A former OpenAI researcher has launched a startup focused on creating high-quality training data, arguing that simply scaling up large language models won't lead to true generalization. Andrew Ho, who left OpenAI after eight months, believes that current AI systems lack the necessary training data to perform reliably in economically valuable tasks. He claims that most skills relevant to real-world applications are underrepresented in existing datasets, which limits the models' ability to generalize effectively. Ho's skepticism is echoed by researchers at the University of Cambridge and Google Deepmind, who note that AI systems are becoming more specialized rather than more versatile. Source: thedecoder

Ho's startup is targeting two key areas: datasets for complex scientific analyses in bioinformatics and datasets for everyday lab work. In bioinformatics, even current models like GPT-5.6 Sol achieve only about a 30 percent success rate, a challenge Ho previously faced at OpenAI. His second focus is on datasets for routine lab tasks, such as when researchers submit photos of experiments to AI models for evaluation. Other areas like chemistry, materials science, healthcare, and broader knowledge work are planned for future development. Ho argues that the current trend of AI development is moving toward extreme specialization, where a few capabilities spike while core skills stagnate or shrink. Source: thedecoder

Cambridge researcher Adam Hunt shares Ho's skepticism, describing how his view of large language models has shifted from optimistic to increasingly pessimistic. Hunt argues that the latest models are not becoming more versatile but more specialized. While programming and complex math capabilities continue to improve, areas like language quality and simple logic are stagnating or even declining. This aligns with Ho's observation of uneven performance across different tasks. Hunt attributes this trend to the nature of reinforcement learning, which works well in domains with clear reward signals and complete training data. Source: thedecoder