Human children have long been the only entities capable of achieving perfect fluency in human languages. Now, machines are catching up, but they still require an inhuman amount of data to match the linguistic prowess of a child. Large language models (LLMs) like GPT, Claude, and DeepSeek can process hundreds of thousands of words more than a child hears in a year, yet they still fall short of human linguistic abilities. 'The progress recently has been amazing,' says Michael C. Frank, a cognitive scientist at Stanford University. 'But we still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year.' This discrepancy is known as the data efficiency gap, which raises critical questions for both AI research and cognitive science. Scientists are exploring how children can learn so much with so little data, hoping to apply these insights to create more efficient AI models. Source: mittr
Children can learn language with significantly less data than AI models. For example, a preteen raised in a linguistically rich home may hear around 100 million words, while modern LLMs are trained on data equivalent to the entire language of a city over a generation. 'Claude has seen the amount of language that an entire city will experience in one generation,' says Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown University. If all the words used to train an LLM were printed on paper, they would form a stack reaching past the International Space Station. In contrast, a child’s 100 million words would stack up just 20 meters. Scientists believe reverse-engineering how children learn could lead to more data-efficient AI models, which could be useful for training AI on video or creating chatbots for minority languages. Source: mittr
The study highlights the contrast between human language acquisition and AI training. While adults struggle with learning new languages, children typically start producing grammatically correct sentences after hearing around 10 million words. 'It’s just totally miraculous,' says Frank. 'If you train GPT-2 on 30 million words, you get a nonsense generator; you don’t get a kid.' The exact mechanisms by which babies achieve this are still unclear, but researchers know a lot about what children learn and how they use language at different developmental stages. The question of why babies can learn language at all remains one of the most enduring mysteries in cognitive science. Source: mittr