Human children can master their native languages flawlessly with a fraction of the input required by large language models (LLMs), which rely on massive datasets—a phenomenon cognitive science refers to as the "data efficiency gap." According to scientists from Stanford and Georgetown universities, while AI systems like ChatGPT and Claude are trained on billions or even trillions of words, a human child can achieve the exact same level of fluency with vastly less exposure to language.
The Data Efficiency Gap: The Difference Between the Human Mind and AI
The recent evolution of AI models has largely relied on a strategy of "bigger models and more data." For instance, Meta’s Llama 3.1 model was trained on 15 trillion word-like tokens, and experts forecast that future frontier models will require multiples of that amount. In contrast, a child growing up in a language-rich environment is exposed to only around 100 million words by the time they reach adolescence.
To put this scale into perspective: if the training data of a modern LLM were printed out on paper, it would form a stack reaching past the International Space Station, whereas the words heard by a child would create a pile just 20 meters high. With raw internet data sources at risk of exhaustion in the coming decades, AI architects are turning to reverse-engineering human learning mechanisms.
Implications of the Research for AI and Cognitive Science
Unlocking the secret behind children's superior learning ability could offer critical advantages for future AI models. Scientists note that developing AI systems capable of mimicking the human brain's data efficiency could be transformative—enabling the creation of chatbots that support low-resource minority languages or training AI more effectively using diverse data types, such as video. At the same time, this research aims to shed light on decades-old cognitive questions, such as the role of innate language instincts in children's mental development.
Frequently Asked Questions
How does the risk of exhausting internet training data impact AI development processes?
Experts predict that if current growth trends continue, easily accessible raw internet data could run out by the 2030s. This situation is forcing companies to move away from gathering more data and instead mimic children's learning models to build smarter systems using far less data.
In which practical areas could adapting children's data efficiency to AI break new ground?
This approach could enable the creation of advanced AI assistants, particularly for minority languages that lack massive datasets, and pave the way for models to process complex data—such as video—with significantly lower costs and energy consumption.
*This news article was prepared based on data published by MIT Tech Review — AI.
💬 Comments
No comments yet. Be the first!
You must be logged in to comment.
🔑 Log In