Skip to content
AnalysisAI-2026-0221

A hundred million words against fifteen trillion

Children reach fluency on about 100 million words. Llama 3.1 took 15 trillion tokens, and nobody can yet say what closes the gap.

3 minMain AI Hub

Elise Cutts set out the data efficiency gap on 24 August, drawing on the BabyLM community and on researchers at Stanford, Princeton, Berkeley, Harvard, UC San Diego and Georgetown. The comparison is stark. A child hears roughly 100 million words by the preteen years and about 300 million by twenty. Llama 3.1 was trained on 15 trillion tokens, and frontier models use an order of magnitude more than that.

The controlled version of the question

BabyLM exists to make the comparison fair: models are trained on a 100 million word budget, with a toddler track at 10 million. Held to that budget, GPT-2 is described as a nonsense generator, while a child on similar exposure produces grammatical sentences. GPT-BERT, the strongest entrant, beat models trained on roughly fifteen thousand times more data on the benchmark.

What has been tried and has not worked

  • Curriculum learning, ordering the data the way a child encounters it, underperformed expectations in the competitions.
  • Adding vision has not helped BabyLM models, despite the obvious argument that children learn words attached to things.
  • Brenden Lake trained a model on 61 hours of headcam footage from a single child and did recover object-word associations, which shows the signal is there without showing how to use it at scale.

Uri Hasson has recorded 1,000 days across seventeen children, twelve hours a day, which will eventually give the field a corpus that is genuinely comparable rather than merely small.

Why this matters outside the lab

The live hypothesis is that children are not passive receivers of a corpus. They experiment, pursue goals and get corrected, an active loop that no static training set contains. If that is the missing ingredient, then data efficiency is not a scaling problem with a scaling answer, and the industry practice of buying the gap down with more tokens has a floor. Nobody in the piece claims to know that it is. That is the honest state of it.

Retold from MIT Technology Review. This is a summary in our own words; follow the link for the original reporting.

Read next

Across the network

Desks that share a zone with this one on the BITBRIEF coverage map.

Terms defined