Context windows stopped being the constraint
Every lab now ships a million tokens or more. The bottleneck moved to retrieval quality, and most teams have not noticed.
9 minMain AI Hub
Two years ago the standard complaint about production language models was that they could not hold enough of the problem in view. A contract, a codebase, a quarter of support tickets — all of it had to be cut down before the model saw any of it. That complaint is now largely obsolete. Every frontier lab ships a context window measured in the hundreds of thousands of tokens, and several exceed a million.
What did not change is that a model reading a million tokens is not the same as a model understanding a million tokens. Retrieval accuracy degrades in the middle of long inputs, and it degrades unevenly across positions. Teams that removed their retrieval layer because the window grew have quietly added it back.
Where the cost actually sits
The economics also moved. Filling a large window on every request is expensive and slow, and prompt caching only helps when the prefix is genuinely stable. Three patterns have settled out of the last year of production work:
- Stable prefix, variable suffix — cache the corpus, vary the question. Cheapest, but only works when the corpus is small enough to fit.
- Retrieval into a medium window — still the default for anything above a few hundred documents. Quality now depends almost entirely on chunking and reranking.
- Agentic search — the model issues its own queries against a store instead of receiving a pre-assembled context. Slower per answer, markedly better on questions the retriever would have got wrong.
What this means for architecture
If you are building now, the useful question is no longer how much context you can afford. It is how much of what you send is load-bearing. A well-ranked ten thousand tokens beats a poorly assembled four hundred thousand on nearly every evaluation we have seen, and it costs a fraction as much.
We cut the context we send by 90% and accuracy went up. That result was not intuitive to anyone on the team.
The teams getting good results treat retrieval as a ranking problem with a measurable target, not as plumbing that feeds the model. They evaluate the retriever separately from the generator, and they keep a held-out set that reflects the questions users actually ask rather than the questions that are easy to score.
What to watch
Position-dependent accuracy is the number to track through the rest of the year. If it flattens across the window, the calculus changes again and a lot of retrieval infrastructure becomes optional. It has not flattened yet.