Recent
- Research noteAI-2026-0223
A budget for how much the teacher says
Self-distillation always aims at the full teacher. Constraining how much teacher information enters at each token improved specialisation in seven of eight settings.
3 minSource: arXiv
- Research noteAI-2026-0222
A scoring rule that punishes guessing
Volume-based accuracy rewards a retrieval system for answering everything. A penalty-aware framework separates the three ways it fails.
3 minSource: arXiv
- AnalysisAI-2026-0221
A hundred million words against fifteen trillion
Children reach fluency on about 100 million words. Llama 3.1 took 15 trillion tokens, and nobody can yet say what closes the gap.
3 minSource: MIT Technology Review
- Research noteAI-2026-0220
Both agent-written papers were rejected
Princeton researchers gave agents six days, $3,000 in credits and unpublished research questions. The original authors turned down the results.
2 minSource: MIT Technology Review
- News briefAI-2026-0219
Gemini 3.7 Flash lands three weeks after 3.6
Google's workhorse model gains sharply on coding and agent benchmarks, with introductory pricing held to the end of the year.
2 minSource: Google
- AnalysisAI-2026-0218
Context windows stopped being the constraint
Every lab now ships a million tokens or more. The bottleneck moved to retrieval quality, and most teams have not noticed.
9 min
- BenchmarkAI-2026-0215
Open weights close the gap on reasoning, not on reliability
Scores converge at the top of the leaderboard. Variance across repeated runs tells a different story.
6 min
- Research noteAI-2026-0211
The quiet standardisation of tool schemas
Four vendors, one shape. Interoperability arrived without an announcement.
7 min
- AnalysisAI-2026-0207
Distillation is doing more work than scale
The interesting releases this quarter were small models trained on the output of large ones.
8 min
- Field reportAI-2026-0203
Evaluation sets are rotting faster than teams replace them
Contamination is no longer an edge case. It is the base rate.
5 min
- Market briefAI-2026-0198
Inference providers compete on latency now, not price
Per-token pricing converged. Time to first token did not.
4 min