Open weights close the gap on reasoning, not on reliability
Scores converge at the top of the leaderboard. Variance across repeated runs tells a different story.
6 minMain AI Hub
On single-pass reasoning benchmarks the distance between the best open-weight models and the best closed ones is now within noise. Run the same prompt twenty times and the picture changes: closed models cluster, open ones spread.
For a chat product that spread is tolerable. For anything that writes to a database, it is the whole problem. We ran a consistency evaluation across nine models and found that leaderboard rank predicted mean accuracy well and predicted worst-case behaviour badly.
The practical advice is unchanged and unglamorous: score the variance, not just the mean, and score it on your own task.