Skip to content
BenchmarkAI-2026-0215

Open weights close the gap on reasoning, not on reliability

Scores converge at the top of the leaderboard. Variance across repeated runs tells a different story.

6 minMain AI Hub

On single-pass reasoning benchmarks the distance between the best open-weight models and the best closed ones is now within noise. Run the same prompt twenty times and the picture changes: closed models cluster, open ones spread.

For a chat product that spread is tolerable. For anything that writes to a database, it is the whole problem. We ran a consistency evaluation across nine models and found that leaderboard rank predicted mean accuracy well and predicted worst-case behaviour badly.

The practical advice is unchanged and unglamorous: score the variance, not just the mean, and score it on your own task.

Read next

Across the network

Desks that share a zone with this one on the BITBRIEF coverage map.

Terms defined