Skip to content
Field reportAI-2026-0203

Evaluation sets are rotting faster than teams replace them

Contamination is no longer an edge case. It is the base rate.

5 minMain AI Hub

Public benchmarks leak into training corpora within months of release. Teams that rely on them to choose models are increasingly measuring memorisation rather than capability.

The teams doing this well maintain a private, versioned set drawn from their own traffic, refresh a portion of it each quarter, and never publish it. That is more work than downloading a benchmark, and it is the only approach that has held up.

Read next

Across the network

Desks that share a zone with this one on the BITBRIEF coverage map.

Terms defined