Evaluation sets are rotting faster than teams replace them
Contamination is no longer an edge case. It is the base rate.
5 minMain AI Hub
Public benchmarks leak into training corpora within months of release. Teams that rely on them to choose models are increasingly measuring memorisation rather than capability.
The teams doing this well maintain a private, versioned set drawn from their own traffic, refresh a portion of it each quarter, and never publish it. That is more work than downloading a benchmark, and it is the only approach that has held up.