Distillation is doing more work than scale
The interesting releases this quarter were small models trained on the output of large ones.
8 minMain AI Hub
Parameter counts stopped being the headline. The releases that changed deployment decisions were mid-size models trained largely on generated data from larger systems, reaching accuracy that would have required an order of magnitude more parameters two years ago.
This shifts cost from inference to training, which suits anyone serving high volume and suits nobody serving a long tail of low-traffic tasks. It also makes the provenance of training data harder to establish, which matters more than it did when licensing terms were simpler.