A vendor case for small models and a set of papers on post-training design and evaluation landed on the same day.
Cohere argued that small AI models can deliver outsized impact for enterprises, and the same day brought papers on post-training science for supervised fine-tuning, SFT-RL annotation budget allocation, the mechanics of LLM-as-a-judge in summarization evaluation, and hypothesis-guided search for research agents — one corporate blog against four arXiv entries.
The axis has moved from scaling up to allocating a fixed budget well. Whether a small model is viable is decided together with how post-training is designed and how its output is judged, and this week's papers work exactly that allocation-and-evaluation side. Vendor claim and research accumulation arrived together, but evidence of actual enterprise adoption is not yet visible.
Whether reproducible guidance emerges for splitting budget between SFT and RL, and whether validation of LLM-as-a-judge reaches standardized evaluation. Also worth tracking is growth in reported small-model deployments.