Inference keeps getting faster and cheaper per call, yet the bills keep growing - both claims landed in the same week.
LiquidAI reported up to 3.2x faster inference with LFM2.5-DSpark, while a trade report unpacked the paradox that falling per-model prices sit alongside rising total AI spend. Of the five items grouped here, four came from vendors and official blogs and one from trade press; academic and community reaction is still thin.
Efficiency does not translate into savings on its own. Cheaper calls invite more calls, and always-on uses like generative recommenders push the total volume up. The remaining two items - a piece on alignment data and a study of multilingual transfer - sit on a different axis and were not announced as part of one current.
Whether the 3.2x figure reproduces outside the vendor's own setup, and whether anyone reports total spend rather than unit price. Disclosures that separate per-call cost from usage volume would help.