Where inference runs became the question this week, moving in three directions at once: in-region, at the edge, and on the desk.
Mistral announced in-region inference and open models as European sovereign-AI infrastructure. The same week, NVIDIA shipped JetPack 7.2.1 for edge devices, Meta released Muse Glimmer, a local-first open model under Apache 2.0, and two community reports covered fast LLM inference on Apple Silicon.
Sourcing is spread across two first-party announcements, one trade report and two community posts, meaning vendors and practitioners are moving at the same time. The motivations differ (regulation, cost, latency), but the direction is consistent: inference is being pulled out of centralized clouds.
Whether in-region inference hardens into a European procurement requirement, and whether local-inference performance claims reproduce and carry real workloads.