Research reframed factual errors in generative AI as a retrieval failure rather than a knowledge gap, alongside a cluster of evaluation and interpretability papers.
Google Research argued that the limit on parametric factuality is not missing knowledge (empty shelves) but the inability to retrieve it (lost keys). Four of the five articles this week were arXiv papers, covering six years of interpretability-to-control work, LLM-generated text detection, and QA reliability.
The emphasis shifts from adding training data toward improving access to knowledge already in the weights. This is research-led, and whether it reaches implementations or products is not established.
Whether recall-oriented techniques appear as inference-time methods, and whether factuality benchmarks adopt the distinction.