Inference × Developer Tools

Apple grades video captions with MCQs

Apple grades video captions with MCQs

✎ Story body

Apple has swapped the yardstick for video descriptions: instead of matching a written-by-hand reference, it asks whether you could answer questions about the video from the description alone.

What happened

Apple published CapQuiz in September, a way of scoring the text that describes what happens in a video. The old approach compared the generated text against a human-written reference and counted overlapping wording, which penalised good descriptions for saying the same thing differently. CapQuiz drops the reference. It builds multiple-choice questions from the video, has people verify them, and then checks whether the description alone carries enough to answer them.

Why it matters

Change the scorecard and you change what writers aim for. Under reference matching, safe and formulaic descriptions won, because rewording cost points even when the content was right. Scoring by whether the questions can be answered shifts the weight onto being accurate and leaving nothing important out. The paper reports that this tracks human judgement more closely than existing measures. The comparison comes from the authors' own experiments, though: no independent replication, and no adoption by anyone else's evaluation setup, has surfaced yet.

What to watch

The first fork is whether the questions and videos are released so other groups can measure on the same ground. The second is whether model rankings move once the reference is gone, which is the real test of the new measure.

▲ Official & Press
Official

Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

Apple Machine Learning Research ・ 2026-09-11 ・ 📌

Apple proposes evaluating video captions via multiple-choice QA

Community

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

Hacker News (Front Page) ・ 2026-09-10

Cognition launches SWE-2, matching Fable 5.1 at 64% lower cost

Community

Quoting Calif Research

Simon Willison's Weblog ・ 2026-09-10

Researchers demo WeWorm, a zero-click worm spreading via WeChat calls

Academic (arxiv etc.) 15 ▾
Academic

GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-10

Academic

Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-10

Academic

Nuha-Speech: Building General-Purpose Arabic Speech-LLMs

arXiv cs.CL (Computation and Language) ・ 2026-09-10

Academic

The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-10

Academic

From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-10

Academic

Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models

arXiv cs.CL (Computation and Language) ・ 2026-09-10

Academic

A Unified Per-Token Gating Family for On-Policy Distillation: FKL/RKL Mixing with Multi-Channel and Bias Coefficients

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-10

Academic

RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety

arXiv cs.CL (Computation and Language) ・ 2026-09-10

Academic

The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challenge

arXiv cs.CL (Computation and Language) ・ 2026-09-10

Academic

When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-10

Academic

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

arXiv cs.CL (Computation and Language) ・ 2026-09-10

Academic

A Dataset and Model for Imputing Water Surface Elevation on a Large and Extremely Sparse Spatiotemporal Graph

arXiv cs.LG (Machine Learning) ・ 2026-09-10

Academic

ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-10

Academic

VikingRAG: Accurate and Token-efficient Retrieval-augmented Generation over Structured Documents

arXiv cs.CL (Computation and Language) ・ 2026-09-10

Academic

SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ

arXiv cs.CL (Computation and Language) ・ 2026-09-10

← Story Archive