推論 (Inference) × 開発者ツール

Apple、動画字幕の品質を選択式QAで評価

Apple、動画字幕の品質を選択式QAで評価

✎ ストーリー本文

動画につける説明文の良し悪しを、Apple は正解文との一致ではなく「その文を読んで中身を当てられるか」で測り直した。

何が起きたか

Apple が 9 月、動画の説明文を採点する方法として CapQuiz を公開した。従来は人が用意した正解文とどれだけ言葉が重なるかで採点していたが、同じ動画でも説明の仕方は何通りもあり、言い回しが違うだけで良い説明文が低く扱われてしまう。CapQuiz は正解文を使わず、動画から作った選択式の設問を人手で確かめたうえで、その設問に説明文だけを読んで答えられるかを見る。

なぜ重要か

採点方法が変わると、作り手が何を目指すかが変わる。正解文との一致で測っていた間は、内容が合っていても言葉を選び直しただけで減点されうるため、無難で似通った説明文が有利だった。設問に答えられるかで測れば、評価されるのは情報が正しく、抜けなく入っているかになる。論文では、この測り方のほうが人間の評価と近い結果になったとしている。※比較は論文内の実験にとどまり、他の研究者による追試や、他社の評価基盤への採用はまだ確認できない。

次に何を見るか

設問と動画のデータが公開され、他の研究チームが同じ土俵で測り直せるようになるかどうかが、最初の分かれ目になる。既存の採点方法で上位だったモデルの順位が入れ替わるかどうかも、この測り方の実力を映す。

▲ 公式・報道
公式

Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

Apple Machine Learning Research ・ 2026-09-11 ・ 📌

Apple、動画キャプションの品質を多肢選択QAで測る評価法を提案

コミュニティ

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

Hacker News (Front Page) ・ 2026-09-10

Cognition、コーディングモデル SWE-2 公開、Fable 5.1 並を 64% 低価格で

コミュニティ

Quoting Calif Research

Simon Willison's Weblog ・ 2026-09-10

WeChat 通話経由のゼロクリック・ワーム、研究チームがデモを公開

学術(arxiv ほか) 15本 ▾
学術

GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-10

学術

Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-10

学術

Nuha-Speech: Building General-Purpose Arabic Speech-LLMs

arXiv cs.CL (Computation and Language) ・ 2026-09-10

学術

The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-10

学術

From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-10

学術

Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models

arXiv cs.CL (Computation and Language) ・ 2026-09-10

学術

A Unified Per-Token Gating Family for On-Policy Distillation: FKL/RKL Mixing with Multi-Channel and Bias Coefficients

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-10

学術

RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety

arXiv cs.CL (Computation and Language) ・ 2026-09-10

学術

The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challenge

arXiv cs.CL (Computation and Language) ・ 2026-09-10

学術

When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-10

学術

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

arXiv cs.CL (Computation and Language) ・ 2026-09-10

学術

A Dataset and Model for Imputing Water Surface Elevation on a Large and Extremely Sparse Spatiotemporal Graph

arXiv cs.LG (Machine Learning) ・ 2026-09-10

学術

ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps

arXiv cs.AI (Artificial Intelligence) ・ 2026-09-10

学術

VikingRAG: Accurate and Token-efficient Retrieval-augmented Generation over Structured Documents

arXiv cs.CL (Computation and Language) ・ 2026-09-10

学術

SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ

arXiv cs.CL (Computation and Language) ・ 2026-09-10

← ストーリー アーカイブ