Apple has swapped the yardstick for video descriptions: instead of matching a written-by-hand reference, it asks whether you could answer questions about the video from the description alone.
Apple published CapQuiz in September, a way of scoring the text that describes what happens in a video. The old approach compared the generated text against a human-written reference and counted overlapping wording, which penalised good descriptions for saying the same thing differently. CapQuiz drops the reference. It builds multiple-choice questions from the video, has people verify them, and then checks whether the description alone carries enough to answer them.
Change the scorecard and you change what writers aim for. Under reference matching, safe and formulaic descriptions won, because rewording cost points even when the content was right. Scoring by whether the questions can be answered shifts the weight onto being accurate and leaving nothing important out. The paper reports that this tracks human judgement more closely than existing measures. The comparison comes from the authors' own experiments, though: no independent replication, and no adoption by anyone else's evaluation setup, has surfaced yet.
The first fork is whether the questions and videos are released so other groups can measure on the same ground. The second is whether model rankings move once the reference is gone, which is the real test of the new measure.