A new architecture for multimodal generation and benchmarks probing cultural and linguistic coverage landed in the same week.
Apple published STARFlow2, which bridges language models and normalizing flows for unified multimodal generation. The same week's arXiv listings included a Cultural Moment Benchmark for video reasoning in Southeast Asia, work on multimodal humor comprehension, and a bibliometric study of Arabic NLP. One official post against four academic papers.
Generation methods and non-English evaluation infrastructure are advancing in parallel. That said, the cluster is grouped under developer tools and its five items span quite different subjects, so a single thread here is not established.
Whether STARFlow2 ships reproducible code, and whether the new cultural and language benchmarks are folded into mainstream multimodal evaluation suites.