Apple presented a large body of computer-vision research at CVPR 2026. The composition is strongly academic—one official Apple ML source plus four arXiv papers—publishing research results rather than a product, in the form of top-conference papers. The focus is presenting foundational work on image/video recognition and generation together at a venue; the through-line is less a flashy product launch than relatively closed Apple's posture of disclosing results to the research community. Apple-characteristic directions—on-device and efficiency—are also visible. But what's shown is research-stage work—reflection into actual products, generalization of performance, and how it stacks up against others can't be concluded from papers alone.
Apple shares CV work at CVPR 2026
Apple shares CV work at CVPR 2026
Academic (arxiv etc.) 6 ▾
DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation
DynaFLIP: dynamics-aware multimodal pre-training for robot perception
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
Qwen-VLA unifies vision-language-action across tasks, environments and robots
LoMo: Local Modality Substitution for Deeper Vision-Language Fusion
LoMo curbs VLM carrier sensitivity via local modality substitution training
CalArena: A Large-Scale Post-Hoc Calibration Benchmark
CalArena: a large-scale benchmark for post-hoc calibration methods
Unveiling the Visual Counting Bottleneck in Vision-Language Models
Unveiling the visual counting bottleneck in vision-language models
PARCEL: pool-anchored resampling reconciles spatial and query token compression