Multimodal A
Showing 31–60 of 124
-
DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation
-
ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programs
-
Same Graph Cross-Task Transfer in GNNs: Protocols and Predictors
-
A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
-
When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence
-
Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
-
HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection
-
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
-
Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners
-
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaborationDeepMind's Gemini Robotics ER 2 adds video understanding, multi-robot teamworkDeepMind introduced Gemini Robotics ER 2, which helps robots reason, collaborate, and solve real-world tasks. The company calls it a step change in video understanding, task orchestration, and multi-robot collaboration for embodied AI.
-
From Japan, Products the World Will Use: An Interview with Sakana AI's Head of Product DevelopmentInterview: Sakana AI's product chief on Japan-born global productsAn interview with Sakana AI's Head of Product Development on building products from Japan that the world will use. The Q&A covers the company's product philosophy and ambitions, offering a look at the strategy of a leading Japanese AI startup.
-
PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?
-
ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding
-
Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation
-
Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory
-
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
-
Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation
-
AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach
-
Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models
-
RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning
-
OPLD: On-Policy Latent Distillation for Multimodal Reasoning
-
Group-Reflective Self-Distillation for Agentic Reinforcement Learning
-
Flux-OPD: On-Policy Distillation with Evolving Contexts
-
TAPO: Transition-Aware Policy Optimization for LLM Agents
-
AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification
-
Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis
-
Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning
-
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
-
DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation
-
The coolest use for the Vision ProApollo dev Christian Selig shares his favorite use for the Vision ProiOS developer Christian Selig (creator of Apollo for Reddit) blogs about what he calls the coolest use for Apple's Vision Pro. The URL slug (vision-pro-house) suggests a home- or spatial-visualization use case, but the raw excerpt is empty, so the specific application, steps, and features are unconfirmed from the source text.