マルチモーダル A
124 件中 31〜60 件目を表示
-
DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation
-
ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programs
-
Same Graph Cross-Task Transfer in GNNs: Protocols and Predictors
-
A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
-
When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence
-
Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
-
HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection
-
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
-
Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners
-
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaborationDeepMind、映像理解と多ロボット連携のGemini Robotics ER 2を発表DeepMindは、ロボット向けモデルGemini Robotics ER 2を発表した。映像理解、タスクの分解・調整、複数ロボットの協調を強化し、ロボットが現実世界の課題を推論しながら協力して解決できるようにする段階的な進歩と位置づける。
-
From Japan, Products the World Will Use: An Interview with Sakana AI's Head of Product DevelopmentSakana AI製品開発責任者、世界で使われる日本発プロダクトを語るSakana AIの製品開発責任者へのインタビュー記事。日本発で世界に使われるプロダクトを生み出す狙いや、同社の製品開発の考え方が語られている。国内AIスタートアップの製品戦略を示す内容となっている。
-
PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?
-
ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding
-
Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation
-
Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory
-
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
-
Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation
-
AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach
-
Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models
-
RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning
-
OPLD: On-Policy Latent Distillation for Multimodal Reasoning
-
Group-Reflective Self-Distillation for Agentic Reinforcement Learning
-
Flux-OPD: On-Policy Distillation with Evolving Contexts
-
TAPO: Transition-Aware Policy Optimization for LLM Agents
-
AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification
-
Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis
-
Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning
-
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
-
DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation
-
The coolest use for the Vision ProiOS開発者C.Selig、Vision Proの最良の使い道を個人ブログで紹介Apollo(Redditクライアント)で知られるiOS開発者Christian Selig氏が、個人ブログでApple Vision Proの「最も優れた使い道」を紹介する記事。URLスラッグ(vision-pro-house)から住宅・空間の可視化に関する活用例とみられるが、raw_excerptが空のため具体的な用途・手順・機能は本文未取得で確認不可。attention 0.5・categories=[multimodal]の個人系トピックとしてexport対象となった。