マルチモーダル
A
81 件中 1〜30 件目を表示
-
PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
-
Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation
-
rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference
-
MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education
-
Rethinking Robot Safety in the Age of AIIEEE Spectrum、物理 AI 時代のロボット安全はセキュリティ問題と指摘IEEE Spectrum が VicOne 提供記事で、マルチモーダルセンサーで知覚し AI で文脈を解釈して動く現代のロボットは、安全性が判断を導くデータの完全性に依存すると論じる。従来は「故障したとき安全か」を問うたが、物理 AI では「何も壊れていないのに攻撃者が知覚や判断を書き換えたとき安全か」が問われる。直接制御なしに挙動を左右できる研究例もあり、既存の安全評価では捉えきれないとする。
-
Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
-
GrainSpeech: Less Context, More Detail for Compact Speech Synthesis
-
ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts
-
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
-
RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection
-
Generalist-Specialist Mixture-of-Experts for Rare Pathology Detection in Multimodal Imaging
-
VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge
-
Learning Array Signal Topologies as Conditional Neural Manifolds
-
Hypothesis-Driven Autonomous Materials Synthesis with Multimodal LLM Agents
-
Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces
-
ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
-
Divide and Conquer: Mixture-of-Bottleneck Experts in Informative Ordinal Space for Video-based Multimodal Sentiment Analysis
-
Emotion Experience, Expression, and Perception: Emotion Analysis on Multimodal Social Media Posts
-
Gemini Live audioSimon Willison、Gemini 3.8 Live を試すブラウザ UI をライブラリ無しで公開Google が音声対話モデル Gemini 3.8 Live と 3.8 Live Extended Thinking を公開したのを受け、Simon Willison がブラウザから両モデルを試せる Web UI を作成した。モデルと音声プリセットを選んで会話でき、モデルの発話中に割り込むこともできる。実装はライブラリを使わず、双方向 WebSocket と Web Audio API だけで録音・再生を処理する。
-
LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs
-
ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation
-
Tables Decoded: DELTA for Structure, TARQA for Understanding
-
CareMirror: Bringing Caregiver Wellbeing into the Dementia Care Ecosystem
-
Type-IV Code Clone Detection via Layer-Wise Non-Contrastive Representation Learning
-
Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs
-
From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic-Geometric Contracts
-
Towards Detecting AI-Assisted Responses in Online Surveys
-
Goal-oriented probabilistic forecasting for dynamic PRB allocation in 5G networks
-
FROD: Feature Matching Residual Denoising Oracle Bone Decipher
-
FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence