Multimodal
A
Showing 1–30 of 81
-
PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
-
Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation
-
rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference
-
MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education
-
Rethinking Robot Safety in the Age of AIIEEE Spectrum: physical AI turns robot safety into a security problemIn a VicOne-sponsored piece, IEEE Spectrum argues robot safety now hinges on the integrity of the data guiding decisions. Classic assessments ask if a machine stays safe when something fails; physical AI asks if it stays safe when an attacker alters what it perceives while nothing looks broken.
-
Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
-
GrainSpeech: Less Context, More Detail for Compact Speech Synthesis
-
ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts
-
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
-
RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection
-
Generalist-Specialist Mixture-of-Experts for Rare Pathology Detection in Multimodal Imaging
-
VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge
-
Learning Array Signal Topologies as Conditional Neural Manifolds
-
Hypothesis-Driven Autonomous Materials Synthesis with Multimodal LLM Agents
-
Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces
-
ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
-
Divide and Conquer: Mixture-of-Bottleneck Experts in Informative Ordinal Space for Video-based Multimodal Sentiment Analysis
-
Emotion Experience, Expression, and Perception: Emotion Analysis on Multimodal Social Media Posts
-
Gemini Live audioSimon Willison builds a library-free browser UI for Gemini 3.8 LiveGoogle shipped two speech-to-speech models, Gemini 3.8 Live and 3.8 Live Extended Thinking. Simon Willison built a browser tool to try them: pick a model and voice preset, add a system prompt, interrupt mid-sentence. No libraries — just a WebSocket and Web Audio.
-
LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs
-
ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation
-
Tables Decoded: DELTA for Structure, TARQA for Understanding
-
CareMirror: Bringing Caregiver Wellbeing into the Dementia Care Ecosystem
-
Type-IV Code Clone Detection via Layer-Wise Non-Contrastive Representation Learning
-
Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs
-
From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic-Geometric Contracts
-
Towards Detecting AI-Assisted Responses in Online Surveys
-
Goal-oriented probabilistic forecasting for dynamic PRB allocation in 5G networks
-
FROD: Feature Matching Residual Denoising Oracle Bone Decipher
-
FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence