Multimodal A
Showing 91–120 of 120
-
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
-
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
-
Face De-Identification: A Domain-Centric Survey from Capture to Processing
-
Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA
-
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
-
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
-
RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement
-
Detecting CSAM Text-to-Image LoRAs From Weights
-
Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography
-
Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction
-
Cognivia: A Cognitive Behavioral Therapy Copilot for Evidence-Based Mental Healthcare
-
DynaBridge: Dynamic Summary-Guided Cross-Task Multimodal Fusion for DASS-Structured Mental Health Assessment
-
MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice
-
Contextual Deconvolution for Variance-Stable Demand Sensing: Kernel-Modulated Operators in Promotional Retail
-
Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications
-
Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs
-
Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact
-
OrthKD: Extracting Generalized Clinical Knowledge from Heterogeneous Teachers for Lightweight Deployment
-
Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering
-
Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control
-
CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
-
MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios
-
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
-
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
-
KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability
-
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
-
ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams
-
Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
-
Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair