Multimodal
A
Showing 61–81 of 81
-
Multi-View Molecular Representation Learning with Hierarchical Graphs and Contextualized Fingerprints
-
Through the Eyes of the Beholder: Biometric and Demographic Conditioning for Multimodal Sexism Detection
-
Don't Count the Edits, Judge by the Outcome Alone: Reward-Based Evaluation for Grammatical Error Correction
-
Adversarial Fashion Confronts Surveillance NormsAdversarial fashion grows into an industry to foil AI surveillance camerasBacklash against face- and plate-reading AI cameras is fueling clothing that confuses object detectors. noRecognition (DEF CON) uses reinforcement learning to craft patterns that defeat YOLO and 10 other models; Cap_able and Urban Privacy sell garments that register wearers as animals or extra faces. Experts warn angles, gait, and model-specific tuning limit the effect.
-
MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding
-
Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings
-
MUSE: A Theory-Harnessed Story Engine for Vibe Narrativizing
-
So you want to use OpenRouter?OpenRouter's auto-routing can vary model behavior by providerSimon Willison flags Mohamed Moustafa's caveats on OpenRouter: its single endpoint auto-routes to the cheapest backend, but providers differ in serving software and settings, so the same model can behave differently. provider.only pins routing.
-
CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation
-
Continue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents
-
Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval
-
Involving before Evolving: A Vision for Trustworthy Enterprise Digital Twin Engineering
-
Anchoring Clinical Events in Time: UID-Preserving Multimodal Reconstruction and Source-Grounded Adjudication
-
Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
-
UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction
-
Tracing and Coordinating Cross-Layer Influence for Multimodal Model Merging
-
Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting
-
Online Video Agent Harness for Long Video Understanding
-
Assisted Spatial Cognition Through Vision-Language Models
-
Calibrated Ambiguity in Multimodal Language Models: Humans reach for cultural references, while models describe the picture
-
Zipbench: Low-Cost Framework for Compressing Comprehensive Benchmarks of Large Language Models