マルチモーダル
A
81 件中 61〜81 件目を表示
-
Multi-View Molecular Representation Learning with Hierarchical Graphs and Contextualized Fingerprints
-
Through the Eyes of the Beholder: Biometric and Demographic Conditioning for Multimodal Sexism Detection
-
Don't Count the Edits, Judge by the Outcome Alone: Reward-Based Evaluation for Grammatical Error Correction
-
Adversarial Fashion Confronts Surveillance Norms監視カメラのAI検知を欺く「敵対的ファッション」、小さな産業に成長顔認識やナンバープレート読取のAIカメラへの反発から、物体検知モデルを誤認させる柄の衣服が商品化されつつある。DEF CONで発表されたnoRecognitionは強化学習でYOLO等11モデルを欺く模様を生成し、Cap_ableやUrban Privacyは人を動物や別の顔として誤検知させる服を販売中。専門家は角度・歩容・モデル依存性から効果は限定的と指摘する。
-
MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding
-
Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings
-
MUSE: A Theory-Harnessed Story Engine for Vibe Narrativizing
-
So you want to use OpenRouter?OpenRouter の自動ルーティング、プロバイダ差で応答が変わる問題Simon Willison が、Mohamed Moustafa による OpenRouter 利用上の注意点を紹介。単一エンドポイントで最もコスト効率の良いバックエンドへ自動ルーティングされるが、プロバイダごとに推論ソフトや設定が異なり、同じモデル指定でも挙動が変わる。vision 非対応や reasoning effort の解釈差もあり、provider.only で routing 先を限定できると指摘する。
-
CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation
-
Continue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents
-
Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval
-
Involving before Evolving: A Vision for Trustworthy Enterprise Digital Twin Engineering
-
Anchoring Clinical Events in Time: UID-Preserving Multimodal Reconstruction and Source-Grounded Adjudication
-
Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
-
UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction
-
Tracing and Coordinating Cross-Layer Influence for Multimodal Model Merging
-
Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting
-
Online Video Agent Harness for Long Video Understanding
-
Assisted Spatial Cognition Through Vision-Language Models
-
Calibrated Ambiguity in Multimodal Language Models: Humans reach for cultural references, while models describe the picture
-
Zipbench: Low-Cost Framework for Compressing Comprehensive Benchmarks of Large Language Models