マルチモーダル

A
81 件中 61〜81 件目を表示
  • arXiv cs.LG (Machine Learning) · EN 学習・ファインチューニング
    Multi-View Molecular Representation Learning with Hierarchical Graphs and Contextualized Fingerprints
    埋め込み (Embeddings) ファインチューニング 検索拡張生成 (RAG)
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.CL (Computation and Language) · EN 推論・効率化
    Through the Eyes of the Beholder: Biometric and Demographic Conditioning for Multimodal Sexism Detection
    推論 (Inference) ニューラルネットワーク
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN マルチモーダル
    Don't Count the Edits, Judge by the Outcome Alone: Reward-Based Evaluation for Grammatical Error Correction
    ニューラルネットワーク 強化学習
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • IEEE Spectrum (AI section) · EN マルチモーダル
    Adversarial Fashion Confronts Surveillance Norms
    監視カメラのAI検知を欺く「敵対的ファッション」、小さな産業に成長
    強化学習 ソフトウェア工学
    顔認識やナンバープレート読取のAIカメラへの反発から、物体検知モデルを誤認させる柄の衣服が商品化されつつある。DEF CONで発表されたnoRecognitionは強化学習でYOLO等11モデルを欺く模様を生成し、Cap_ableやUrban Privacyは人を動物や別の顔として誤検知させる服を販売中。専門家は角度・歩容・モデル依存性から効果は限定的と指摘する。
    元記事を読む (IEEE Spectrum (AI section)) ↗
  • arXiv cs.CL (Computation and Language) · EN マルチモーダル
    MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding
    推論 (Inference) 機械学習 ニューラルネットワーク 検索拡張生成 (RAG)
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN 推論・効率化
    Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings
    埋め込み (Embeddings) 推論 (Inference) 強化学習
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN マルチモーダル
    MUSE: A Theory-Harnessed Story Engine for Vibe Narrativizing
    ニューラルネットワーク
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • Simon Willison's Weblog · EN 開発者ツール
    So you want to use OpenRouter?
    OpenRouter の自動ルーティング、プロバイダ差で応答が変わる問題
    深層学習 人間のフィードバックによる強化学習 (RLHF)
    Simon Willison が、Mohamed Moustafa による OpenRouter 利用上の注意点を紹介。単一エンドポイントで最もコスト効率の良いバックエンドへ自動ルーティングされるが、プロバイダごとに推論ソフトや設定が異なり、同じモデル指定でも挙動が変わる。vision 非対応や reasoning effort の解釈差もあり、provider.only で routing 先を限定できると指摘する。
    元記事を読む (Simon Willison's Weblog) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 新モデル・リリース
    CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation
    検索拡張生成 (RAG) 強化学習
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN 新モデル・リリース
    Continue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents
    AI エージェント ニューラルネットワーク 強化学習 音声処理
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 開発者ツール
    Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval
    AI エージェント 機械学習
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN マルチモーダル
    Involving before Evolving: A Vision for Trustworthy Enterprise Digital Twin Engineering
    深層学習 強化学習
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN マルチモーダル
    Anchoring Clinical Events in Time: UID-Preserving Multimodal Reconstruction and Source-Grounded Adjudication
    DeepSeek ニューラルネットワーク
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN マルチモーダル
    Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
    コンピュータビジョン ロボティクス
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 新モデル・リリース
    UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction
    埋め込み (Embeddings) ファインチューニング 強化学習 Transformer
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN マルチモーダル
    Tracing and Coordinating Cross-Layer Influence for Multimodal Model Merging
    強化学習
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 新モデル・リリース
    Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN マルチモーダル
    Online Video Agent Harness for Long Video Understanding
    AI エージェント 深層学習 ニューラルネットワーク ソフトウェア工学 音声処理
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN マルチモーダル
    Assisted Spatial Cognition Through Vision-Language Models
    コンピュータビジョン 強化学習
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN マルチモーダル
    Calibrated Ambiguity in Multimodal Language Models: Humans reach for cultural references, while models describe the picture
    強化学習
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN マルチモーダル
    Zipbench: Low-Cost Framework for Compressing Comprehensive Benchmarks of Large Language Models
    ニューラルネットワーク 検索拡張生成 (RAG)
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗