マルチモーダル

A
81 件中 1〜30 件目を表示
  • arXiv cs.CL (Computation and Language) · EN マルチモーダル
    PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
    コンピュータビジョン ニューラルネットワーク 検索拡張生成 (RAG) 強化学習
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN マルチモーダル
    Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation
    ニューラルネットワーク 検索拡張生成 (RAG)
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 推論・効率化
    rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference
    コンピュータビジョン 推論 (Inference) 強化学習
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN マルチモーダル
    MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education
    コンピュータビジョン ニューラルネットワーク 検索拡張生成 (RAG) 強化学習
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • IEEE Spectrum (AI section) · EN マルチモーダル
    Rethinking Robot Safety in the Age of AI
    IEEE Spectrum、物理 AI 時代のロボット安全はセキュリティ問題と指摘
    ロボティクス
    IEEE Spectrum が VicOne 提供記事で、マルチモーダルセンサーで知覚し AI で文脈を解釈して動く現代のロボットは、安全性が判断を導くデータの完全性に依存すると論じる。従来は「故障したとき安全か」を問うたが、物理 AI では「何も壊れていないのに攻撃者が知覚や判断を書き換えたとき安全か」が問われる。直接制御なしに挙動を左右できる研究例もあり、既存の安全評価では捉えきれないとする。
    元記事を読む (IEEE Spectrum (AI section)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN マルチモーダル
    Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
    コンピュータビジョン 検索拡張生成 (RAG)
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 新モデル・リリース
    GrainSpeech: Less Context, More Detail for Compact Speech Synthesis
    強化学習 音声処理
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN 開発者ツール
    ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts
    AI エージェント ニューラルネットワーク 強化学習
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 新モデル・リリース
    Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
    強化学習
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN マルチモーダル
    RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection
    AI エージェント コンピュータビジョン 推論 (Inference) 検索拡張生成 (RAG)
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 新モデル・リリース
    Generalist-Specialist Mixture-of-Experts for Rare Pathology Detection in Multimodal Imaging
    推論 (Inference) Mixture of Experts (MoE) 強化学習
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN 推論・効率化
    VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge
    推論 (Inference) 強化学習
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.LG (Machine Learning) · EN 開発者ツール
    Learning Array Signal Topologies as Conditional Neural Manifolds
    ニューラルネットワーク
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 資金・M&A
    Hypothesis-Driven Autonomous Materials Synthesis with Multimodal LLM Agents
    AI エージェント
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN マルチモーダル
    Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces
    コンピュータビジョン 強化学習 ソフトウェア工学
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN 推論・効率化
    ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
    コンピュータビジョン 量子化
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN 新モデル・リリース
    Divide and Conquer: Mixture-of-Bottleneck Experts in Informative Ordinal Space for Video-based Multimodal Sentiment Analysis
    深層学習 検索拡張生成 (RAG) 強化学習
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN 新モデル・リリース
    Emotion Experience, Expression, and Perception: Emotion Analysis on Multimodal Social Media Posts
    ニューラルネットワーク 強化学習
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • Simon Willison's Weblog · EN 新モデル・リリース
    Gemini Live audio
    Simon Willison、Gemini 3.8 Live を試すブラウザ UI をライブラリ無しで公開
    Gemini Google GPT OpenAI 音声処理
    Google が音声対話モデル Gemini 3.8 Live と 3.8 Live Extended Thinking を公開したのを受け、Simon Willison がブラウザから両モデルを試せる Web UI を作成した。モデルと音声プリセットを選んで会話でき、モデルの発話中に割り込むこともできる。実装はライブラリを使わず、双方向 WebSocket と Web Audio API だけで録音・再生を処理する。
    元記事を読む (Simon Willison's Weblog) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 推論・効率化
    LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs
    埋め込み (Embeddings) 推論 (Inference) 量子化 音声処理
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN マルチモーダル
    ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation
    AI エージェント コンピュータビジョン 検索拡張生成 (RAG)
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN 新モデル・リリース
    Tables Decoded: DELTA for Structure, TARQA for Understanding
    強化学習 ソフトウェア工学
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN マルチモーダル
    CareMirror: Bringing Caregiver Wellbeing into the Dementia Care Ecosystem
    ニューラルネットワーク
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN 新モデル・リリース
    Type-IV Code Clone Detection via Layer-Wise Non-Contrastive Representation Learning
    GPT 機械学習 検索拡張生成 (RAG) 強化学習 Transformer
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN マルチモーダル
    Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs
    コンピュータビジョン ニューラルネットワーク
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 新モデル・リリース
    From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic-Geometric Contracts
    機械学習 検索拡張生成 (RAG)
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN 新モデル・リリース
    Towards Detecting AI-Assisted Responses in Online Surveys
    AI エージェント ニューラルネットワーク 強化学習
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.LG (Machine Learning) · EN マルチモーダル
    Goal-oriented probabilistic forecasting for dynamic PRB allocation in 5G networks
    Transformer
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN マルチモーダル
    FROD: Feature Matching Residual Denoising Oracle Bone Decipher
    検索拡張生成 (RAG) 強化学習
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN マルチモーダル
    FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence
    アルゴリズム・理論 コンピュータビジョン 推論 (Inference) 検索拡張生成 (RAG) 強化学習
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗