NVIDIA が音声モデル「Nemotron Speech」で臨床領域の音声認識(ASR)評価を高速化したと示した。発生元は5件すべて NVIDIA 公式(一部 Hugging Face・DeepMind 連携)で、一社主導の技術提示――発表主導だが、医療という応用ドメインに踏み込む動きだ。診療記録の音声起こしなど、専門用語が多く精度要求の高い臨床 ASR を対象に、評価と処理を速める点が主眼で、本流はモデルの汎用性能より、特定ドメインでの実用精度と処理効率へ焦点が移っている点にある。医療現場の文書作業を軽くする応用として位置づけられる。ただし現段階は技術提示が中心で、実臨床での精度検証や規制対応、採用の広がりは今後の確認点だ。
NVIDIA、Nemotron Speechで臨床ASRを高速化
NVIDIA、Nemotron Speechで臨床ASRを高速化
Evaluate Clinical ASR Models Faster with Agent Skills and NVIDIA Nemotron Speech
NVIDIA、Agent SkillsとNemotron Speechで臨床ASR評価を高速化
AIエージェントもフィッシング詐欺に引っかかる? 米セキュリティ企業がOpenClawで検証 結果は……
AIエージェントもフィッシングに引っかかる VaronisがOpenClawで検証
Building a persistent cognitive architecture for LLM agents using Elixir and OTP
Elixir と OTP で LLM エージェントの永続的認知アーキテクチャを構築
Can Voice Agents Handle Bilingual Customers? Benchmarking Frontier ASR on Code-Switched Speech
ServiceNow AI、コードスイッチ音声で最先端ASRをベンチマーク評価
NVIDIA、DGX SparkでAIインフラのライフサイクル管理を強化
NVIDIA、FP8チェックポイントをTensorRTで高性能推論エンジンに変換する手法を解説
Accelerating Federated Learning Research with AI Agents and NVIDIA FLARE Auto-FL
NVIDIA、AIエージェントとFLARE Auto-FLで連合学習研究を加速
Fluid, natural voice translation with Gemini 3.5 Live Translate
Google、Gemini 3.5 Live Translateで自然な音声翻訳を提供
Introducing North Mini Code: Cohere’s first model for developers
Cohere、初の agentic コーディングモデル「North Mini Code」を OSS 公開
Ubuntu、サンドボックス化された開発環境をコマンド一発で構築。新機能「Workshop」リリース
Canonical、AIエージェント向けサンドボックス開発環境「Workshop」を公開
学術(arxiv ほか) 99本 ▾
Operadic consistency: a label-free signal for compositional reasoning failures in LLMs
LLMの合成推論の失敗をラベル不要で検出する「オペラド整合性」を提案
Valid Inference with Synthetic Data via Task Exchangeability
合成データで妥当な推論を保証するtask exchangeabilityを提案
Beyond Uniform Tokens: Adaptive Compression for Time Series Language Models
時系列LLMの非対称トークン圧縮で効率化、予測から異常検知まで有効
Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models
思考連鎖に『コミットメント境界』、以降の手順は無影響と判明
単体制約スパースバギングSCSB、アンサンブルを最大96%圧縮し較正改善
Timeflies、未来の観測有無と値を同時推定する予測枠組み
A2D2: Fine-Tuning Any-Length Discrete Diffusion for Adaptive Decoding
A2D2、可変長離散拡散モデルの報酬誘導ファインチューニングを統一
NetCause: Counterfactual Learning for Root Cause Analysis in Large-Scale Networks
NetCause、反実仮想学習でクラウド網障害の根本原因を順位付け
Graphical Causal Reasoning for Root Cause Analysis in Cloud Networks
クラウド網障害の根本原因分析、因果グラフ探索で85.7%の再現率
GF-DiT: Scheduling Parallelism for Diffusion Transformer Serving
GF-DiT、拡散Transformer配信のGPU並列度を動的スケジューリング
Optical Implementation of Equilibrium Propagation Using Spatial Photonic Ising Machines
空間フォトニックイジングマシンで平衡伝播を光学実装
Accelerating Speculative Diffusions via Block Verification
拡散モデルに投機的復号のブロック検証を導入、受理率を理論的に改善
PolyFlow、多面体制約を埋込む射影不要のflow matching
MiniMax、超長文脈向けブロック疎注意MSAを発表
SmartFont: Dynamic Condition Allocation for Few-Shot Font Generation
SmartFont、大域と局所条件を多段配分する少数事例フォント生成
Hölder++: Improving the Quality-Coherence Trade-off in Multimodal VAEs
Hölder++、マルチモーダルVAEの品質と整合性の両立を改善
VideoMDM: Towards 3D Human Motion Generation From 2D Supervision
VideoMDM、3D正解データなしで動画の2D姿勢から3D動作生成を学習
SkillCAT: Contrastive Assessment and Topology-Aware Skill Self-Evolution for LLM Agents
SkillCAT、LLMエージェントのスキル自己進化を検証付き3段階に分離
SICI: A Semantic-Pragmatic Complexity Index Reveals Regime Shifts in LLM Stance Detection
意味・語用論的複雑性指標SICI、LLM立場検出のレジーム変化を解明
MiniPIC: Flexible Position-Independent Caching in <100LOC
MiniPIC、100行未満でvLLMに位置非依存KVキャッシュを実装
Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models
Reroute、視覚トークンを除去でなく再ルーティングしVLM推論を効率化
Context-Driven Incremental Compression for Multi-Turn Dialogue Generation
C-DIC、多ターン対話の文脈を逐次圧縮し長対話を安定化
DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?
DIRECT、身体エージェントのテスト時計算をプロンプト別に配分
Doc-to-Atom: Learning to Compile and Compose Memory Atoms
Doc2Atom、文書を知識アトムに分解し合成的パラメトリック記憶を構築
System Report for CCL25-Eval Task 5: New Dataset and LoRA-Fine-Tuned Qwen2.5
PoetryQwen、CCPoetry-49Kで古典中国詩の鑑賞を専門化
TAHOE: Text-to-SQL with Automated Hint Optimization from Experience
Tahoe、経験からヒントを学びText-to-SQLを本番最適化
ATLAS: Active Theory Learning for Automated Science
ATLAS、能動学習で解釈可能な行動モデルを自動的に発見
APPO: Agentic Procedural Policy Optimization
APPO、細粒度の決定点で分岐と信用割当を行うエージェントRL
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
Bebop、棄却サンプリングでMTP受理率を改善しRL学習を加速
Latent World Recovery for Multimodal Learning with Missing Modalities
LWR、欠損モダリティ下で潜在世界を復元しマルチモーダル学習
CHORUS: Decentralized Multi-Embodiment Collaboration with One VLA Policy
CHORUS、単一VLA方策で分散的な多体協調を実現
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks
Claw-SWE-Bench、OpenClaw型エージェントのコード能力を多言語評価
ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing
ALIGNBEAM、語彙を越えて安全アンカーのlogitを移植
Fourier Features Let Agents Learn High Precision Policies with Imitation Learning
Fourier 特徴で点群方策が高精度ロボット操作を獲得
Measuring Semantic Progress in Multi-turn Dialogue via Information Gain
多ターン対話の意味的進展を情報利得で測る指標を提案
PROJECTMEM: A Local-First, Event-Sourced Memory and Judgment Layer for AI Coding Agents
projectmem、コーディングエージェントに局所優先の記憶層を追加
A Five-Plane Reference Architecture for Runtime Governance of Production AI Agents
本番AIエージェントの実行時統治に5プレーン参照アーキテクチャ
Harness In-Context Operator Learning with Chain of Operators
CHOP、演算子連鎖で凍結ICONを分布外タスクへ汎化
CCKS: Consensus-based Communication and Knowledge Sharing
CCKS、合意ベースの通信と知識共有で協調的MARLを改善
Ideogram 4.0 を INT8 量子化、民生 GPU で FP8 品質を維持
Mathematical perspective on genetic algorithms with optimization guided operators
最適化誘導演算子を持つ遺伝的アルゴリズムの数理的視座
The Impossibility of Eliciting Latent Knowledge
潜在知識の引き出し(ELK)は不可能と因果影響図で形式化
VIA-SD: Verification via Intra-Model Routing for Speculative Decoding
VIA-SD、モデル内ルーティングで投機的デコードを高速化
Re-evaluating Confidence Remasking in Masked Diffusion Language Models
拡散言語モデルの remasking 手法 WINO、再評価で効果は限定的
Can News Predict the Market? Limits of Zero-Shot Financial NLP and the Role of Explainable AI
ゼロショット金融NLPの限界、ニュースで株価予測は困難と示す
Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models
SKIM、手続き的スキルを適応的多解像度で圧縮しコスト削減
Implicit Neural Representations of Individual Behavior
Behavioral INR、無ラベルの多方策行動から方策表現を学習
A Resource for Enthymeme Detection in Controversial Political Discourse
論争的政治言説のエンチメーム検出データセットを公開
Towards Responsibly Non-Compliant Machines
責任ある「不服従」が可能なAIエージェントの工学を提起
FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents
FORT、近道耐性のある探索課題を合成し深層探索エージェントを訓練
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
Arbor、仮説ツリー改良で長期間の自律研究ループを運用
Notes2Skills: From Lab Notebooks to Certainty-Aware Scientific Agent Skills
Notes2Skills、実験ノートを確信度付きの科学エージェントスキルに変換
Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training
ART、視覚入力のみ最適化で凍結MLLMをソフトトークン微調整
WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning
WorldReasoner、エージェントの事象予測の推論妥当性を評価
MultiToP、視覚トークンを修復し動画LMMの幻覚を緩和
Fast Speech Foundation Model Distillation Using Interleaved Stacking
交互スタッキングで音声基盤モデルの蒸留学習を高速化
EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents
EEVEE、複数データセット対応のテスト時プロンプト学習を実現
Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories
Data2Story、データから検証可能なマルチモーダル記事を自動生成
Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models
全二重音声対話モデルの対話性をRLで多面的に改善
ReasonAlloc: Hierarchical Decoding-Time KV Cache Budget Allocation for Reasoning Models
ReasonAlloc、推論モデルのKVキャッシュ予算を階層的に配分
Itoマップ、確率動力学の任意ステップ流れマップを単一パスで予測
ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity
ABC-Bench、LLMエージェントのバイオセキュリティ関連能力を評価
RoboNaldo、動作誘導カリキュラムRLでヒューマノイドの強シュート
A History-Aware Visually Grounded Critic for Computer Use Agents
HiViG、計算機操作エージェント向けの履歴認識・視覚接地批評
T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains
T1-Bench、25ドメインの多シナリオエージェントを高忠実評価
What Fits (Into Few Tokens) Doesn't Overfit: Compression and Generalization in ML Research Agents
ML研究エージェントの圧縮性が過学習の少なさを説明
Workflow-GYM、専門ソフトの長期GUIタスクを評価
AuRA: Internalizing Audio Understanding into LLMs as LoRA
AuRA、音声理解をLoRAとしてLLMに内在化
DFP、履歴誘導の拡散計画で自動運転の軌跡を安定化
OpenClawの危険を非専門ユーザ向けに7分類し平易に解説
Mind the Gap: Can Frontier LLMs Pass a Standardized Office Proficiency Exam?
最先端LLM、Office操作試験で最高36.6%にとどまる
CLP: Collocation-Length Prediction for Zero-Loss Adaptive Multi-Token Inference
CLP、バックボーンを設計者とし零損失の適応的多トークン推論
Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages
最先端コーディングエージェント、難解言語へメタプログラミングで適応
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
Role-Agent、単一LLMがエージェントと環境を兼ね共進化
Range Penalization: Theoretical Insights with Applications in Federated Learning
範囲正則化、連合学習で量子化に資する重みの極性クラスタ化
What Do Deepfake Speech Detectors Actually Hear?
ディープフェイク音声検出器が実際に何を聴くかを可視化
Ethical and Technical Limits of Deepfake Speech Datasets
ディープフェイク音声39データセットを監査、公平性評価は困難
RAT: Reference-Augmented Training for ASV Anti-Spoofing
RAT、参照拡張学習でASV成りすまし検知を最先端化
Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation
経験的知識の統合と活性化でLLMのツール呼び出しを強化
ConvMemory v2: A Recall-Preserving Top-10 Evidence Reranker for Conversational Memory Retrieval
ConvMemory v2、会話記憶検索の想起保持型Top-10再ランカー
Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
ADAS、マスク拡散言語モデルの並列復号を再ランクで安全化
K-Forcing: Joint Next-K-Token Decoding via Push-Forward Language Modeling
K-Forcing、押し出し言語モデルで次kトークンを同時復号
Recovering the Zipfian Distribution in Unsupervised Term Discovery
教師なし語彙発見、グラフクラスタリングがK-means等を上回る
密から疎へ、継続学習でLLMをスパース化するレシピを提示
Attention Expansion、長文書のキーフレーズ抽出を強化
REAL: A Reasoning-Enhanced Graph Framework for Long-Term Memory Management of LLMs
REAL提案、推論強化グラフでLLMの長期記憶を管理
Infini Memory: Maintainable Topic Documents for Long-Term LLM Agent Memory
Infini Memory、トピック文書でエージェントの長期記憶を保守
自己教師あり表現で多言語の単語強制アラインメントを高精度化
Speaker Group Encoding in Self-supervised Speech Recognition Models
自己教師あり音声モデルの話者グループ情報の符号化を解析
ParaBridge: Bridging Paralinguistic Perception and Dialogue Behavior in Speech Language Models
ParaBridge、音声LLMのパラ言語知覚と対話行動の溝を埋める