NVIDIA が AI 工場(大規模データセンター)向けの蓄電池・電力設計の指針を提示した。発生元は NVIDIA・Anthropic の公式に simon_willison と arXiv が薄く絡む構成で、ベンダーの技術指針が起点――発表主導だが電力という物理制約に踏み込む動きだ。AI 需要の急増で膨らむ消費電力を、蓄電池を含む電源設計でどう支えるかという主眼で、本流はモデルやソフトより、AI を動かす電力・インフラ側がボトルネックとして前面に出てきた点にある。計算資源の議論が「電気をどう賄うか」まで降りてきた形だ。ただし内容は設計指針の提示が中心で、実際の導入効果や標準化は今後の確認点になる。
NVIDIA、AI工場向け蓄電池の設計指針を提示
NVIDIA、AI工場向け蓄電池の設計指針を提示
Designing Production-Ready Battery Energy Storage Systems for AI Factories
NVIDIA、AIファクトリー向けの実用的なバッテリー蓄電システム設計を解説
Results from the first Anthropic Public Record
Anthropic、米国民5.2万人のAI意識調査「Public Record」初回結果を公表
NVIDIA、MiniMax M3の長文脈推論とagenticワークフロー展開手法を解説
Claude Fable is relentlessly proactive
Simon Willison氏、Claude Fable 5を「徹底的にproactive」と評価
学術(arxiv ほか) 119本 ▾
AdaSR: Adaptive Streaming Reasoning with Hierarchical Relative Policy Optimization
AdaSR、入力を逐次処理する「ストリーミング推論」を提案
CORA、マルチモーダルRLVRの「思考と回答のずれ」を是正
Beyond task performance: Decoding bioacoustic embeddings with speech features
生物音響の埋め込みを音声特徴で解読し透明性を検証
Which Directions Matter? Sparse Design for Affine Robust Optimization
アフィンロバスト最適化で重要な不確実性方向を疎に選択
Graph Diffusion Residuals for Control-Function Instrumental Variables
制御関数IV推定にグラフ拡散の残差を活用
Characterizing Cultural Localization in AI-Generated Stories
AI生成の物語における文化的ローカライズを分析
Regulating the Machine Contributor: Governance and Policy Alignment in Open Source
オープンソースへのAIエージェント貢献を統治する枠組みを論じる
AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models
音声言語モデルの事後学習向け推論データセット「AudioDER」
SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model
世界モデルでLLM実行計画の潜在的失敗を測る「SIMMER」
ORCA: A Platform for Open-Source Dexterity Research
ORCA、オープンソースの器用さ研究プラットフォーム
Rethinking Global Average Pooling: Your Classifier Is Secretly a Multi-Instance Learner
大域平均プーリングの分類器は実は多重インスタンス学習器と指摘
Provably Safe, Yet Scalable Reinforcement Learning
証明可能な安全性とスケール性を両立する強化学習
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
LLMが「チャットボット」から永続自律AIへ移行する転換を概念化
音声モデルの説明は予測を変えず操作できるその脆弱性を検証
Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR
言いよどみを聞き取る継続学習でASRを頑健化
A Low-Rank Subspace Analysis of LLM Interventions
LLM介入を低ランク部分空間で分析、副作用を抑制
Discovery under Hypothesis Redundancy: A Geometric Theory of Discovery Bottlenecks
仮説の冗長性で生じる発見の頭打ちを幾何学的に理論化
Can Deep Neural Networks Improve Compression of Very Large Scientific Data?
深層NNで大規模科学データの誤差有界圧縮を改善できるか
Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation
知識グラフ接地のデータ生成で高精度Text-to-Cypherを実現
ScoreGate、RAGの取得チャンク数を適応的に選択
Decoupled Mixture-of-Experts for Parametric Knowledge Injection
分離型MoEでパラメトリックな知識注入を改善
Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources
動的な音源の時空間を扱う音声言語モデル
Knowledge Graph Enhanced Memory-Augmented Retrieval for Long Context Modeling
知識グラフ強化の記憶検索で長文脈を一貫処理
The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models
句動詞の全体的記憶を文字・音声モデルで分析
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
動的環境でのLLMエージェント評価基盤EvoArena、パッチ型記憶EvoMemも提案
Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning
類推による推論をRAGとRL微調整で学習するRA-RFT
Influcoder: Distilling Decoders' Gradient Influence Rankings into an Encoder for Data Attribution
Influcoder、勾配影響ランキングを蒸留しデータ帰属を高速化
HyperTool: Beyond Step-Wise Tool Calls for Tool-Augmented Agents
HyperTool、ツール呼び出しをコードに畳み込みMCP精度を15.7%から35.3%へ
Operadic consistency: a label-free signal for compositional reasoning failures in LLMs
LLMの合成推論の失敗をラベル不要で検出する「オペラド整合性」を提案
SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation
SkMTEB、スロバキア語の埋め込みベンチと小型モデル公開
From Tokens to Faces: Investigating Discrete Speech Representations for 3D Facial Animation
音声駆動3D顔アニメに最適な音声表現を比較、音素情報が鍵と判明
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
AgentBeats、judge agentで評価を標準化するAAAを提唱
Distribution-Agnostic Robust Trajectory Optimization via Chance-Constrained Reinforcement Learning
機会制約RLで分布非依存のロバスト軌道最適化、地球-火星遷移で検証
Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models
思考連鎖に『コミットメント境界』、以降の手順は無影響と判明
Reward Modeling for Multi-Agent Orchestration
OrchRM、多agent統括の品質を自己教師で評価し効率10倍
EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution
EvTexture++、イベント信号で動画超解像のテクスチャ強化
Uncertainty-Aware Hybrid Retrieval for Long-Document RAG
UMG-RAG、粒度を信頼度推定として扱う訓練不要のRAG検索
AgentRivet: an automated system for producing Rivet routines from journal publications
AgentRivet、論文から欠落Rivetルーチンを自動生成するLLM系
Graphical Causal Reasoning for Root Cause Analysis in Cloud Networks
クラウド網障害の根本原因分析、因果グラフ探索で85.7%の再現率
Measurement-Calibrated Multi-Camera Fusion for Vision-Based Indoor Localization
単眼誤差の定量化で多カメラ融合を校正する屋内測位手法
Leveraging Audio-LLMs to Filter Speech-to-Speech Training Data
音声LLMで音声翻訳の学習データを選別、ASR-BLEU最大+1.4
Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations
オントロジー記憶で長い対話のASR誤り訂正を文脈接地
Clustering Node Attributed Networks with Graph Neural Networks and Self Learning
GNNと自己学習の反復でノード属性グラフをクラスタリング
SmartFont: Dynamic Condition Allocation for Few-Shot Font Generation
SmartFont、大域と局所条件を多段配分する少数事例フォント生成
An LLM System for Autonomous Variational Quantum Circuit Design
LLMで量子回路を自律設計する閉ループagent枠組み
From Passive Generation to Investigation: A Proactive Scientific Peer Review Agent
能動的に論文を調査する査読エージェントProReviewer、5品質次元で最高評価
追加学習なしで拡散モデルの低密度領域を探索、希少サンプル網羅を改善
電力系統向け確率予測ベンチマークPowerPhase、最大3.7万チャネル
SkillCAT: Contrastive Assessment and Topology-Aware Skill Self-Evolution for LLM Agents
SkillCAT、LLMエージェントのスキル自己進化を検証付き3段階に分離
Physics-Guided Spatiotemporal Learning for Coastal Wave Peak Period Estimation from Video
物理誘導の時空間学習で沿岸映像から波のピーク周期を推定
Clipping Makes Distributed and Federated Asynchronous SGD Robust to Stragglers
勾配クリッピングが非同期SGDを遅延耐性にすると理論的に正当化
TimeLens、大エジプト博物館向けオンデバイス文化財認識とRAG型QA
ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm
ComAct、COM経由の決定論的操作で産業CADソフトをエージェント制御
PolyAlign: Conditional Human-Distribution Alignment
PolyAlign、文脈に応じた人間の応答分布へLLMを整合
SICI: A Semantic-Pragmatic Complexity Index Reveals Regime Shifts in LLM Stance Detection
意味・語用論的複雑性指標SICI、LLM立場検出のレジーム変化を解明
MemRefine: LLM-Guided Compression for Long-Term Agent Memory
MemRefine、LLM判断で予算内にエージェント長期記憶を圧縮
NTS-CoT、思考連鎖でニュース年表要約の幻覚を抑制
Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models
Reroute、視覚トークンを除去でなく再ルーティングしVLM推論を効率化
Context-Driven Incremental Compression for Multi-Turn Dialogue Generation
C-DIC、多ターン対話の文脈を逐次圧縮し長対話を安定化
DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?
DIRECT、身体エージェントのテスト時計算をプロンプト別に配分
Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs
ModSleuth、公開成果物からLLMの依存グラフを再構成し監査
Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization
RACES、検証可能環境を再帰合成しRL推論の汎化を拡張
Adjoint Method versus Physics-Informed Neural Networks in PDE-Constrained Inverse Problems
PDE 制約逆問題で随伴法と PINN を同条件で公平比較
Fourier Features Let Agents Learn High Precision Policies with Imitation Learning
Fourier 特徴で点群方策が高精度ロボット操作を獲得
PROJECTMEM: A Local-First, Event-Sourced Memory and Judgment Layer for AI Coding Agents
projectmem、コーディングエージェントに局所優先の記憶層を追加
Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
MedMisBench、誤誘導文脈下でのLLM医療判断の脆弱性を測定
Ideogram 4.0 を INT8 量子化、民生 GPU で FP8 品質を維持
DiffCold: A Diffusion-based Generative Model for Cold-Start Item Recommendation
DiffCold、拡散生成でコールドスタート推薦のシーソー難を解消
Re-evaluating Confidence Remasking in Masked Diffusion Language Models
拡散言語モデルの remasking 手法 WINO、再評価で効果は限定的
MLT-Dedup、多層表現で大規模オンライン動画の重複排除を効率化
OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models
OpenMedReason、約45万件の医療多モーダル推論コーパスを公開
A Controlled Study of Decoding-Time Truthfulness Methods on Instruction-Tuned LLMs
CHAIR、層ごとのlogitから幻覚を検出する教師あり手法
nD-RoPE: A Generalized RoPE for n-Dimensional Position Embedding
nD-RoPE、n次元位置埋め込みへRoPEを分解なく一般化
Augmenting Molecular Language Models with Local $n$-gram Memory
MolGram、局所n-gram記憶で分子言語モデルを効率改善
InDex、意味継承でVLAモデルを器用な多関節ハンドへ適応
DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model
DAM-VLA、モダリティ別の非同期処理で実機操作成功率を倍増
FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents
FORT、近道耐性のある探索課題を合成し深層探索エージェントを訓練
On the Limits of LLM-as-Judge for Scientific Novelty Assessment
LLM審判の科学的新規性評価の限界、RQ-Benchで検証
Simplicity Suffices for Parameter Noise Injection in Stochastic Gradient Descent
SGD へのノイズ注入、単純な等方性手法で十分と実証
Decoding Multimodal Cues: Unveiling the Implicit Meaning Behind Hateful Videos
IARE、根拠付きで有害動画を説明可能に検出する枠組み
SemEval-2026 Task 8、疎検索とリスト再順位で多ターンRAG
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
Arbor、仮説ツリー改良で長期間の自律研究ループを運用
An Ontology-Guided Multi-Anchor Graph Retrieval Framework for Traffic Legal Liability Determination
OMAGR、オントロジー誘導の多アンカーグラフ検索で交通法責任を判定
Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code
CodeSpear、文法制約デコードを悪用しLLMに悪性コードを生成させる
Lius、継続的指示調整で低資源クパン・マレー語の翻訳を改善
EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents
EEVEE、複数データセット対応のテスト時プロンプト学習を実現
Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories
Data2Story、データから検証可能なマルチモーダル記事を自動生成
Flaws in the LLM Automation Narrative
LLMの専門家代替論に疑義、人間専門家が平均で上回ると報告
ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity
ABC-Bench、LLMエージェントのバイオセキュリティ関連能力を評価
First-Order Trajectory Matching: Fast Ensemble Predictions of Chaotic, Turbulent, Stochastic Systems
FTM、軌跡から確率流速を学習し安価なアンサンブル予測
DMT、PPGから人口統計条件付きTransformerでカフレス血圧推定
TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning
TRACE、エージェントRLのロールアウト予算をツリー状に配分
Towards Autonomous Accelerator Design: FPGA Accelerator Generation with SECDA
SECDA-DSE、LLMでFPGAアクセラレータの設計空間探索を誘導
PhantomBench: Benchmarking the Non-existential Threat of Language Models
PhantomBench、存在しない概念へのLLMのハルシネーションを評価
RoboNaldo、動作誘導カリキュラムRLでヒューマノイドの強シュート
VISTA: A Versatile Interactive User Simulation Toolkit for Agent Evaluation
VISTA、エージェント評価向けの多用途ユーザシミュレーション基盤
A History-Aware Visually Grounded Critic for Computer Use Agents
HiViG、計算機操作エージェント向けの履歴認識・視覚接地批評
T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains
T1-Bench、25ドメインの多シナリオエージェントを高忠実評価
Generative Archetype-Grounded Item Representations for Sequential Recommendation
GenAIR、対象顧客像で接地した生成的アイテム表現で推薦改善
LLM拡張XAIで次世代ネットワークの説明を自然言語生成
Democratising Camera Trap AI: An Open-Source Model for Detecting UK Mammals
英国哺乳類検出の開源モデルを公開、mAP 0.984を達成
DocTrace、長文書QA向けの構造認識・オンデマンド超グラフ記憶
Task Robustness via Re-Labelling Vision-Action Robot Data
TREAD、VLMでロボットの視覚行動データを再ラベルし頑健化
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
Role-Agent、単一LLMがエージェントと環境を兼ね共進化
Range Penalization: Theoretical Insights with Applications in Federated Learning
範囲正則化、連合学習で量子化に資する重みの極性クラスタ化
Ethical and Technical Limits of Deepfake Speech Datasets
ディープフェイク音声39データセットを監査、公平性評価は困難
Pose-ICL: 3D-Aware In-Context Learning for Pose-Controllable Subject Customization
Pose-ICL、3D認識の文脈内学習でポーズ制御可能な被写体生成
Janus: A Benchmark for Goal-Conditioned Information Distortion in LLMs
JANUS、真実の事実を選択的に扱うLLMの情報歪曲を測定
Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
ADAS、マスク拡散言語モデルの並列復号を再ランクで安全化
Recovering the Zipfian Distribution in Unsupervised Term Discovery
教師なし語彙発見、グラフクラスタリングがK-means等を上回る
N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization
N-GRPO提案、埋め込み近傍の混合でGRPOの探索を強化
Infini Memory: Maintainable Topic Documents for Long-Term LLM Agent Memory
Infini Memory、トピック文書でエージェントの長期記憶を保守
How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMs
FlowTracer、attention情報流でRLのトークン信用割当を精緻化
ParaBridge: Bridging Paralinguistic Perception and Dialogue Behavior in Speech Language Models
ParaBridge、音声LLMのパラ言語知覚と対話行動の溝を埋める