新モデル・リリース
A
219 件中 181〜210 件目を表示
-
Continue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents
-
MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
-
MAxBench: A Multinomial Concept Recovery Benchmark
-
Expert-Space Exploration in MoE Reinforcement Learning
-
DynSHAP: Towards Explainable Dynamic Survival Analysis
-
How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
-
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
-
Investigating Temporal Motion Features for Pose-to-Text Indian Sign Language Translation
-
SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading
-
Generative Retrieval for Unsupervised Text-Based Person Search
-
EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics
-
英国王立協会特集号に見る、世界モデルの最前線とAIの未来英国王立協会が世界モデル特集号、Sakana AI の David Ha も寄稿英国王立協会の『Philosophical Transactions A』が、世界モデルを主題とする特集号「World Models in Natural and Artificial Intelligence」を公開した。Sakana AI CEO の David Ha が巻頭記事の共著者として参加。大規模モデルの「できること」と「わかっていること」の差は計算資源だけでは埋まらない、自己の内部状態を予測する学習が内部表現を整理する、世界モデルの問いは人工生命へ通じる、の 3 点を軸に論じる。
-
ARC: Autonomous Robotics Compliance A Three-Layer Governance Architecture for Deployed Autonomous Systems
-
UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction
-
Behavior Quotient Learning for Low-Rank Adaptation of LLM Agents
-
Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting
-
MedSNIP: Building and Benchmarking Snippet-Level Granularity for Medical Fact Verification
-
Don't sleep on wraptureGraham Dumpleton、テストと可観測性を兼ねる wrapture を公開Graham Dumpleton 氏が 8 月 31 日に公開した Python 向け monkey patching ライブラリ wrapture が、テストと可観測性 (New Relic 的トレーシング) を同時に担う設計として注目されている。公開以降ほぼ毎日チュートリアルが追加され、呼び出しの記録とツリー表示、複数回の呼び出しで挙動を変えるフェーズ動作、属性・辞書・ジェネレータへのパッチまでを扱う。Simon Willison 氏は「ほとんど話題になっていないのが不思議」と評した。
-
MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations
-
Evaluating Context Segmentation in Locally Deployable SLMs for Cybersecurity CTF Tasks
-
Scaling Clinical Judgment to Evaluate Medical AI
-
4D Parallelism Unlocks Exascale Bayesian Neural Networks for High-Fidelity Atmospheric Modeling
-
RunningTensor: Generalizing Linear Attention to Higher-Order Recurrent States
-
K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments
-
Interpreting the predictions of neural network classification based on a Taylor Coefficient Analysis (TCA)
-
SCQ: Stabilizing Conservative Q-Learning with Sigmoid-Bounded Entropy
-
Assisted Spatial Cognition Through Vision-Language Models
-
Doc2FRC: Length-Consistent Document-Level Machine Translation via Fixed-Range Chunking
-
SWARM: A Multilingual Human-Annotated Dataset for Russian Propaganda Detection in Search Engine Results
-
SteerDuplex: Steerable Duplex Speech Dialogue Models