New Model Releases
A
Showing 181–210 of 218
-
Continue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents
-
MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
-
MAxBench: A Multinomial Concept Recovery Benchmark
-
Expert-Space Exploration in MoE Reinforcement Learning
-
DynSHAP: Towards Explainable Dynamic Survival Analysis
-
How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
-
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
-
Investigating Temporal Motion Features for Pose-to-Text Indian Sign Language Translation
-
SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading
-
Generative Retrieval for Unsupervised Text-Based Person Search
-
EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics
-
英国王立協会特集号に見る、世界モデルの最前線とAIの未来Royal Society issue maps world models; Sakana AI's David Ha co-authorsThe Royal Society's Philosophical Transactions A published a special issue, "World Models in Natural and Artificial Intelligence," co-authored in its opening article by Sakana AI CEO David Ha. Contributors argue scaling compute alone will not close the gap between what large models can do and what they understand, and tie world models to artificial life.
-
ARC: Autonomous Robotics Compliance A Three-Layer Governance Architecture for Deployed Autonomous Systems
-
UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction
-
Behavior Quotient Learning for Low-Rank Adaptation of LLM Agents
-
Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting
-
MedSNIP: Building and Benchmarking Snippet-Level Granularity for Medical Fact Verification
-
Don't sleep on wraptureDumpleton's wrapture unifies Python mocking and tracingGraham Dumpleton released wrapture, a Python monkey patching library serving both unit testing and New Relic-style observability. Near-daily tutorials since its August 31 debut cover call recording as trees, phased behaviour across successive calls, and patching attributes, dictionaries and generators.
-
MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations
-
Evaluating Context Segmentation in Locally Deployable SLMs for Cybersecurity CTF Tasks
-
Scaling Clinical Judgment to Evaluate Medical AI
-
4D Parallelism Unlocks Exascale Bayesian Neural Networks for High-Fidelity Atmospheric Modeling
-
RunningTensor: Generalizing Linear Attention to Higher-Order Recurrent States
-
K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments
-
Interpreting the predictions of neural network classification based on a Taylor Coefficient Analysis (TCA)
-
SCQ: Stabilizing Conservative Q-Learning with Sigmoid-Bounded Entropy
-
Assisted Spatial Cognition Through Vision-Language Models
-
Doc2FRC: Length-Consistent Document-Level Machine Translation via Fixed-Range Chunking
-
SWARM: A Multilingual Human-Annotated Dataset for Russian Propaganda Detection in Search Engine Results
-
SteerDuplex: Steerable Duplex Speech Dialogue Models