Developer Tools B
Showing 121–150 of 430
-
Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
-
Cybersecurity Detection Classification with Reasoning-enabled Language Models
-
Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation
-
Oracle-Budgeted Molecular Optimization with Short-Term Graph Memory
-
Metaphor Tracer: A Theory-Informed Analysis of Hidden States
-
Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image Classification
-
Negative controls reveal volume-driven confounding in radiomics and imaging foundation model features
-
QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Quantum Error Correction
-
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning
-
QQWorld: Quantile-Quantile Matching for World Model Regularization
-
NVIDIA Exemplar Cloud: Lessons for Unlocking Full Performance on AI InfrastructureNVIDIA shares Exemplar Cloud lessons for unlocking AI infra performanceNVIDIA shared lessons from its Exemplar Cloud, noting that two clusters built from identical H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different performance. The guidance focuses on tuning and operations to unlock full AI infrastructure performance.
-
Windowed thinning and query complexity for the bouncy particle and Zigzag samplers
-
Can Large Language Models Execute Parent Orders?
-
Hierarchical Multilevel Monte Carlo for Order-Optimal Neural Actor-Critic in Average-Reward CMDPs
-
When Specifications Conflict: A Symmetry-Based Framework for Measuring LLM Preferences
-
HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection
-
How Benchmarks Mis-Score Computer-Use Agents
-
ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
-
Teffic-Audio: Tell Fact from Fiction
-
LLMs struggle to simulate human belief updates in controlled environments
-
Reflected diffusion, no-flux continuity equations and confined Lagrangian flows in bounded domains
-
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaborationDeepMind's Gemini Robotics ER 2 adds video understanding, multi-robot teamworkDeepMind introduced Gemini Robotics ER 2, which helps robots reason, collaborate, and solve real-world tasks. The company calls it a step change in video understanding, task orchestration, and multi-robot collaboration for embodied AI.
-
From Japan, Products the World Will Use: An Interview with Sakana AI's Head of Product DevelopmentInterview: Sakana AI's product chief on Japan-born global productsAn interview with Sakana AI's Head of Product Development on building products from Japan that the world will use. The Q&A covers the company's product philosophy and ambitions, offering a look at the strategy of a leading Japanese AI startup.
-
Investigating three real-world incidents in our cybersecurity evaluationsAnthropic's Frontier Red Team probes three cybersecurity-eval incidentsAnthropic's Frontier Red Team published a review of three real-world incidents tied to its cybersecurity evaluations. The investigation examines potential misuse and the validity of its evaluation methods to strengthen the safety of frontier models.
-
Paying for Honesty Without Knowing the Truth: Reputation-Penalty Design for LLM Marketplace Agents
-
Measuring Distortion in the Empty Regions of Dimensionality Reduction Scatterplots with the Gap Index
-
PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?
-
One Human, $N$ Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence
-
Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing
-
From Textual Requirements to Microservice Architectures - A Comprehensive Evaluation of LLM-Based Design Synthesis