NVIDIA × Safety & Evaluation

NVIDIA outlines AI factory storage

NVIDIA outlines AI factory storage

✎ Story body

NVIDIA laid out guidance for energy storage and power design for AI factories (large data centers). The signal pairs NVIDIA and Anthropic official sources with thin involvement from simon_willison and arXiv—vendor technical guidance, announcement-led but stepping into the physical constraint of power. The focus is how to support surging AI power consumption through power design including batteries; the through-line is the power and infrastructure side—rather than models or software—coming to the fore as the bottleneck. The compute conversation has descended to 'how to supply the electricity.' But the content is mainly design guidance—real deployment impact and standardization are to confirm.

▲ Official & Press
Official

Designing Production-Ready Battery Energy Storage Systems for AI Factories

NVIDIA Developer Blog ・ 2026-06-10 ・ 📌

NVIDIA details production-ready battery energy storage for AI factories

Official

Results from the first Anthropic Public Record

Anthropic News ・ 2026-06-12

Anthropic shares first Public Record survey of 52,000 Americans on AI

Official

Deploy Long-Context Reasoning and Agentic Workflows with MiniMax M3 on NVIDIA Accelerated Infrastructure

NVIDIA Developer Blog ・ 2026-06-12

NVIDIA details deploying MiniMax M3 for long-context agentic workflows

Community

Claude Fable is relentlessly proactive

Simon Willison's Weblog ・ 2026-06-11

Simon Willison: Claude Fable 5 is 'relentlessly proactive'

Academic (arxiv etc.) 119 ▾
Academic

AdaSR: Adaptive Streaming Reasoning with Hierarchical Relative Policy Optimization

arXiv cs.CL (Computation and Language) ・ 2026-06-12

AdaSR enables adaptive streaming reasoning for reasoning models

Academic

CORA: Analyzing and bridging thinking-answer gap in Multimodal RLVR via Consistency-Oriented Reasoning Alignment

arXiv cs.CL (Computation and Language) ・ 2026-06-12

CORA aligns reasoning and answers in multimodal RLVR

Academic

Flood and Harvest: The Provable Necessity of Trivia for Generating Valuable Mathematics via the Lens of Language Generation in the Limit

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

Why generating 'trivia' is provably necessary for valuable mathematics

Academic

Beyond task performance: Decoding bioacoustic embeddings with speech features

arXiv cs.LG (Machine Learning) ・ 2026-06-12

Decoding what bioacoustic embeddings encode via speech features

Academic

Graph Structured Combinatorial Semi-Bandit with Nonlinear Reward Associations through Separable Signals

arXiv cs.LG (Machine Learning) ・ 2026-06-12

Graph-structured combinatorial semi-bandits with nonlinear rewards

Academic

Which Directions Matter? Sparse Design for Affine Robust Optimization

arXiv cs.LG (Machine Learning) ・ 2026-06-12

Sparse design identifies which directions matter in robust optimization

Academic

Graph Diffusion Residuals for Control-Function Instrumental Variables

arXiv cs.LG (Machine Learning) ・ 2026-06-12

Graph diffusion residuals for control-function instrumental variables

Academic

Characterizing Cultural Localization in AI-Generated Stories

arXiv cs.CL (Computation and Language) ・ 2026-06-12

Characterizing cultural localization in AI-generated stories

Academic

Regulating the Machine Contributor: Governance and Policy Alignment in Open Source

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

Governance and policy alignment for AI contributors in open source

Academic

AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

AudioDER: a deduplication-enhanced reasoning dataset for audio LLMs

Academic

SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

SIMMER: benchmarking latent failures in LLM executable planning

Academic

ORCA: A Platform for Open-Source Dexterity Research

arXiv cs.LG (Machine Learning) ・ 2026-06-12

ORCA: an open-source platform for dexterity research

Academic

Rethinking Global Average Pooling: Your Classifier Is Secretly a Multi-Instance Learner

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

Rethinking GAP: your classifier is secretly a multi-instance learner

Academic

Provably Safe, Yet Scalable Reinforcement Learning

arXiv cs.LG (Machine Learning) ・ 2026-06-12

Provably safe yet scalable reinforcement learning

Academic

From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

From chatbot to digital colleague: the shift to persistent autonomous AI

Academic

The Perceived Fragility of Explanations in Audio Models: Manipulation of Attribution with Unchanged Predictions

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

Manipulating audio-model explanations while predictions stay unchanged

Academic

Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

Learning to hear hesitation: continual learning for disfluency-aware ASR

Academic

A Low-Rank Subspace Analysis of LLM Interventions

arXiv cs.LG (Machine Learning) ・ 2026-06-12

A low-rank subspace analysis of LLM behavioral interventions

Academic

Discovery under Hypothesis Redundancy: A Geometric Theory of Discovery Bottlenecks

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

A geometric theory of discovery bottlenecks under hypothesis redundancy

Academic

Can Deep Neural Networks Improve Compression of Very Large Scientific Data?

arXiv cs.LG (Machine Learning) ・ 2026-06-12

Can deep neural networks improve compression of very large scientific data?

Academic

Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-12

Precise Text-to-Cypher via grounded knowledge graph data generation

Academic

ScoreGate: Adaptive Chunk Selection for Retrieval-Augmented Generation via Dual-Score Statistical Fusion

arXiv cs.CL (Computation and Language) ・ 2026-06-12

ScoreGate: adaptive chunk selection for RAG via dual-score fusion

Academic

Decoupled Mixture-of-Experts for Parametric Knowledge Injection

arXiv cs.CL (Computation and Language) ・ 2026-06-12

Decoupled mixture-of-experts for parametric knowledge injection

Academic

Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources

arXiv cs.CL (Computation and Language) ・ 2026-06-12

Spatio-temporal audio language modeling for dynamic sound sources

Academic

Knowledge Graph Enhanced Memory-Augmented Retrieval for Long Context Modeling

arXiv cs.CL (Computation and Language) ・ 2026-06-12

Knowledge-graph-enhanced memory-augmented retrieval for long context

Academic

The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-12

Holistic storage of verb-up phrases in text and audio language models

Academic

EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments

arXiv cs.CL (Computation and Language) ・ 2026-06-11

EvoArena tests LLM agents in evolving environments with patch-based memory

Academic

Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

RA-RFT teaches LLMs to reason by analogy via retrieval-augmented RL fine-tuning

Academic

Influcoder: Distilling Decoders' Gradient Influence Rankings into an Encoder for Data Attribution

arXiv cs.CL (Computation and Language) ・ 2026-06-11

Influcoder distills gradient influence rankings for fast data attribution

Academic

HyperTool: Beyond Step-Wise Tool Calls for Tool-Augmented Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-11

HyperTool folds tool workflows into code, lifting MCP-Universe to 35.29%

Academic

Operadic consistency: a label-free signal for compositional reasoning failures in LLMs

arXiv cs.CL (Computation and Language) ・ 2026-06-11

Operadic consistency flags LLM compositional reasoning errors label-free

Academic

SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

SkMTEB: first comprehensive Slovak text embedding benchmark and compact models

Academic

From Tokens to Faces: Investigating Discrete Speech Representations for 3D Facial Animation

arXiv cs.CL (Computation and Language) ・ 2026-06-11

Phonetic encoding key for speech-driven 3D facial animation, study finds

Academic

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

AgentBeats proposes agentified, protocol-standardized agent assessment (AAA)

Academic

Distribution-Agnostic Robust Trajectory Optimization via Chance-Constrained Reinforcement Learning

arXiv cs.LG (Machine Learning) ・ 2026-06-11

Chance-constrained RL robustifies trajectories for Earth-Mars transfer

Academic

Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

Chain-of-thought reasoning crosses a 'commitment boundary,' study shows

Academic

Reward Modeling for Multi-Agent Orchestration

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

OrchRM: self-supervised reward modeling for multi-agent orchestration

Academic

EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

EvTexture++ uses event signals for texture enhancement in video super-resolution

Academic

Uncertainty-Aware Hybrid Retrieval for Long-Document RAG

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

UMG-RAG treats chunk granularity as query-specific reliability for RAG

Academic

AgentRivet: an automated system for producing Rivet routines from journal publications

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

AgentRivet auto-generates missing Rivet routines from physics papers

Academic

Graphical Causal Reasoning for Root Cause Analysis in Cloud Networks

arXiv cs.LG (Machine Learning) ・ 2026-06-11

Causal graph traversal recalls 85.7% of cloud incident root causes

Academic

Measurement-Calibrated Multi-Camera Fusion for Vision-Based Indoor Localization

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

Measurement-calibrated multi-camera fusion improves indoor visual localization

Academic

Leveraging Audio-LLMs to Filter Speech-to-Speech Training Data

arXiv cs.CL (Computation and Language) ・ 2026-06-11

Audio-LLM filters speech translation data, gaining up to +1.4 ASR-BLEU

Academic

Heterogeneous LiDAR Early Fusion and Learned Re-Ranking Strategy for Robust Long-Term Place Recognition in Unstructured Environments

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

Heterogeneous LiDAR fusion and re-ranking boost place recognition in fields

Academic

Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

Ontology memory grounds ASR correction in long text-speech conversations

Academic

Clustering Node Attributed Networks with Graph Neural Networks and Self Learning

arXiv cs.LG (Machine Learning) ・ 2026-06-11

Self-learning rounds couple GNN embeddings and graph clustering

Academic

Mod-Guide: An LLM-based Content Moderation Feedback System to Address Insensitive Speech toward Indigenous Ethnic and Religious Minority Communities

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

Mod-Guide integrates minority perspectives into LLM content moderation

Academic

SmartFont: Dynamic Condition Allocation for Few-Shot Font Generation

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

SmartFont allocates global and local conditions for few-shot font generation

Academic

An LLM System for Autonomous Variational Quantum Circuit Design

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-11

LLM agent framework autonomously designs variational quantum circuits

Academic

From Passive Generation to Investigation: A Proactive Scientific Peer Review Agent

arXiv cs.CL (Computation and Language) ・ 2026-06-11

ProReviewer agent proactively investigates papers for deeper peer review

Academic

Enhanced Low-Density Region Exploration in Classifier-Guided Diffusion Models Through Modified Reverse Diffusion Sampling

arXiv cs.LG (Machine Learning) ・ 2026-06-11

Training-free sampling tweak helps diffusion models reach rare samples

Academic

Navigating the Safety-Fidelity Trade-off: Massive-Variate Time Series Forecasting for Power Systems via Probabilistic Scenarios

arXiv cs.LG (Machine Learning) ・ 2026-06-11

PowerPhase benchmarks probabilistic forecasting on 37K grid channels

Academic

SkillCAT: Contrastive Assessment and Topology-Aware Skill Self-Evolution for LLM Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-11

SkillCAT verifies and routes self-evolved skills for LLM agents

Academic

Physics-Guided Spatiotemporal Learning for Coastal Wave Peak Period Estimation from Video

arXiv cs.LG (Machine Learning) ・ 2026-06-11

Physics-guided video learning estimates coastal wave peak periods

Academic

Clipping Makes Distributed and Federated Asynchronous SGD Robust to Stragglers

arXiv cs.LG (Machine Learning) ・ 2026-06-11

Theory shows clipping makes asynchronous SGD robust to stragglers

Academic

TimeLens: On-Device Artifact Recognition with Retrieval-Augmented Question Answering for the Grand Egyptian Museum

arXiv cs.CL (Computation and Language) ・ 2026-06-11

TimeLens: on-device artifact recognition with RAG for Egyptian museum

Academic

ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm

arXiv cs.CL (Computation and Language) ・ 2026-06-11

ComAct drives industrial CAD software via COM-as-Action paradigm

Academic

PolyAlign: Conditional Human-Distribution Alignment

arXiv cs.CL (Computation and Language) ・ 2026-06-11

PolyAlign matches LLMs to context-specific human response distributions

Academic

SICI: A Semantic-Pragmatic Complexity Index Reveals Regime Shifts in LLM Stance Detection

arXiv cs.CL (Computation and Language) ・ 2026-06-11

SICI complexity index reveals regime shifts in LLM stance detection

Academic

MemRefine: LLM-Guided Compression for Long-Term Agent Memory

arXiv cs.CL (Computation and Language) ・ 2026-06-11

MemRefine compresses agent memory to budget with LLM-judged merging

Academic

NTS-CoT: Mitigating Hallucinations in LLM-based News Timeline Summarization with Chain-of-Thought Reasoning

arXiv cs.CL (Computation and Language) ・ 2026-06-11

NTS-CoT curbs hallucinations in LLM news timeline summarization

Academic

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

Reroute replaces visual-token removal with recoverable routing in VLMs

Academic

Context-Driven Incremental Compression for Multi-Turn Dialogue Generation

arXiv cs.CL (Computation and Language) ・ 2026-06-10

C-DIC compresses multi-turn dialogue context incrementally for stability

Academic

DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

DIRECT allocates test-time compute per prompt for embodied planners

Academic

Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs

arXiv cs.CL (Computation and Language) ・ 2026-06-10

ModSleuth audits invisible recursive dependencies in modern LLMs

Academic

Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization

arXiv cs.CL (Computation and Language) ・ 2026-06-10

RACES recursively composes verifiable RL environments to scale reasoning

Academic

Adjoint Method versus Physics-Informed Neural Networks in PDE-Constrained Inverse Problems

arXiv cs.LG (Machine Learning) ・ 2026-06-10

Adjoint optimization vs PINNs: a fair test on PDE inverse problems

Academic

Fourier Features Let Agents Learn High Precision Policies with Imitation Learning

arXiv cs.LG (Machine Learning) ・ 2026-06-10

Fourier features give point-cloud policies high-precision control

Academic

PROJECTMEM: A Local-First, Event-Sourced Memory and Judgment Layer for AI Coding Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

projectmem adds a local-first memory layer for AI coding agents

Academic

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

arXiv cs.CL (Computation and Language) ・ 2026-06-10

MedMisBench shows LLM medical judgment collapses under misleading context

Academic

The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

Standard Interpretable Model designs interpretability via Lagrangian mechanics

Academic

Holding the FP8 Quality Ceiling at 8-Bit Weights and Activations: INT8 and GGUF Post-Training Quantization of Ideogram 4.0 for Consumer GPUs

arXiv cs.LG (Machine Learning) ・ 2026-06-10

INT8 quantization of Ideogram 4.0 holds FP8 quality on consumer GPUs

Academic

DiffCold: A Diffusion-based Generative Model for Cold-Start Item Recommendation

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

DiffCold uses diffusion to resolve the cold-start recommendation seesaw

Academic

Re-evaluating Confidence Remasking in Masked Diffusion Language Models

arXiv cs.LG (Machine Learning) ・ 2026-06-10

Re-evaluation: WINO remasking adds little in masked diffusion LLMs

Academic

MLT-Dedup: Efficient Large-Scale Online Video Deduplication via Multi-Level Representations and Spatial-Temporal Matching

arXiv cs.LG (Machine Learning) ・ 2026-06-10

MLT-Dedup speeds large-scale online video deduplication

Academic

OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

OpenMedReason releases a ~450K medical multimodal reasoning corpus

Academic

A Controlled Study of Decoding-Time Truthfulness Methods on Instruction-Tuned LLMs

arXiv cs.CL (Computation and Language) ・ 2026-06-10

CHAIR detects hallucinations from per-layer token logits

Academic

nD-RoPE: A Generalized RoPE for n-Dimensional Position Embedding

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

nD-RoPE generalizes RoPE to n-dimensional position embedding

Academic

Augmenting Molecular Language Models with Local $n$-gram Memory

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

MolGram adds local n-gram memory to boost molecular language models

Academic

Bridging the Morphology Gap: Adapting VLA Models to Dexterous Manipulation via Intent-Conditioned Fine-Tuning

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

InDex adapts VLA models to dexterous hands via semantic inheritance

Academic

DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model

arXiv cs.LG (Machine Learning) ・ 2026-06-10

DAM-VLA's asynchronous modality buffers double manipulation success

Academic

FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-10

FORT synthesizes shortcut-resistant tasks to train deep search agents

Academic

On the Limits of LLM-as-Judge for Scientific Novelty Assessment

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-10

RQ-Bench exposes limits of LLM-as-judge for scientific novelty

Academic

Simplicity Suffices for Parameter Noise Injection in Stochastic Gradient Descent

arXiv cs.LG (Machine Learning) ・ 2026-06-10

Simple isotropic noise injection suffices for SGD, study finds

Academic

Decoding Multimodal Cues: Unveiling the Implicit Meaning Behind Hateful Videos

arXiv cs.CL (Computation and Language) ・ 2026-06-10

IARE enables explainable hateful-video detection with rationales

Academic

uva-irlab-conv at SemEval-2026 Task 8: Multi-Turn RAG with Learned Sparse Retrieval and Listwise Reranking

arXiv cs.CL (Computation and Language) ・ 2026-06-10

Multi-turn RAG with sparse retrieval and listwise reranking for SemEval-2026

Academic

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement

arXiv cs.CL (Computation and Language) ・ 2026-06-10

Arbor runs autonomous research via hypothesis-tree refinement

Academic

An Ontology-Guided Multi-Anchor Graph Retrieval Framework for Traffic Legal Liability Determination

arXiv cs.CL (Computation and Language) ・ 2026-06-10

OMAGR uses ontology-guided multi-anchor retrieval for traffic-law liability

Academic

Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code

arXiv cs.CL (Computation and Language) ・ 2026-06-10

CodeSpear jailbreaks LLMs into malicious code via grammar-constrained decoding

Academic

Lius: Translation Model Based Instructional Lingustic Using Continual Instruction Tuning In Kupang Malay

arXiv cs.CL (Computation and Language) ・ 2026-06-10

Lius improves low-resource Kupang Malay translation via continual tuning

Academic

EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-09

EEVEE: multi-dataset test-time prompt learning for self-improving agents

Academic

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories

arXiv cs.CL (Computation and Language) ・ 2026-06-09

Data2Story: a multi-agent framework for verifiable, multimodal data stories

Academic

Flaws in the LLM Automation Narrative

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-09

Human experts still outperform a frontier LLM on a coding data-analysis task

Academic

ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-09

ABC-Bench evaluates LLM agents' biosecurity-relevant capabilities

Academic

First-Order Trajectory Matching: Fast Ensemble Predictions of Chaotic, Turbulent, Stochastic Systems

arXiv cs.LG (Machine Learning) ・ 2026-06-09

FTM learns probability-current velocity from trajectories for cheap forecasts

Academic

DMT: Demographic Conditioning, Morphology-Enhanced Transformer for Cuffless Blood Pressure Estimation from PPG Signals

arXiv cs.LG (Machine Learning) ・ 2026-06-09

DMT: demographic-conditioned Transformer for cuffless BP from PPG

Academic

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-09

TRACE: tree-structured rollout budget allocation for agentic RL

Academic

Towards Autonomous Accelerator Design: FPGA Accelerator Generation with SECDA

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-09

SECDA-DSE: LLM-guided design space exploration for FPGA accelerators

Academic

PhantomBench: Benchmarking the Non-existential Threat of Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-09

PhantomBench: testing LLM hallucination on non-existent concepts

Academic

RoboNaldo: Accurate, Stable and Powerful Humanoid Soccer Shooting via Motion-Guided Curriculum Reinforcement Learning

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-09

RoboNaldo: motion-guided curriculum RL for powerful humanoid soccer shots

Academic

VISTA: A Versatile Interactive User Simulation Toolkit for Agent Evaluation

arXiv cs.CL (Computation and Language) ・ 2026-06-09

VISTA: a versatile user-simulation toolkit for agent evaluation

Academic

A History-Aware Visually Grounded Critic for Computer Use Agents

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-09

HiViG: a history-aware, visually grounded critic for computer-use agents

Academic

T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-09

T1-Bench: high-fidelity evaluation of multi-domain agents across 25 domains

Academic

Generative Archetype-Grounded Item Representations for Sequential Recommendation

arXiv cs.CL (Computation and Language) ・ 2026-06-09

GenAIR: archetype-grounded item representations for sequential recommendation

Academic

Generative Explainability for Next-Generation Networks: LLM-Augmented XAI with Mutual Feature Interactions

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-09

LLM-augmented XAI generates natural-language explanations for networks

Academic

Democratising Camera Trap AI: An Open-Source Model for Detecting UK Mammals

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-09

Open-source UK mammal detector reaches 0.984 [email protected]

Academic

Trace Only What You Need: Structure-Aware On-Demand Hypergraph Memory for Long-Document Question Answering

arXiv cs.CL (Computation and Language) ・ 2026-06-09

DocTrace: structure-aware on-demand hypergraph memory for long-document QA

Academic

Task Robustness via Re-Labelling Vision-Action Robot Data

arXiv cs.LG (Machine Learning) ・ 2026-06-09

TREAD: re-labelling robot vision-action data with VLMs for task robustness

Academic

Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-09

Role-Agent: a single LLM as both agent and environment for co-evolution

Academic

Range Penalization: Theoretical Insights with Applications in Federated Learning

arXiv cs.LG (Machine Learning) ・ 2026-06-09

Range penalization for quantization-friendly federated learning

Academic

Ethical and Technical Limits of Deepfake Speech Datasets

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-09

Auditing 39 deepfake speech datasets reveals fairness gaps and overlap

Academic

Pose-ICL: 3D-Aware In-Context Learning for Pose-Controllable Subject Customization

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-09

Pose-ICL: 3D-aware in-context learning for pose-controllable generation

Academic

Janus: A Benchmark for Goal-Conditioned Information Distortion in LLMs

arXiv cs.CL (Computation and Language) ・ 2026-06-09

JANUS: measuring goal-conditioned information distortion in LLMs

Academic

Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-09

ADAS: attention-discounted reranking for masked diffusion decoding

Academic

Recovering the Zipfian Distribution in Unsupervised Term Discovery

arXiv cs.CL (Computation and Language) ・ 2026-06-09

Graph clustering beats K-means in unsupervised term discovery

Academic

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization

arXiv cs.CL (Computation and Language) ・ 2026-06-09

N-GRPO mixes neighbor embeddings to boost policy optimization

Academic

Infini Memory: Maintainable Topic Documents for Long-Term LLM Agent Memory

arXiv cs.CL (Computation and Language) ・ 2026-06-09

Infini Memory: topic documents for long-term LLM agent memory

Academic

How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMs

arXiv cs.CL (Computation and Language) ・ 2026-06-09

FlowTracer traces attention flow for token-level RL credit

Academic

ParaBridge: Bridging Paralinguistic Perception and Dialogue Behavior in Speech Language Models

arXiv cs.CL (Computation and Language) ・ 2026-06-09

ParaBridge links paralinguistic cues to dialogue behavior

← Story Archive