Inference & Efficiency A
Showing 61–90 of 170
-
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning
-
QuantWAMs: Calibrating at the Right Granularity for World Action Models
-
Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
-
Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners
-
Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
-
Fully Inductive Cardinality Estimation
-
Semi-Supervised Learning for Molecular Graphs via Ensemble Consensus
-
CACHE-UK: A Stability-Aware Memory Editor for Sequentially Updated Quantized LLMs in Finance
-
Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3
-
Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation
-
Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory
-
Operationally Guided Placement-Aware Learning for Industrial Online 3D Bin Packing
-
Are AI Models Working Harder Than They Need to?Are AI models working harder than they need to?IEEE Spectrum examines how much of modern AI relies on massive amounts of multiplication. Questioning whether the neural networks behind everything from generated answers to photo organization really need all that computation, the piece explores the potential to make AI inference far more efficient.
-
AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach
-
OPLD: On-Policy Latent Distillation for Multimodal Reasoning
-
Information Bottleneck Learning for Faithful Time Series Forecasting Explanations
-
MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck
-
From Expert Reduction to Behavioral Divergence: Tracing Numerical State through Sparse MoE Inference
-
Distilling Answer Set Programming Theories from Large Language Models
-
GGC: Selective Query Correction for Reliable Text-to-SPARQL Generation
-
Group-Reflective Self-Distillation for Agentic Reinforcement Learning
-
SemPIC: Learning Semantic Position-Independent KV Caches
-
Stimulus-Evoked Network Dynamics in Human Cortical Organoids: From a Graph-Computational Framework to Repeated-Stimulation Depression
-
A Query-Efficient Stochastic Volume Rendering Framework for Time-Varying Implicit Neural Volumes
-
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
-
Flux-OPD: On-Policy Distillation with Evolving Contexts
-
Driving up Inference Energy on SNNs: Per-Sample and Universal Sponge Attacks
-
Generalization Bounds on Optimal Control for Transformer Training and Wasserstein Distributional Robustness
-
TAPO: Transition-Aware Policy Optimization for LLM Agents
-
Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning