Inference & Efficiency A
Showing 91–120 of 158
-
A Sparse Glimpse of the Whole: Train-Free Self-Speculative Decoding
-
Recall Before You Rank: Similarity-Guided Top-$K$ Reuse for Efficient Long-Context Attention
-
Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories
-
Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs
-
From Classification to Regression: Using a Fruitfly to Solve Equations
-
Improving Item Discoverability in e-Commerce Search via Related Intent Generation
-
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
-
Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes
-
InferScale: GPU-Native KV Injection for Personalized LLM Serving
-
On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
-
CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation
-
Mitigating Compounding Error via Video Representation Regularization
-
Generation or Judgement? A Paradigm Perspective on LLM-Based Emotion-Cause Pair Extraction in Conversation
-
Belief-Guided Decision Making with Uncertainty Gating in the Game of Go
-
DIRECT: Direct Decoding for Efficient and Aligned Sequence Labeling with Large Language Models
-
SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning
-
No Data Is Not No Risk: Visibility Aware Graph-Based Inference of Business Conduct Risk
-
Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
-
From Found to Designed: Concepts as a Design Axis for Large Language Models
-
FedTopo: Relation-Level Topology Sharing for Model-Heterogeneous Federated Learning
-
See2Think: Do Multimodal Models Really Use Intermediate Visual States?
-
MediaWiki Code2Code Search: Neural Retrieval for the Semantic Discovery of Open-Source Software Entities
-
Metis: Memory Foundation Model
-
AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution
-
Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability
-
Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes
-
FedWeave: Rethinking the Unit of Specialization in Heterogeneous Federated MoE-LoRA
-
Where Detectors Fail: Closing the Tail-Domain Gap with Expert-Guided Mutual Distillation
-
Which RAG Paradigm Wins at Scale? A Scaling Study of Retrieval-Augmented Generation Paradigms
-
Voice Memory for Agentic Speech Recognition