Inference & Efficiency
A
Showing 31–60 of 93
-
Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback
-
Dense vs. MoE Models: Active Parameters, Throughput, and When to Choose EachNVIDIA compares dense and MoE models on active parameters and throughputNVIDIA published a guide to choosing between dense and Mixture-of-Experts models. Using Nemotron 3.5 Lightning, it shows how a 30B model can activate only 3B parameters per token while keeping the larger model's capacity, and maps active parameters to throughput.
-
How NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera RubinNVIDIA on Groq 3 LPX deterministic execution on Vera RubinNVIDIA described how deterministic execution in Groq 3 LPX delivers power-efficient, high-interactivity inference on Vera Rubin. With power the defining constraint for AI factories, fixed execution timing cuts idle waits while preserving responsiveness.
-
FlashVector: Agent for Hierarchical Model Serving Stack Optimization
-
Personalized Federated Learning through Global Knowledge Distillation and Local Head Adaptation
-
ECHO: Early-layer Collaborative Hierarchical Orchestration with Bonus Logits in Speculative Decoding
-
Cross-Domain Inference for Human Localization: Applying Wi-Fi RSSI Data to CSI-Trained Models
-
End-to-End Latency-Minimizing and Load-Balanced Request Scheduling for Edge LLM Inference in Agentic AI Services
-
LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers
-
An Empirical Study of Counterfactual Self-Explanations in LLMs
-
Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs
-
Sample-Conditioned Representation Selection for Audio Few-Shot Learning
-
Neuro-Symbolic Hierarchical Intention Anticipation in Human Behavior
-
Autoformalizing Argumentative Material Inferences
-
Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation
-
MedPCFM-TED: One-Step Point Cloud Flow Matching for Implant Generation via Teacher-Guided Endpoint Distillation
-
Verbalizing Subliminal Learning Effects Using Text Optimization
-
Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science
-
Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States
-
Privacy-Aligned Personalized Federated Learning with Compact Adaptation and Variable-Length Gaussian Communication
-
SlipSense: Multimodal Tactile Learning for Low-Latency and Generalized Slip Detection
-
Bridging Control, Inference, Transport, and Thermodynamics: From Theory to Applications in Learning
-
Anatomical Grounding and Leakage-Aware Multimodal Contrastive Learning for Alzheimer's Disease Classification from Structural MRI
-
Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport
-
Learning to Coach for Experiential Learning
-
Proportional-Fair Resource Allocation and Dual-Threshold Early-Exit Inference for Secure Cooperative Multi-Layer Edge Intelligence
-
Event-Native Symbolic-Temporal Spike Encoding Framework for Heterogeneous Cyber Streams
-
Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models
-
Look Before You Leap: Factual Decoding with Internal Attribution Signals
-
Merging the Knowledge of LLMs for Automatic Speech Recognition