Inference & Efficiency

A
Showing 31–60 of 93
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback
    Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • NVIDIA Developer Blog · EN Inference & Efficiency
    Dense vs. MoE Models: Active Parameters, Throughput, and When to Choose Each
    NVIDIA compares dense and MoE models on active parameters and throughput
    Generative AI Inference Mixture of Experts (MoE) NVIDIA
    NVIDIA published a guide to choosing between dense and Mixture-of-Experts models. Using Nemotron 3.5 Lightning, it shows how a 30B model can activate only 3B parameters per token while keeping the larger model's capacity, and maps active parameters to throughput.
    Read original (NVIDIA Developer Blog) ↗
  • NVIDIA Developer Blog · EN Infrastructure & Hardware
    How NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera Rubin
    NVIDIA on Groq 3 LPX deterministic execution on Vera Rubin
    Generative AI Inference NVIDIA
    NVIDIA described how deterministic execution in Groq 3 LPX delivers power-efficient, high-interactivity inference on Vera Rubin. With power the defining constraint for AI factories, fixed execution timing cuts idle waits while preserving responsiveness.
    Read original (NVIDIA Developer Blog) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    FlashVector: Agent for Hierarchical Model Serving Stack Optimization
    AI Agents Machine Learning NVIDIA
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    Personalized Federated Learning through Global Knowledge Distillation and Local Head Adaptation
    Neural Network
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    ECHO: Early-layer Collaborative Hierarchical Orchestration with Bonus Logits in Speculative Decoding
    Fine-tuning Inference Neural Network Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    Cross-Domain Inference for Human Localization: Applying Wi-Fi RSSI Data to CSI-Trained Models
    Inference Neural Network
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    End-to-End Latency-Minimizing and Load-Balanced Request Scheduling for Edge LLM Inference in Agentic AI Services
    Inference Neural Network Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers
    Deep Learning Inference Neural Network Reinforcement Learning Transformer
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN Infrastructure & Hardware
    An Empirical Study of Counterfactual Self-Explanations in LLMs
    Inference Llama Neural Network
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs
    Software Engineering
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Sample-Conditioned Representation Selection for Audio Few-Shot Learning
    Inference Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Neuro-Symbolic Hierarchical Intention Anticipation in Human Behavior
    Deep Learning Inference Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    Autoformalizing Argumentative Material Inferences
    Inference
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation
    Neural Network
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    MedPCFM-TED: One-Step Point Cloud Flow Matching for Implant Generation via Teacher-Guided Endpoint Distillation
    Inference Reinforcement Learning Reinforcement Learning from Human Feedback (RLHF)
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    Verbalizing Subliminal Learning Effects Using Text Optimization
    Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science
    Gemini Google Inference Machine Learning Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States
    Llama Reinforcement Learning
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.LG (Machine Learning) · EN New Model Releases
    Privacy-Aligned Personalized Federated Learning with Compact Adaptation and Variable-Length Gaussian Communication
    Deep Learning Quantization
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    SlipSense: Multimodal Tactile Learning for Low-Latency and Generalized Slip Detection
    Computer Vision Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    Bridging Control, Inference, Transport, and Thermodynamics: From Theory to Applications in Learning
    Inference Machine Learning Neural Network Reinforcement Learning
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Anatomical Grounding and Leakage-Aware Multimodal Contrastive Learning for Alzheimer's Disease Classification from Structural MRI
    Inference Transformer
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport
    Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    Learning to Coach for Experiential Learning
    Inference Neural Network
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    Proportional-Fair Resource Allocation and Dual-Threshold Early-Exit Inference for Secure Cooperative Multi-Layer Edge Intelligence
    Inference Meta Neural Network Reinforcement Learning
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    Event-Native Symbolic-Temporal Spike Encoding Framework for Heterogeneous Cyber Streams
    Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models
    Deep Learning Inference Retrieval-Augmented Generation (RAG) Speech Processing
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Look Before You Leap: Factual Decoding with Internal Attribution Signals
    Inference Machine Learning Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    Merging the Knowledge of LLMs for Automatic Speech Recognition
    Inference Retrieval-Augmented Generation (RAG) Reinforcement Learning Speech Processing
    Read original (arXiv cs.CL (Computation and Language)) ↗