Inference & Efficiency A

Showing 151–171 of 171
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    Where Steering Signals Come From: Activation Source Selection in Activation Steering
    Inference Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    TabRank: Chain-of-Thought Distillation for Table Re-Rankers
    Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • Apple Machine Learning Research · EN Inference & Efficiency extract
    Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers
    Apple details memory-efficient on-device audio synthesis for Siri voices
    Quantization Speech Processing Transformer
    Apple ML published the memory-efficient audio synthesis architecture behind Siri Expressive Voices, which generate configurable speech in real time entirely on device. Powered by its AFM 3 Core Advanced on-device foundation model, a detokenizer converts semantic audio tokens into high-fidelity audio within the Apple Matrix Coprocessor (AMX) budget using a residual vector quantization (RVQ), streaming design. Details past the streaming component are truncated in the excerpt.
    Read original (Apple Machine Learning Research) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation
    AI Agents Deep Learning Neural Network
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
    Inference Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Denial of Deadline: Network-Driven Accuracy Collapse in Distributed Inference Pipelines
    Deep Learning Inference Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    Co-Learning for Missing Arbitrary Modalities in Multi-modal Classification
    Deep Learning Inference Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    Explainable Reinforcement Learning via Physics-Aware Policy Distillation
    Neural Network Retrieval-Augmented Generation (RAG) Reinforcement Learning Robotics
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention
    DeepSeek Inference Reinforcement Learning
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents
    AI Agents Reinforcement Learning Software Engineering
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Infrastructure & Hardware
    From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference
    Deep Learning Inference Neural Network
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    PYPM-GGD: Pitman-Yor Process Mixture with Generalized Gaussian Density using ADAM
    Deep Learning Inference Reinforcement Learning
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Multimodal
    CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding
    Computer Vision Deep Learning Inference Machine Learning Software Engineering
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    Evaluating Fuzz Testing for Reinforcement Learning Agents
    AI Agents Retrieval-Augmented Generation (RAG) Reinforcement Learning Robotics
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    Bit-Accurate FPGA Evaluation of Learned Feature Gating in a Fixed-Point Fourier-Feature Automatic Modulation Classifier
    Quantization Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Hierarchical Group-Conditional Conformal Risk Control for Selective Prediction in Language Models
    Llama Machine Learning Quantization
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    EgoPlay: Event-Triggered Video Editing for Egocentric Streams
    Deep Learning Fine-tuning Inference Neural Network Transformer
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding
    Inference Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings
    Inference
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages
    Embeddings Reinforcement Learning Unsupervised Learning
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    The K-SCAN Clustering Algorithm
    Algorithms & Theory Quantization
    Read original (arXiv cs.LG (Machine Learning)) ↗