Inference & Efficiency A
Showing 151–171 of 171
-
Where Steering Signals Come From: Activation Source Selection in Activation Steering
-
TabRank: Chain-of-Thought Distillation for Table Re-Rankers
-
Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion TransformersApple details memory-efficient on-device audio synthesis for Siri voicesApple ML published the memory-efficient audio synthesis architecture behind Siri Expressive Voices, which generate configurable speech in real time entirely on device. Powered by its AFM 3 Core Advanced on-device foundation model, a detokenizer converts semantic audio tokens into high-fidelity audio within the Apple Matrix Coprocessor (AMX) budget using a residual vector quantization (RVQ), streaming design. Details past the streaming component are truncated in the excerpt.
-
How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation
-
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
-
Denial of Deadline: Network-Driven Accuracy Collapse in Distributed Inference Pipelines
-
Co-Learning for Missing Arbitrary Modalities in Multi-modal Classification
-
Explainable Reinforcement Learning via Physics-Aware Policy Distillation
-
PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention
-
SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents
-
From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference
-
PYPM-GGD: Pitman-Yor Process Mixture with Generalized Gaussian Density using ADAM
-
CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding
-
Evaluating Fuzz Testing for Reinforcement Learning Agents
-
Bit-Accurate FPGA Evaluation of Learned Feature Gating in a Fixed-Point Fourier-Feature Automatic Modulation Classifier
-
Hierarchical Group-Conditional Conformal Risk Control for Selective Prediction in Language Models
-
EgoPlay: Event-Triggered Video Editing for Egocentric Streams
-
LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding
-
EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings
-
From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages
-
The K-SCAN Clustering Algorithm