Inference & Efficiency A
Showing 121–150 of 158
-
Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification
-
How GPT-5.6 fuses frontier intelligence with frontier efficiencyOpenAI: GPT-5.6 fuses frontier intelligence with frontier efficiencyOpenAI explained how GPT-5.6 improves efficiency across models, inference, and agentic workflows while retaining top-tier capability. The company frames it as fusing frontier intelligence with frontier efficiency to deliver more useful AI at lower cost.
-
Pass the Baton: Trajectory-Relayed On-Policy Distillation
-
$π\mathbf{R}^2$: Reactive Real-time Flow Policies
-
Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
-
UniMem: Complementary Episodic-to-Parametric Memory for Boundary-Agnostic Task Streams
-
MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar
-
Parallel Decoding Distillation for Fast Image and Video Generation
-
MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents
-
Quasi-SVD: Learning a Lie-constrained matrix factorisation for real-time imaging
-
A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series
-
The OlmoEarth Platform: Geospatial inference at planetary scaleAllen AI unveils OlmoEarth for planetary-scale geospatial inferenceAllen Institute for AI (Ai2) published OlmoEarth on Hugging Face, an infrastructure platform for geospatial inference at planetary scale over satellite and Earth-observation data. Model size, supported tasks and benchmarks are not verifiable as the body was not retrieved; the 'moe' tag hints at a possible Mixture-of-Experts design, unconfirmed.
-
Penelope: Localized Latent Recurrence for Efficient Structured Reasoning
-
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
-
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
-
Stemma: Induced Decision Regions Reveal LLM Provenance
-
VAD to the Bone: Ultra-Tiny Speech Activity Detection for Edge Deployment
-
HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs
-
AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding
-
Prototype Adaptation for Zero-Shot sEMG Movement Classification
-
LFM2.5-Encoders for Fast Long-Context Inference on CPULiquid AI releases LFM2.5-Encoders for fast long-context inference on CPULiquid AI published LFM2.5-Encoders on Hugging Face, a family of encoder models optimized for fast long-context inference on CPU. The release targets efficient CPU-side embedding and retrieval over long inputs without relying on a GPU. Note: the article body was not available, so model sizes, benchmark figures, supported tasks, and comparisons to existing encoders could not be confirmed; this is a neutral summary synthesized from the title and source (Hugging Face / LiquidAI).
-
Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL
-
Detecting CSAM Text-to-Image LoRAs From Weights
-
IRIS: Reusable Identity Representations from Frozen LLMs for Entity Alignment
-
OrthKD: Extracting Generalized Clinical Knowledge from Heterogeneous Teachers for Lightweight Deployment
-
Mind the Missing Split: Resolving Feature Heterogeneity in Swarm Learning with Random Forests
-
Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization
-
Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors
-
CAST: Game Solvers as Turn-Level Teachers for LLM Agents
-
CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention