Inference & Efficiency A

Showing 121–150 of 158
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification
    Inference Machine Learning Retrieval-Augmented Generation (RAG) Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • OpenAI Blog · EN New Model Releases extract
    How GPT-5.6 fuses frontier intelligence with frontier efficiency
    OpenAI: GPT-5.6 fuses frontier intelligence with frontier efficiency
    GPT Inference
    OpenAI explained how GPT-5.6 improves efficiency across models, inference, and agentic workflows while retaining top-tier capability. The company frames it as fusing frontier intelligence with frontier efficiency to deliver more useful AI at lower cost.
    Read original (OpenAI Blog) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Pass the Baton: Trajectory-Relayed On-Policy Distillation
    Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Multimodal
    $π\mathbf{R}^2$: Reactive Real-time Flow Policies
    Computer Vision Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
    AI Agents Inference Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Developer Tools
    UniMem: Complementary Episodic-to-Parametric Memory for Boundary-Agnostic Task Streams
    AI Agents Deep Learning Inference Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar
    Inference Quantization Retrieval-Augmented Generation (RAG) Transformer
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    Parallel Decoding Distillation for Fast Image and Video Generation
    Inference
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Developer Tools
    MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents
    AI Agents Machine Learning Neural Network Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN New Model Releases
    Quasi-SVD: Learning a Lie-constrained matrix factorisation for real-time imaging
    Algorithms & Theory Neural Network Reinforcement Learning
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series
    Embeddings Inference Neural Network Retrieval-Augmented Generation (RAG) Software Engineering
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • Hugging Face Blog · EN Inference & Efficiency extract
    The OlmoEarth Platform: Geospatial inference at planetary scale
    Allen AI unveils OlmoEarth for planetary-scale geospatial inference
    Inference Mixture of Experts (MoE)
    Allen Institute for AI (Ai2) published OlmoEarth on Hugging Face, an infrastructure platform for geospatial inference at planetary scale over satellite and Earth-observation data. Model size, supported tasks and benchmarks are not verifiable as the body was not retrieved; the 'moe' tag hints at a possible Mixture-of-Experts design, unconfirmed.
    Read original (Hugging Face Blog) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Penelope: Localized Latent Recurrence for Efficient Structured Reasoning
    Deep Learning Inference Software Engineering Transformer
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Multimodal
    SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
    Computer Vision Inference Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
    Inference Llama Machine Learning Neural Network Software Engineering
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    Stemma: Induced Decision Regions Reveal LLM Provenance
    Inference Neural Network Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    VAD to the Bone: Ultra-Tiny Speech Activity Detection for Edge Deployment
    Inference Neural Network Quantization Speech Processing
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs
    AI Agents Inference Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding
    Inference Reinforcement Learning
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    Prototype Adaptation for Zero-Shot sEMG Movement Classification
    Embeddings Inference Reinforcement Learning
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • Hugging Face Blog · EN Inference & Efficiency extract
    LFM2.5-Encoders for Fast Long-Context Inference on CPU
    Liquid AI releases LFM2.5-Encoders for fast long-context inference on CPU
    Inference
    Liquid AI published LFM2.5-Encoders on Hugging Face, a family of encoder models optimized for fast long-context inference on CPU. The release targets efficient CPU-side embedding and retrieval over long inputs without relying on a GPU. Note: the article body was not available, so model sizes, benchmark figures, supported tasks, and comparisons to existing encoders could not be confirmed; this is a neutral summary synthesized from the title and source (Hugging Face / LiquidAI).
    Read original (Hugging Face Blog) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL
    AI Agents Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN Training & Fine-tuning
    Detecting CSAM Text-to-Image LoRAs From Weights
    Fine-tuning Inference Meta
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    IRIS: Reusable Identity Representations from Frozen LLMs for Entity Alignment
    Inference Reinforcement Learning
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    OrthKD: Extracting Generalized Clinical Knowledge from Heterogeneous Teachers for Lightweight Deployment
    Neural Network Retrieval-Augmented Generation (RAG) Transformer
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    Mind the Missing Split: Resolving Feature Heterogeneity in Swarm Learning with Random Forests
    Algorithms & Theory Inference Machine Learning Reinforcement Learning
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization
    Algorithms & Theory Inference Quantization Reinforcement Learning
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors
    Inference
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    CAST: Game Solvers as Turn-Level Teachers for LLM Agents
    AI Agents Retrieval-Augmented Generation (RAG) Reinforcement Learning Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention
    Inference Neural Network
    Read original (arXiv cs.CL (Computation and Language)) ↗