Inference & Efficiency

A
Showing 1–30 of 93
  • Hacker News (Front Page) · EN Inference & Efficiency
    DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression
    DeepSeek-V4.1 Flash pushes KV cache compression 4x for agent workloads
    DeepSeek
    A close read of the DeepSeek-V4.1 Flash technical report, whose stated aim is pushing KV cache compression to its limit for long-horizon agent workflows. Borrowing from YOCO, only 20 of the 40 layers run during prefill, cutting prefill active parameters to 8B against 16B at decode. GQA-style head reduction, cross-layer CSA2 block compression and an FP4 KV cache together shrink the KV cache roughly 4x while keeping task quality.
    Read original (Hacker News (Front Page)) ↗
  • ITmedia AI+ · JA New Model Releases
    「中国AIがClaudeやGPTから数十億トークン抽出」 米当局が暴いた「知識蒸留」の実態
    US NSA, CISA and FBI warn of Chinese distillation of US frontier models
    Claude GPT Inference
    The NSA, CISA and FBI issued a joint advisory alleging Chinese AI companies are systematically distilling US frontier models, extracting reasoning capabilities and other behaviour through proxy access. The agencies recommend countermeasures.
    Read original (ITmedia AI+) ↗
  • NVIDIA Developer Blog · EN Inference & Efficiency
    TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor
    NVIDIA runs MLPerf Edge Agentic 6.4x faster on one Jetson AGX Thor
    AI Agents Generative AI Inference NVIDIA Software Engineering
    In MLPerf Inference v6.1's Edge Agentic benchmark, NVIDIA's TensorRT Edge-LLM ran Qwen3.6-27B on one Jetson AGX Thor, finishing 1,007 turns in 24m36s — 6.4x faster than the llama.cpp reference (2h37m). NVFP4 quantization, tree-based MTP and KV cache reuse drive the gain.
    Read original (NVIDIA Developer Blog) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference
    Computer Vision Inference Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    Long-Lived Characters, Local Inference: Incremental Memory Maintenance for Game NPCs
    Inference Quantization
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    Structured Claim-Level Discourse Representations for Dense Health Narratives
    Inference Neural Network Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Multimodal
    Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
    Computer Vision Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Taming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN
    AI Agents Neural Network OpenAI
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It
    Neural Network Software Engineering
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection
    Fine-tuning Inference Reinforcement Learning Speech Processing
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages
    Inference Neural Network Natural Language Processing (NLP) Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    Clueing up LLMs with Tool-Augmented Deductive Reasoning
    AI Agents Gemini GPT Inference
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning
    Embeddings Inference Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN New Model Releases
    Toward Composable Network Digital Twins: A Subgraph-Based Latency Prediction Study
    Machine Learning Reinforcement Learning
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.CL (Computation and Language) · EN Multimodal
    RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection
    AI Agents Computer Vision Inference Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    Generalist-Specialist Mixture-of-Experts for Rare Pathology Detection in Multimodal Imaging
    Inference Mixture of Experts (MoE) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    Voice of Reason: Reinforcement Learning for Spoken Math
    Fine-tuning Reinforcement Learning Software Engineering Speech Processing
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification
    Embeddings Inference Neural Network Speech Processing
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Infrastructure & Hardware
    Selection Is Retrieval, Abstention Is Not: On-Device Tool Routing over 70 Korean-English Actions
    Neural Network Reinforcement Learning Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge
    Inference Reinforcement Learning
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.LG (Machine Learning) · EN Inference & Efficiency
    Peak-Aware Short-Term Load Forecasting Across Distribution Grid Aggregation Levels
    Inference Machine Learning Reinforcement Learning
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.LG (Machine Learning) · EN New Model Releases
    TTM-Bench: A Framework for Text-to-Music System Performance Benchmarking
    Machine Learning Neural Network Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Policy & Regulation
    On-the-Fly Homographies Calibration for Multi-Camera Tracking
    Meta
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    Align, Integrate, and Fire: Efficient Token-Level Alignment for Zero-Shot SpeechLLMs
    Fine-tuning Natural Language Processing (NLP) Speech Processing
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    Size Matters: Foundation Model for Czech HTML documents
    Machine Learning Reinforcement Learning
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
    Computer Vision Quantization
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN Training & Fine-tuning
    Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning
    AI Agents Fine-tuning Inference Neural Network Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • Google Research Blog · EN Inference & Efficiency
    Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train
    Google unveils Retrieve-for-Train, skipping inference-time reasoning
    Algorithms & Theory Data Mining Generative AI Inference
    Google Research introduced Retrieve-for-Train, which returns a coherent slate of results rather than one best match. Instead of inference-time reasoning, it trains a lightweight diffusion model once via reinforcement learning to generate the slate instantly.
    Read original (Google Research Blog) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs
    Embeddings Inference Quantization Speech Processing
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    JustFit: 200K-Token LLM Serving on a 24 GiB Laptop with Just-in-Time State Management
    Inference Machine Learning Neural Network Quantization Software Engineering
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗