Inference & Efficiency A
Showing 1–30 of 173
-
Sakana AI、日本語特化のLLM API「Sakana Namazu」を提供開始Sakana AI launches Namazu, a Japanese-focused OpenAI-compatible LLM APISakana AI released Namazu, an LLM API tuned for Japanese and local business use. Built on Moonshot AI's open Kimi K2.6 and refined with in-house data, it adds built-in web search and code execution. Being OpenAI-compatible, existing code works by swapping the base_url, filling the gap between costly frontier models and raw open ones.
-
OpenAI、アクティブユーザー10億人超に 導入企業は200万社超OpenAI passes 1 billion active users and 2 million business customersOpenAI said it surpassed one billion active users and two million business customers. It cited efficiency gains from retained reasoning, better context management, and production optimization that cut costs and improved token throughput, alongside price cuts on some GPT-5.6 models.
-
Co-Designing AI Model Attention for Fast, Interactive Long-Context InferenceNVIDIA details co-designed attention for fast long-context inferenceNVIDIA describes co-designing model attention with hardware to speed up interactive long-context inference. As agentic and long-context workloads grow, attention takes a larger share of inference time, and the approach targets that bottleneck for faster serving.
-
GQ-FSL: Green Quantized Federated Split Learning
-
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning
-
QASP: Query-Adaptive Robust Vector Search Policy
-
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
-
The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs
-
ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression
-
Adaptive FastOPD: Progress-Aware Rollout Horizon Expansion for Efficient On-Policy Distillation
-
Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?
-
Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember
-
TFGformer: Multivariate Time Series Forecasting via Time-Frequency Graph Learning and Covariate Fusion
-
Analytical and Bootstrap Confidence Intervals of Double Machine Learning: Simulation studies and an application to rural-urban difference in obesity prevalence
-
Beyond Retrieval: Analytic Memory for Multimodal Agents
-
Bridging the Question-Answer Gap in Retrieval-Augmented Generation: Hypothetical Prompt Embeddings
-
OnlineCache: Learning Dynamic Caching Policies with Error Correction for Efficient Diffusion Inference
-
Studying quantization trade-offs for efficient inference deployment in machine translation
-
Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning
-
MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation
-
Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation
-
OsteoCAD: A Human-in-the-Loop Cloud-Edge Framework for Bone Tumor Segmentation
-
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
-
Data Turnstile: A Scalable Open Framework for Function-Calling Data Generation
-
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
-
Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters
-
FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution
-
MOSAIC: Masked Outsourcing of Secure AI Computations
-
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
-
SERUM: State Extraction and Refinement for User Modeling