推論・効率化 A
173 件中 1〜30 件目を表示
-
Sakana AI、日本語特化のLLM API「Sakana Namazu」を提供開始Sakana AI、日本語特化LLM「Namazu」をOpenAI互換APIで提供開始Sakana AIが、日本語と日本の商習慣に特化したLLM API「Sakana Namazu」の提供を開始した。Sakana Chat搭載モデルを更新したもので、Moonshot AIのオープンモデル「Kimi K2.6」をベースに社内データで日本語・業務文脈への適合を進めた。Web検索とコード実行のビルトインツールを備え、OpenAI互換のためbase_urlの変更だけで既存コードから利用できる。高コストなフロンティアモデルと素のオープンモデルの中間を埋める選択肢として位置づける。
-
OpenAI、アクティブユーザー10億人超に 導入企業は200万社超OpenAI、アクティブユーザー10億人・導入企業200万社を突破OpenAIは、アクティブユーザーが10億人、導入企業が200万社を超えたと公表した。推論の保持やコンテキスト管理の改善、本番ソフトウェアの最適化によりコスト削減とトークン生成効率の向上を実現し、GPT-5.6の一部モデルは値下げした。
-
Co-Designing AI Model Attention for Fast, Interactive Long-Context InferenceNVIDIA、長文脈推論を高速化するattention協調設計手法を解説NVIDIAは、エージェント型・長文脈ワークロードの増加でattentionが推論時間の大きな割合を占める課題に対し、モデルのattention機構をハードウェアと協調設計して高速かつ対話的な長文脈推論を実現する手法を紹介した。
-
GQ-FSL: Green Quantized Federated Split Learning
-
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning
-
QASP: Query-Adaptive Robust Vector Search Policy
-
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
-
The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs
-
ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression
-
Adaptive FastOPD: Progress-Aware Rollout Horizon Expansion for Efficient On-Policy Distillation
-
Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?
-
Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember
-
TFGformer: Multivariate Time Series Forecasting via Time-Frequency Graph Learning and Covariate Fusion
-
Analytical and Bootstrap Confidence Intervals of Double Machine Learning: Simulation studies and an application to rural-urban difference in obesity prevalence
-
Beyond Retrieval: Analytic Memory for Multimodal Agents
-
Bridging the Question-Answer Gap in Retrieval-Augmented Generation: Hypothetical Prompt Embeddings
-
OnlineCache: Learning Dynamic Caching Policies with Error Correction for Efficient Diffusion Inference
-
Studying quantization trade-offs for efficient inference deployment in machine translation
-
Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning
-
MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation
-
Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation
-
OsteoCAD: A Human-in-the-Loop Cloud-Edge Framework for Bone Tumor Segmentation
-
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
-
Data Turnstile: A Scalable Open Framework for Function-Calling Data Generation
-
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
-
Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters
-
FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution
-
MOSAIC: Masked Outsourcing of Secure AI Computations
-
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
-
SERUM: State Extraction and Refinement for User Modeling