推論・効率化
A
94 件中 1〜30 件目を表示
-
DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache CompressionDeepSeek-V4.1 Flash、KV キャッシュを 4 倍圧縮し推論を高速化DeepSeek-V4.1 Flash の技術報告を読み解いた解説記事。長時間稼働するエージェント用途で肥大化する KV キャッシュと prefill 計算負荷を抑えることが狙いで、YOCO に倣い全 40 層のうち prefill では 20 層のみを使い活性パラメータを 8B (decode は 16B) に抑える。さらに GQA 的なヘッド数圧縮、層をまたぐ CSA2 のブロック圧縮、FP4 KV キャッシュを組み合わせ、タスク品質を保ったまま KV キャッシュを約 4 倍圧縮したと整理する。
-
「中国AIがClaudeやGPTから数十億トークン抽出」 米当局が暴いた「知識蒸留」の実態米NSA・CISA・FBI、中国AI企業による米モデルからの知識蒸留で共同勧告米国のNSA、CISA、FBIは、中国のAI企業が米国のフロンティアモデルから組織的に「知識蒸留」を行っているとして共同勧告を公開した。プロキシ経由で推論機能などを抽出していると指摘し、対抗措置を推奨している。
-
TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX ThorNVIDIA、MLPerf エッジ agent ベンチを Jetson Thor で 6.4 倍高速化NVIDIA は MLPerf Inference v6.1 の Edge Agentic ベンチで、Qwen3.6-27B を Jetson AGX Thor 1 台で走らせ 1,007 ターンを 24 分 36 秒で完走。llama.cpp 参照実装 (2 時間 37 分) の 6.4 倍で、BFCL 精度は 87.94%。NVFP4 量子化、ツリー型 MTP、KV キャッシュ再利用の併用による。
-
rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference
-
Long-Lived Characters, Local Inference: Incremental Memory Maintenance for Game NPCs
-
Structured Claim-Level Discourse Representations for Dense Health Narratives
-
Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
-
Taming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN
-
Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It
-
FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection
-
A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages
-
Clueing up LLMs with Tool-Augmented Deductive Reasoning
-
Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning
-
Toward Composable Network Digital Twins: A Subgraph-Based Latency Prediction Study
-
RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection
-
Generalist-Specialist Mixture-of-Experts for Rare Pathology Detection in Multimodal Imaging
-
Voice of Reason: Reinforcement Learning for Spoken Math
-
Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification
-
Selection Is Retrieval, Abstention Is Not: On-Device Tool Routing over 70 Korean-English Actions
-
VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge
-
Peak-Aware Short-Term Load Forecasting Across Distribution Grid Aggregation Levels
-
TTM-Bench: A Framework for Text-to-Music System Performance Benchmarking
-
On-the-Fly Homographies Calibration for Multi-Camera Tracking
-
Align, Integrate, and Fire: Efficient Token-Level Alignment for Zero-Shot SpeechLLMs
-
Size Matters: Foundation Model for Czech HTML documents
-
ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
-
Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning
-
Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-TrainGoogle、推論時思考を回避する検索枠組み Retrieve-for-TrainGoogle Research が、AI 検索・推薦で一貫した結果セットを返す枠組み Retrieve-for-Train を公開した。推論時に高コストな自己回帰的推論を重ねる代わりに、強化学習で軽量な拡散モデルを一度学習させ、補完関係にある結果群をまとめて即座に生成する。キャンプ用品の検索でテント・寝袋・コンロを揃えて返すような用途を想定する。
-
LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs
-
JustFit: 200K-Token LLM Serving on a 24 GiB Laptop with Just-in-Time State Management