推論・効率化
A
93 件中 31〜60 件目を表示
-
Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback
-
Dense vs. MoE Models: Active Parameters, Throughput, and When to Choose EachNVIDIA、Dense と MoE の使い分けを活性パラメータ視点で解説NVIDIA が、密結合 (Dense) モデルと Mixture-of-Experts (MoE) モデルの選び分けを解説した。総パラメータ 30B のうちトークンあたり 3B だけを活性化しながら大規模モデル並みの容量を使える仕組みを Nemotron 3.5 Lightning を例に示し、活性パラメータ数とスループットの関係から用途ごとの最適解を整理している。
-
How NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera RubinNVIDIA、Groq 3 LPX の決定的実行で電力効率と応答性を両立NVIDIA が、Vera Rubin 上で動く Groq 3 LPX の決定的実行方式を解説した。AI ファクトリでは電力が最大の制約になるため、実行タイミングを確定させて無駄な待ちを減らし、高い対話性を保ちながら電力効率を高める狙い。プラットフォームを構成する各要素の性能を最大化する設計思想の一環と位置づけている。
-
FlashVector: Agent for Hierarchical Model Serving Stack Optimization
-
Personalized Federated Learning through Global Knowledge Distillation and Local Head Adaptation
-
ECHO: Early-layer Collaborative Hierarchical Orchestration with Bonus Logits in Speculative Decoding
-
Cross-Domain Inference for Human Localization: Applying Wi-Fi RSSI Data to CSI-Trained Models
-
End-to-End Latency-Minimizing and Load-Balanced Request Scheduling for Edge LLM Inference in Agentic AI Services
-
LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers
-
An Empirical Study of Counterfactual Self-Explanations in LLMs
-
Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs
-
Sample-Conditioned Representation Selection for Audio Few-Shot Learning
-
Neuro-Symbolic Hierarchical Intention Anticipation in Human Behavior
-
Autoformalizing Argumentative Material Inferences
-
Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation
-
MedPCFM-TED: One-Step Point Cloud Flow Matching for Implant Generation via Teacher-Guided Endpoint Distillation
-
Verbalizing Subliminal Learning Effects Using Text Optimization
-
Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science
-
Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States
-
Privacy-Aligned Personalized Federated Learning with Compact Adaptation and Variable-Length Gaussian Communication
-
SlipSense: Multimodal Tactile Learning for Low-Latency and Generalized Slip Detection
-
Bridging Control, Inference, Transport, and Thermodynamics: From Theory to Applications in Learning
-
Anatomical Grounding and Leakage-Aware Multimodal Contrastive Learning for Alzheimer's Disease Classification from Structural MRI
-
Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport
-
Learning to Coach for Experiential Learning
-
Proportional-Fair Resource Allocation and Dual-Threshold Early-Exit Inference for Secure Cooperative Multi-Layer Edge Intelligence
-
Event-Native Symbolic-Temporal Spike Encoding Framework for Heterogeneous Cyber Streams
-
Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models
-
Look Before You Leap: Factual Decoding with Internal Attribution Signals
-
Merging the Knowledge of LLMs for Automatic Speech Recognition