推論・効率化
A
93 件中 61〜90 件目を表示
-
Beyond Accuracy: Robustness, Cost, and Governance Trade-offs for Vision-Language Models in Templated Document Extraction
-
Backward SDEs-based Diffusion for Physics-Constrained Generation
-
NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities
-
Don't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models
-
Circuit-MLLM: Topological Logic-Guided Latent-Space Visual Reasoning for Circuit Schematic Understanding
-
IROH: Insightful Ranking Of Humor using Multi-Stage Hybrid Retrieval with Rationale-Distilled LLM Judges for JOKER 2026 Track Task 1 English
-
How OpenAI Used Its Own LLMs to Design Its Jalapeño ChipOpenAI、自社LLMで初の自社チップJalapeñoを設計、RTLからテープアウトまで9カ月IEEE Spectrumは、OpenAIが8月25日に公開した初の自社AIアクセラレータ「Jalapeño」の設計プロセスを報じた。4bit演算13.4PFLOPS、232GBメモリを備え、NVIDIA GB300比で最大3.6倍のレイテンシ短縮を主張する。100人弱のチームがGoogle発の高位合成ツールXLSに自社LLMを組み合わせ、構想から初回シリコンまで20カ月未満、RTLからテープアウトまで9カ月で完了。初回シリコン到着後のカーネル最適化はAIにより約40時間で理論性能の0.31%から88.94%へ到達した。物理設計はBroadcomが担当した。
-
Through the Eyes of the Beholder: Biometric and Demographic Conditioning for Multimodal Sexism Detection
-
Option-Aware Retrieval and Task-Specific VLM Adaptation for Medical VQA
-
To Each Language Its Tokenizer: Modular Tokenizers for Efficient Multilingual LLMs
-
Beyond Noise: Understanding and Overcoming Temperature Effects in Analog DNN Inference
-
How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus
-
MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding
-
Dynamic Semantic Compression for Efficient Latent-Space Inference in Large Language Models
-
Reducing the Output-Mode Gap in Speech Language Models via Joint-Output On-Policy Distillation
-
Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings
-
Semiotic Relations and Proof Methods: A Cross-Genre Study of Argument Structure with Large Language Models
-
EMR: Self-Evolving Medical Multi-Agent System via Experience Mining and Reuse
-
Rethinking Heterogeneous System Disaggregation for Subquadratic Attention
-
Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning
-
Attention Quantization for Tabular Foundation Models
-
Label-Guided Knowledge Distillation for 3D-CNNs in Action Recognition
-
TileNet: Tile-Based CNN-SVM Architecture for Autonomous Unmanned Aerial Systems Inspection of Flat Roofs
-
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
-
Parameter-Efficient Retrievers for Polish and European Languages
-
Behavior Quotient Learning for Low-Rank Adaptation of LLM Agents
-
A Graph-Based Approach for Mapping Kernel-Level Telemetry to MITRE ATT&CK
-
GraphAHA: Graph-Based Adaptive Search with Heterogeneous Actions for Test-Time Code Generation
-
Doc2FRC: Length-Consistent Document-Level Machine Translation via Fixed-Range Chunking
-
LifeMem: Enabling Lifelong Experience Reuse for LLM Agents