Inference & Efficiency A
Showing 31–60 of 170
-
vLLM for Baidu KunlunBaidu open-sources a vLLM port for its Kunlun AI chipsBaidu published vLLM-Kunlun on GitHub, a port of the high-throughput vLLM inference engine targeting its in-house Kunlun AI accelerators. The project expands the options for running LLM inference on non-NVIDIA hardware.
-
Can Zero-Shot LLMs Predict Child Malnutrition? A Fairness and Temporal Robustness Study
-
TransMem: Transforming Hidden States into Memory for Large Language Models
-
GoldenRetriever: Non-Interactive Homomorphic Encrypted Retrieval for Privacy-Preserving RAG
-
Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning
-
Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models
-
BLADE: Boundary-Expanded and Layer-Adaptive Dynamic Exit for Efficient LLM Reasoning
-
Thinking Machines、軽量モデル「Inkling-Small」正式公開 サイズ4分の1で「Inkling」に匹敵する性能Thinking Machines releases Inkling-Small, matching Inkling at 1/4 the sizeThinking Machines Lab released the final version of Inkling-Small, an open-weight AI model. At a quarter the size of its predecessor, the company says data improvements and reinforcement learning let it match the larger Inkling on tasks such as code generation.
-
Chromeに13年以上潜んでいた脆弱性、AIで発見 直近2回のアプデで過去23回分を上回るバグ修正Google's Gemini agent finds 13-year-old Chrome flaw; tests twice-weekly updatesGoogle detailed its use of AI for Chrome security, saying a Gemini-based agent uncovered a vulnerability hidden for over 13 years and that its last two updates fixed more bugs than the previous 23 combined. To counter faster AI-driven attacks, Google is trialing twice-weekly security updates.
-
Google、ロボット向けAI「Gemini Robotics 2」発表 ヒューマノイドの全身制御や指先作業を実現Google unveils Gemini Robotics 2 for whole-body and fine fingertip controlGoogle and Google DeepMind announced Gemini Robotics 2, a family of robotics AI models supporting humanoid whole-body control, fine fingertip manipulation, and multi-robot collaboration. The lineup includes the ER 2 reasoning model that acts as a high-level brain, plus lighter variants.
-
Token-Level Diagnosis of Sycophancy in LLMs with Attribution-Guided Steering
-
ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
-
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
-
MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision Transformers
-
$β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
-
Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs
-
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
-
Doubly Robust Functional Representation Learning for Longitudinal Causal Inference with Irregular Histories
-
APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems
-
MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems
-
Stage-Replay Divergence Follows the KV Cache: Fixed-Prefix Precision Controls and Bidirectional Cache Transplantation
-
A Fuzzy Rule-based Neuro-Symbolic Approach for Pipe Severity Prediction in Sewer Networks
-
Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
-
Improving Mental Health Screening and Early Risk Detection in Spanish
-
Towards Autonomous Aircraft Surveillance from Nanosatellites through On-Board Inference and Generative Data Augmentation
-
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute
-
Machines that know they are aging: a framework for hardware-aware autonomous intelligence
-
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models
-
QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Quantum Error Correction
-
When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence