開発者ツール B
430 件中 121〜150 件目を表示
-
Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
-
Cybersecurity Detection Classification with Reasoning-enabled Language Models
-
Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation
-
Oracle-Budgeted Molecular Optimization with Short-Term Graph Memory
-
Metaphor Tracer: A Theory-Informed Analysis of Hidden States
-
Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image Classification
-
Negative controls reveal volume-driven confounding in radiomics and imaging foundation model features
-
QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Quantum Error Correction
-
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning
-
QQWorld: Quantile-Quantile Matching for World Model Regularization
-
NVIDIA Exemplar Cloud: Lessons for Unlocking Full Performance on AI InfrastructureNVIDIA、AI基盤の性能を引き出すExemplar Cloudの知見を公開NVIDIAは、H100やGB200 NVL72、GB300 NVL72など同一構成のAIクラスタでも性能が大きく異なりうる点に着目し、Exemplar Cloudの取り組みからインフラの実力を最大限引き出すための知見を紹介した。
-
Windowed thinning and query complexity for the bouncy particle and Zigzag samplers
-
Can Large Language Models Execute Parent Orders?
-
Hierarchical Multilevel Monte Carlo for Order-Optimal Neural Actor-Critic in Average-Reward CMDPs
-
When Specifications Conflict: A Symmetry-Based Framework for Measuring LLM Preferences
-
HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection
-
How Benchmarks Mis-Score Computer-Use Agents
-
ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
-
Teffic-Audio: Tell Fact from Fiction
-
LLMs struggle to simulate human belief updates in controlled environments
-
Reflected diffusion, no-flux continuity equations and confined Lagrangian flows in bounded domains
-
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaborationDeepMind、映像理解と多ロボット連携のGemini Robotics ER 2を発表DeepMindは、ロボット向けモデルGemini Robotics ER 2を発表した。映像理解、タスクの分解・調整、複数ロボットの協調を強化し、ロボットが現実世界の課題を推論しながら協力して解決できるようにする段階的な進歩と位置づける。
-
From Japan, Products the World Will Use: An Interview with Sakana AI's Head of Product DevelopmentSakana AI製品開発責任者、世界で使われる日本発プロダクトを語るSakana AIの製品開発責任者へのインタビュー記事。日本発で世界に使われるプロダクトを生み出す狙いや、同社の製品開発の考え方が語られている。国内AIスタートアップの製品戦略を示す内容となっている。
-
Investigating three real-world incidents in our cybersecurity evaluationsAnthropic、サイバーセキュリティ評価で実世界3件の事例を調査AnthropicのFrontier Red Teamは、自社のサイバーセキュリティ評価に関連する実世界の3件のインシデントを調査した結果を公表した。モデルの悪用リスクや評価手法の妥当性を検証し、フロンティアモデルの安全性向上に役立てる。
-
Paying for Honesty Without Knowing the Truth: Reputation-Penalty Design for LLM Marketplace Agents
-
Measuring Distortion in the Empty Regions of Dimensionality Reduction Scatterplots with the Gap Index
-
PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?
-
One Human, $N$ Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence
-
Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing
-
From Textual Requirements to Microservice Architectures - A Comprehensive Evaluation of LLM-Based Design Synthesis