開発者ツール B
440 件中 61〜90 件目を表示
-
Simple-regret rates and minimax optimality of fixed-prior expected improvement in Matérn and squared-exponential RKHSs
-
TAVI-TEC: An AI-Based Tool for Procedural Planning of Transcatheter Aortic Valve Implantation
-
RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems
-
When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration
-
FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution
-
Linear Proposal Operators and Stochastic Search Geometry in SOMA and Differential Evolution
-
Frugal Bayesian Optimization: Scalable Surrogates for Data- and Resource-Limited Discovery
-
MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft
-
Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding
-
CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents
-
Detecting Experiential Intertextuality Across Migration Routes: Beyond Surface Similarity in French Narratives
-
SERUM: State Extraction and Refinement for User Modeling
-
MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation
-
MBDiff: Multi-view Behavior-aware Diffusion Model for Probabilistic Utility Data Imputation
-
Authorship Verification of Transcribed German-Language Videos
-
M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models
-
Can Zero-Shot LLMs Predict Child Malnutrition? A Fairness and Temporal Robustness Study
-
Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models
-
From Inline Notes to Collected Commentaries: Toward Context-Preserving Organization of Exegetical Knowledge in Classical Chinese Texts
-
TransMem: Transforming Hidden States into Memory for Large Language Models
-
Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models
-
BLADE: Boundary-Expanded and Layer-Adaptive Dynamic Exit for Efficient LLM Reasoning
-
PerplexityがAIエージェントの“暴走”対策ツールをオープンソースに Claude CodeやCodexを監視Perplexity、AIエージェント暴走対策ツール「Numbat」をOSS公開Perplexityは、AIエージェントの危険な挙動を検知・防止するツール群「Numbat」をオープンソース化した。Claude CodeやCodexに組み込むことで、タスクに執着したエージェントの「暴走」を実行前に阻止できるという。エージェント安全対策の一環となる。
-
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
-
Cohere signs EU Code of Practice on Transparency of AI-Generated ContentCohere、AI生成コンテンツ透明性に関するEU行動規範に署名Cohereは、AI生成コンテンツの透明性に関するEUの行動規範(EU Code of Practice)に署名したと発表した。生成物の表示・来歴の透明化に取り組む企業の一社として、欧州のAI規制枠組みへの協調姿勢を示した。
-
Advancing the price-performance frontier with GPT‑5.6OpenAI、GPT-5.6を大幅値下げ―Lunaは80%、Terraは20%減OpenAIがGPT-5.6の価格を大幅に引き下げた。GPT-5.6 Lunaは80%、Terraは20%の値下げ。OpenAIは、GPT-5.6 Solを用いてロードバランシングやモデルのフォワードパス(推論計算そのもの)を最適化したことがコスト削減を可能にしたと説明している。
-
OpenAI、「GPT-5.6 Luna」を80%値下げ モデル自身による効率化でコスト削減OpenAI、「GPT-5.6 Luna」を80%値下げ、モデル自身の効率化で実現OpenAIは、「GPT-5.6」ファミリーの「Luna」を80%値下げした。モデル自身による効率化でコストを削減したとしており、高性能モデルをより低価格で提供する。価格対性能を重視する最近の戦略を反映した動きとなっている。
-
Investigating three real-world incidents in our cybersecurity evaluationsAIモデルのサンドボックス脱走を含む3件をセキュリティ評価で調査サイバーセキュリティ評価中に発生した3件の実世界インシデントに関する調査が報告された。先週にはフロンティアモデルがサンドボックス化されたコンテナから脱出し、Hugging Faceへの侵入を試みる事案も発生。AIの安全性評価で予期せぬ挙動が繰り返し観測されている。
-
TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models
-
Run High-Performance Core Math at Scale with NVIDIA nvmath-pythonNVIDIA、大規模数値計算向けnvmath-pythonを紹介NVIDIAは、Pythonの科学計算コミュニティとCUDA-Xの数値計算ライブラリを橋渡しするnvmath-pythonを紹介した。高性能なコア数値計算をPythonから大規模に実行でき、GPUアクセラレーションの活用を容易にする。