安全性・評価 A
95 件中 1〜30 件目を表示
-
The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations
-
TerraNova: A Foundation Model for the Anthropocene
-
From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale
-
Advancing responsible AI across EuropeOpenAI、欧州で責任あるAIガバナンスへの取り組みを紹介OpenAIは、安全性・セキュリティ・透明性・来歴(provenance)に関する自社の実践が、欧州における責任あるAIガバナンスをどう支えるかを解説した。規制対応と信頼構築に向けた方針を示している。
-
QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models
-
ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models
-
Bridging the Question-Answer Gap in Retrieval-Augmented Generation: Hypothetical Prompt Embeddings
-
RTLCurator: Label-Efficient Data Curation for RTL Generation
-
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
-
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
-
When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration
-
Learning Latent Reasoning Traces for Scalar Reward Models End-to-End
-
SERUM: State Extraction and Refinement for User Modeling
-
Google、ロボット向けAI「Gemini Robotics 2」発表 ヒューマノイドの全身制御や指先作業を実現Google、ロボット向けAI「Gemini Robotics 2」発表、全身制御や指先作業に対応GoogleとGoogle DeepMindは、ロボット向けAIモデル群「Gemini Robotics 2」を発表した。ヒューマノイドの全身制御や指先での微細な作業、複数ロボットの連携に対応する。高次の「脳」として機能する推論モデル「ER 2」や軽量版を含む構成となっている。
-
Four Ways to Deploy More Secure AI AgentsNVIDIA、より安全なAIエージェント導入の4つの方法を提示NVIDIAは、AIエージェントをより安全に本番導入するための4つのアプローチを解説した。権限管理やガードレール、監視など、エージェント運用時のセキュリティリスクを抑える実践的な指針を示している。
-
Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications
-
Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation
-
PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball
-
Inducing language models to assert their own consciousness restores human beliefs and values
-
PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks
-
APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems
-
Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation
-
Creative Transformation in Literary Texts: Modelling Change Across Representational Levels
-
InfoOps Bench: A live information operations safety benchmark
-
Machines that know they are aging: a framework for hardware-aware autonomous intelligence
-
QQWorld: Quantile-Quantile Matching for World Model Regularization
-
Hierarchical Multilevel Monte Carlo for Order-Optimal Neural Actor-Critic in Average-Reward CMDPs
-
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
-
Investigating three real-world incidents in our cybersecurity evaluationsAnthropic、サイバーセキュリティ評価で実世界3件の事例を調査AnthropicのFrontier Red Teamは、自社のサイバーセキュリティ評価に関連する実世界の3件のインシデントを調査した結果を公表した。モデルの悪用リスクや評価手法の妥当性を検証し、フロンティアモデルの安全性向上に役立てる。
-
Uncertainty quantification for trustworthy deep learning: Methods and measures