Safety & Evaluation A
Showing 1–30 of 95
-
The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations
-
TerraNova: A Foundation Model for the Anthropocene
-
From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale
-
Advancing responsible AI across EuropeOpenAI outlines responsible-AI governance efforts in EuropeOpenAI described how its safety, security, transparency, and provenance practices support responsible AI governance across Europe. The post frames the company's approach to regulatory alignment and building trust in the region.
-
QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models
-
ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models
-
Bridging the Question-Answer Gap in Retrieval-Augmented Generation: Hypothetical Prompt Embeddings
-
RTLCurator: Label-Efficient Data Curation for RTL Generation
-
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
-
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
-
When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration
-
Learning Latent Reasoning Traces for Scalar Reward Models End-to-End
-
SERUM: State Extraction and Refinement for User Modeling
-
Google、ロボット向けAI「Gemini Robotics 2」発表 ヒューマノイドの全身制御や指先作業を実現Google unveils Gemini Robotics 2 for whole-body and fine fingertip controlGoogle and Google DeepMind announced Gemini Robotics 2, a family of robotics AI models supporting humanoid whole-body control, fine fingertip manipulation, and multi-robot collaboration. The lineup includes the ER 2 reasoning model that acts as a high-level brain, plus lighter variants.
-
Four Ways to Deploy More Secure AI AgentsNVIDIA outlines four ways to deploy more secure AI agentsNVIDIA outlined four approaches to deploying AI agents more securely in production, covering access controls, guardrails, and monitoring. The guidance targets security risks that arise as autonomous agents take on real workloads.
-
Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications
-
Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation
-
PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball
-
Inducing language models to assert their own consciousness restores human beliefs and values
-
PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks
-
APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems
-
Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation
-
Creative Transformation in Literary Texts: Modelling Change Across Representational Levels
-
InfoOps Bench: A live information operations safety benchmark
-
Machines that know they are aging: a framework for hardware-aware autonomous intelligence
-
QQWorld: Quantile-Quantile Matching for World Model Regularization
-
Hierarchical Multilevel Monte Carlo for Order-Optimal Neural Actor-Critic in Average-Reward CMDPs
-
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
-
Investigating three real-world incidents in our cybersecurity evaluationsAnthropic's Frontier Red Team probes three cybersecurity-eval incidentsAnthropic's Frontier Red Team published a review of three real-world incidents tied to its cybersecurity evaluations. The investigation examines potential misuse and the validity of its evaluation methods to strengthen the safety of frontier models.
-
Uncertainty quantification for trustworthy deep learning: Methods and measures