Safety & Evaluation
A
Showing 31–58 of 58
-
Privacy-Aligned Personalized Federated Learning with Compact Adaptation and Variable-Length Gaussian Communication
-
Recurrent GraphNeural NetworkswithSet-BasedAggregation
-
Safe Meta-Reinforcement Learning via Information Space Reachability
-
Inoculation Midtraining with Learned Neologisms
-
K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
-
CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering
-
Sharp Rates and a One-Line Correction for Spectral Representation Learning
-
Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control
-
KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI
-
Look Before You Leap: Factual Decoding with Internal Attribution Signals
-
Principal-timestep Restricted Init via Sparse Matrix-decomposition in Flow-matching
-
The Misery of Mechanistic Interpretability: A Formal Perspective
-
Beyond Safe Answers: Segment-Aware Listwise Alignment for Reasoning Safety in Large Reasoning Models
-
Why Andon Labs Puts AI Agents in Charge of Real BusinessesAndon Labs puts AI agents in charge of real stores to measure autonomyAI safety firm Andon Labs runs a real San Francisco store and cafe managed by AI agents to measure how much real-world responsibility they can handle. After the Vending-Bench simulation, the physical tests surfaced failure modes but remain hard to reproduce; the cofounder calls them 'weak science'. Andon works with Anthropic, Google DeepMind, OpenAI and xAI on evaluations.
-
Who Gets to Define the Rules for AI?Cohere CEO calls Anthropic's antitrust-waiver AI safety plan a 'cartel'Aidan Gomez argues Dario Amodei's antitrust-waiver plan letting top labs set AI safety rules would entrench incumbents, as rating agencies once did. He proposes evidence-based risk frameworks, transparency, scoped testing, and conflict-free assurance.
-
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
-
Quoting Boris ChernyBoris Cherny: Claude-written production code needs a higher barSimon Willison quotes Anthropic's Boris Cherny arguing that production code written by Claude should be held to a higher bar than human-written code. Anthropic layers on guardrails: extensive lint rules and tests, Claude-driven end-to-end tests, Claude-powered fuzzers run daily, automated code and security reviews, and automated refactoring. Without them, he warns, the result is hard to maintain.
-
CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation
-
ASTRIL-MPC: Autonomous Traversal Framework of Articulated Tracked Robots with Language-Guided Neural-Kinematic MPC
-
MAxBench: A Multinomial Concept Recovery Benchmark
-
Unified CT and MRI Pancreas Segmentation for Label-Efficient Cross-Modality Subregion Transfer
-
Comfort by Construction: Adaptive, Comfort-Bounded Action Spaces for Learned Driving Policies
-
ARC: Autonomous Robotics Compliance A Three-Layer Governance Architecture for Deployed Autonomous Systems
-
3D CT-to-PET Translation via Latent Brownian Bridge Diffusion
-
A Graph-Based Approach for Mapping Kernel-Level Telemetry to MITRE ATT&CK
-
Online Video Agent Harness for Long Video Understanding
-
Assisted Spatial Cognition Through Vision-Language Models
-
When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation