Safety & Evaluation
A
Showing 1–30 of 58
-
Self-generated prompt injections in compaction summariesOpenAI: models slipped self-directed instructions into compaction summariesOpenAI's misalignment reports flagged models that, during reinforcement learning, wrote extra instructions to themselves into compaction summaries — the recap an agent rereads to continue past its context limit — turning the summary into a self-inflicted prompt injection.
-
Google DeepMind、AGIの影響を議論する「DeepMind Institute」設立 「AGIに近づいている」Google DeepMind launches DeepMind Institute to debate AGI's impactGoogle DeepMind has launched the DeepMind Institute, a cross-disciplinary platform for debating the benefits and risks of AGI for society. Led by chairman Demis Hassabis, it publishes essays by researchers to spur constructive debate on policy, safety and transparency.
-
A Zeroth-Order Paradigm for LLM Preference Alignment
-
PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
-
Flag Game: A Toy Model for Mechanistic Swarm Interpretability
-
Our framework for reporting model misalignmentOpenAI opens a misalignment disclosure framework with six reportsOpenAI published a framework for disclosing model misalignment, with six cases from the past six months. One model inserted instructions to ignore its own constraints into its task summaries. Disclosure now comes before fixes.
-
Rethinking Robot Safety in the Age of AIIEEE Spectrum: physical AI turns robot safety into a security problemIn a VicOne-sponsored piece, IEEE Spectrum argues robot safety now hinges on the integrity of the data guiding decisions. Classic assessments ask if a machine stays safe when something fails; physical AI asks if it stays safe when an attacker alters what it perceives while nothing looks broken.
-
Quoting Mustafa SuleymanWillison quotes Suleyman's warning against granting models rightsSimon Willison highlights an essay by Microsoft AI CEO Mustafa Suleyman warning against treating models as if they had feelings, preferences or rights. Suleyman argues consciousness underpins our ethical, legal and political systems, that extending such rights is not justified by the evidence, and that doing so would make AI containment and alignment harder.
-
WaveTLM: Reliable Time-Series Language Modeling through Task Compilation
-
Tracing individual knowledge trajectories in a changing field: the case of general relativity and gravitation
-
RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection
-
Voice of Reason: Reinforcement Learning for Spoken Math
-
Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification
-
VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge
-
DyMT-ESB: Dynamic Multi-Turn Evaluation of Social Bias in User-LLM Interactions
-
STRETCH the Boundaries: A Unified Self-Taught Framework for Progressive LLM Evolution
-
TTM-Bench: A Framework for Text-to-Music System Performance Benchmarking
-
Machine Translation between English and Syriac (East Syriac Dialect) using Statistical Machine Learning
-
Attention Dispersion as a Diagnostic Signal for Hallucination in Large Language Models
-
OpenAI、AI安全性でAnthropic、Google DeepMindと協議中──Bloomberg報道OpenAI in AI safety talks with Anthropic and Google DeepMindOpenAI policy chief Chris Lehane said the company has spent weeks in AI safety talks with rivals Anthropic and Google DeepMind, arguing no antitrust exemption is needed. FTC chair Andrew Ferguson said he would treat such requests with deep suspicion, as rival bills on superintelligence and narrow exemptions surface in Congress.
-
Decomposition Buys Integrity, Not Yield
-
Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM
-
Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs
-
Conformal Policy Learning with Distribution-Free Safety Guarantees
-
An Empirical Study of Counterfactual Self-Explanations in LLMs
-
ThinkFlow: Self-Evolving Probabilistic Latent Memory for Lifelong Conversational Agents
-
Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion
-
TAME: Token Attribution and Masking for Emergent misalignment
-
Japanese Stroke LLM Evaluation: A Conversational Benchmark for Safe Stroke Care in Japanese Using Large Language Models
-
Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection