Safety & Evaluation A
Showing 61–89 of 89
-
Constitutional Midtraining: Content Presence Drives Alignment Gains
-
Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis
-
Misalignment Has a Personality: A Big Five Account of Emergent Misalignment
-
Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs
-
Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach
-
VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening
-
Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment
-
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
-
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
-
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
-
How Do LLMs Read Bug Reports? An Empirical Study of Attention in LLMs for Automated Program Repair
-
Evaluation of Adversarial Robustness in Arabic Language Models
-
Rashomon Alignment
-
AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review
-
Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models
-
MemSFT: Mitigating Alignment Tax with an External Parametric Memory
-
Evaluation of forced alignment of code-mixed speech: the case of Hindi-English
-
IRIS: Reusable Identity Representations from Frozen LLMs for Entity Alignment
-
AMPBench-MT: A Homology-Controlled Benchmark for Antimicrobial Peptide Potency, Spectrum, and Safety Prediction
-
Phase Structure in Rotary Attention: A Spectral Framework for Semantic Continuity and Execution-Boundary Governance
-
PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
-
Data-Dependent Regret and Polyak Corrections for Constrained Online Convex Optimization
-
Emergent Latent-State Computation under Stochastic Volatility
-
Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context
-
MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios
-
NVIDIAやMicrosoftなど30社超、オープンAIの防御ツール共同開発の「Open Secure AI Alliance」設立30+ firms including NVIDIA, Microsoft form Open Secure AI AllianceNVIDIA, Microsoft, SpaceX AI and 30-plus others launched the Open Secure AI Alliance, an initiative to improve the safety of open AI models and jointly develop cybersecurity tools. Using open technologies to patch software vulnerabilities and co-build defensive tooling, the group frames open models as a defensive asset against excessive regulation.
-
Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
-
Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels