安全性・評価 A
94 件中 61〜90 件目を表示
-
Constitutional Midtraining: Content Presence Drives Alignment Gains
-
Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis
-
Misalignment Has a Personality: A Big Five Account of Emergent Misalignment
-
Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs
-
Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach
-
VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening
-
Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment
-
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
-
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
-
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
-
How Do LLMs Read Bug Reports? An Empirical Study of Attention in LLMs for Automated Program Repair
-
Evaluation of Adversarial Robustness in Arabic Language Models
-
Rashomon Alignment
-
AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review
-
Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models
-
MemSFT: Mitigating Alignment Tax with an External Parametric Memory
-
Evaluation of forced alignment of code-mixed speech: the case of Hindi-English
-
IRIS: Reusable Identity Representations from Frozen LLMs for Entity Alignment
-
AMPBench-MT: A Homology-Controlled Benchmark for Antimicrobial Peptide Potency, Spectrum, and Safety Prediction
-
Phase Structure in Rotary Attention: A Spectral Framework for Semantic Continuity and Execution-Boundary Governance
-
PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
-
Data-Dependent Regret and Polyak Corrections for Constrained Online Convex Optimization
-
Emergent Latent-State Computation under Stochastic Volatility
-
Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context
-
MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios
-
NVIDIAやMicrosoftなど30社超、オープンAIの防御ツール共同開発の「Open Secure AI Alliance」設立NVIDIAやMicrosoftなど30社超、AI防御ツール共同開発の「Open Secure AI Alliance」設立NVIDIA、Microsoft、SpaceX AIなど30社超が、AIオープンモデルの安全性向上とサイバーセキュリティツールの共同開発を目指すイニシアチブ「Open Secure AI Alliance」を設立した。オープンな技術を活用してソフトウェアの脆弱性修正や防御ツールの共同開発を推進し、オープンモデルを過度な規制に対する防御資産として位置づけ、その重要性を主張する。
-
Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
-
Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
-
D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models