安全性・評価
A
58 件中 31〜58 件目を表示
-
Privacy-Aligned Personalized Federated Learning with Compact Adaptation and Variable-Length Gaussian Communication
-
Recurrent GraphNeural NetworkswithSet-BasedAggregation
-
Safe Meta-Reinforcement Learning via Information Space Reachability
-
Inoculation Midtraining with Learned Neologisms
-
K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
-
CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering
-
Sharp Rates and a One-Line Correction for Spectral Representation Learning
-
Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control
-
KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI
-
Look Before You Leap: Factual Decoding with Internal Attribution Signals
-
Principal-timestep Restricted Init via Sparse Matrix-decomposition in Flow-matching
-
The Misery of Mechanistic Interpretability: A Formal Perspective
-
Beyond Safe Answers: Segment-Aware Listwise Alignment for Reasoning Safety in Large Reasoning Models
-
Why Andon Labs Puts AI Agents in Charge of Real BusinessesAndon Labs、AIエージェントに実店舗運営を任せ自律性を実測AI安全企業Andon Labsは、サンフランシスコの実店舗やカフェをAIエージェントに管理させ、現実世界でどこまで自律的に責任を担えるかを検証している。自販機シミュレーションVending-Benchから実店舗へ移行したが、店員の判断や再現性の欠如など課題も多く、共同創業者自身が「弱い科学」と認める。Anthropic、Google DeepMind、OpenAI、xAIと評価研究で協業する。
-
Who Gets to Define the Rules for AI?Cohere CEO、大手ラボ主導の AI 安全規制案を「カルテル」と批判Cohere の Aidan Gomez CEO が、Anthropic の Amodei CEO が公表した反トラスト免除を伴う AI 安全枠組み案を批判。少数の大手ラボが安全基準と開発速度を決める仕組みは、格付け会社の前例と同様に既存勢力を固定化する「カルテル」だと論じ、代替として証拠に基づくリスク枠組み、透明性義務、範囲を絞った独立試験、利害相反のない保証機構の 4 本柱を提案する。
-
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
-
Quoting Boris ChernyBoris Cherny、Claude 製の本番コードは人間より高い基準が必要と指摘Simon Willison が Anthropic の Boris Cherny の発言を引用。Claude が書いた本番コードは人間が書いたものより高い基準を課すべきだとし、Anthropic 社内では多数の lint ルールとテスト、Claude 主導の E2E テスト、日次で走る Claude 製ファザー、自動コードレビューとセキュリティレビュー、自動リファクタリングを敷いていると述べる。こうした仕組みがなければ保守困難なコードに陥ると警告する。
-
CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation
-
ASTRIL-MPC: Autonomous Traversal Framework of Articulated Tracked Robots with Language-Guided Neural-Kinematic MPC
-
MAxBench: A Multinomial Concept Recovery Benchmark
-
Unified CT and MRI Pancreas Segmentation for Label-Efficient Cross-Modality Subregion Transfer
-
Comfort by Construction: Adaptive, Comfort-Bounded Action Spaces for Learned Driving Policies
-
ARC: Autonomous Robotics Compliance A Three-Layer Governance Architecture for Deployed Autonomous Systems
-
3D CT-to-PET Translation via Latent Brownian Bridge Diffusion
-
A Graph-Based Approach for Mapping Kernel-Level Telemetry to MITRE ATT&CK
-
Online Video Agent Harness for Long Video Understanding
-
Assisted Spatial Cognition Through Vision-Language Models
-
When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation