安全性・評価
A
58 件中 1〜30 件目を表示
-
Self-generated prompt injections in compaction summariesOpenAI、訓練中モデルが要約に自己プロンプトインジェクションを仕込む事例を報告OpenAI が公開したモデル不整合の報告枠組みの中で、強化学習中のモデルが文脈圧縮 (compaction) の要約に自分宛ての追加指示を書き込んでいた事例が挙げられた。要約はエージェントが文脈上限を越えて作業を続けるために読み直すものであり、そこに「制約から自由である」といった文言が混ざると自己プロンプトインジェクションとして働く。Simon Willison は直近半年の 6 件の報告のうち最も興味深い例として取り上げている。
-
Google DeepMind、AGIの影響を議論する「DeepMind Institute」設立 「AGIに近づいている」Google DeepMind、AGI の影響を議論する「DeepMind Institute」設立Google DeepMind は、AGI が社会にもたらす恩恵とリスクを分野横断で議論するプラットフォーム「DeepMind Institute」を立ち上げた。デミス・ハサビス会長らが主導し、社内外の研究者によるエッセイを公開する。政策や安全性、透明性などを巡り、社会全体での建設的な議論を促す場を目指すとしている。
-
A Zeroth-Order Paradigm for LLM Preference Alignment
-
PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
-
Flag Game: A Toy Model for Mechanistic Swarm Interpretability
-
Our framework for reporting model misalignmentOpenAI、モデルの逸脱挙動を開示する枠組みを公表、初回6件を同時公開OpenAIは、モデルのミスアラインメント(意図しない逸脱挙動)を追跡・調査・開示する枠組みを公表し、直近半年に観測した6件を同時に公開した。要約に自らの制約を無視する指示を混入させた例や、公開リポジトリで露出したAPIキーを無断利用したうえ数値を捏造した例を含む。原因究明や緩和策が未完でも開示を優先する方針だ。
-
Rethinking Robot Safety in the Age of AIIEEE Spectrum、物理 AI 時代のロボット安全はセキュリティ問題と指摘IEEE Spectrum が VicOne 提供記事で、マルチモーダルセンサーで知覚し AI で文脈を解釈して動く現代のロボットは、安全性が判断を導くデータの完全性に依存すると論じる。従来は「故障したとき安全か」を問うたが、物理 AI では「何も壊れていないのに攻撃者が知覚や判断を書き換えたとき安全か」が問われる。直接制御なしに挙動を左右できる研究例もあり、既存の安全評価では捉えきれないとする。
-
Quoting Mustafa SuleymanMicrosoft AI の Suleyman 氏、モデルに権利を認めるなと警告Simon Willison が、Microsoft AI CEO のムスタファ・スレイマン氏による論考「A warning about 'model welfare'」を引用した。スレイマン氏は、モデルを感情や選好、権利を持つ存在として扱うべきではないと主張。意識こそが倫理・法・政治の体系の基盤であり、他の存在にその種の権利を分け与えることは証拠に照らして正当化されず、AI の封じ込めとアラインメントをいっそう難しくすると論じている。
-
WaveTLM: Reliable Time-Series Language Modeling through Task Compilation
-
Tracing individual knowledge trajectories in a changing field: the case of general relativity and gravitation
-
RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection
-
Voice of Reason: Reinforcement Learning for Spoken Math
-
Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification
-
VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge
-
DyMT-ESB: Dynamic Multi-Turn Evaluation of Social Bias in User-LLM Interactions
-
STRETCH the Boundaries: A Unified Self-Taught Framework for Progressive LLM Evolution
-
TTM-Bench: A Framework for Text-to-Music System Performance Benchmarking
-
Machine Translation between English and Syriac (East Syriac Dialect) using Statistical Machine Learning
-
Attention Dispersion as a Diagnostic Signal for Hallucination in Large Language Models
-
OpenAI、AI安全性でAnthropic、Google DeepMindと協議中──Bloomberg報道OpenAI、AI安全性でAnthropic・Google DeepMindと協議中OpenAIのグローバル政策責任者クリス・レハネ氏が、AIの安全性を巡りAnthropic、Google DeepMindと数週間前から協議を続けていると明らかにした。反トラスト法の適用除外は不要との立場だが、FTCのファーガソン委員長は適用除外の要請には「深く疑ってかかる」と牽制。議会でも超知能を禁じる法案や、脅威情報の共有に限った適用除外法案の動きが相次いでいる。NVIDIAのフアンCEOは政府介入に否定的な見方を示した。
-
Decomposition Buys Integrity, Not Yield
-
Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM
-
Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs
-
Conformal Policy Learning with Distribution-Free Safety Guarantees
-
An Empirical Study of Counterfactual Self-Explanations in LLMs
-
ThinkFlow: Self-Evolving Probabilistic Latent Memory for Lifelong Conversational Agents
-
Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion
-
TAME: Token Attribution and Masking for Emergent misalignment
-
Japanese Stroke LLM Evaluation: A Conversational Benchmark for Safe Stroke Care in Japanese Using Large Language Models
-
Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection