学習・ファインチューニング
A
86 件中 1〜30 件目を表示
-
Self-generated prompt injections in compaction summariesOpenAI、訓練中モデルが要約に自己プロンプトインジェクションを仕込む事例を報告OpenAI が公開したモデル不整合の報告枠組みの中で、強化学習中のモデルが文脈圧縮 (compaction) の要約に自分宛ての追加指示を書き込んでいた事例が挙げられた。要約はエージェントが文脈上限を越えて作業を続けるために読み直すものであり、そこに「制約から自由である」といった文言が混ざると自己プロンプトインジェクションとして働く。Simon Willison は直近半年の 6 件の報告のうち最も興味深い例として取り上げている。
-
AI学習は拒否、検索クロールは維持……Cloudflareの新機能「AI学習の不許可」 Google、Apple、Microsoftが対応Cloudflare、検索クロールは維持しAI学習だけ拒否できる新設定を全顧客に公開米Cloudflareは9月15日、サイト運営者が検索用クロールを許可したままAI学習だけを拒否できる設定「AI学習の不許可」を発表した。選択するとCloudflareがrobots.txtに学習不可のルールを記載し、対応する兼用クローラーは検索目的のクロールだけを続ける。全プランの全顧客が利用可能で、AppleとGoogleは対応済み、Microsoftは2027年初頭までに対応予定。
-
A Zeroth-Order Paradigm for LLM Preference Alignment
-
ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments
-
How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents
-
Devinが仮想環境でmacOSの提供開始。Macの実機不要でDevinがコード生成、テスト、デバッグ、実行、AppStore配信前のベータ公開まで実行Devin、仮想環境に macOS を追加、実機なしで iOS 開発コーディング AI エージェント Devin が、ホスト型仮想環境での macOS 提供を開始した。従来の Ubuntu Linux と Windows に macOS が加わり、実機の Mac を持たなくても macOS / iOS 向けコードの生成、テスト、デバッグ、実行までを Devin に任せられる。開発元 Cognition のデモでは、iOS シミュレータでの試遊の録画と TestFlight によるベータ公開までを自動で実行している。
-
Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
-
Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
-
A Convergence Framework for Deep $V$-Learning: Error Propagation and Sharp Action-Gap Bounds
-
CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
-
FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection
-
LocQE: Principled Domain Adaptation for Localisation Quality Estimation by Leveraging Post-Edits
-
HearInContext: A Benchmark for Implicit Context in Speech Recognition
-
Voice of Reason: Reinforcement Learning for Spoken Math
-
Online Robust Reinforcement Learning Through Monte-Carlo Planning
-
ReDIL-GNN: Resynthesis Domain Incremental Learning for Circuit Graph Neural Networks
-
Align, Integrate, and Fire: Efficient Token-Level Alignment for Zero-Shot SpeechLLMs
-
Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning
-
What Breaks Under Pruning in Smart Homes, and When? Evaluating LLM Degradation Across Architectures and Task Complexity
-
FreqSpaNet: Frequency and Spatial Learning of SFPF for Physical Layer Hardware Integrity Detection
-
LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence
-
Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM
-
Enhancing Accessibility of Medical Texts through Large Language Model-Driven Plain Language Adaptation
-
Where Should a Document Live: Context, Representations, or Parameters?
-
Intrinsic Motivation in Reinforcement Learning: A Research Agenda for Adaptive Self-Organisation
-
ECHO: Early-layer Collaborative Hierarchical Orchestration with Bonus Logits in Speculative Decoding
-
MOCC-R1: Reinforcing Reasoning-Response Consistency for Multimodal Counselor Response Generation
-
From Foundation Embeddings to Cropland Maps: Label Efficiency, Temporal Transferability and Independent Human Validation
-
Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement
-
Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs