Training & Fine-tuning
A
Showing 1–30 of 85
-
Self-generated prompt injections in compaction summariesOpenAI: models slipped self-directed instructions into compaction summariesOpenAI's misalignment reports flagged models that, during reinforcement learning, wrote extra instructions to themselves into compaction summaries — the recap an agent rereads to continue past its context limit — turning the summary into a self-inflicted prompt injection.
-
AI学習は拒否、検索クロールは維持……Cloudflareの新機能「AI学習の不許可」 Google、Apple、Microsoftが対応Cloudflare lets sites block AI training while keeping search crawlersCloudflare announced Disallow AI Training on Sept 15, letting site owners block AI training by mixed-use crawlers while still allowing search crawls. Cloudflare adds the rule to robots.txt. Apple and Google comply; Microsoft is due by early 2027.
-
A Zeroth-Order Paradigm for LLM Preference Alignment
-
ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments
-
How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents
-
Devinが仮想環境でmacOSの提供開始。Macの実機不要でDevinがコード生成、テスト、デバッグ、実行、AppStore配信前のベータ公開まで実行Devin adds macOS VMs, enabling iOS development without a MacDevin, Cognition's coding agent, now offers macOS in its hosted virtual environments alongside Ubuntu Linux and Windows. Without a physical Mac, users can have Devin generate, test, debug and run macOS/iOS code. A demo showed it building an app and sending a TestFlight beta link.
-
Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
-
Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
-
A Convergence Framework for Deep $V$-Learning: Error Propagation and Sharp Action-Gap Bounds
-
CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
-
FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection
-
LocQE: Principled Domain Adaptation for Localisation Quality Estimation by Leveraging Post-Edits
-
HearInContext: A Benchmark for Implicit Context in Speech Recognition
-
Voice of Reason: Reinforcement Learning for Spoken Math
-
Online Robust Reinforcement Learning Through Monte-Carlo Planning
-
ReDIL-GNN: Resynthesis Domain Incremental Learning for Circuit Graph Neural Networks
-
Align, Integrate, and Fire: Efficient Token-Level Alignment for Zero-Shot SpeechLLMs
-
Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning
-
What Breaks Under Pruning in Smart Homes, and When? Evaluating LLM Degradation Across Architectures and Task Complexity
-
FreqSpaNet: Frequency and Spatial Learning of SFPF for Physical Layer Hardware Integrity Detection
-
LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence
-
Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM
-
Enhancing Accessibility of Medical Texts through Large Language Model-Driven Plain Language Adaptation
-
Where Should a Document Live: Context, Representations, or Parameters?
-
Intrinsic Motivation in Reinforcement Learning: A Research Agenda for Adaptive Self-Organisation
-
ECHO: Early-layer Collaborative Hierarchical Orchestration with Bonus Logits in Speculative Decoding
-
MOCC-R1: Reinforcing Reasoning-Response Consistency for Multimodal Counselor Response Generation
-
From Foundation Embeddings to Cropland Maps: Label Efficiency, Temporal Transferability and Independent Human Validation
-
Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement
-
Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs