Training & Fine-tuning

A
Showing 61–85 of 85
  • arXiv cs.CL (Computation and Language) · EN Safety & Evaluation
    Beyond Safe Answers: Segment-Aware Listwise Alignment for Reasoning Safety in Large Reasoning Models
    Reinforcement Learning from Human Feedback (RLHF) Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    Authorship attribution and aesthetic evaluation of AI poetry: a case study with Haiku
    Gemini GPT Llama Neural Network
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • OpenAI Blog · EN Training & Fine-tuning
    How Fyxer built an AI executive assistant people trust
    OpenAI showcases Fyxer's AI EA: 53% of drafts sent unedited, 90% retention
    Fine-tuning OpenAI
    OpenAI profiles Fyxer, a UK startup whose AI executive assistant is trained on 500,000+ hours of human EA workflows. Email handling is split across 30–50 specialized fine-tuned models, and user edits feed a DPO self-training loop. 53% of drafts are accepted as written, 90-day retention tops 90%, and ARR grew from $1M to $32M in 2025.
    Read original (OpenAI Blog) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
    AI Agents GPT Reinforcement Learning
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN Training & Fine-tuning
    Parameter-Efficient Adaptation of Pretrained Language Models for Time-Series Forecasting
    Embeddings Fine-tuning GPT Machine Learning Transformer
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN Infrastructure & Hardware
    Reducing the Output-Mode Gap in Speech Language Models via Joint-Output On-Policy Distillation
    Deep Learning Fine-tuning Reinforcement Learning Software Engineering Speech Processing
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • Hacker News (Front Page) · EN Training & Fine-tuning
    EuroBirdPortal – Live bird movements across Europe
    EuroBirdPortal maps live weekly bird movements across Europe
    Reinforcement Learning from Human Feedback (RLHF)
    EuroBirdPortal aggregates bird observation records from across Europe and animates weekly changes in species counts and distribution on an interactive map. It offers a double-map view for comparing two species and a 52-week timeline playback, making seasonal migration patterns easy to follow. The citizen-science visualization drew attention on Hacker News.
    Read original (Hacker News (Front Page)) ↗
  • Publickey · JA Agents & Tool Use
    TailscaleのVPNにAIエージェントを組み込める「Aperture」正式リリース。AIによるTailscaleやノードの操作も可能に
    Tailscale ships Aperture AI gateway GA, adds MCPs for agent-driven network ops
    AI Agents Machine Learning
    Tailscale announced general availability of Aperture, an AI gateway that lets any node in a tailnet reach AI services without distributing API keys. It now bundles cost controls, guardrails, audit logs, MCP/API proxies, and in-gateway token purchasing. New Tailscale MCP and Tailscale SSH MCP let AI agents add nodes, SSH into machines, and deploy services via Aperture's chat UI, with existing ACLs enforced, human approval required for new nodes, and every action logged.
    Read original (Publickey) ↗
  • Simon Willison's Weblog · EN Developer Tools
    So you want to use OpenRouter?
    OpenRouter's auto-routing can vary model behavior by provider
    Deep Learning Reinforcement Learning from Human Feedback (RLHF)
    Simon Willison flags Mohamed Moustafa's caveats on OpenRouter: its single endpoint auto-routes to the cheapest backend, but providers differ in serving software and settings, so the same model can behave differently. provider.only pins routing.
    Read original (Simon Willison's Weblog) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    Expert-Space Exploration in MoE Reinforcement Learning
    Mixture of Experts (MoE) Reinforcement Learning
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Multimodal
    Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
    Computer Vision Robotics
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    Generative Retrieval for Unsupervised Text-Based Person Search
    Neural Network Retrieval-Augmented Generation (RAG) Software Engineering
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • Sakana AI Blog (ja) · JA New Model Releases
    英国王立協会特集号に見る、世界モデルの最前線とAIの未来
    Royal Society issue maps world models; Sakana AI's David Ha co-authors
    Reinforcement Learning Robotics
    The Royal Society's Philosophical Transactions A published a special issue, "World Models in Natural and Artificial Intelligence," co-authored in its opening article by Sakana AI CEO David Ha. Contributors argue scaling compute alone will not close the gap between what large models can do and what they understand, and tie world models to artificial life.
    Read original (Sakana AI Blog (ja)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    Parameter-Efficient Retrievers for Polish and European Languages
    Embeddings Fine-tuning Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction
    Embeddings Fine-tuning Reinforcement Learning Transformer
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    Behavior Quotient Learning for Low-Rank Adaptation of LLM Agents
    AI Agents Inference Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Developer Tools
    Evaluating Context Segmentation in Locally Deployable SLMs for Cybersecurity CTF Tasks
    AI Agents
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Training & Fine-tuning
    Scaling Clinical Judgment to Evaluate Medical AI
    Deep Learning Fine-tuning Health & Bio Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    RunningTensor: Generalizing Linear Attention to Higher-Order Recurrent States
    Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Training & Fine-tuning
    Cognition on Graph: Navigating Massive Knowledge Space via Cognitive Cycles and Bidirectional Graph-Text Synergy
    Deep Learning Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    SCQ: Stabilizing Conservative Q-Learning with Sigmoid-Bounded Entropy
    Reinforcement Learning Robotics
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Infrastructure & Hardware
    Residual Vector-based Reconstruction as Long-Context Recall Regardless of Context Window Size
    Deep Learning Fine-tuning Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    Doc2FRC: Length-Consistent Document-Level Machine Translation via Fixed-Range Chunking
    Deep Learning Fine-tuning Inference
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    Earth-Agent-Pro: Towards Real-World Full-Chain Earth Observation with Agents
    AI Agents Fine-tuning GPT Neural Network Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • ITmedia AI+ · JA Training & Fine-tuning
    フィジカルAI向け「データ収集工場」に"潜入" 実機写真で学習データ生成の流れを解説
    J-HRTI opens a physical-AI robot data factory in Narashino, Japan
    Robotics
    A five-company cross-industry consortium, J-HRTI, has started its Kanto Data Factory in Narashino, Chiba, producing training data for physical AI and robotics. Member firms share the collected data; ITmedia details the capture process with on-site photos.
    Read original (ITmedia AI+) ↗