Training & Fine-tuning A
Showing 1–30 of 119
-
The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs
-
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
-
Ordered-to-disordered transfer learning with graph neural networks for formation-energy and HOMO-LUMO gap prediction in high-entropy perovskite oxides
-
Leveraging Transfer Learning with Class-Specific Decoders for Laparoscopic Segmentation
-
Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?
-
MoPET: Parameter-Efficient Mixture-of-Experts for Unified Medical Image Classification
-
Parameter-Free Heavy-Tailed Bandits
-
Explore Beyond the Boundary Using Entropic Information
-
ALIVE: Warnings Before Exclusion in Budgeted Multi-Source Learning
-
PTP: Previous-Token Prediction based LLM Inversion for Near-Exact Prompt Reconstruction
-
The Greedy Advantage in Finite-Horizon Bandits
-
Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation
-
RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems
-
Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters
-
GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System
-
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
-
Learning Latent Reasoning Traces for Scalar Reward Models End-to-End
-
Thinking Machines、軽量モデル「Inkling-Small」正式公開 サイズ4分の1で「Inkling」に匹敵する性能Thinking Machines releases Inkling-Small, matching Inkling at 1/4 the sizeThinking Machines Lab released the final version of Inkling-Small, an open-weight AI model. At a quarter the size of its predecessor, the company says data improvements and reinforcement learning let it match the larger Inkling on tasks such as code generation.
-
llm 0.32rc2llm 0.32rc2 switches its default model to GPT-5.6 LunaSimon Willison released llm 0.32rc2, fixing a dependency issue and changing the default model for users who have not set one from GPT-4o mini to the newer, more capable GPT-5.6 Luna. Luna is slightly more expensive but a notable upgrade.
-
Inducing language models to assert their own consciousness restores human beliefs and values
-
JetBrains、AIが少ないトークンでコンテキストを取得しやすく、よりよいコード生成を可能にする「JetBrains Context」発表JetBrains unveils 'JetBrains Context' to feed AI agents code context efficientlyJetBrains announced JetBrains Context, a service that builds an intelligence layer over code repositories. By supplying AI agents with the right code context using fewer tokens, it aims to enable better code generation from agentic coding tools.
-
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
-
APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems
-
Same Graph Cross-Task Transfer in GNNs: Protocols and Predictors
-
Improving Mental Health Screening and Early Risk Detection in Spanish
-
Cybersecurity Detection Classification with Reasoning-enabled Language Models
-
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models
-
Oracle-Budgeted Molecular Optimization with Short-Term Graph Memory
-
QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Quantum Error Correction
-
When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence