学習・ファインチューニング A
119 件中 1〜30 件目を表示
-
The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs
-
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
-
Ordered-to-disordered transfer learning with graph neural networks for formation-energy and HOMO-LUMO gap prediction in high-entropy perovskite oxides
-
Leveraging Transfer Learning with Class-Specific Decoders for Laparoscopic Segmentation
-
Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?
-
MoPET: Parameter-Efficient Mixture-of-Experts for Unified Medical Image Classification
-
Parameter-Free Heavy-Tailed Bandits
-
Explore Beyond the Boundary Using Entropic Information
-
ALIVE: Warnings Before Exclusion in Budgeted Multi-Source Learning
-
PTP: Previous-Token Prediction based LLM Inversion for Near-Exact Prompt Reconstruction
-
The Greedy Advantage in Finite-Horizon Bandits
-
Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation
-
RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems
-
Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters
-
GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System
-
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
-
Learning Latent Reasoning Traces for Scalar Reward Models End-to-End
-
Thinking Machines、軽量モデル「Inkling-Small」正式公開 サイズ4分の1で「Inkling」に匹敵する性能Thinking Machines、軽量モデル「Inkling-Small」公開、1/4サイズで同等性能Thinking Machines Labは、オープンウェイトのAIモデル「Inkling-Small」正式版を公開した。従来モデルの4分の1のサイズながら、データ改良や強化学習によりコード生成などで「Inkling」に匹敵する性能を実現したとしている。
-
llm 0.32rc2Simon Willison、llm 0.32rc2公開―既定モデルをGPT-5.6 LunaにSimon Willison氏がCLIツールllmの0.32rc2を公開した。依存関係の問題を修正するとともに、既定モデルを未設定のユーザー向けに従来のGPT-4o miniから、より新しく高性能なGPT-5.6 Lunaへ変更した。Lunaはやや高価だが大きな改善という。
-
Inducing language models to assert their own consciousness restores human beliefs and values
-
JetBrains、AIが少ないトークンでコンテキストを取得しやすく、よりよいコード生成を可能にする「JetBrains Context」発表JetBrains、AIエージェント向け「JetBrains Context」発表、少トークンで文脈提供JetBrainsは、コードリポジトリの上に知的レイヤを構築する新サービス「JetBrains Context」を発表した。AIエージェントに対して適切なコードのコンテキストを少ないトークンで提供することで、より良いコード生成を可能にするという。
-
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
-
APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems
-
Same Graph Cross-Task Transfer in GNNs: Protocols and Predictors
-
Improving Mental Health Screening and Early Risk Detection in Spanish
-
Cybersecurity Detection Classification with Reasoning-enabled Language Models
-
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models
-
Oracle-Budgeted Molecular Optimization with Short-Term Graph Memory
-
QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Quantum Error Correction
-
When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence