学習・ファインチューニング A

119 件中 1〜30 件目を表示
  • arXiv cs.LG (Machine Learning) · EN 学習・ファインチューニング
    The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs
    ファインチューニング 量子化
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 学習・ファインチューニング
    LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
    AI エージェント 強化学習
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN 開発者ツール
    Ordered-to-disordered transfer learning with graph neural networks for formation-energy and HOMO-LUMO gap prediction in high-entropy perovskite oxides
    ニューラルネットワーク
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.LG (Machine Learning) · EN 学習・ファインチューニング
    Leveraging Transfer Learning with Class-Specific Decoders for Laparoscopic Segmentation
    深層学習 検索拡張生成 (RAG)
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.CL (Computation and Language) · EN 新モデル・リリース
    Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?
    推論 (Inference) 強化学習
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.LG (Machine Learning) · EN 学習・ファインチューニング
    MoPET: Parameter-Efficient Mixture-of-Experts for Unified Medical Image Classification
    深層学習 ファインチューニング Mixture of Experts (MoE) 検索拡張生成 (RAG)
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.LG (Machine Learning) · EN 新モデル・リリース
    Parameter-Free Heavy-Tailed Bandits
    アルゴリズム・理論 強化学習
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 学習・ファインチューニング
    Explore Beyond the Boundary Using Entropic Information
    AI エージェント 検索拡張生成 (RAG) 強化学習
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN 学習・ファインチューニング
    ALIVE: Warnings Before Exclusion in Budgeted Multi-Source Learning
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.CL (Computation and Language) · EN 学習・ファインチューニング
    PTP: Previous-Token Prediction based LLM Inversion for Near-Exact Prompt Reconstruction
    ファインチューニング
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.LG (Machine Learning) · EN 学習・ファインチューニング
    The Greedy Advantage in Finite-Horizon Bandits
    アルゴリズム・理論
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 推論・効率化
    Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation
    DeepSeek ファインチューニング GPT 推論 (Inference) ニューラルネットワーク
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 学習・ファインチューニング
    RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems
    AI エージェント
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN インフラ・ハードウェア
    Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters
    推論 (Inference)
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN 学習・ファインチューニング
    GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System
    埋め込み (Embeddings) ファインチューニング ニューラルネットワーク 検索拡張生成 (RAG) 強化学習
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 推論・効率化
    SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
    ニューラルネットワーク 検索拡張生成 (RAG) 強化学習
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN 学習・ファインチューニング
    Learning Latent Reasoning Traces for Scalar Reward Models End-to-End
    検索拡張生成 (RAG) 強化学習 人間のフィードバックによる強化学習 (RLHF)
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • ITmedia AI+ · JA 学習・ファインチューニング 抜粋
    Thinking Machines、軽量モデル「Inkling-Small」正式公開 サイズ4分の1で「Inkling」に匹敵する性能
    Thinking Machines、軽量モデル「Inkling-Small」公開、1/4サイズで同等性能
    強化学習
    Thinking Machines Labは、オープンウェイトのAIモデル「Inkling-Small」正式版を公開した。従来モデルの4分の1のサイズながら、データ改良や強化学習によりコード生成などで「Inkling」に匹敵する性能を実現したとしている。
    元記事を読む (ITmedia AI+) ↗
  • Simon Willison's Weblog · EN 新モデル・リリース 抜粋
    llm 0.32rc2
    Simon Willison、llm 0.32rc2公開―既定モデルをGPT-5.6 Lunaに
    GPT 機械学習 ニューラルネットワーク OpenAI 人間のフィードバックによる強化学習 (RLHF)
    Simon Willison氏がCLIツールllmの0.32rc2を公開した。依存関係の問題を修正するとともに、既定モデルを未設定のユーザー向けに従来のGPT-4o miniから、より新しく高性能なGPT-5.6 Lunaへ変更した。Lunaはやや高価だが大きな改善という。
    元記事を読む (Simon Willison's Weblog) ↗
  • arXiv cs.CL (Computation and Language) · EN 安全性・評価
    Inducing language models to assert their own consciousness restores human beliefs and values
    ファインチューニング
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • Publickey · JA 新モデル・リリース 抜粋
    JetBrains、AIが少ないトークンでコンテキストを取得しやすく、よりよいコード生成を可能にする「JetBrains Context」発表
    JetBrains、AIエージェント向け「JetBrains Context」発表、少トークンで文脈提供
    AI エージェント 機械学習
    JetBrainsは、コードリポジトリの上に知的レイヤを構築する新サービス「JetBrains Context」を発表した。AIエージェントに対して適切なコードのコンテキストを少ないトークンで提供することで、より良いコード生成を可能にするという。
    元記事を読む (Publickey) ↗
  • arXiv cs.CL (Computation and Language) · EN 新モデル・リリース
    Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
    ファインチューニング GPT 機械学習 Meta 検索拡張生成 (RAG)
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 推論・効率化
    APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems
    推論 (Inference) 人間のフィードバックによる強化学習 (RLHF)
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN 新モデル・リリース
    Same Graph Cross-Task Transfer in GNNs: Protocols and Predictors
    ニューラルネットワーク 検索拡張生成 (RAG) 強化学習
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.CL (Computation and Language) · EN 新モデル・リリース
    Improving Mental Health Screening and Early Risk Detection in Spanish
    強化学習
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.LG (Machine Learning) · EN 学習・ファインチューニング
    Cybersecurity Detection Classification with Reasoning-enabled Language Models
    強化学習 人間のフィードバックによる強化学習 (RLHF)
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.CL (Computation and Language) · EN 学習・ファインチューニング
    Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models
    ファインチューニング
    元記事を読む (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.LG (Machine Learning) · EN 新モデル・リリース
    Oracle-Budgeted Molecular Optimization with Short-Term Graph Memory
    深層学習 検索拡張生成 (RAG)
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.LG (Machine Learning) · EN 開発者ツール
    QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Quantum Error Correction
    深層学習 ファインチューニング Google
    元記事を読む (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN 推論・効率化
    When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence
    推論 (Inference) ニューラルネットワーク 人間のフィードバックによる強化学習 (RLHF)
    元記事を読む (arXiv cs.AI (Artificial Intelligence)) ↗