学習・ファインチューニング A
114 件中 31〜60 件目を表示
-
On-Policy and Off-Policy Learning for Large Action Spaces
-
Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
-
ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
-
HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks
-
Filling the Pareto-Optimal Front for Affordance Segmentation on Embedded Devices Using RGB-D Cameras
-
CACHE-UK: A Stability-Aware Memory Editor for Sequentially Updated Quantized LLMs in Finance
-
Agentic Method for Deterministic Validation of Legacy Code Migration
-
Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory
-
LM-GRASP: Instance-Specific Language Models for Combinatorial Construction via Online Imitation Learning
-
Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training
-
From Expert Reduction to Behavioral Divergence: Tracing Numerical State through Sparse MoE Inference
-
LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models
-
GVR-Coder: A Visual-Feedback Framework for Structured SVG Generation in Complex Document and Meeting Scenarios
-
SemPIC: Learning Semantic Position-Independent KV Caches
-
A Query-Efficient Stochastic Volume Rendering Framework for Time-Varying Implicit Neural Volumes
-
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
-
Building a User Foundation Model for the Open Web
-
TriShield: Zero-Utility-Loss Defense Against Privacy Backdoors in Federated Language Model Fine-Tuning via Orthogonal Gradient Projection and Optimizer State Entanglement
-
Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting
-
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
-
FinanceHarness: Autonomous Financial Deep Research Framework
-
Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation
-
Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO
-
Baikal: Structured Search for Deep Research over Data Lakes
-
Tight Sample Complexity for Low-Rank Adaptation: Matching Bounds and Rank Selection
-
Training Skills Like Parameters via Self-Supervised Semantic Diffusion
-
Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?
-
DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search
-
Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork
-
Improving Item Discoverability in e-Commerce Search via Related Intent Generation