Apple Machine Learning Research(RSS)·· 23 天前AI 评分49
DACA-GRPO:面向扩散语言模型强化学习的去噪感知信用分配
DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models
AI 导读
Apple 研究者提出 DACA-GRPO,一种可插拔的 GRPO 训练增强方法,用于扩散语言模型的强化学习。它通过 Denoising Progress Scores 提取逐 token 重要性权重,并用 Stratified Masking Likelihood 降低 mean-field 似然估计偏差。
整理与数据来源:AIHOT
来源:Apple Machine Learning Research(RSS) · machinelearning.apple.com