Hacker News· mrkn1·· 4 小时前AI 评分37
ProVer:如何不评估每一步即可修复 GRPO 的信用分配问题
Fixing GRPO's credit assignment problem without evaluating every step
AI 导读
ProVer 框架针对 GRPO 在智能体强化学习中给所有策略 token 均匀分配优势的信用分配问题,通过智能体判官对比成功与失败轨迹来定位关键决策片段,并用前后续写成功率差验证其优势。
正文
Abstract:Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.36178 [cs.CL] |
| (or arXiv:2609.36178v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.36178 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Dongwon Jung [view email]
[v1]
Mon, 28 Sep 2026 19:48:05 UTC (132 KB)
来源:Hacker News · arxiv.org