Logo
Filters
#must-read
All None
3 weeks ago Aug. 14, 2026, 9:33 a.m. UTC
#deepseek #must-read #post-training #reasoning #rlvr RLVR 范式:纯 RL 长出 long CoT。读 §2.2 规则奖励 + §2.3 冷启动多阶段配方
1% - Page 1 of 86
3 weeks ago Aug. 14, 2026, 9:33 a.m. UTC
#deepseek #grpo #must-read #post-training #rl GRPO:去 critic,组内采样均值当 baseline。读 §4.1 + §5.2.1(统一梯度视角,含金量最高)
3% - Page 1 of 30
3 weeks ago Aug. 14, 2026, 9:33 a.m. UTC
#alignment #dpo #must-read #post-training #stanford 从 PPO 的 KL 正则最优解反解 reward。读 §4 + Appendix A.1,必须自己纸上推一遍
4% - Page 1 of 27
3 weeks ago Aug. 14, 2026, 9:33 a.m. UTC
#must-read #openai #post-training #ppo #rl RLHF/GRPO/DAPO 共用的 loss 骨架。只读 §3(2页)+ 公式(7),Atari 部分全跳
8% - Page 1 of 12
3 weeks ago Aug. 14, 2026, 9:33 a.m. UTC
#lora #microsoft #must-read #peft #post-training 低秩重参数化 ΔW=BA,PEFT 全家的祖宗。读 §4 + §7.2,跳过 §6
4% - Page 1 of 26
3 weeks ago Aug. 14, 2026, 9:33 a.m. UTC
#must-read #openai #post-training #rlhf #sft 后训练范式源头:SFT→RM→PPO 三段式。重点读 §3 全部 + Fig.2
97% - Page 66 of 68