PdfDing
Bulk Edit
Layout
Sort
Compact
Grid
List
Minimal
Newest
Oldest
A - Z
Z - A
Most Viewed
Least Viewed
Recently Viewed
Add
M
M
mail.dark4scope@gmail.com
User ID: 1
Settings
Admin
Sign out
Personal
Personal
✓
Create Workspace
Workspace Settings
Default
All
Default
✓
Create Collection
Collection Settings
PDFs
share
Shared
Starred
Annotations
Archive
Tags
Tree Mode
PdfDing
v1.9.0
Filters
#must-read
Archive
Delete
Set Collection
Set Tags
Star
No
Yes
Default
Execute
All
None
3 weeks ago
Aug. 14, 2026, 9:33 a.m. UTC
06 · DeepSeek-R1 — Incentivizing Reasoning via RL
#
deepseek
#
must-read
#
post-training
#
reasoning
#
rlvr
RLVR 范式:纯 RL 长出 long CoT。读 §2.2 规则奖励 + §2.3 冷启动多阶段配方
Details
share
Share
Star
Archive
Delete
1% - Page 1 of 86
3 weeks ago
Aug. 14, 2026, 9:33 a.m. UTC
05 · DeepSeekMath — GRPO 出处
#
deepseek
#
grpo
#
must-read
#
post-training
#
rl
GRPO:去 critic,组内采样均值当 baseline。读 §4.1 + §5.2.1(统一梯度视角,含金量最高)
Details
share
Share
Star
Archive
Delete
3% - Page 1 of 30
3 weeks ago
Aug. 14, 2026, 9:33 a.m. UTC
04 · DPO — Your Language Model is Secretly a Reward Model
#
alignment
#
dpo
#
must-read
#
post-training
#
stanford
从 PPO 的 KL 正则最优解反解 reward。读 §4 + Appendix A.1,必须自己纸上推一遍
Details
share
Share
Star
Archive
Delete
4% - Page 1 of 27
3 weeks ago
Aug. 14, 2026, 9:33 a.m. UTC
03 · PPO — Proximal Policy Optimization Algorithms
#
must-read
#
openai
#
post-training
#
ppo
#
rl
RLHF/GRPO/DAPO 共用的 loss 骨架。只读 §3(2页)+ 公式(7),Atari 部分全跳
Details
share
Share
Star
Archive
Delete
8% - Page 1 of 12
3 weeks ago
Aug. 14, 2026, 9:33 a.m. UTC
02 · LoRA — Low-Rank Adaptation of Large Language Models
#
lora
#
microsoft
#
must-read
#
peft
#
post-training
低秩重参数化 ΔW=BA,PEFT 全家的祖宗。读 §4 + §7.2,跳过 §6
Details
share
Share
Star
Archive
Delete
4% - Page 1 of 26
3 weeks ago
Aug. 14, 2026, 9:33 a.m. UTC
01 · InstructGPT — Training LMs to Follow Instructions with Human Feedback
#
must-read
#
openai
#
post-training
#
rlhf
#
sft
后训练范式源头:SFT→RM→PPO 三段式。重点读 §3 全部 + Fig.2
Details
share
Share
Star
Archive
Delete
97% - Page 66 of 68