PdfDing
Bulk Edit
Layout
Sort
Compact
Grid
List
Minimal
Newest
Oldest
A - Z
Z - A
Most Viewed
Least Viewed
Recently Viewed
Add
M
M
mail.dark4scope@gmail.com
User ID: 1
Settings
Admin
Sign out
Personal
Personal
✓
Create Workspace
Workspace Settings
Default
All
Default
✓
Create Collection
Collection Settings
PDFs
share
Shared
Starred
Annotations
Archive
Tags
Tree Mode
PdfDing
v1.9.0
Filters
#post-training
Archive
Delete
Set Collection
Set Tags
Star
No
Yes
Default
Execute
All
None
3 weeks ago
Aug. 14, 2026, 9:33 a.m. UTC
09 · Tülu 3 — Pushing Frontiers in Open Post-Training
#
ai2
#
data
#
post-training
#
recipe
完整开源后训练配方,数据混合比例与消融全公开。当参考答案用,不当教材读
Details
share
Share
Star
Archive
Delete
26% - Page 21 of 82
3 weeks ago
Aug. 14, 2026, 9:33 a.m. UTC
08 · Llama 2 — Open Foundation and Fine-Tuned Chat Models
#
meta
#
post-training
#
reward-model
#
rlhf
唯一把 RM 工程细节写清楚的公开材料。只读 §3.2.2:margin loss + 双 RM
Details
share
Share
Star
Archive
Delete
1% - Page 1 of 77
3 weeks ago
Aug. 14, 2026, 9:33 a.m. UTC
07 · DAPO — Open-Source LLM RL System at Scale
#
bytedance
#
grpo
#
post-training
#
rlvr
#
verl
GRPO 规模化会崩,这里给 4 个修法:clip-higher/动态采样/token-level loss/超长奖励整形。跑 verl 时读
Details
share
Share
Star
Archive
Delete
6% - Page 1 of 16
3 weeks ago
Aug. 14, 2026, 9:33 a.m. UTC
06 · DeepSeek-R1 — Incentivizing Reasoning via RL
#
deepseek
#
must-read
#
post-training
#
reasoning
#
rlvr
RLVR 范式:纯 RL 长出 long CoT。读 §2.2 规则奖励 + §2.3 冷启动多阶段配方
Details
share
Share
Star
Archive
Delete
1% - Page 1 of 86
3 weeks ago
Aug. 14, 2026, 9:33 a.m. UTC
05 · DeepSeekMath — GRPO 出处
#
deepseek
#
grpo
#
must-read
#
post-training
#
rl
GRPO:去 critic,组内采样均值当 baseline。读 §4.1 + §5.2.1(统一梯度视角,含金量最高)
Details
share
Share
Star
Archive
Delete
3% - Page 1 of 30
3 weeks ago
Aug. 14, 2026, 9:33 a.m. UTC
04 · DPO — Your Language Model is Secretly a Reward Model
#
alignment
#
dpo
#
must-read
#
post-training
#
stanford
从 PPO 的 KL 正则最优解反解 reward。读 §4 + Appendix A.1,必须自己纸上推一遍
Details
share
Share
Star
Archive
Delete
4% - Page 1 of 27
3 weeks ago
Aug. 14, 2026, 9:33 a.m. UTC
03 · PPO — Proximal Policy Optimization Algorithms
#
must-read
#
openai
#
post-training
#
ppo
#
rl
RLHF/GRPO/DAPO 共用的 loss 骨架。只读 §3(2页)+ 公式(7),Atari 部分全跳
Details
share
Share
Star
Archive
Delete
8% - Page 1 of 12
3 weeks ago
Aug. 14, 2026, 9:33 a.m. UTC
02 · LoRA — Low-Rank Adaptation of Large Language Models
#
lora
#
microsoft
#
must-read
#
peft
#
post-training
低秩重参数化 ΔW=BA,PEFT 全家的祖宗。读 §4 + §7.2,跳过 §6
Details
share
Share
Star
Archive
Delete
4% - Page 1 of 26
3 weeks ago
Aug. 14, 2026, 9:33 a.m. UTC
01 · InstructGPT — Training LMs to Follow Instructions with Human Feedback
#
must-read
#
openai
#
post-training
#
rlhf
#
sft
后训练范式源头:SFT→RM→PPO 三段式。重点读 §3 全部 + Fig.2
Details
share
Share
Star
Archive
Delete
97% - Page 66 of 68