Sebastian Raschka

@rasbt

Reasoning From Scratch: Reinforcement Learning with Verifiable Rewards (RLVR) round 2. Covering clipped policy ratios, KL loss term, format rewards, and other GRPO tips & tricks. 00:00 Introduction and recap 01:52 Interpreting basic GRPO training metrics 06:34 Planned improvements to GRPO 08:58 Running longer training jobs with Python scripts 13:39 Running the baseline GRPO training script 17:29 Loading and plotting training logs 19:29 Diagnosing unstable training 23:55 Evaluating checkpoints on MATH-500 26:26 Downloading existing checkpoints 30:09 Tracking advantage statistics 34:53 Understanding entropy 40:32 Computing entropy in PyTorch 44:17 Interpreting entropy values 48:58 Adding entropy tracking to GRPO 53:36 Analyzing advantage and entropy metrics 56:18 Stabilizing GRPO with clipped policy ratios 1:03:27 Implementing the clipped policy loss 1:09:39 Analyzing clipped policy training results 1:11:25 KL divergence and reward hacking 1:15:12 Adding a KL loss term 1:20:34 Limitations of the simplified KL loss 1:23:04 Format rewards and think tags 1:25:47 Adding special tokens to the tokenizer 1:30:29 Implementing the format reward 1:35:56 Analyzing format reward training 1:38:25 Rewarding format only for correct answers 1:40:48 Further GRPO improvements from research 1:45:43 Next steps and distillation
打开原帖#511482
  1. Research

    Yann LeCun: @Helios_Hua Nice work
  2. Industry

    Mustafa Suleyman: Super Intelligence must be contained... Today, this is an…
  3. Industry

    Databricks: Admins can now configure coding agents in one place with Unity Gateway