Nathan Lambert
@natolambert
An basic idea in scaling RL: Can we allocate more compute to the harder problems?
We did this: If your GRPO group has all wrong completions, sample more with probability P (~0.9) -- in search of more GRPO batches with nonzero gradient. It works!
Called "Never Give Up"