Nathan Lambert

@natolambert

An basic idea in scaling RL: Can we allocate more compute to the harder problems? We did this: If your GRPO group has all wrong completions, sample more with probability P (~0.9) -- in search of more GRPO batches with nonzero gradient. It works! Called "Never Give Up"
打开原帖#511482
  1. Industry

    Alexandr Wang: We believe strongly in the necessity to invest into alignment
  2. Industry

    Alexandr Wang: We believe strongly in the necessity to invest into alignment
  3. Industry

    Mark Zuckerberg: Last month I wrote about how we can build a positive and safe future…