Sasha Rush

@srush_nlp

Here's a fun one with @preston_fu et al -- how should we reward RL policies in a way that scales to longer and more difficult tasks? Our answer lies in-between RL and imitation learning, and provides a simple way to assign dense per-token credit to long trajectories.
打开原帖#511482
  1. Research

    OpenAI's Chief Scientist Issues a Rare Warning: We Grew an Alien Mind, and Our Ability to Monitor It Is Fading
  2. Industry

    Alexandr Wang: if this were an mma match, only one guy would be left standing (hint…
  3. Industry

    Alexandr Wang: 3/ Muse operates with the principle of least privilege, so you can de…