Nathan Lambert

@natolambert

The biggest disagreement JSD and I had in this podcast was on the impact of distillation. In our research for Interconnects, @xeophon and I agree that there's no hard evidence that distillation is a massive impact for the Chinese labs. At the same time, the gossip mill in SF has been doubling down on the Chinese labs getting massive gains from distilltion. The argument where distillation is a huge impact is something along the lines of distilled traces go into mid training and make RL work far more easily. My argument, that distillation is a 1-2 month pull ahead in capabilities closer to the frontier, is that Chinese labs already have sufficiently strong models, where this mid-training setup can be done on their own models, and most of the capabilities gains are from scaling RL environment training, which doesn't link cleanly to the distillation data pathway. We feel like the recent RL dashboard from @XiaomiMiMo supports this claim. A lot of it comes down to a gut call on if you can believe it that the Chinese labs are really great at building LLMs, maybe even better than focusing than OpenAI/Ant etc, as their current ambitions are a bit narrower (catching up), rather than transformative products/inventions
打开原帖#511482
  1. Industry

    Clément Delangue: Feels like I’m the diversity pick on this one 😅
  2. Industry

    Alexandr Wang: there it is a real career highlight for me
  3. Industry

    Aravind Srinivas: /btw to initiate side chats inside a Computer session with a shared s…