Sebastian Raschka

@rasbt

@BiploveYadav The model (training) should work fine without format reward, it’s just an optional stylistic thing. If the model never gets the answers right then you’d need to start with a simpler dataset.
打开原帖#511482
  1. Industry

    Rohan Paul: Another beautiful robotic hand
  2. Industry

    Rohan Paul: FT just reported Nvidia is in early talks to acquire, or invest furth…
  3. Industry

    Rohan Paul: New Stanford paper finds that harness changes fix agents that loop or…