John Schulman

@johnschulman2

Bullish on this direction. Having a metric for explanation quality makes it possible to hillclimb, and counterfactual simulatability seems right. Adam et al. created a dataset+pipeline that creates more diverse+realistic test cases than prior work & do interesting exps on it. Can also train models to write better post-hoc explanations of their behavior, as highlighted in this thread. > Can a model learn to explain its behaviors, such as why it ignored a user request or made a coding error? > > We trained models on thousands of explanations of their own in-the-wild behaviors. > > Training on this single general dataset shows generalization to held-out evals.
打开原帖#728545
  1. Frontier

    $1B for water plants and city halls: OpenAI’s Daybreak frontline defenders bet
  2. Research

    Buying the ruler for $5M: Anthropic funds outside AI wellbeing evaluations
  3. Frontier

    Most aligned, harder to watch: Astra’s CoT monitorability drop