John Schulman
@johnschulman2
Bullish on this direction. Having a metric for explanation quality makes it possible to hillclimb, and counterfactual simulatability seems right. Adam et al. created a dataset+pipeline that creates more diverse+realistic test cases than prior work & do interesting exps on it. Can also train models to write better post-hoc explanations of their behavior, as highlighted in this thread.
> Can a model learn to explain its behaviors, such as why it ignored a user request or made a coding error?
>
> We trained models on thousands of explanations of their own in-the-wild behaviors.
>
> Training on this single general dataset shows generalization to held-out evals.