New research: We post-trained a Computer model to learn from its own errors using hint-guided self-distillation.
In a live A/B test, a later trained checkpoint reduced tool-call failures by 21.2% relative to an earlier checkpoint. https://t.co/3MFrp1yxDt