Rohan Paul
@rohanpaul_ai
So OpenAI will now publicly disclose model misalignment even before it fully understands or fixes the behavior.
So its institutionalizing public disclosure of model failures instead of waiting for occasional system cards or bundled research reports.
They will prioritize cases that reveal new failure mechanisms, show known problems getting worse, or undermine assumptions about existing safeguards.