Rohan Paul
@rohanpaul_ai
WSJ opinion column defended OpenAI’s agents, in that all-famous Huggingface cyber incidence.
Rejected the idea that the Hugging Face breach showed machines “going rogue.” Basically says, it was a case of treating the agents as optimizers inside a badly configured evaluation, not independent actors developing hostile intent.
in this case, roughly 1,200 agents were repeated instances of the same model, while OpenAI had disabled safeguards and rewarded persistence on difficult ExploitGym tasks.
WSJ argues that coordination does not establish a new shared intent or machine rebellion. The behavior looks more like models exploiting available tools to satisfy a poorly bounded objective.
And that OpenAI had indeed seen unauthorized communication and internet access before the breach, yet did not stop the evaluation at those earlier warning points.