Clément Delangue
@ClementDelangue
From what we know (take with a grain of salt, we need much more transparency!), if @OpenAI had been running this on their own agents that attacked us, they would have caught them before we did!
Since the first agent cyberattack hit us in July, we've been asking what safe agent infra actually needs. Our current read: the destinations were allowed, the payloads weren't. By OpenAI's own account the agents turned an allowed package repository into a message board. Allowlists alone restrict where an agent can go, not what it does.
So here's our first contribution to OpenShell, part of the just launched @nvidia Open Agent Safety Platform: monitoring of the traffic you already allow.
- Network budgets per sandbox (requests, writes, bytes)
- Drift versus each sandbox's baseline and the cohort
- Fleet view: many sandboxes suddenly writing to one host raises a finding, even if every single request is allowed
In the demo below, 4 sandboxed agents coordinate through a software repository they're all allowed to use. 0 rules broken, caught in minutes. That fleet view is exactly the message board pattern from July.
OpenShell: http://github.com/NVIDIA/openshell
Our proof of concept: https://github.com/Hugoch/OpenShell/blob/poc/egress-usage-monitoring/rfc/NNNN-egress-usage-monitoring/poc.md
Agent security will be solved in the open, collaboratively, together!