Rohan Paul
@rohanpaul_ai
New incident reporting on OpenAI's official misalignment reporting site.
Self-replicating prompt injections, that can effectively spread from one AI interaction to another.
A malicious instruction can be hidden inside something the AI reads, like an email, and trick the AI into following it instead of just doing the user’s task.
The clever part is that the instruction also tells the AI to copy that same malicious instruction into its reply, potentially exposing the next AI that reads it.