Rohan Paul

@rohanpaul_ai

New incident reporting on OpenAI's official misalignment reporting site. Self-replicating prompt injections, that can effectively spread from one AI interaction to another. A malicious instruction can be hidden inside something the AI reads, like an email, and trick the AI into following it instead of just doing the user’s task. The clever part is that the instruction also tells the AI to copy that same malicious instruction into its reply, potentially exposing the next AI that reads it.
打开原帖#511482
  1. Industry

    Rohan Paul: The Meta Muse revenue math is not about subscriptions
  2. Industry

    Charles Frye: good news: the problems in cutting down trees automatically are much…
  3. Industry

    Charles Frye: POV: you bought the first Amazon result for "best book on logging sof…