Simon Willison

@simonw

"Deterministic code checks the result" sounds like they might be implementing a variant of the DeepMind CaMeL paper https://simonwillison.net/2025/Apr/11/camel/
打开原帖#511482
  1. Products

    Grok Bot for Enterprise: the news is audit and allowlist, not another AI teammate pitch
  2. Research

    Anthropic’s Model Hardware Standard: the news is the shared driver, not the robot
  3. Research

    DeepMind’s double-blind Gemini eval is an enclave pattern—not a public safety scorecard