Hamel Husain

@hamelhusain

Recently had the chance to interview @benhylak on how he is thinking about simulations in our AI Evals Course 😁 Simulations are incredibly useful for evals when you do them well. Ben discusses what he learned from implementing simulations for evals at several companies. Some fun topics came up, my favorite being "simulation awareness" - when models know they are in a simulation. Below is a chapter summary: Why input/output tests miss agent failures Replay production traces to see what a change breaks Recreate the data and tools around the agent How agents detect simulations and change their behavior Choose scenarios that test the change you made Check answers, costs, and tool behavior across runs Replay known failures to verify a fix Common mistakes that make simulations misleading Check that the simulation reproduces production behavior When to start using simulations You can also watch this on YT: https://www.youtube.com/watch?v=eLkESCLvvAs
打开原帖#511482
  1. Research

    Yann LeCun: @andrewgwils A generalisation theory that applies to self-supervised…
  2. Research

    Yann LeCun: @tdietterich @andrewgwils We would have to insist that we prefer mode…
  3. Frontier

    Odyssey-3 opens a world-model preview