Hamel Husain
@hamelhusain
Put this through its paces yesterday. Some thoughts:
The Bad:
1. The workflow creates evals before looking at data, and has you correct its mistakes. I believe you should be looking at data first to inform your understanding and to prioritize what to work on before trying to write evals.
2. It created markdown files for looking at data and asked us to "tell it" which labels were off. You should create your own annotation app instead. You are using coding agents after all (we did this in our livestream)!
3. The evaluators it scoped were too broad, and bundled too many types of failures at once. This can be avoided with looking at the data first. I'm afraid this might steer people in the wrong direction.
4. The workflow put us DEEP into the rabbit hole of a specific eval right away. It also asks "does it look right" at several steps without really giving you enough information to make that determination. I think this will steer people towards the wrong workflow in many cases.
The Good:
1. I was impressed by out of the box ability for issue discovery that other auto-eval approaches haven't been able to find! It found issues with human handoff, formatting, voice agents, and more. It's still better to look at your data iteratively with an agent, but this was the strongest performance I've seen with a more "one-shot" issue discovery approach.
2. This is a big improvement in terms of UX from their prior eval plugin thing https://x.com/ClaudeDevs/status/2098500999656923145?s=20
3. The blog post they released conveys thinking that I agree with, such as the importance of looking at data, sampling intelligently, not saturating your own evals, etc. I'm really happy more people are thinking about evals this way.
Video of our attempt at using this here: https://x.com/i/broadcasts/1nKOLQOwQEEGR?s=20