Hamel Husain

@hamelhusain

Q: Should I build automated evaluators for every failure mode I find? A: No. Different types of evaluators have different costs (code vs. LLM judge), which you need to weigh before building an eval. https://hamel.dev/blog/posts/evals-faq/should-i-build-automated-evaluators-for-every-failure-mode-i-find.html
打开原帖#511482
  1. Research

    François Chollet: The critical distinction between base LLMs (2024 and earlier) and mod…
  2. Industry

    Michael Truell: Grok Bot gets useful work done, without you needing to ask.
  3. Industry

    Aravind Srinivas: We're open-sourcing our multimodal Decision model and offering it thr…