
Early drug work often stalls not on naming a molecule, but on designing a protein binder that latches a target — then surviving wet-lab validation. Anthropic’s August 18 research post puts hard numbers on the table: Claude (Mythos Preview and Opus 4.8) designed minibinders against 15 targets and, per external evaluation, succeeded on 14. Per-design hit rates landed between 22% and 35%, versus a typical 10%–15% in today’s campaigns. Some top designs beat previously published affinities.
[1]The official setup is concrete. Multi-target mode: Opus 4.8 and Mythos Preview ran 48 wall-clock hours with up to ~12,500 H100-hours and hit about 22.6% and 26.7%. Single-target mode: Mythos Preview used 24 hours and up to ~2,500 H100-hours per target, reaching 35.1%. Claude chose epitopes, orchestrated structure/sequence/co-folding models, ran optimization loops, and shipped designs to Adaptyv Bio and Twist Bioscience; humans mostly approved network access and watched infra. Second track: Opus 5 got raw NMR and LC-MS files plus a two-sentence prompt, finishing in 23 and 19 minutes with hydrogen counts and purity (96.4% vs 96.33%) matching the contract lab.
[1]The story is orchestration, not “AI invents drugs.” Specialist models still fold and score; Claude runs the campaign — site selection, loops, diversity filters. On Adaptyv contest targets, Mythos Preview hit 40% on RBX1 versus ~3.7% among participants; Opus 4.8 produced cross-species TNFα binders where Mythos Preview failed — Anthropic says it does not know why. Author’s read: this compresses weeks of computational-biology operations into an auditable agent session; it does not replace the wet bench.
[1]Pharma will ask whether hit rates and affinities reproduce under different prompts and budgets. Safety watchers will note Mythos-class life-science work remains gated and the scientist access program is still rolling out. Anthropic also says policy and operational bottlenecks dominate end-to-end drug timelines. Wet-lab results come from partners; treat “14/15 targets” as the published collaborative evaluation, not a public raw-plate dump.
[1]Math formalization offers second-scale checks; protein design puts verification back on a weekly clock. Publishing comparable targets and hit rates is how Anthropic argues agent orchestration belongs in experimental science. Next contest: who wires orchestration logs, wet-lab feedback, and access gates into one pipeline.
[1]