AI21 Labs
@AI21Labs
We analyzed a year of agent research across 4 tasks: deep research, RAG indexing, agentic search and coding.
Found that first asking ‘Is this task verifiable?’ exposed inefficiencies in the agent architecture that were worth optimizing before reaching for the scaling lever.
That’s our verifiability litmus test:
> If you can't verify the task → aggregate
> If you can verify it → generate cheaply and spend your budget on the check
Full write up here: https://t.co/HKr3Q5UcWl