I like papers like this because it starts to communicate this differential in terms of return for compute.
It’s actually a harder space to measure comparisons on because the experimental setup for test time compute is more complex and less easy to replicate.