Dylan Patel
@dylan522p
We are excited to bring the first open benchmarking of Google's TPUs to the world
Running every day, on many models + scenarios
$/token is better than B200 and B300
Huge shout-out to Google @inferact and the InferenceX team at SemiAnalysis to this effort that's taken many months
Quoted: TPU Inference Externalization — Full Steam Ahead - InferenceX; up to 50% better performance per dollar; rapid externalization of TPU stack; Ironwood, TPUv8i, reducing CUDA moat. https://newsletter.semianalysis.com/p/tpu-inferencex-full-steam