Charles Frye
@charles_irl
New on the @modal blog: a write-up by me and Shreya on our work to speed up AI-SQL queries
- why AI-SQL? why high-throughput inference?
- how can we possibly be >10x faster than vLLM?
- what does the future of inference engineering look like?
Read here: https://modal.com/blog/quail-billion-tpm