Rohan Paul

@rohanpaul_ai

New Tsinghua paper builds TokenRouter, a serving system that runs per-token small-and-large model routing at upto 64.15X the throughput of existing setups. Current popular serving frameworks (like vLLM and SGLang) run one model per request, so when two models share an answer, every step waits for the slower one. TokenRouter gives each model its own server and lets them pass work back and forth. So, it hands a half-written answer between them while keeping the model’s memory of the text so far (the KV cache), and holds requests for a moment so each model works on bigger batches. Across 5 routing methods, throughput rose 2.01 to 64.15 times over the stronger existing setup. – arxiv. org/abs/2610.12242 Title: "TokenRouter: Efficient Serving System for Token-Level LLM Routing"
打开原帖#511482
  1. Research

    Yann LeCun: @Helios_Hua Nice work
  2. Industry

    Mustafa Suleyman: Super Intelligence must be contained... Today, this is an…
  3. Industry

    Databricks: Admins can now configure coding agents in one place with Unity Gateway