Rohan Paul

@rohanpaul_ai

New Google paper shows for small quantized LLMs on serverless CPUs, 55–70% of cold-start latency is just model loading. i.e. the bottleneck is often moving model weights, not generating tokens. The inference itself is less of the problem than getting the model into memory. Giving the same model 8 GB of Cloud Run memory instead of 4 GB unlocks roughly 2× the CPU, cutting warm inference time almost in half.
打开原帖#511482
  1. Industry

    Alexandr Wang: MUSE FEATURE ALERT: muse for mac now has computer use! queue up your…
  2. Industry

    Alexandr Wang: MUSE FEATURE ALERT: y'all have been asking for this one
  3. Industry

    Clément Delangue: Full transcript: Minister Barrot, members of the Security Council, th…