Rohan Paul

@rohanpaul_ai

Andrew Feldman, co-founder and CEO of Cerebras gives the best explanation of why Cerebras' wafer-scale architecture is 2,500X faster than a GPU during LLM inference. During inference, there are 2 stages: - pre-fill, where the model first processes the user's prompt, and - decode, where it generates the answer 1 token at a time in sequence. During that sequencial Decode phase, before each token is calculated the model weights have to be moved from memory into compute. On a GPU those weights are moved from HBM, while Cerebras keeps them in much faster SRAM spread across its very large wafer-scale processor, so the memory-to-compute movement that must happen for every token is about 2,500× faster ---- From The MAD Podcast with Matt Turck and Cerebras YouTube channel, (link in comment)
打开原帖#511482
  1. Industry

    Alexandr Wang: you can use a muse gadget to track your wine collection
  2. Industry

    Alexandr Wang: mathematicians and muse spark collaborated to solve 6 open problems i…
  3. Industry

    Alexandr Wang: eink muse gadget we're basically one click away from the talking hogw…