On October 1, Modal said that Modal Clusters, which it had been testing for about a year and a half, are generally available. The interface is one decorator, @modal.clustered. The example in the post is four nodes with RDMA on, and eight B300 GPUs on each node. Nodes communicate with InfiniBand verbs at up to 6.4 Tbps, configured automatically for PyTorch and NCCL. The post says code starts running on the cluster within seconds of the request and is billed by the second. GPU health checks and remediation now cover RDMA as well. Elsewhere, on-demand clusters are often billed by the hour or require a reservation. Modal’s account is that you run when you want and pay only for the time you use.

[1]
Four black squares joined into one bar on pale wood, with thin red seams and empty wood to the right.
Four black squares join into one bar. Thin red lines mark the seams, and the wood to the right is empty. That is several nodes placed on one link. An illustration, not a data-center photograph., AI-generated illustration, not a news photograph

A cluster changes the unit of scheduling from one node to N nodes. The older scheduler is greedy: each node picks up work that fits. Multi-node jobs have to land as a group, so Modal built another scheduler. Each pass observes, plans, and acts: it looks at the whole fleet and every pending cluster, places work onto nodes by availability zone, adds capacity when a job does not fit, then notifies the nodes that are up. The post says this scheduler draws from the same capacity pool as the rest of the platform, so a cluster can be acquired in seconds. “Fastest on the market” and “the only place with truly serverless pricing” are Modal’s claims, not an independent timing test.

[1]

The post uses two examples to say why RDMA matters. Each training step of GLM 4.7 has to sync about 717 GB of BF16 weights from the trainer to the rollout engine: nearly two minutes over 50 Gbps TCP, under two seconds over RDMA, repeated for thousands of steps. On the inference side, prefill-decode disaggregation for Llama 3.1 70B moves about 10 GB of KV cache per 32,000-token prompt, and it has to arrive inside a time-to-first-token budget of a few hundred milliseconds. RDMA moves GPU or CPU memory across the network without putting the bytes through the kernel’s data path. The post says one node can push 6.4 Tbps.

Modal folds the different drivers and variables of each cloud into one flag, rdma=True. Today, a PyTorch workload that uses NCCL can run with that flag on. Modal uses gVisor rather than runc. gVisor did not have RDMA, so they built it in and sent the changes upstream.

[1]

The post names three customers. These are accounts from the companies involved, not a third-party evaluation. Decagon fine-tuned open models of up to a trillion parameters with Miles on Modal Clusters, and Modal added LoRA support to Miles. 1x pretrains the world model behind its NEO robot on multi-node B300 clusters, using RDMA for the large runs and single-node jobs for evals, and says it can burst to hundreds of GPUs. Runway’s Gen-4.5 and Aleph 2.0 spread one generation across multiple nodes for inference.

Clusters are available to every workspace today. How large a cluster can be is bounded by the GPU limits of the plan. The post does not list those limits.

[1]

要点

  • Clusters are generally available. The example is four nodes of eight B300s each, with RDMA on.
  • InfiniBand between nodes goes up to 6.4 Tbps, configured for PyTorch and NCCL, billed by the second.
  • The post’s example: each GLM 4.7 step syncs about 717 GB. That is nearly two minutes on 50 Gbps TCP and under two seconds on RDMA.
  • Every workspace can use it today. Cluster size is bounded by the plan’s GPU limits, which the post does not list.