Modal Opens Multi-Machine AI Computing to All Workspaces With Per-Second Billing
The company says clusters start within seconds and configure high-speed networking automatically. Access is broad, but each plan’s GPU limits still bound cluster size.
Loading page…
The company says clusters start within seconds and configure high-speed networking automatically. Access is broad, but each plan’s GPU limits still bound cluster size.
Listen to this story
After 1.5 years of testing, Modal has made Clusters generally available for coordinated multi-machine AI training and inference. Its gang scheduler places nodes together, and clusters can scale as a unit; RDMA support enables high-speed GPU communication for training and inference workloads. Per-second billing makes short runs possible, but access is not unlimited: cluster size remains subject to each plan’s GPU limits, and Modal asks customers to contact it about larger runs.
Decagon used Modal to fine-tune open models with up to one trillion parameters using its Miles training software.
1X uses multi-node B300 clusters with RDMA to pre-train the world model for its NEO home robot.
Runway distributes inference across multiple nodes, including for low-latency, real-time video conversations in Runway Characters.
Developers across Modal’s workspaces can now request multiple machines working together for AI training and inference. Modal has made Modal Clusters generally available after 1.5 years of testing. The company says workloads can start within seconds and are billed by the second.
Access does not mean unlimited capacity: cluster size is bounded by each plan’s GPU limits. Modal directs customers planning larger runs to contact it. Clusters also connect to existing Modal tools for storing training checkpoints, loading data and coordinating jobs.
Modal names Decagon, 1X and Runway as users. It says Decagon fine-tuned open models with up to a trillion parameters using Miles training software. 1X uses multi-node B300 clusters with RDMA networking to pre-train the world model behind its NEO home robot, alongside single-node evaluation jobs.
Runway spreads inference for its video generation and editing models across multiple nodes, according to Modal. Its autoscaler grows or shrinks each cluster as a unit. In a March 26 announcement, Modal described Runway Characters using multi-node clusters for low-latency, real-time video conversations.
Modal Clusters let us scale up to hundreds of B300s with RDMA, then spin them down when we’re done, with minimal overhead.
Sam Sinha, Head of World Model Lab at 1X, in Modal’s announcement
Developers request a cluster through @modal.clustered, a Python decorator attached to a function. Modal’s example requests four containers—the environments running customers’ code—with eight B300 GPUs each:
@app.function(gpu="B300:8")
@modal.clustered(size=4, rdma=True)Behind that request is a new “gang scheduler.” Rather than assigning work machine by machine, it examines the fleet and pending clusters together. It chooses placements, adds capacity where needed and groups nodes by availability zone and network.
RDMA, or Remote Direct Memory Access, transfers data between machines’ memory without routing it through the operating-system kernel. Modal says its InfiniBand networking supports up to 6.4 terabits per second per node. The rdma=True flag automatically configures it for PyTorch workloads using NCCL, software that coordinates GPU communication.
The connection supports training synchronization and transfers of cached model context during inference. Modal also extended automated health checks to RDMA. Its gVisor container sandbox lacked RDMA support, so the company says it built that capability and contributed the changes to the open-source project.
Loading discussion...
Join the conversation
Explain when flexibility should outweigh fixed capacity limits.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.