CoreWeave Launches Forge to Connect AI Deployment, Monitoring and Improvement
The platform brings production feedback into training and evaluation. Pro starts at $60 a month, but teams with 50 or more employees must use Enterprise.
Loading page…
The platform brings production feedback into training and evaluation. Pro starts at $60 a month, but teams with 50 or more employees must use Enterprise.
Listen to this story
Forge packages CoreWeave’s former Weights & Biases, OpenPipe and Marimo products with new services in a shared environment for improving deployed AI systems. Its core workflow carries production traces into datasets, training and release evaluations, so teams can investigate failures and test fixes without relying on disconnected handoffs. The offer is integrated with CoreWeave’s cloud but does not require a cloud move; pricing and eligibility are tiered, with Pro limited to teams under 50 employees.
Agent Lens records agent steps, decisions and tool calls; flagged failures become versioned datasets that retain production context.
Teams can improve prompts, tools or retrieval without retraining; training options include supervised fine-tuning, reinforcement learning and distillation.
Sandboxes can run agents and evaluations on infrastructure teams already use, while Forge also supports workloads on other clouds.
CoreWeave launched Forge on September 30, 2026, bringing AI deployment, monitoring, data preparation, training and evaluation into one development environment. Available starting today, the platform is designed to turn problems found in live models and agents into tested improvements. CoreWeave says teams can connect workloads on other clouds or their own infrastructure rather than move everything onto its cloud.
Announced at Fully Connected 2026, Forge addresses a problem CoreWeave describes as a disconnected improvement cycle. A production team may see whether a service is running, while researchers lack the conversations needed to judge answer quality. Evaluation tests can fall behind changing user behavior. The company’s pitch is to connect those records and workflows, reducing the manual handoffs between finding an issue and checking a fix.
The connection starts with traces: records of an agent’s steps, decisions and tool calls. Agent Lens captures those records, and monitors score live behavior against baselines set by the team. Flagged failures become versioned datasets in Weights & Biases Models, retaining the production context. Those examples can then inform both training and the tests used to decide whether a replacement is ready to ship.
An improvement need not mean retraining a model. Forge’s workflow can begin with changes to a prompt, a tool or retrieval—the process of finding information for a model to use. When training is needed, CoreWeave offers supervised fine-tuning on examples, reinforcement learning using reward signals, and distillation. The company says teams can use these services without setting up a training cluster.
Forge is not a collection of entirely new products. It brings together the former Weights & Biases, OpenPipe and Marimo products, alongside CoreWeave services. The launch introduces Agent Lens, Model Distillation and Notebooks as new services; ARIA and Sandboxes are now generally available. Three components illustrate the different jobs within that shared environment:
ARIA adds assistance across that work: it analyzes experiment history, proposes experiments and recommends code changes. CoreWeave describes it as advisory, with the user’s judgment remaining in charge. Evaluations compare a candidate with the current version using production traces. Registry records the dataset and model checkpoint that passed, plus the version to roll back to. The workflow therefore includes a check before deployment, not just suggestions for changes.
CoreWeave says Forge works with teams’ existing models, frameworks and clouds, while keeping generated improvement data in portable formats. That openness sits alongside tighter integration under one account and navigation. CoreWeave also says the platform runs best on its own cloud. Teams can connect external workloads and choose CoreWeave’s training or inference services when they want the infrastructure as well.
The Forge pricing page lists Free at $0 a month and Pro starting at $60 a month, billed monthly, with a 30-day trial. Pro is restricted to early-stage teams with fewer than 50 employees; customers exceeding that limit must transition to custom-priced Enterprise plans. Existing Weights & Biases users keep their credentials, projects and wandb.ai links, with the other Forge tools accessible through the same account.
Loading discussion...
Join the conversation
Explain whether portability would change your preference for separate tools.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.