Harell Data Picks CoreWeave to Run AI Training Without Sending Out Raw Data

The multi-year deal would add cloud capacity for Harell’s marketplace, where dataset owners and model builders can each earn from their work.

By 3 min read
Harell Data Picks CoreWeave to Run AI Training Without Sending Out Raw Data
Harell Data Picks CoreWeave to Run AI Training Without Sending Out Raw Data

Listen to this story

The audio brief

About 1:21
0:001:21
Read transcript
CoreWeave will provide cloud capacity for Harell Data’s planned AI platform, where companies can train models on proprietary scientific data without taking away the raw files. The multi-year agreement covers training, fine-tuning and inference, using NVIDIA A100 and Hopper processors inside Harell’s platform. The design flips the usual arrangement: instead of handing sensitive research data to a model builder, Harell says the model goes to the data. Builders can take away the trained model and its intellectual property, then sell access to that model through Harell’s marketplace. Dataset owners, meanwhile, are meant to earn a share whenever a training job uses their data. Harell has not disclosed the revenue-share rate, and the announcement offers no earnings figures from the new deployment. Harell says its evaluation process also keeps test questions hidden from builders until results are published against a common standard. That could help make comparisons without exposing the test material, but it is a separate part of the platform’s approach. There’s an important caveat: the deployment is still planned. The agreement adds computing capacity, but it does not yet show how well the platform performs or independently verify Harell’s claims about keeping data protected. The next concrete test is whether researchers can run real workloads on the CoreWeave-backed service while the underlying files remain inside Harell’s platform.

Story brief

3 key points

Harell Data has signed a multi-year agreement to run its data-local AI platform on CoreWeave Cloud, adding capacity for training, fine-tuning and inference on proprietary scientific datasets. The design lets builders take away trained models and related IP while raw files remain inside Harell’s platform; runs will use NVIDIA A100 and Hopper GPUs. Dataset owners are promised usage-based revenue shares, though Harell...

  1. 01

    Harell says dataset owners earn a share each time a training job uses their data; builders can charge for later model use through its marketplace.

  2. 02

    For evaluations, Harell says test questions stay hidden from builders until results are published against a common standard.

  3. 03

    The agreement does not yet demonstrate performance or independently validate the platform’s data-handling claims.

Harell Data plans to put CoreWeave’s cloud behind a way to train AI on proprietary scientific data without sending the raw datasets to model builders. Under a multi-year agreement, Harell will deploy its platform on CoreWeave Cloud for training, fine-tuning and inference. The deal puts computing capacity behind Harell’s central promise: builders can work with valuable data without taking possession of it.

The model travels to the data

Harell’s platform is built around a separation: model builders can run training against a dataset without receiving the underlying data. CoreWeave says those runs will execute inside Harell’s platform on NVIDIA A100 and Hopper GPUs. The raw data stays there; what a builder takes away is the trained model, its intellectual property and the right to sell it. That is Harell’s description of the design, not an independent security assessment.

Harell argues that handing valuable research data to another party for commercial model training can destroy its value. Keeping the files inside its platform is its proposed answer to that risk. The CoreWeave agreement is intended to supply the cloud infrastructure as Harell scales the service for researchers.

We created Harell Data to connect researchers to proprietary scientific datasets for model training and inference without requiring them to view or extract the underlying data.

Harlan Robins, founder and president of Harell Data

Two ways to earn from one dataset

The agreement also backs a marketplace with payments on both sides of a model’s life. Harell says it meters computing use for each training run and gives the dataset owner a revenue share whenever a job uses its data. That ties payment to training activity rather than requiring the owner to hand over a copy. Harell has not specified the owner’s share.

Builders have a separate route to revenue. They can list models trained through Harell on its marketplace under a Models as a Service arrangement and collect a fee each time someone runs one. A data partner earns when its dataset is used for training; a builder can earn from later use of the resulting model. Those are Harell’s stated payment terms, not earnings figures from the CoreWeave deployment.

The test data stays hidden, too

Harell applies a similar separation when models are evaluated. Its data partners keep test questions hidden from builders, so a builder has not seen the answers before a result is published. Harell says results are then published against one standard. Unlike keeping training data in place, this step is meant to make model evaluations comparable without exposing the test material to builders.

The deployment is still planned, and the announcement includes no results from jobs run under the new agreement. For researchers, the next test is how the CoreWeave-backed platform performs when they use it—not only how its data-handling design is described.

Sources

  1. coreweave.comHarell Data Selects CoreWeave for Secure AI Training

Loading discussion...

YOUR READING SPACE

Notifications