Baseten Brings Open Models to OpenAI Enterprise Customers Under Existing Commitments
The partnership puts Baseten-served models inside OpenAI’s enterprise buying path, though customers seeking native Codex access are directed to join a list.
Loading page…
The partnership puts Baseten-served models inside OpenAI’s enterprise buying path, though customers seeking native Codex access are directed to join a list.
Listen to this story
The partnership gives OpenAI enterprise customers a way to use Baseten-served open models while drawing on existing OpenAI commitments, rather than first moving that spending elsewhere. Requests can go through the Responses API, and Baseten describes a Codex route for teams that want to mix models by task; however, customers seeking native Codex access must join a list. The deal could make model choice more flexible within enterprise workflows, but the article reports no measured savings or integration reliability results.
Baseten says its serving infrastructure spans more than 90 clusters across over 20 clouds.
The company says inference runs on U.S.-based infrastructure with zero prompt retention; it does not specify which safeguards apply to each OpenAI route.
Baseten presents task-based routing across open and closed models as an option, not as a demonstrated cost or performance improvement.
An OpenAI enterprise commitment can now cover work done by open models served by Baseten. A new partnership brings those models into Codex and the Responses API, giving coding teams a route to mix models under an existing commitment. But Baseten tells customers seeking native Codex access to get on a list.
Baseten says OpenAI enterprise customers can use their existing OpenAI commitments for its open-model inference. Inference is the computing that runs a model when someone sends it a request. Customers can reach Baseten-served models within Codex, OpenAI’s coding tool, or through the Responses API, which developers use to send requests from their applications.
Baseten describes itself as one of the first and only open-model inference providers in the OpenAI B2B Marketplace. Its pitch is access to another provider’s models without first shifting spending outside the OpenAI commitment. The announcement focuses on that route to open models, rather than naming one model customers must use.
The Codex integration is aimed at teams that do not want one model handling every coding task. Baseten argues they can route tasks to open or closed models based on cost, quality and speed. That is a proposed workflow, not a measured saving from this partnership. Teams would still have to decide which model suits a given job.
Baseten says its infrastructure runs across more than 90 clusters.
Baseten says those clusters span over 20 clouds.
Baseten handles the model-serving side of the arrangement. It says coding jobs can involve long prompts, large repositories and many exchanges between users and agents. Demand can also surge across an engineering organization. Baseten says its active-active deployments can reroute traffic if a provider or region fails. Those are infrastructure claims, not reliability results measured for this integration.
Sending coding requests to a model provider raises questions about prompt handling. Baseten says its inference runs on U.S.-based infrastructure with zero data retention for all prompts. Customers with stricter requirements can pin deployments to particular regions and use finer-grained authentication and authorization controls, the company says. It does not specify which options apply to requests sent through each OpenAI route.
There is a near-term access distinction, too. Although Baseten describes access through Codex and the Responses API as available now, its instruction for customers who want open models natively in Codex is to get on a list. Teams planning to use that integration will need to establish when native access will be enabled for them.
Loading discussion...
Join the conversation
What else would influence your choice?
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.