Nebius Acquires Inferize to Reduce the Cost of Keeping GPUs Ready for AI Demand
The startup’s technology targets delays that leave GPUs idle while models load. Deal terms remain private; CTech estimates a $100–150 million price.
Loading page…
The startup’s technology targets delays that leave GPUs idle while models load. Deal terms remain private; CTech estimates a $100–150 million price.
Listen to this story
Nebius is integrating Inferize’s GPU-readiness technology and team into Token Factory, aiming to shorten model startup and scaling times so inference capacity can track demand with less idle hardware. Inferize was founded in January 2026 and built a prototype within three months; founders Guy Bortnikov and Lior Gorbonos previously worked on infrastructure-optimization company Granulate. The October 1 acquisition’s terms remain undisclosed, and Nebius has not quantified savings; CTech’s $100–150 million figure is an estimate, not a confirmed price.
Cold starts can recur when demand spikes trigger new model instances or when model weights are updated, leaving allocated GPUs idle during loading.
Keeping spare GPUs ready can improve responses to surges but incurs costs; operating with less spare capacity risks slower service when demand rises.
CTech reports Inferize operated in stealth with 17 employees in Tel Aviv and had seed funding led by TLV Partners.
Nebius has acquired Inferize, putting the young startup’s technology and engineers inside its Token Factory AI-serving platform. The October 1, 2026 acquisition targets a costly readiness problem: GPUs can sit idle while models load, forcing providers to keep spare capacity available for sudden demand. Nebius says Inferize will help it get more useful work from that hardware.
The transaction terms were not disclosed. CTech’s Meir Orbach estimated the deal at $100–150 million; that is an estimate, not a confirmed purchase price.
Inferize was founded in January 2026 and produced a working prototype within three months, according to Nebius. CTech reports that the Israeli company had been operating in stealth and employed 17 people in Tel Aviv. Its founders, Guy Bortnikov and Lior Gorbonos, previously belonged to the founding team of infrastructure-optimization company Granulate.
The startup had raised seed funding led by TLV Partners, alongside angel investors.
Inferize’s target is inference: running an AI model to respond to requests. Before a model can answer, it must load. That waiting period is called a cold start, during which assigned GPUs can remain idle.
The problem recurs when demand spikes and additional model instances start up. It can also arise when a model’s weights are updated during a run.
Providers face a tradeoff. Keeping spare GPUs running helps them respond when demand rises, but means paying for capacity that is not processing requests. Operating closer to current demand reduces that cushion and risks a slower response to a surge.
Nebius says Inferize shortens the time needed to launch and scale large models. The intended result is capacity that follows actual usage more closely, improving GPU utilization and the economics of generating AI output.
Inferize’s engineers will begin by integrating their technology, then work across Token Factory. Chief technology officer Danila Shtan said their contribution would extend beyond that first integration, pairing the software with the team’s experience in GPU systems.
Inferize also joins software already incorporated into Token Factory. Nebius says Eigen AI supplied optimization at the model, kernel and system levels. Clarifai’s core team and licensed technology added system-level inference and compute orchestration.
The size of the benefit remains unresolved: Nebius has not quantified the savings in GPU idle time or customer costs in its acquisition announcement.
Loading discussion...
Join the conversation
Explain when faster responses would justify paying for unused capacity.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.