Flower Labs Starts Endeavor 1.0 Preview With a Private Deployment Option
The model’s differentiator is not just Flower’s frontier-performance claim. Customers can run it as a managed service or place selected workloads inside their own infrastructure, though the launch remains limited to early users.
Listen to this story
The audio brief
Story brief
3 key pointsFlower Labs is previewing Endeavor 1.0 as a production-oriented generalist model, but access remains limited to selected organizations and partners while the company expands compute. Its commercial distinction is deployment flexibility: Flower can operate Endeavor as a managed service, or customers can keep it inside their own infrastructure and potentially migrate between modes. Flower reports strong results on...
- 01
Endeavor scored 92.0 on GPQA, 98.2 on HumanEval, 99.9 on AIME 2026, and 94.1 on IFEval in Flower’s comparison.
- 02
The preview is not self-service; Flower is onboarding selected organizations and partners before a broader rollout.
- 03
Private deployment targets sensitive workloads, with Flower claiming customer data can remain outside a centralized database.
Flower Labs has begun previewing Endeavor 1.0, a generalist AI model the company says can handle reasoning, coding, and long-horizon agent work. The release offers a managed service and an option to deploy the model within an organization’s own infrastructure.
Access is restricted for now. Flower is onboarding a select group of organizations and partners rather than opening self-service access, while it increases compute availability for a broader rollout. It describes Endeavor as its first production-ready release in the model line.
Two operating paths
Teams can have Flower handle deployment, scaling, and model operations through its managed service, which the company presents as a faster path to production. Private deployment keeps the model in an organization’s environment and gives it more control over sensitive workloads. Flower says customers can begin with the managed option and later move more of the deployment into their own systems.
A performance claim with narrow evidence
Flower’s case for frontier status rests on its own four-test launch comparison. It reports Endeavor scored 92.0 on GPQA, 98.2 on HumanEval, 99.9 on AIME 2026, and 94.1 on IFEval; its table places Endeavor first on HumanEval and level with GPT-5.6 Sol and Claude Fable 5 on AIME 2026. Flower also says small benchmark collections cannot fully capture practical usefulness.
The company says performance also depends on the system around the trained model: how it assigns reasoning effort, maintains context, uses tools, and checks or recovers from failed steps. Flower says its FlowerBench enterprise evaluation runs opt-in organizations’ tasks inside their own environments, retaining proprietary data and internal context there while producing sanitized results.
Built from an earlier sovereign model
Endeavor follows Lizzy, Flower’s sovereign 7B model for UK use, by four months. Flower says Endeavor combines capabilities from open-weight models with its own specialist knowledge, model behaviors, and training advances. Tech.eu reports that Endeavor is licensed and can be trained on customer data without moving that data to a centralized database.