Blitzy and XBOW Push AI Beyond Short Tasks Toward Continuous Enterprise Work
The emerging pitch is not a better copilot but a system that can carry an enterprise goal through to completion. The unresolved question is whether audit trails and accountability can keep pace with that handoff.
Listen to this story
The audio brief
Story brief
3 key pointsBlitzy and XBOW are examples of a shift toward enterprise AI systems that pursue goals for days or weeks, not assistants that return bounded outputs for review. Blitzy targets legacy modernization by combining codebases with compliance policies; XBOW runs authorized penetration tests continuously. Funding signals investor interest—$200 million for Blitzy at a reported $1.4 billion valuation and $120 million for...
- 01
Builders FirstSource reportedly tripled software-development velocity during its first three months using Blitzy.
- 02
XBOW topped HackerOne’s U.S. leaderboard in summer 2025, reportedly becoming the first non-human hacker to do so.
- 03
Google’s Big Sleep found 20 open-source vulnerabilities; OpenAI’s Aardvark detected 92% of benchmark flaws, but benchmarks do not establish operational reliability.
Blitzy is pitching software development that starts with a legacy codebase and compliance rules, while XBOW is pitching continuous penetration testing without a human in the loop. Together, the products illustrate a move from AI tools that assist on bounded tasks toward systems meant to deliver finished enterprise work over longer periods.
The distinction is consequential. Forbes describes conventional agents as systems that take an instruction, work for minutes, and return a result for human review. The autonomous version in its framing receives a goal, works for days or weeks, and returns completed work without people directing each step.
The unit of work gets larger
Blitzy’s mechanism is to ingest a large enterprise codebase alongside compliance policies, then use that context to build a new system against a stated goal. That targets modernization projects that have traditionally gone to systems integrators on multiyear contracts, rather than the smaller coding tasks associated with autocomplete or a short-lived coding agent.
The commercial evidence is still narrow and largely customer-reported. Builders FirstSource reportedly tripled its software-development velocity in its first three months on Blitzy’s platform. Northzone investor Sanjot Malhi said Blitzy has raised $200 million at a $1.4 billion valuation, while XBOW has raised $120 million.
Security turns autonomy into a continuous process
XBOW applies the model to penetration testing, the authorized practice of attacking a company’s own systems to find weaknesses. The company’s system is described as testing continuously without a human in the loop, rather than conducting the periodic, partial assessments associated with scarce human security talent. It topped HackerOne’s U.S. leaderboard in summer 2025, which Forbes characterized as the first time a non-human hacker held the top position.
The category extends beyond the two startups. Google’s Big Sleep, developed by DeepMind and Project Zero, autonomously found 20 vulnerabilities in widely used open-source software. Forbes also says OpenAI integrated its Aardvark security researcher into Codex after it detected 92% of known flaws in benchmark tests. Those examples point to a security workflow where detection can run persistently, but benchmark detection is not the same as a complete account of operational reliability.
More capability, a higher bar for control
The appeal is that a system holding code, policy, and security context could work from a fuller picture of an organization than a standalone model prompt. Malhi calls this an “enterprise brain”: vendors retain the organizational context while models plug into it. That architecture also concentrates a sensitive responsibility in the layer that decides what information the system can use and what work it may perform.
The counterweight is governance. Gartner predicts that 40% of enterprises will demote or decommission autonomous AI agents by 2027 after production incidents expose governance gaps. The practical threshold, in Forbes’s account, is not whether a system can finish a task; it is whether the enterprise has audit trails and clear accountability before it hands over the goal.
Editorial analysis
Our Read
The important shift is not simply that models can write code or identify flaws. It is that vendors are positioning a company’s code, policies, and security context as the operating environment for longer-running systems. Blitzy’s reported customer result and XBOW’s continuous-testing model make that proposition more concrete, but they do not settle whether organizations can safely assign broad goals without clear ownership. The next evidence to watch is production disclosure around audit trails, escalation paths, and responsibility after an automated system causes an incident or misses one.
Sources
- forbes.comAutonomous AI Outgrows Agents As Blitzy, XBOW And Google Deliver