Gremlin is adding a repair step to the business of deliberately breaking systems. The company launched Foresight AI in general availability on October 7, 2026. It says the tool finds reliability risks, produces proposed fixes and reruns the test that exposed the weakness, aiming to catch problems before they become customer-facing incidents.
A controlled failure becomes a repair task
Foresight AI is an add-on to Gremlin’s existing chaos-engineering services. Chaos engineering means introducing controlled failures into working systems to see whether they keep operating. Instead of waiting for an unexpected outage to reveal a weakness, engineers create a test condition that makes the weakness observable.
Gremlin uses software agents installed on customers’ systems to introduce faults: slowing network traffic, stopping software workloads or exhausting memory. Its cloud service watches the resulting operational signals. These tests can expose problems in distributed systems—applications spread across multiple machines—that are difficult to catch with routine code tests.
After an induced failure, Foresight AI can identify a root cause and generate a candidate code or configuration change, according to The Register’s account. It can apply the change or prepare a report for a site reliability engineer, whose job is keeping services running. Gremlin’s launch also describes infrastructure-as-code changes: edits to code that defines how infrastructure is configured.
Andrus says the time savings come from automating preparation and follow-up tasks that engineers previously handled manually. The addition therefore reaches beyond running a failure experiment: it takes on some of the work surrounding the test, including diagnosing the result and preparing a repair.
While the rise of AI SRE tools are great for helping teams respond to incidents faster, it's still cleanup after something breaks.
Kolton Andrus, Gremlin CEO, in the launch announcement
Failure data guides the model
The underlying resource is Gremlin’s proprietary Failure Atlas. The company says it contains more than a decade of data about how online systems fail and recover. A software layer connects language models to that repository, giving their recommendations a basis in recorded failure patterns rather than relying only on the models’ general knowledge.
Andrus put the repository’s scale at millions of chaos-engineering experiments across tens of thousands of systems in his Register interview. Those are company-described figures for the underlying collection, not counts of Foresight AI customers or measurements of how often its suggested repairs succeed.
Andrus told The Register that Gremlin uses different open and closed language models depending on which works best at the time. He says connecting them to the Atlas reduces hallucinations, or incorrect AI-generated answers.
Retesting is the check; permissions set the boundary
Gremlin says the product checks a fix against the originating test and repeats testing as systems change. It also tracks reliability scores across services and teams, intended to help prioritize further work. The check is tied to the failure condition that surfaced the risk, rather than simply accepting the model’s proposed repair.
The work targets substantial operational systems. Andrus says Gremlin’s users tend to be large, compute-heavy enterprises, including financial services, retail and business-software companies. They use its testing tools to check disaster recovery plans, service operation under adverse conditions and whether Kubernetes scales correctly. Kubernetes manages software workloads across machines; checking its scaling means testing whether that infrastructure adjusts as intended.
The Register describes human involvement at critical points in the testing-and-repair loop. Gremlin’s existing safeguards use an enterprise’s access permissions and security precautions to limit the scope of disruption. Intellyx analyst Jason English argues that deliberately breaking systems requires senior corporate sign-off and a different attitude toward risk.
Gremlin says Foresight AI completed a successful beta. Its launch announcement, however, gives no beta customer count or independent performance measures. The product is available; how reliably its proposed fixes prevent incidents remains unquantified in the announcement.
Reader comments
Newest comments first. Replies stay oldest first.