Goodfire launches AI agent monitors that check internal signals, not just output
The system reserves a second AI model’s review for flagged activity. Goodfire’s cost and detection figures come from its own Kimi K3 tests.
Loading page…
The system reserves a second AI model’s review for flagged activity. Goodfire’s cost and detection figures come from its own Kimi K3 tests.
Listen to this story
Goodfire launched internal-signal monitors on October 8, making probes available to Baseten customers as a lower-cost way to flag risky agent activity for further review. In Goodfire’s Kimi K3 tests, monitoring cost about $51 across 1,500 sessions, but the results are specific to those tests: the probes caught 94% of malicious hacking sessions, and alerts only trigger the customer’s chosen response. The approach’s broader reliability—especially on long or unfamiliar inputs—remains an open deployment question.
Goodfire compared its roughly $51 monitoring cost with $233 for a cheaper model reviewing every step and about $10,000 for a top-tier reviewer.
The probes sent 8.7% of harmless hacking sessions for a second look; that escalation rate is not an automatic-block rate.
Running four probes together added less than 2% to the time before the model began responding, not to total session time.
Baseten customers can now check an AI agent’s internal signals for risky activity, reserving closer AI review for steps that trigger an alert. Goodfire launched the monitors on October 8, offering an alternative to having another model reread every step. The company says its Kimi K3 tests showed substantially lower monitoring costs.
The monitors use small detectors called probes to inspect signals inside the model at each step of an agent’s work. A probe’s alert can prompt a separate AI model to take a closer look. That creates two stages: a lightweight check throughout the session, followed by more detailed scrutiny of flagged activity.
Goodfire CEO Eric Ho attributes the low running cost to reusing calculations the model already performs while producing its response. Rather than making a second model process all the same material, the probes classify intermediate internal signals. The proposed saving comes from reading existing computations, not duplicating the agent’s entire workload.
Customers choose which risks to watch, including offensive hacking, chemical and biological weapons misuse, and reward hacking. They also choose how the system responds. An alert does not have to mean an automatic refusal; the available responses span recordkeeping, human intervention and blocking.
Goodfire reported this monitoring cost across about 1,500 Kimi K3 sessions in its tests.
A cheaper AI model checking every step cost $233 in the company’s comparison.
A top-tier model checking every step cost about $10,000 in the same comparison.
Goodfire says the probes caught 94% of malicious hacking sessions and sent 8.7% of harmless sessions for a second look. That second figure describes escalation, not automatic rejection. Whether a flagged request is blocked depends on the customer’s chosen response. The detection result concerns hacking sessions, not every risk category customers can select.
The company also says running four probes together added less than 2% to the time the model took to start responding. That measures the delay before a response begins, rather than the time required to finish a whole agent session.
Goodfire is not the first to deploy this approach. A Google DeepMind paper submitted in January says its research informed misuse-detection probes deployed in user-facing Gemini systems. It also identifies a central limitation: probes can fail when production inputs differ from their training examples, especially when moving from short inputs to long ones.
Those researchers found that handling longer inputs was not enough by itself. Broad reliability also required suitable probe designs and training on varied examples. Their results support the idea of pairing efficient probes with a second classifier, but they do not validate Goodfire’s product. They show why low cost and strong results on one test are only part of the deployment question.
Loading discussion...
Join the conversation
Explain when preventing harm outweighs the disruption of a false alarm.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.