Scale AI Details a Human-Routing Rule for Credit-Recovery Agents
The company says its pre-deployment system found more recoverable vendor credits by reserving analyst time for decisions where a mistake would be costly and hard to undo.
Listen to this story
The audio brief
Story brief
3 key pointsScale AI’s case study proposes an economic rule for deciding when an accounts-payable agent should hand work to analysts: escalate when expected loss—estimated error probability multiplied by credit value and severity—exceeds review cost. In one month of pre-deployment testing with an unnamed global media technology company, the system found several times more recoverable vendor credits than the existing process and...
- 01
Testing covered classification, extraction, reconciliation, and ERP posting—not just an isolated model step.
- 02
Higher-value or harder-to-reverse mistakes receive stricter autonomy thresholds than errors likely to be caught downstream.
- 03
Scale’s evidence is limited to one month of pre-deployment testing, not production performance.
Scale AI says an agentic accounts-payable system built with a global media technology company could recover millions of dollars in vendor credits that would otherwise go unclaimed. But its evidence so far is a month of pre-deployment testing: Scale says the system surfaced several times more recoverable credits than the existing process and reached at least 92% end-to-end disposition accuracy.
The more consequential part of Scale’s newly published case study is not the automation itself. It is a proposed rule for deciding when automation should stop. Rather than sending every low-confidence result to a person, Scale proposes review when the estimated chance of an agent mistake, multiplied by the impact of that mistake, exceeds the cost of an analyst’s time.
A workflow built around unclaimed credits
The workflow handles a familiar accounts-payable gap. Vendor statements can contain credits, overpayments, or duplicate charges, yet manual reconciliation limits how many statements a finance team can examine. Scale says its system classifies vendor documents, identifies credit items, obtains credit memos, and posts credits into an enterprise resource planning system.
Straightforward cases are routed toward automated credit posting; other cases go to analysts for deeper review. The reported test result spans the whole process—classification, extraction, and reconciliation—rather than isolating the performance of a single model step.
Confidence alone is not the proposed gate
A fixed confidence cutoff treats a mistake on a small credit much like a mistake on a large one. Scale’s framework instead increases the confidence required for autonomous action as the credit value and potential loss rise. Under a fixed review budget, it ranks cases by expected loss: error probability multiplied by a severity factor and the credit value.
The three inputs to escalation
- Estimated error probability, based on runtime uncertainty signals calibrated against observed outcomes rather than a model’s self-reported certainty.
- Credit value, which determines how much money is at stake in the decision.
- Severity, which accounts for whether an error is likely to be detected and corrected later or may silently leave money unrecovered.
The key assumption is how errors are priced
That severity adjustment is central to the argument. Scale distinguishes a credit wrongly closed without escalation, which may be difficult to discover, from an invalid credit that is posted and later reversed through a future reconciliation. The first can justify review at a lower value because the potential loss is less reversible; the second can tolerate a higher threshold if downstream processes are likely to catch it.
The approach turns human review from a blanket safety layer into a scarce resource assigned to the cases where it is expected to prevent the most loss. It also leaves open a practical question for any deployment: whether the calibrated error estimates and severity assumptions continue to hold once the system moves beyond the reported pre-deployment test.
Sources
- scale.comIn An Agentic World Where Automation Gets Cheap, Which Work Is Worth Routing to a Human?
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.