Every AI vendor counts its work delivered the moment it closes it. Actuals takes that work out of your own systems, follows each claim through its window, and counts it delivered only if it holds — measured against what your own people achieve on the same work.
Support is the first instance: a ticket counts resolved only if the same customer does not come back with the same issue — on any channel, in any thread — for 72 hours. The vendor figure is a floor, not a measurement.
A customer whose refund never arrives does not reopen the old ticket. They start a fresh one, often on a different channel, often with a different subject line. The vendor's reopen rate stays clean. The work happened twice anyway.
We match on what the two tickets have in common: an order ID, an error code, an account, an invoice, an amount. Every match is published as an evidence row you can inspect and overrule.
A measurement is only worth what its rules are worth, so the rules are published first and they do not move for a customer. Every one of them is written down in the spec this product is built from, and the report prints them next to the finding.
MATCH-V1 returns one of three verdicts on every candidate pair. When it cannot decide, the verdict is unclear, and an unclear is recorded as a new issue — not as a return. A return we cannot prove is a return the vendor keeps.
The rule runs one way only. We undercount ourselves before we overcount the vendor, so the durable rate we publish is the kindest reading of your data the evidence allows. A finding that survives its own handicap is a finding a vendor has to answer.
Order numbers, invoice numbers, error codes, product SKUs and tracking numbers are pulled out of ticket text by deterministic pattern rules — no model involved — and stored against the ticket. The patterns are configured per tenant and printed in the report's method section, because a printed rule is one you can argue with. A candidate that shares a hard identifier with the claim is never dropped from the shortlist.
Same-issue detection therefore works across threads and across channels. It never reads the vendor's own thread IDs or reopen counters, which is the whole reason it can see the returns those counters miss.
Your own people's tickets run through identical machinery: the same period, the same contact reasons, the same t0, the same window, the same truth events. Nothing about the control group is measured more gently than the AI is.
Only the excess above that human baseline is attributed to the AI. Humans produce repeat contacts too; those are not the vendor's to carry, and the report never hands them over.
Three gates stand between the pipeline and any report. The thresholds are floors: they never lower, and the labelled sets only grow. A red gate blocks the release of any report.
| Gate | Threshold · standing | Measured 2026-08-09 · labelled eval set |
|---|---|---|
| Match precision | ≥ 90% | 96.1% n=215 |
| Requester-join recall | ≥ 85% | 100% |
| Resolver-attribution disagreement | ≤ 5% | 3.2% |
These values are measured on the labelled evaluation sets — seeded-sandbox ground truth, synthetic by construction, plus hand-written pairs — not on customer data. The first real dataset re-verifies every gate with the customer in the loop (M6). The measured column is an observation on a dated run, not a property of the product. Every prompt or model change re-runs the whole labelled set and appends a new row to the log; the log, not this page, is the record.
Fifty matched pairs are read through with you before your first report is final — your data, your judgement, one call. Where you overrule us, the correction is ingested as labelled data under your name as the labeller, the gates are re-run on the grown set, and the result is appended to the accuracy log whatever it says.
Every match row in the finished report carries the same affordance. The report is not only evidence out; it is labelled data in.
These print in every report, verbatim, next to the finding they qualify.
A customer who gives up files nothing to say so. Some share of what we count as durable is a customer who churned or solved it themselves. Direction: pushes the durable rate up — kinder to the vendor.
A customer who returns by phone is matched only where the phone record carries a shared clue. Where telephony does not write back to the helpdesk, those returns are invisible to us. Direction: the durable rate we report is a ceiling — kinder to the vendor.
Where the AI resolves by sending the customer elsewhere, the real work happens off the thread. If the redirected contact lands back in the helpdesk we count it; if it lands anywhere else we cannot see it. Direction: the claimed rate keeps credit for work it signposted away — kinder to the vendor.
Every limitation on this list moves the number the same way: it makes our reading kinder to the vendor, never harsher. That is by construction, and it is why a finding here is a floor.
Three claims from the public record, quoted as published, each with its source. We are not disputing any of them. We could not — and neither can you, and that is the point of the section.
Klarna reported that its AI assistant, one month after launch, had handled 2.3 million conversations — two-thirds of the company's customer service chats — and was “doing the equivalent work of 700 full-time agents”. The same release reported the assistant “more accurate in errand resolution, leading to a 25% drop in repeat inquiries”, customer satisfaction “on par with human agents”, and an estimated $40 million USD profit improvement in 2024.
The company later said it was recruiting human customer service agents again, in what its chief executive described as an Uber-type setup. Sebastian Siemiatkowski, quoted by Bloomberg:
“As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality.”
“Really investing in the quality of the human support is the way of the future for us.”
CNBC reported the same week that the company had shrunk its workforce by about 40%, which its chief executive attributed in part to its investments in AI and in part to natural attrition after a hiring freeze.
The vendor publishes a resolution rate on its own front page: “Fin has industry-leading resolution rates, averaging 76% across 12,000+ customers, with many seeing over 85%”, alongside “2 million weekly resolutions”.
A marketing page is a live document; this is what it said on the date above.
A resolution rate is a count of closures. Published on its own it cannot tell you what share of those closures came back — under a new subject line, on another channel, two days later — or what the same company's own people achieved on the same work in the same period.
One figure above does speak to returns: the 25% drop in repeat inquiries. What the record does not say is how a repeat inquiry was identified — whether a customer who came back on a fresh thread, or through a different channel, was counted as one at all.
Nothing here is hidden. It is simply not measured in public, by anyone, in either direction. It is measurable — from the buyer's own data, on the buyer's own instrument. That gap is the entire reason this exists.
This is not your number. It applies the gap measured in our calibrated sample — a durable rate {{ estRatioPct }} of the claimed rate, and a vendor fee of {{ estFeeLabel }} on every resolution billed — to the volume you entered. Your queue mix moves it by more than your vendor's marketing does. The only way to get your figure is to measure it.
Every headline number in this product anchors to evidence rows you can open, inspect, and dispute. That is the whole product.
Measuring AI on code, not support? The same meter is in pilot — see the code tab of the sample audit.