Loading Actuals…
Actuals — the independent meter for AI-delivered work
actuals Independent meter
The independent meter for AI-delivered work

Vendors report claimed.
We report actuals.

Every AI vendor counts its work delivered the moment it closes it. Actuals takes that work out of your own systems, follows each claim through its window, and counts it delivered only if it holds — measured against what your own people achieve on the same work.

Support is the first instance: a ticket counts resolved only if the same customer does not come back with the same issue — on any channel, in any thread — for 72 hours. The vendor figure is a floor, not a measurement.

Walk the sample audit
Calibrated simulation Figures shown here are simulated, not a live client — "Solvo AI" is a fictional vendor, invented for this walkthrough. Your number is different.
Solvo AI · weekly summary {{ phaseLabel }}
Claimed resolved by AI · 78.2%
{{ figure }}%
{{ figCaption }}
{{ deltaText }}
{{ deltaLabel }}
DurableCame back
Same-issue returns surfacing
{{ r.pair }} {{ r.clue }}
{{ r.gap }}
{{ footNote }}
The new-thread problem · in support

Same issue, new ticket — invisible to reopen counters.

A customer whose refund never arrives does not reopen the old ticket. They start a fresh one, often on a different channel, often with a different subject line. The vendor's reopen rate stays clean. The work happened twice anyway.

We match on what the two tickets have in common: an order ID, an error code, an account, an invoice, an amount. Every match is published as an evidence row you can inspect and overrule.

64%
of returns arrive as a brand-new ticket
2.1×
the human return rate, on the same queues
ZD-418822 Closed by AI
“Refund for order FB-8841-2207 has been approved and will appear in 3–5 days.”
31 HOURS LATER · NEW THREAD · NO REOPEN RECORDED
ZD-419540 Same issue
“Still no refund on FB-8841-2207 — this is my second time asking.”
Shared clue: order ID. The vendor counted this as one resolution and one new contact. It is one failure, billed twice.
The method

The rules, before the numbers.

A measurement is only worth what its rules are worth, so the rules are published first and they do not move for a customer. Every one of them is written down in the spec this product is built from, and the report prints them next to the finding.

01

Ambiguity resolves in the vendor's favour.

same_issue  →  the claim failed
new_issue   →  the claim held
unclear    →  the claim held

MATCH-V1 returns one of three verdicts on every candidate pair. When it cannot decide, the verdict is unclear, and an unclear is recorded as a new issue — not as a return. A return we cannot prove is a return the vendor keeps.

The rule runs one way only. We undercount ourselves before we overcount the vendor, so the durable rate we publish is the kindest reading of your data the evidence allows. A finding that survives its own handicap is a finding a vendor has to answer.

02

Matching runs on shared hard identifiers.

Order numbers, invoice numbers, error codes, product SKUs and tracking numbers are pulled out of ticket text by deterministic pattern rules — no model involved — and stored against the ticket. The patterns are configured per tenant and printed in the report's method section, because a printed rule is one you can argue with. A candidate that shares a hard identifier with the claim is never dropped from the shortlist.

Same-issue detection therefore works across threads and across channels. It never reads the vendor's own thread IDs or reopen counters, which is the whole reason it can see the returns those counters miss.

03

The human baseline is measured by the same instrument.

Your own people's tickets run through identical machinery: the same period, the same contact reasons, the same t0, the same window, the same truth events. Nothing about the control group is measured more gently than the AI is.

Only the excess above that human baseline is attributed to the AI. Humans produce repeat contacts too; those are not the vendor's to carry, and the report never hands them over.

04

Gates before numbers.

Three gates stand between the pipeline and any report. The thresholds are floors: they never lower, and the labelled sets only grow. A red gate blocks the release of any report.

Gate Threshold · standing Measured 2026-08-09 · labelled eval set
Match precision ≥ 90% 96.1% n=215
Requester-join recall ≥ 85% 100%
Resolver-attribution disagreement ≤ 5% 3.2%

These values are measured on the labelled evaluation sets — seeded-sandbox ground truth, synthetic by construction, plus hand-written pairs — not on customer data. The first real dataset re-verifies every gate with the customer in the loop (M6). The measured column is an observation on a dated run, not a property of the product. Every prompt or model change re-runs the whole labelled set and appends a new row to the log; the log, not this page, is the record.

05

The first report ships only after you have checked us.

Fifty matched pairs are read through with you before your first report is final — your data, your judgement, one call. Where you overrule us, the correction is ingested as labelled data under your name as the labeller, the gates are re-run on the grown set, and the result is appended to the accuracy log whatever it says.

Every match row in the finished report carries the same affordance. The report is not only evidence out; it is labelled data in.

06

Named limitations, before anyone asks.

These print in every report, verbatim, next to the finding they qualify.

Silent quit

A customer who gives up files nothing to say so. Some share of what we count as durable is a customer who churned or solved it themselves. Direction: pushes the durable rate up — kinder to the vendor.

Channel leakage

A customer who returns by phone is matched only where the phone record carries a shared clue. Where telephony does not write back to the helpdesk, those returns are invisible to us. Direction: the durable rate we report is a ceiling — kinder to the vendor.

Signpost attribution

Where the AI resolves by sending the customer elsewhere, the real work happens off the thread. If the redirected contact lands back in the helpdesk we count it; if it lands anywhere else we cannot see it. Direction: the claimed rate keeps credit for work it signposted away — kinder to the vendor.

Every limitation on this list moves the number the same way: it makes our reading kinder to the vendor, never harsher. That is by construction, and it is why a finding here is a floor.

The public record

Published in full. Checkable by no one.

Three claims from the public record, quoted as published, each with its source. We are not disputing any of them. We could not — and neither can you, and that is the point of the section.

27 February 2024
Klarna, company newsroom
klarna.com — press release ↗

Klarna reported that its AI assistant, one month after launch, had handled 2.3 million conversations — two-thirds of the company's customer service chats — and was “doing the equivalent work of 700 full-time agents”. The same release reported the assistant “more accurate in errand resolution, leading to a 25% drop in repeat inquiries”, customer satisfaction “on par with human agents”, and an estimated $40 million USD profit improvement in 2024.

The company later said it was recruiting human customer service agents again, in what its chief executive described as an Uber-type setup. Sebastian Siemiatkowski, quoted by Bloomberg:

“As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality.”
“Really investing in the quality of the human support is the way of the future for us.”

CNBC reported the same week that the company had shrunk its workforce by about 40%, which its chief executive attributed in part to its investments in AI and in part to natural attrition after a hiring freeze.

Retrieved 11 August 2026
Fin by Intercom, product marketing
fin.ai ↗

The vendor publishes a resolution rate on its own front page: “Fin has industry-leading resolution rates, averaging 76% across 12,000+ customers, with many seeing over 85%”, alongside “2 million weekly resolutions”.

A marketing page is a live document; this is what it said on the date above.

What none of them contains

A resolution rate is a count of closures. Published on its own it cannot tell you what share of those closures came back — under a new subject line, on another channel, two days later — or what the same company's own people achieved on the same work in the same period.

One figure above does speak to returns: the 25% drop in repeat inquiries. What the record does not say is how a repeat inquiry was identified — whether a customer who came back on a fresh thread, or through a different channel, was counted as one at all.

Nothing here is hidden. It is simply not measured in public, by anyone, in either direction. It is measurable — from the buyer's own data, on the buyer's own instrument. That gap is the entire reason this exists.

Next: the same arithmetic on your volume, then a sample audit you can walk end to end.
Your number

One figure you already have.

What you already know
Costed at the sample's own entered cost per contact ({{ estCpcLabel }}) and its vendor's published rate ({{ estClaimedLabel }}).
What the benchmark implies
Durable rate at 72 hours
{{ estDurablePct }}%
{{ estDeltaPts }} against the rate you are shown
HeldCame back
Annual exposure
{{ estExposureRange }}
Excess repeat contacts{{ estExcessContacts }}
Resolutions that did not hold{{ estNonDurable }}
Find out what it actually is →

This is not your number. It applies the gap measured in our calibrated sample — a durable rate {{ estRatioPct }} of the claimed rate, and a vendor fee of {{ estFeeLabel }} on every resolution billed — to the volume you entered. Your queue mix moves it by more than your vendor's marketing does. The only way to get your figure is to measure it.

Published audits
Nothing here yet, and we will not invent it.
Slot 01 · reserved
First cohort audit, publishable with client permission
Headline pair, method, and named limitations. Verbatim or not at all.
Slot 02 · reserved
Vendor response, printed unedited
Where a vendor disputes a finding, their reply runs next to it.
Slot 03 · reserved
Method review by an outside statistician
Commissioned, not solicited. Published whatever it says.

You can't coach a bot.
You can measure it.

Every headline number in this product anchors to evidence rows you can open, inspect, and dispute. That is the whole product.

Walk the sample audit Start your subscription

Measuring AI on code, not support? The same meter is in pilot — see the code tab of the sample audit.