Programme

The AI Ledger Benchmark

Every accuracy claim in this category comes from the vendor making it. None is accompanied by a test set, an error taxonomy or independent verification. We intend to fix that, and this page publishes the method before the results.

Status: not yet run

No results have been published, and no product on this site carries a benchmark score. We are publishing the protocol first so it can be criticised while it can still change. If you think the design below is wrong, say so — that is the point of putting it up early.

Why this is the gap worth filling

The whole category rests on an unmeasured claim. Vendors say “up to 98% auto-categorization” or “automates ~85%”. Those figures may well be accurate. They are also unfalsifiable as published: no shared test set, no definition of what counts as a transaction, no statement of whether exceptions are in the denominator, no independent audit.

Meanwhile a buyer choosing between a $129/month layer and a $494/month service is making a decision that hinges entirely on which one produces books that are right. There is currently no evidence available to them that is not marketing.

Test set design

The value of a benchmark is almost entirely in the test set. Ours will be:

What gets measured

MetricDefinitionWhy it is in the set
Categorization correct rateCorrect ÷ attemptedThe headline number, reported but not led with
Error rateIncorrect ÷ attemptedWhat actually costs money
Dollar-weighted errorValue incorrectly categorized ÷ total valueOne wrong $40,000 item outweighs twenty wrong coffees
Referral rateReferred to a human ÷ totalHow much work remains with you
Ambiguity handlingShare of designed-ambiguous items referred rather than guessedThe best single predictor of trustworthiness
Reconciliation accuracyCorrect matches ÷ matchable itemsThe second core task
Intervention timeMinutes to clear the exception queueConverts referral rate into the hours you are buying back
Time to first passElapsed hours from connection to complete draftMatters for close deadlines

Results will be led by error rate and ambiguity handling, not by correct rate. A product that refers every hard case and gets the rest right is more useful than one with a higher headline number and confident mistakes.

Error taxonomy

“Wrong” is not one thing. Errors will be classified as:

The last two are reported separately and prominently, because they are the failure modes a monthly review is least likely to catch.

Vendor policy

Likely first round

Practical constraint: a product needs a free tier or a trial for a first round to be affordable. 3 of the 21 products we track currently offer one.

Products without a trial will be included where budget allows, and their absence noted rather than glossed over.

How results will appear

This is already built. Every product record carries a benchmark object with a status field, currently not-tested for all 21. When a round completes, results attach to the record and appear automatically on the product page, in category comparison tables, in the software finder, and in structured data — with the round identifier and date. No page will need rewriting.

Until then, every page that touches accuracy says plainly that it has not been independently measured. That is the honest position, and it is the one this whole site is built on.

Questions

Has the AI Ledger Benchmark published results yet?

No. No product has been tested yet. This page publishes the protocol before the results deliberately, so the method can be criticised while it can still be changed. Any page on this site that mentions a product's accuracy says explicitly that it has not been independently measured.

Why does an independent benchmark not already exist?

Because it is expensive and awkward. It needs a realistic labelled transaction set, paid subscriptions to every product, weeks of setup, and a willingness to publish findings that will annoy vendors you might later want a commercial relationship with. Those incentives explain why every accuracy figure in this category comes from the company selling the product.

Will vendors be able to influence the results?

They will be shown their own results before publication and given a right of reply, which will be published alongside. They will not see other vendors' results, will not be able to alter the test set, and no commercial relationship will affect participation or presentation. A vendor that declines to participate will be listed as having declined.

What will you actually measure?

Categorization correctness and error rate, dollar-weighted error, referral rate, how ambiguous items are handled, reconciliation accuracy, exception-queue resolution time, and total elapsed time to a first complete pass. Error rate and ambiguity handling matter more than headline accuracy, and the results will be presented in that order.

Where will the data live?

Product records on this site already carry a benchmark field with a status of "not tested". When a round completes, results attach to each product record and appear automatically on that product's page, in category comparison tables, and in the software finder. The data model was built for this from the start.

Want to help or object? If you are an accountant willing to review the labelled test set, a vendor who wants to participate, or someone who thinks the design above is flawed, get in touch. Criticism now is more useful than criticism after publication.

Next steps