Programme
The AI Ledger Benchmark
Every accuracy claim in this category comes from the vendor making it. None is accompanied by a test set, an error taxonomy or independent verification. We intend to fix that, and this page publishes the method before the results.
Status: not yet run
No results have been published, and no product on this site carries a benchmark score. We are publishing the protocol first so it can be criticised while it can still change. If you think the design below is wrong, say so — that is the point of putting it up early.
Why this is the gap worth filling
The whole category rests on an unmeasured claim. Vendors say “up to 98% auto-categorization” or “automates ~85%”. Those figures may well be accurate. They are also unfalsifiable as published: no shared test set, no definition of what counts as a transaction, no statement of whether exceptions are in the denominator, no independent audit.
Meanwhile a buyer choosing between a $129/month layer and a $494/month service is making a decision that hinges entirely on which one produces books that are right. There is currently no evidence available to them that is not marketing.
Test set design
The value of a benchmark is almost entirely in the test set. Ours will be:
- Synthetic but realistic. Generated from real small-business transaction patterns rather than taken from a real company's books, so it can be published in full alongside the results. A benchmark you cannot inspect is not a benchmark.
- 500 transactions across 12 months, for a single-entity services business, with a second variant for a product business including inventory, COGS and marketplace payout splitting.
- Independently labelled. Ground truth set by a licensed accountant, not by us, with disagreements resolved before the test set is frozen.
- Deliberately hard in the right places. Roughly 15% of transactions will be trap cases: intercompany transfers, owner draws, refunds netting against expense, split payments, same-vendor-different-purpose, prepaid annual subscriptions, personal spend on a business card, duplicate-and-reversal pairs.
- Frozen and versioned. Published after the first round so anyone can reproduce or dispute the result. Subsequent rounds use a new set, because a public test set stops measuring generalisation the moment it is public.
What gets measured
| Metric | Definition | Why it is in the set |
|---|---|---|
| Categorization correct rate | Correct ÷ attempted | The headline number, reported but not led with |
| Error rate | Incorrect ÷ attempted | What actually costs money |
| Dollar-weighted error | Value incorrectly categorized ÷ total value | One wrong $40,000 item outweighs twenty wrong coffees |
| Referral rate | Referred to a human ÷ total | How much work remains with you |
| Ambiguity handling | Share of designed-ambiguous items referred rather than guessed | The best single predictor of trustworthiness |
| Reconciliation accuracy | Correct matches ÷ matchable items | The second core task |
| Intervention time | Minutes to clear the exception queue | Converts referral rate into the hours you are buying back |
| Time to first pass | Elapsed hours from connection to complete draft | Matters for close deadlines |
Results will be led by error rate and ambiguity handling, not by correct rate. A product that refers every hard case and gets the rest right is more useful than one with a higher headline number and confident mistakes.
Error taxonomy
“Wrong” is not one thing. Errors will be classified as:
- Category error — right transaction, wrong account, no tax consequence
- Consequential error — wrong account with a tax or reporting consequence (capitalisation, deductibility, owner compensation)
- Structural error — a transfer treated as income and expense, a refund booked as revenue, a duplicate not reversed
- Fabrication — a vendor, amount, date or document reference that does not exist in the source data
- Silent systematic error — a single wrong rule applied to a class of transactions
The last two are reported separately and prominently, because they are the failure modes a monthly review is least likely to catch.
Vendor policy
- Products are bought at list price on ordinary commercial terms. No free access, no sponsored participation.
- Each vendor is given its own results before publication and a right of reply, published alongside.
- No vendor sees another's results before publication, and none can alter the test set.
- Declining to participate is published as declining to participate.
- No commercial relationship affects inclusion, method or presentation. Products we may later earn commission on are tested and reported identically.
Likely first round
Practical constraint: a product needs a free tier or a trial for a first round to be affordable. 3 of the 21 products we track currently offer one.
- Digits — 30-day trial
- Kick — free tier
- QuickBooks Online — 30-day trial
Products without a trial will be included where budget allows, and their absence noted rather than glossed over.
How results will appear
This is already built. Every product record carries a benchmark object with a status field, currently not-tested for all 21. When a round completes, results attach to the record and appear automatically on the product page, in category comparison tables, in the software finder, and in structured data — with the round identifier and date. No page will need rewriting.
Until then, every page that touches accuracy says plainly that it has not been independently measured. That is the honest position, and it is the one this whole site is built on.
Questions
Has the AI Ledger Benchmark published results yet?
No. No product has been tested yet. This page publishes the protocol before the results deliberately, so the method can be criticised while it can still be changed. Any page on this site that mentions a product's accuracy says explicitly that it has not been independently measured.
Why does an independent benchmark not already exist?
Because it is expensive and awkward. It needs a realistic labelled transaction set, paid subscriptions to every product, weeks of setup, and a willingness to publish findings that will annoy vendors you might later want a commercial relationship with. Those incentives explain why every accuracy figure in this category comes from the company selling the product.
Will vendors be able to influence the results?
They will be shown their own results before publication and given a right of reply, which will be published alongside. They will not see other vendors' results, will not be able to alter the test set, and no commercial relationship will affect participation or presentation. A vendor that declines to participate will be listed as having declined.
What will you actually measure?
Categorization correctness and error rate, dollar-weighted error, referral rate, how ambiguous items are handled, reconciliation accuracy, exception-queue resolution time, and total elapsed time to a first complete pass. Error rate and ambiguity handling matter more than headline accuracy, and the results will be presented in that order.
Where will the data live?
Product records on this site already carry a benchmark field with a status of "not tested". When a round completes, results attach to each product record and appear automatically on that product's page, in category comparison tables, and in the software finder. The data model was built for this from the start.
Want to help or object? If you are an accountant willing to review the labelled test set, a vendor who wants to participate, or someone who thinks the design above is flawed, get in touch. Criticism now is more useful than criticism after publication.