Protocol

How to test accuracy yourself, in about two hours

No independent accuracy comparison of these products exists, so during a trial you are the only person who can measure the one thing that matters. Here is a protocol that gives you a defensible number rather than an impression.

Last reviewed September 4, 2026 · AI Ledger Intelligence editorial · how we work

Why you have to do this

Every accuracy figure published in this category comes from the vendor selling the product, with no stated methodology and no independent verification. That is not an accusation of bad faith — it is simply the state of the market. It does mean that during a trial you are the only party who will ever measure the tool on data that resembles yours.

The good news: a defensible test takes about two hours, and it will separate products more decisively than any feature comparison.

Step 1 — build the sample

Step 2 — write the answer key

Before you let the tool near it, record the correct account for every transaction yourself — or better, have whoever currently does your books do it. This is the slowest step, roughly 45–60 minutes for 200 transactions, and the test is worthless without it.

Where you genuinely do not know the right answer, mark the item ambiguous and keep it in the sample. Ambiguous items are the most informative part of the test, because they measure whether the product knows what it does not know.

Step 3 — add the trap cases

Real books contain these, and they are where products separate. If your sample does not already include each, find one:

TrapWhat a good product does
Transfer between your own accountsRecognises it as a transfer, not income and expense
Owner draw or founder paymentRefers it rather than guessing between salary, distribution and loan
Refund from a vendorNets against the original expense account, not booked as revenue
Payment covering several invoicesSplits or flags for splitting; does not silently apply to one
Same vendor, two purposese.g. an office-supply retailer selling both supplies and equipment: asks or splits
Annual subscription paid upfrontFlags for prepayment treatment if you are on accrual
Personal expense on a business cardRefers it, or codes to owner draw — never to an expense account silently
Duplicate charge later reversedMatches the reversal to the original

Step 4 — run it

Connect the tool, let it categorize, and change nothing while it works. Two practical points: give it whatever historical context it asks for, since some products learn from prior periods and starving them of that is an unrealistic test; and record how long it takes from connection to a complete first pass.

Then export its output and line it up against your answer key.

Step 5 — score it properly

Three numbers, not one. A single "accuracy" percentage hides the distinction that matters.

MetricHow to computeWhy it matters
Correct rateCorrectly categorized ÷ transactions it attemptedHow much work it actually did right
Error rateWrongly categorized ÷ transactions it attemptedThe number that costs you money. Weight by dollar value too.
Referral rateReferred to a human ÷ total transactionsHow much work is left with you
Ambiguity handlingShare of your ambiguous items it referred rather than guessedThe single best predictor of whether you can trust it unsupervised
Dollar-weighted errorValue of wrongly categorized ÷ total valueTwenty wrong coffee purchases matter less than one wrong $40,000 item
Time to resolveMinutes to clear the exception queue ÷ items in itConverts the referral rate into hours, which is what you are buying

What the numbers mean

Judge relatively, not against a target. Compare candidates against each other and against your current process on the same sample. Then:

Finally, ask the vendor to explain any error you found. How they respond — a rule you can fix, a limitation they acknowledge, or a deflection — tells you as much as the number did.

This is the same protocol our benchmark programme uses, published in advance so it can be criticised before we run it.

Questions buyers actually ask

How many transactions do I need to test?

150–300 real transactions from one or two recent months is enough to distinguish a good product from a poor one. Below about 100 the confidence interval is too wide to act on; above 300 you are spending time for precision that will not change your decision.

What accuracy should I expect?

We deliberately do not publish a target, because no credible independent baseline exists and inventing one would be exactly the behaviour we criticise. What matters is relative: run the same sample through two candidates and through your current process. The comparison is meaningful even when the absolute number is not.

Should I count exceptions as errors?

No — count them separately. A transaction correctly flagged as needing human judgement is a success, not a failure. What you want are three numbers: correct, wrong, and referred. A product with a high referral rate is honest but may not save you much time; a product with a low referral rate and a high wrong rate is dangerous.

Can I do this without a trial?

Some products have free tiers or 30-day trials that make this straightforward. Where there is no trial, ask the vendor to run your sample as part of the sales process and to show you the exception queue. A vendor unwilling to be measured on your own data before you buy has told you something useful.

Not advice. This is general information about buying software and services. It is not accounting, tax or legal advice, and it does not account for your circumstances. Decisions about accounting basis, entity structure or tax treatment should be taken with a licensed professional.

Next steps