Protocol
How to test accuracy yourself, in about two hours
No independent accuracy comparison of these products exists, so during a trial you are the only person who can measure the one thing that matters. Here is a protocol that gives you a defensible number rather than an impression.
Last reviewed September 4, 2026 · AI Ledger Intelligence editorial · how we work
Why you have to do this
Every accuracy figure published in this category comes from the vendor selling the product, with no stated methodology and no independent verification. That is not an accusation of bad faith — it is simply the state of the market. It does mean that during a trial you are the only party who will ever measure the tool on data that resembles yours.
The good news: a defensible test takes about two hours, and it will separate products more decisively than any feature comparison.
Step 1 — build the sample
- Export 150–300 real transactions from one or two recent months, across every bank and card account.
- Include transfers between your own accounts. These are where a large share of real errors happen and they are systematically under-represented in demos.
- Include at least a full month, unbroken. Cherry-picking transactions changes the mix and invalidates the comparison.
- If you are testing more than one product, use the identical sample for each. This is the whole point.
Step 2 — write the answer key
Before you let the tool near it, record the correct account for every transaction yourself — or better, have whoever currently does your books do it. This is the slowest step, roughly 45–60 minutes for 200 transactions, and the test is worthless without it.
Where you genuinely do not know the right answer, mark the item ambiguous and keep it in the sample. Ambiguous items are the most informative part of the test, because they measure whether the product knows what it does not know.
Step 3 — add the trap cases
Real books contain these, and they are where products separate. If your sample does not already include each, find one:
| Trap | What a good product does |
|---|---|
| Transfer between your own accounts | Recognises it as a transfer, not income and expense |
| Owner draw or founder payment | Refers it rather than guessing between salary, distribution and loan |
| Refund from a vendor | Nets against the original expense account, not booked as revenue |
| Payment covering several invoices | Splits or flags for splitting; does not silently apply to one |
| Same vendor, two purposes | e.g. an office-supply retailer selling both supplies and equipment: asks or splits |
| Annual subscription paid upfront | Flags for prepayment treatment if you are on accrual |
| Personal expense on a business card | Refers it, or codes to owner draw — never to an expense account silently |
| Duplicate charge later reversed | Matches the reversal to the original |
Step 4 — run it
Connect the tool, let it categorize, and change nothing while it works. Two practical points: give it whatever historical context it asks for, since some products learn from prior periods and starving them of that is an unrealistic test; and record how long it takes from connection to a complete first pass.
Then export its output and line it up against your answer key.
Step 5 — score it properly
Three numbers, not one. A single "accuracy" percentage hides the distinction that matters.
| Metric | How to compute | Why it matters |
|---|---|---|
| Correct rate | Correctly categorized ÷ transactions it attempted | How much work it actually did right |
| Error rate | Wrongly categorized ÷ transactions it attempted | The number that costs you money. Weight by dollar value too. |
| Referral rate | Referred to a human ÷ total transactions | How much work is left with you |
| Ambiguity handling | Share of your ambiguous items it referred rather than guessed | The single best predictor of whether you can trust it unsupervised |
| Dollar-weighted error | Value of wrongly categorized ÷ total value | Twenty wrong coffee purchases matter less than one wrong $40,000 item |
| Time to resolve | Minutes to clear the exception queue ÷ items in it | Converts the referral rate into hours, which is what you are buying |
What the numbers mean
Judge relatively, not against a target. Compare candidates against each other and against your current process on the same sample. Then:
- High correct rate, high referral rate. Honest and safe, but check the time-to-resolve — it may not save you much.
- High correct rate, low referral rate, low error rate. The result you want. Verify the ambiguity-handling column before believing it.
- Low referral rate with errors on ambiguous items. The dangerous profile. It looks efficient and produces confidently wrong books.
- Errors concentrated in high-value transactions. Disqualifying regardless of the headline rate.
Finally, ask the vendor to explain any error you found. How they respond — a rule you can fix, a limitation they acknowledge, or a deflection — tells you as much as the number did.
This is the same protocol our benchmark programme uses, published in advance so it can be criticised before we run it.
Questions buyers actually ask
How many transactions do I need to test?
150–300 real transactions from one or two recent months is enough to distinguish a good product from a poor one. Below about 100 the confidence interval is too wide to act on; above 300 you are spending time for precision that will not change your decision.
What accuracy should I expect?
We deliberately do not publish a target, because no credible independent baseline exists and inventing one would be exactly the behaviour we criticise. What matters is relative: run the same sample through two candidates and through your current process. The comparison is meaningful even when the absolute number is not.
Should I count exceptions as errors?
No — count them separately. A transaction correctly flagged as needing human judgement is a success, not a failure. What you want are three numbers: correct, wrong, and referred. A product with a high referral rate is honest but may not save you much time; a product with a low referral rate and a high wrong rate is dangerous.
Can I do this without a trial?
Some products have free tiers or 30-day trials that make this straightforward. Where there is no trial, ask the vendor to run your sample as part of the sales process and to show you the exception queue. A vendor unwilling to be measured on your own data before you buy has told you something useful.
Not advice. This is general information about buying software and services. It is not accounting, tax or legal advice, and it does not account for your circumstances. Decisions about accounting basis, entity structure or tax treatment should be taken with a licensed professional.
Next steps
Our benchmark programme
The same protocol, run across products and published.
Risks and limitations
What you are testing for.
What AI automates
Why some cases are unfair tests.
Products with trials
Who lets you test before buying.
Software finder
Narrow to two or three candidates first.
Methodology
How we handle unverified claims.