Guide

Risks and limitations worth pricing in

This category is genuinely useful and genuinely oversold. These are the failure modes we would want a buyer to understand before signing anything — each with the specific mitigation, not a general warning.

Last reviewed September 4, 2026 · AI Ledger Intelligence editorial · how we work

Errors are systematic, not random

This is the most important thing to understand and the least discussed. When a human bookkeeper makes a mistake, it is usually one transaction. When a rule-learning system makes a mistake, it applies the same mistake to every transaction matching that pattern — for as long as the pattern holds.

A software subscription miscoded as cost of goods sold rather than operating expense does not look wrong in a transaction list. It looks wrong in a gross margin that has drifted three points, which nobody notices for a quarter.

Mitigation: review account balances month over month, not transaction lists. A five-minute comparison of this month's P&L by account against the last three months catches a whole miscategorised class at once. Spot-checking twenty transactions catches almost nothing.

Confident answers to ambiguous questions

Some transactions have no correct answer derivable from the data. A $12,000 transfer to the founder could be salary, a distribution, a loan repayment or an expense reimbursement. They are identical in the bank feed and have entirely different tax treatments.

A well-built product routes these to a human. A poorly built one picks the most likely answer and moves on, which is worse than leaving it uncategorised, because a wrong classification looks like completed work.

Mitigation: in a trial, deliberately introduce ambiguous transactions and see what the tool does. Agree with your accountant a list of categories the automation is not permitted to decide alone — anything touching equity, intercompany, owner compensation, or tax classification — and check the tool can enforce that.

Accuracy claims are unverifiable

Every accuracy figure in this category comes from the vendor. “Up to 98% auto-categorization.” “Automates ~85%.” “95% of transactions.” None is accompanied by a published test set, a defined error taxonomy, or third-party verification. We could not find a single independent, reproducible accuracy comparison of these products.

These numbers are also easy to state truthfully while being misleading. If exceptions are excluded from the denominator, automation rate measures only the easy transactions. If a wrongly categorized transaction that nobody caught counts as automated, the rate measures throughput rather than correctness.

Mitigation: ask three questions of any automation figure — what is the denominator, do exceptions count as failures, and does a wrong-but-unreviewed categorization count as automated? Then run your own test during the trial. This is also why we are building an independent benchmark.

Vendor continuity

This is a young, heavily funded category, and funding is not the same as durability. The concrete precedent: Bench, an established provider with tens of thousands of customers, ceased operations without notice on 27 December 2024. It was acquired days later by Employer.com; the brand now appears as part of Mainstreet. Customers were locked out of their books in between.

The exposure is not uniform. If your books live in your own QuickBooks or Xero file, a vendor failure costs you a provider. If the vendor holds the ledger, it costs you access to your accounting records.

Mitigation: before signing with any vendor that holds your ledger, get written answers to: what export formats are available, does the export include full transaction-level detail and attachments, how quickly can it be produced, and what happens to data access if the company ceases trading. Then export once, early, to confirm the answer is real. Treat multi-year prepayment to a small vendor as a credit decision.

Data portability and lock-in

A trial balance is not a data export. Moving books between systems needs transaction-level detail with dates, accounts, memos, attachments and the reconciliation state. Products that offer a summary export are offering you a report, not your data.

Lock-in also comes in a softer form: the rules, categorizations and learned behaviour your team has built up over a year do not transfer. That is a real switching cost even where the raw data moves cleanly.

Mitigation: run one export during the trial, not at the end of the relationship, and confirm it imports into a standard accounting system. If it does not, you have learned something important while it still costs you nothing.

The review that never happens

Every software-only product assumes somebody works the exception queue. The most common failure in practice is not a model error at all — it is that nobody does. The books look up to date because transactions are categorized, and the queue quietly grows.

Mitigation: put a recurring calendar block on it, in someone's name, with a defined output — queue cleared, balances compared to prior months, anything unresolved escalated. If you cannot name the person and the date, you are buying a service whether or not you have paid for one. See AI + human services.

A pre-purchase checklist

  1. Who works the exception queue, by name, and when?
  2. Which categories is the automation not allowed to decide alone, and can the product enforce that?
  3. What is the audit trail for an automated entry, and can a reviewer reverse it in one action?
  4. What export formats exist, do they include attachments, and how fast can you get one?
  5. Is our data used for model training, and can we opt out? Get it in writing.
  6. What security certifications does the vendor hold? (Some publish SOC 2; many say nothing.)
  7. If we cancel mid-term, what happens to data access, and what is the notice period?
  8. What does the vendor's automation figure actually measure — and what is the denominator?

Questions buyers actually ask

What is the biggest risk of AI bookkeeping?

That errors are systematic rather than random. A human miscodes one transaction; a rule-learning system miscodes every transaction matching a pattern, silently, for months. The mitigation is not spot-checking transactions — it is reviewing account balances month over month, which surfaces a whole miscategorised class at once.

Can I trust AI-generated books for a tax return or a loan?

Only after a competent human has reviewed and signed off. That is true of manually prepared books too, but the failure mode is different: automated books tend to be internally consistent and confidently wrong, which is harder to spot than the obvious errors humans make. If a lender, investor or auditor is going to rely on the numbers, budget for professional review.

What happens if my AI bookkeeping vendor shuts down?

It depends entirely on where your ledger lives. If the vendor works inside your QuickBooks or Xero file, a shutdown is an inconvenience. If the vendor holds the ledger, it is a data-recovery problem — and there is precedent: Bench ceased operations without notice in December 2024 and customers lost access before an acquisition completed days later.

Is my financial data used to train AI models?

Vendors vary and most do not address it prominently. Ask directly, get the answer in writing, and ask specifically whether data is used for training, whether it is de-identified, and whether you can opt out. This is a standard question in any software procurement and there is no reason to accept a vague answer.

Not advice. This is general information about buying software and services. It is not accounting, tax or legal advice, and it does not account for your circumstances. Decisions about accounting basis, entity structure or tax treatment should be taken with a licensed professional.

Next steps