AI for Accounting Firms
Supplier invoices, receipts and statements read into the ledger, and an answer desk over the firm’s own files and procedures.
The documents this gets built for
Four recurring documents, and what makes each one awkward to handle by hand.
Forty photographs from one client
Taken on a phone across a month, some duplicated, some blurred, one of a receipt that should not be in there. Sorting, deduplicating and rejecting is as much of the build as reading the amounts.
An invoice that runs onto page two
The header values are the easy half. Line items continue overleaf, credits appear as negatives, and the coding depends on which client the invoice belongs to.
A statement that is a picture of a page
Bank and card statements arriving as scans rather than as data, read into dated rows that then have to reconcile against a running balance. In most bookkeeping lines this is where the hours sit.
A tax agency notice, scanned at the front desk
Correspondence read for the client, the period and the amount, and filed against the right client with a task raised. It arrives in volume in the same few weeks every year.
What this is for an accounting firm
The documents an accounting firm processes are produced by other people: supplier invoices, bank statements, the client folder that arrives as forty photographs of receipts, the payroll register, last year’s workpapers. Almost none of it arrives in a form the ledger can take.
The work here points a model at that material and at the firm’s own procedures, so the recurring keying and the recurring questions are handled by the system rather than by staff. It runs in a Microsoft or AWS account the firm owns, and every request is logged against a named user. The client files never leave that account.
None of it learns from client files. Documents are read at the moment they are processed, and a person approves everything that reaches the ledger. A first build takes weeks and needs no data science team.
The document type sets the scope. Invoices, receipts and statements read into fields, with a person approving every record.
What gets built
One piece at a time, starting with whichever costs the most staff hours. Most firms find that is a single document type.
Supplier invoices and receipts into the ledger
Header values, line items, tax and coding read into fields and posted after a person approves the batch. The client’s folder of photographed receipts is the case it gets built for, not the clean PDF.
Bank and card statements as transactions
Statement pages turned into dated rows with amounts and descriptions ready for coding. That includes the statements that arrive as scanned images rather than data, which is where the hours usually are.
An answer desk over the firm’s own material
Staff ask a question in Teams — how a treatment was handled last year, what the engagement letter says, which procedure applies — and get an answer naming the file it came from. It reads what is already in SharePoint or the practice system rather than a new copy of it.
Missing documents chased
What each client still owes for a period, worked out from what has actually arrived rather than a checklist somebody maintains. The reminder goes out and the reply is filed against the right client.
Client email into the practice system
Enquiries and attachments arriving by email filed against the right client, with the task created and anything ambiguous flagged for a person to look at.
The month-end pack, assembled
The recurring extract pulled from the ledger and the practice system, formatted and distributed on schedule. It says so when a source has not updated instead of sending a pack full of last month’s numbers.
First drafts from your own precedent
Engagement letters, management letters and recurring client correspondence drafted from the firm’s approved wording and prior work, for a partner to edit and sign.
Somebody who owns it afterwards
Model versions, extraction accuracy and spend reviewed on a set cadence once the build is done. Accuracy is re-measured after a supplier changes its template, which happens without notice.
How the work runs
Pick the document that costs the most
Usually supplier invoices or the client receipt bundle. The hours it takes now is the number the build gets measured against, so it gets counted before anything is built.
Collect a real sample
A few hundred actual documents from actual clients, including the photographed, the sideways and the half-legible. The awkward ones decide the project, so they go in the sample deliberately.
Build and measure
Extraction runs against documents whose correct values are already known, so accuracy is a measured figure rather than an assertion. That set stays in place afterwards to test every later change.
Put it in the workflow
The review screen goes to the people who key the documents now, and approved records start landing in the ledger. Corrections they make feed back into the build.
What decides whether it still works next season
The documents belong to other people and the load is seasonal. These are the parts of the build those two facts decide.
-
Accuracy is measured against known answers
A set of documents whose correct values are recorded gets built first. Without it there is no way to tell whether a change improved extraction or quietly made it worse, which is how these systems drift.
-
Line items are harder than totals
A single value on an invoice is straightforward. Line items that run across a page break with a subtotal in the middle are where the engineering time goes.
-
Suppliers redesign their paperwork
Templates change without notice, and a supplier that moves the invoice number does not send word that it has. The system is built to show that a field has moved rather than post a confidently wrong number.
-
The write-back is where the care goes
Reading an invoice wrong costs a correction. Posting it to the ledger twice costs an afternoon of reconciliation, so anything that writes is built to be safe when it runs again.
-
Client files are separated by permission
An answer desk returns what the person asking is entitled to see, so client and engagement boundaries are set in the identity system before anything is indexed. Content that accumulated over a decade usually has broader access than anybody intends.
-
The load arrives in a few weeks of the year
January to April is not the same system as August. Review queues, capacity and spend reporting are sized against the peak week rather than the annual average.
Who this is for
Firms with a bookkeeping or outsourced-finance line
Where document volume is the product, and coding and keying are most of the delivery cost.
Tax and audit practices with a document estate
Years of workpapers, engagement files and correspondence, where two people know how to navigate it and the rest of the firm asks them.
Firms already on Microsoft 365
Identity and much of the licensing are usually in place already, which is most of the groundwork for anything built here.
Firms whose staff already paste work into chat tools
Client material is going into consumer tools and the firm wants a supported route with logging and a named account behind it.
Frequently asked questions
-
Does client data get used to train the model?
No. The model reads a document at the moment it is processed and does not learn from it. On Azure, prompts and completions are not used to train foundation models without your instruction.
-
How accurate is invoice extraction?
That depends on the documents and gets measured during the build against invoices whose correct values are known. A number quoted before anybody has seen your client paperwork would be made up, which is why the measurement is part of the work.
-
Does someone still review every posting?
Yes, and that is the design. The system produces a queue of extracted records beside the original page, and a person approves what goes into the ledger. Records it is confident about move quickly; the rest are queued for a closer look.
-
Will it work with our accounting software?
The write-back is built against whichever system holds the ledger. Where a supported interface exists it is used; where one does not, approved records are delivered in the format the system imports.
-
Can it read a client’s photographed receipts?
That is the case worth testing first, because it is where the hours are. Photographs at an angle, in poor light, with a thumb across the corner are part of the sample the accuracy figure is measured on.
-
How long before it is in use?
A single document type is normally three to six weeks, most of which is spent on the awkward documents rather than the common ones. An answer desk over material already in SharePoint is usually quicker.
Name the document that gets keyed
The one that arrives every week and costs the most hours. That is where this starts.