Home/ Journal/ Memo

AI for CPA firms in 2026, sorted by what actually returns hours.

In a CPA firm the review step is not overhead, it is the product: the client is paying for a return a licensed professional examined and will stand behind. That single fact sorts the AI market better than any feature comparison. Applications that compress preparation time, meaning they get a preparer to a reviewable first draft faster, pay reliably, because they leave the product untouched and lower the cost of producing it. Four land consistently: retrieval over the firm's own prior-year treatment and accumulated guidance, where the answer already exists in the firm and is simply not findable; classification and routing of the annual client document flood; drafting the repetitive client correspondence such as missing-item chases and status updates with a preparer sending; and workpaper support that starts a reviewer from a draft rather than a blank page. A middle category returns real but smaller value, and only when the firm designs and staffs the verification step: figure extraction from source documents, tax research orientation, call summaries. The category that keeps failing is anything positioned to produce a tax position or a client-facing answer with no licensed reviewer attached, general-purpose chat used against client data outside a sanctioned system, and any tool bought to fix a process the firm has never written down. The calendar decides more outcomes than the technology. Nothing requiring new staff behavior can be introduced between late January and April 15, and anything not adopted before the season will not be adopted during it, which leaves May through November as the only real window to build and train.

A managing partner deciding where to spend on AI this year is really deciding which part of the firm's work is safe to hand over, and the two outcomes are not symmetric. Choose well and a meaningful share of preparation hours comes back, most of it at the staff level that is hardest to hire and hold. Choose badly and the firm has paid for something that sits unused past April, plus a cost that is harder to write off: a team that now believes the whole category does not work and will resist the next attempt for two years. The sorting is not difficult, but it does not follow the categories the software market sells in, and it does not turn on which model is underneath. What separates the applications that return hours from the ones that get quietly abandoned is a structural feature of the profession itself, and once a partner sees it the shortlist gets short fast.

MemoJuly 2026
Read time11 minutes
AudienceManaging Partners, Tax Partners, Firm Administrators

The test that sorts every application, and why it has nothing to do with the technology.

A client paying a CPA firm is not buying a completed form; competent bookkeeping produces forms. What the client buys is a return or a statement a licensed professional examined and will put a name to, along with the judgment behind it and the willingness to defend it later. That is what the fee attaches to, and why the profession is licensed at all.

Which makes the review step something other than friction. In most businesses a checking step is waste to be engineered out. In a CPA firm it is the product. Preparation builds the artifact; review attaches the judgment, and the judgment is what was sold. Once a partner sees the work that way the AI market sorts itself, because every product on offer either shortens the road to a reviewable draft or tries to take the review out of the middle.

The first category compounds. Preparation hours are the firm's largest controllable cost and the hardest to staff, and compressing them changes nothing about what the client receives. The reviewer still reviews and the partner still signs. The firm sells the same product at a lower cost to produce, which is the only efficiency a licensed practice can bank without spending down what makes it worth hiring.

The second category is where the money goes to die, and not because the technology is immature. A firm cannot accept an output that establishes a tax position without review, so the review happens regardless. Now the reviewer is checking work that nobody in the building did, with no preparer to question. That is slower than reviewing a staff draft. The vendor sold a saving the firm's own obligations make it unable to take, and a better model next year does not change that.

The test has a second half that catches the near misses. A faster draft only pays if a reviewer can check it. An output the reviewer has to re-derive from scratch in order to trust is not a first draft, it is a second thing to verify. So the question to ask about any product here is whether it gets a preparer to a checkable draft faster than the preparer could alone, without making the checking harder.

What returns hours now, in the order it tends to land.

Four applications clear that test consistently inside a mid-market firm, listed here by reliability rather than by how well they demonstrate, which is roughly the reverse order.

Retrieval over the firm's own accumulated guidance and prior-year treatment. This is the highest-return application in the category and the least discussed, because it does not look like AI in the way the market advertises. The situation it fixes: somebody needs to know how the firm handled this exact issue, for this client or one like it, in a prior year. The answer exists. A partner wrote a memo, or the call was made in an email thread in 2023. In practice it is unfindable, so a staff member interrupts a senior person, who reconstructs it from memory in four minutes, and the firm pays its most expensive hour to retrieve something it already owned.

Retrieval handles this well because the task is finding rather than deciding. A well-built system returns the passage and the source document rather than an assertion, so the preparer reads the original and forms a view. It also fails safely: when it finds nothing, the preparer asks the partner exactly as they do today. No version of this produces a wrong filing position, because it does not produce positions. The constraint is that the firm's own material has to be reachable and organized enough to search, which is why data readiness is a property of a specific workflow rather than of a company.

Intake, classification, and routing of the annual document flood. Between January and April a firm receives client documents in every form a human being can produce: portal uploads, emailed scans, a photograph of a W-2 taken on a kitchen counter, a shoebox arriving as forty unnamed PDFs. Someone has to work out what each one is, whose it is, and whether it satisfies an outstanding request. This works because the output is checkable at a glance by whoever opens the file next, and because an error costs a misfiled document rather than a wrong return. It also moves information forward in time: a firm that classifies as documents arrive knows in late January which clients are missing which items, instead of in March when the calendar has no room left.

The repetitive client correspondence, with a preparer in the loop. Missing-item chases, status updates, extension notices, engagement letter variants. These are already written from templates; what varies is client-specific fact, which the firm holds. Drafting them is the task staff resent most and defer longest, which is much of why documents arrive late. Draft-and-send works here because the person sending has enough context to catch an error and reading a paragraph takes seconds. It stops working the moment somebody removes that person to save the seconds, because output that leaves the firm with nobody attached is the failure pattern below.

Workpaper support that starts the reviewer from a draft rather than a blank page. Assembling support, pulling prior-year comparatives, flagging variances above a threshold, laying out the tie-out so a preparer works through it instead of building it. Nothing here decides anything. The reviewer reviews what they always reviewed, and the preparer arrives with more of the mechanical assembly done. The gain is unglamorous and it is where the largest number of hours sit. It is also the most sensitive to how a particular firm works, which is why it tends to be what a custom engagement inside a CPA firm centers on.

What half-works, and the human step that never shows up in the pricing.

A middle category returns genuine value and is sold at roughly double what it delivers, because the vendor quotes the gross saving and the firm absorbs the verification. These are not bad purchases; they are bad purchases at the advertised number.

Extraction of figures from client source documents. Accuracy is high and improving, and the failure mode is the problem rather than the rate. An extraction error is silent and plausible. A transposed figure does not announce itself the way a missing document does, so the firm has to design and staff a verification pass side by side against the source. Do that and the net gain is real, smaller than advertised, and durable. Skip it and the firm has not saved time, it has moved a risk somewhere nobody is watching. Firms that get burned here rarely decided to skip verification; they assumed somebody downstream was doing it.

Tax research assistance. A research tool orients a preparer quickly on an unfamiliar issue and produces a reasonable starting list of authority. It does not conclude, and every citation still has to be pulled and read, because a citation that looks right and is not is the most expensive object in the building. Firms that use this well treat the output as a table of contents. Firms that get hurt treat the summary as the research, and the position rests on something nobody read.

Call and meeting summarization. Useful for internal continuity, weak as a record of what a client agreed to. Brief the next person on the engagement with it, but do not let it become the file's account of a client decision.

Client-facing status answers. Telling a client where their return stands is a bounded, low-risk question and a reasonable candidate, conditional on the status being accurate in a system somewhere. In most firms it is not: status lives in the preparer's head and in the last email, so the tool becomes a fast way to tell clients something wrong. That is a process problem, and fixing it pays whether or not a tool gets bought.

What consistently fails, and why the reason is structural.

Three patterns account for most of the abandoned projects in this profession, and none of them improve with a stronger model next year.

Output positioned as a tax position or a client-facing answer with no licensed reviewer attached. This never arrives labeled that way. It arrives as a pilot producing draft returns for review, where review turns out to mean re-deriving the work, or as a client chat answering substantive questions because restricting it to status felt like underusing the investment. The firm cannot delegate the judgment, so the judgment stays where it was, and the cost of the tool now sits on top of a review that got harder. When one of these stalls, what to do after a failed AI pilot is a diagnostic exercise rather than a search for somebody new to try again with.

General-purpose chat used against client data outside a sanctioned system. IRC Section 7216 and the regulations under Treasury Reg 301.7216 restrict the use and disclosure of tax return information, and the AICPA Code of Professional Conduct imposes its own duty regarding confidential client information. What matters more for a managing partner is what the behavior indicates. The person pasting client figures into a consumer tool is usually one of your stronger preparers, solving a real problem the firm has not solved for them. A policy memo ends the visible behavior and leaves the problem, and the next version happens on a personal phone where nothing is visible at all.

A tool bought to fix a process the firm has never defined. The document chase is slow, so the firm buys something to accelerate the chase. The chase was slow because no single person owned it, the request list differed by preparer, and nobody agreed what complete meant. The purchase automates the ambiguity and delivers it faster, now inside a system where it is harder to argue with. This one is easy to detect in advance: if two people in the firm describe the process differently, there is no process to automate yet, and the month spent writing it down is the higher-return project.

The compression-season constraint, which kills more of these than the technology does.

Two facts about the calendar decide more CPA firm AI outcomes than any decision about tooling. Anything that requires staff to learn a new behavior cannot be introduced between late January and April 15. And anything not adopted before the season will not be adopted during it.

The second is the one firms underestimate. Under deadline pressure people revert to the method they trust, and that is correct professional behavior rather than resistance. A preparer holding forty returns and a hard date is not going to experiment with a new intake flow. So the season does not test a half-adopted system, it discards it, and the firm reads the discarding as the tool having failed.

What follows is a cadence rather than a project plan. The window for building and training runs roughly May through November, with a smaller slot around the extension deadlines for anything narrow. The season is the proof window, the only time a firm finds out whether something holds at real volume with real client behavior. The review belongs in May, while people still remember what broke and can say so specifically. A firm that judges a system on its first season is grading it during the year it was learning.

The trap is that the decision usually gets made at the wrong point in that cycle. The firm feels the pain most acutely in March, starts looking in April, and scopes something in June, by which point the partner who felt it is thinking about something else. Momentum decays, the work drifts toward the season it was meant to serve, and lands too late to train anybody. That decay is the same mechanism behind why mid-market rollouts stall around month four. The tax calendar makes it sharper, because there is a hard date after which the window closes for a full year.

One consequence is worth acting on before anything else. Measure the current state during a season you are already running, before anything changes. A firm that never wrote down what the document chase consumed will spend the following spring arguing about whether the system helped, and that argument gets settled by whoever is most invested in the answer.

When to commission nothing at all, and what a partner can do this week.

We build custom systems for a living, so read this section with the appropriate suspicion. The firms that commission at the wrong moment are the ones who end up believing the whole category is a waste of money.

The software you already run may ship it. The tax and practice management platforms most mid-market firms own have added much of the capability described above, usually in a tier you already have or can buy this quarter. A custom version of a shipped feature costs more up front, has to be maintained, and ages against a product with a roadmap behind it. If the need is generic to the profession and the documents already live in that system, buy the tier. The build versus buy decision turns on how specific the workflow is to your firm, and the calculation worth running first is what the platform costs across your full headcount at renewal, since per-seat pricing at scale is where firms discover they picked the expensive option by accident.

Doing nothing this year is a legitimate answer. It is right if the firm is mid-migration on practice management, if the partner group is not aligned, since a build with a dissenting partner becomes a tool nobody enforces, or if the volume behind the workflow is too low for automation to return anything. And it is right if the process does not exist in written form: a year spent standardizing document naming, fixing intake, and writing down how a return moves through the office returns more than a build laid over the disorder, at a fraction of the cost.

Custom is the right instrument in a narrower set of cases than the market implies: when the workflow is specific to how your firm works, when no vendor will ship it, and when the volume behind it is real. If you are in that set, the shortlist of AI consultants working with CPA firms is worth reading for the questions it raises.

Four things a managing partner can do this week without involving a vendor. Name the question that most often interrupts your most senior tax person and count how many times it got asked last week; if the answer already existed in the firm both times, you have found the retrieval case. Ask your three heaviest preparers, with no consequence attached, which tools they already use on their own, and write down the problem each was solving rather than the tool. Count what share of last season's client email was chasing missing documents. Then put each finding in a column headed preparation or a column headed review.

Two columns and an afternoon. Firms that run it honestly come out with two candidates worth funding and one purchase they were about to make for the wrong reason. If something real lands in the preparation column and survives a full season, the signs that a firm is ready to commission a custom build will already be visible, and the first vendor conversation starts from a measured problem instead of a demonstration.

Field-note context

What we notice inside an accounting practice.

The answer the staff member needed was already in the building.

In most firms, the questions that interrupt partners most often have documented answers sitting in the firm's own files. The treatment was decided, somebody wrote it down, and the writing is in a folder nobody remembers, or in an email thread, or in a binder from a year when the client had a different entity structure. Firms read this as a training gap and respond with more onboarding, which does not touch it. The tell that it is a retrieval problem instead is that the same senior person answers the same question for different staff members inside the same month. Counting that is a two-week exercise for a firm administrator with a notepad, and it usually settles the argument about where to start better than any assessment.

Review time expands to fit the quality of what the preparer hands over.

The number partners watch is preparation hours. The number that moves quietly is review hours. A draft assembled faster but less carefully costs the reviewer more, and because review sits with the most expensive people in the practice, a change that looks like a win in the preparation column can be a net loss and stay invisible for a full season. Before and after numbers are worth capturing for both, which is often the first time anyone has looked at the second one. The firms that track only preparation are usually the ones who conclude an earlier tool did not really change anything, and frequently that is exactly what happened. The hours moved rather than disappeared.

Somebody in the firm is already using AI, and nobody has asked them what for.

Nearly every practice has at least one person who has quietly adopted a consumer tool for part of their work, usually a strong performer, usually for drafting or for orienting on an unfamiliar issue. The instinct is to treat it as a policy breach, and the confidentiality exposure is real enough that it has to be addressed. The behavior is also the cheapest research the firm will ever get for free. It identifies the exact task where the current process is worst, selected by the person who feels it every day, at their own inconvenience. Firms that ask before they prohibit tend to arrive at a better first project than firms that ask a vendor what to do.

Extended questions

The questions partners ask about AI in a tax practice.

Why does some AI work in a CPA firm while other tools fail?

Because of what the firm is actually selling. A client pays for a return or a statement that a licensed professional has examined and will stand behind, which makes the review step the product rather than overhead. That single fact sorts the market. Tools that shorten the path to a reviewable first draft leave the product untouched and lower the cost of producing it, which is why retrieval over the firm's own prior-year treatment, document intake and classification, drafted client correspondence, and workpaper assembly return hours reliably. Tools sold on removing or shortcutting the review ask the firm to sell less of what it sells while carrying the same exposure. The review then happens anyway, and the reviewer is checking work that nobody in the building did and cannot be asked about, which is slower than reviewing a staff draft rather than faster. The failure is structural rather than a matter of model quality, which is why waiting for a better model next year does not fix it.

Can AI replace the review step on a tax return?

No, and the reason is not caution about accuracy. The signature, the professional judgment behind it, and the willingness to defend a position are what the client is buying and what the license exists to make meaningful. Peer review assumes a named professional stood behind the work. A firm cannot accept an output that establishes a tax position without a reviewer, so removing the reviewer removes nothing. It relocates the same work to someone with less context and no preparer to question. What AI can change is what the reviewer receives: a draft with the support assembled, prior-year comparatives pulled, and variances above a threshold flagged, arriving in less time than a preparer needed to build it from nothing. The review itself is identical, and it simply starts later in the process. That is a real saving, and it does not require the profession to pretend the review is optional.

Is it a problem if staff use general AI chat tools with client data?

Yes, and it is worth separating the exposure from the signal. IRC Section 7216 and the regulations under Treasury Reg 301.7216 restrict the use and disclosure of tax return information, and the AICPA Code of Professional Conduct imposes its own duty regarding confidential client information. Neither depends on the staff member's intent, and consumer tools generally sit outside anything the firm has sanctioned, controls, or can audit. The signal matters as much as the exposure. The person doing this is usually a capable preparer who found a task the firm's current process handles badly and solved it themselves on their own time. A policy memo ends the visible behavior and leaves the underlying problem, which reappears somewhere less visible. The stronger response is to ask what they were trying to accomplish, then decide whether the firm should provide a sanctioned way to do the same thing inside a system it controls.

When in the year should a CPA firm start an AI project?

Build and train between roughly May and November, with a smaller usable window around the extension deadlines. Two calendar facts govern this. Anything that requires staff to learn a new behavior cannot be introduced between late January and April 15, and anything not adopted before the season will not be adopted during it, because people working against a hard deadline revert to the method they trust, which is correct professional behavior rather than resistance. The season is the proof window rather than the adoption window. It tells you whether the system holds at real volume, and the review belongs in May while people still remember what broke. The common failure here is timing rather than tooling. A firm feels the pain most acutely in March, starts looking in April, scopes in June, and by then the urgency has faded and the work drifts past the training window into the next season, where it arrives unadopted and gets discarded.

When should a CPA firm skip a custom AI build entirely?

In four situations, and they are more common than the market suggests. When the tax or practice management platform the firm already runs ships the capability in a tier you own or can buy, since building a custom version of a shipped feature costs more up front, needs maintaining, and ages against a product that has a roadmap. When the firm is mid-migration on practice management, because you would be building against a system that is about to change underneath you. When the partner group is not aligned, since a build with a dissenting partner becomes a tool nobody enforces and quietly dies after one season. And when the process itself has never been defined, which you can test in an afternoon by asking two people to describe it and seeing whether the descriptions match. A year spent standardizing intake, document naming, and the actual sequence a return moves through returns more than a build laid over the disorder, at a fraction of the cost.

How does the AI Maturity Index help a CPA firm sort this?

It runs the two-column exercise in this memo without requiring a call or a vendor. The Index asks which process is worth investing in first, who touches it and how often, and where that process gets its inputs, which is most of what a firm needs in order to tell a preparation problem from a review problem. For an accounting practice the most useful output is usually the elimination rather than the recommendation: it surfaces the candidates that depend on removing a professional judgment, which are the ones to close the file on this year, and it sizes the ones that only compress preparation. Ten minutes, no call, and the result is specific enough to take into a partner meeting in May, which is the month when a decision about the next season can still be acted on.

Not sure which of these your firm should fund first?

Start with the AI Maturity Index. Ten minutes, no call, and it names the one workflow worth investing in first, sizes who touches it and how often, and tells you whether what you are looking at is a preparation problem or a review problem before you talk to anyone.