The test that sorts every application, and why it has nothing to do with the technology.
A client paying a CPA firm is not buying a completed form; competent bookkeeping produces forms. What the client buys is a return or a statement a licensed professional examined and will put a name to, along with the judgment behind it and the willingness to defend it later. That is what the fee attaches to, and why the profession is licensed at all.
Which makes the review step something other than friction. In most businesses a checking step is waste to be engineered out. In a CPA firm it is the product. Preparation builds the artifact; review attaches the judgment, and the judgment is what was sold. Once a partner sees the work that way the AI market sorts itself, because every product on offer either shortens the road to a reviewable draft or tries to take the review out of the middle.
The first category compounds. Preparation hours are the firm's largest controllable cost and the hardest to staff, and compressing them changes nothing about what the client receives. The reviewer still reviews and the partner still signs. The firm sells the same product at a lower cost to produce, which is the only efficiency a licensed practice can bank without spending down what makes it worth hiring.
The second category is where the money goes to die, and not because the technology is immature. A firm cannot accept an output that establishes a tax position without review, so the review happens regardless. Now the reviewer is checking work that nobody in the building did, with no preparer to question. That is slower than reviewing a staff draft. The vendor sold a saving the firm's own obligations make it unable to take, and a better model next year does not change that.
The test has a second half that catches the near misses. A faster draft only pays if a reviewer can check it. An output the reviewer has to re-derive from scratch in order to trust is not a first draft, it is a second thing to verify. So the question to ask about any product here is whether it gets a preparer to a checkable draft faster than the preparer could alone, without making the checking harder.
What returns hours now, in the order it tends to land.
Four applications clear that test consistently inside a mid-market firm, listed here by reliability rather than by how well they demonstrate, which is roughly the reverse order.
Retrieval over the firm's own accumulated guidance and prior-year treatment. This is the highest-return application in the category and the least discussed, because it does not look like AI in the way the market advertises. The situation it fixes: somebody needs to know how the firm handled this exact issue, for this client or one like it, in a prior year. The answer exists. A partner wrote a memo, or the call was made in an email thread in 2023. In practice it is unfindable, so a staff member interrupts a senior person, who reconstructs it from memory in four minutes, and the firm pays its most expensive hour to retrieve something it already owned.
Retrieval handles this well because the task is finding rather than deciding. A well-built system returns the passage and the source document rather than an assertion, so the preparer reads the original and forms a view. It also fails safely: when it finds nothing, the preparer asks the partner exactly as they do today. No version of this produces a wrong filing position, because it does not produce positions. The constraint is that the firm's own material has to be reachable and organized enough to search, which is why data readiness is a property of a specific workflow rather than of a company.
Intake, classification, and routing of the annual document flood. Between January and April a firm receives client documents in every form a human being can produce: portal uploads, emailed scans, a photograph of a W-2 taken on a kitchen counter, a shoebox arriving as forty unnamed PDFs. Someone has to work out what each one is, whose it is, and whether it satisfies an outstanding request. This works because the output is checkable at a glance by whoever opens the file next, and because an error costs a misfiled document rather than a wrong return. It also moves information forward in time: a firm that classifies as documents arrive knows in late January which clients are missing which items, instead of in March when the calendar has no room left.
The repetitive client correspondence, with a preparer in the loop. Missing-item chases, status updates, extension notices, engagement letter variants. These are already written from templates; what varies is client-specific fact, which the firm holds. Drafting them is the task staff resent most and defer longest, which is much of why documents arrive late. Draft-and-send works here because the person sending has enough context to catch an error and reading a paragraph takes seconds. It stops working the moment somebody removes that person to save the seconds, because output that leaves the firm with nobody attached is the failure pattern below.
Workpaper support that starts the reviewer from a draft rather than a blank page. Assembling support, pulling prior-year comparatives, flagging variances above a threshold, laying out the tie-out so a preparer works through it instead of building it. Nothing here decides anything. The reviewer reviews what they always reviewed, and the preparer arrives with more of the mechanical assembly done. The gain is unglamorous and it is where the largest number of hours sit. It is also the most sensitive to how a particular firm works, which is why it tends to be what a custom engagement inside a CPA firm centers on.
What half-works, and the human step that never shows up in the pricing.
A middle category returns genuine value and is sold at roughly double what it delivers, because the vendor quotes the gross saving and the firm absorbs the verification. These are not bad purchases; they are bad purchases at the advertised number.
Extraction of figures from client source documents. Accuracy is high and improving, and the failure mode is the problem rather than the rate. An extraction error is silent and plausible. A transposed figure does not announce itself the way a missing document does, so the firm has to design and staff a verification pass side by side against the source. Do that and the net gain is real, smaller than advertised, and durable. Skip it and the firm has not saved time, it has moved a risk somewhere nobody is watching. Firms that get burned here rarely decided to skip verification; they assumed somebody downstream was doing it.
Tax research assistance. A research tool orients a preparer quickly on an unfamiliar issue and produces a reasonable starting list of authority. It does not conclude, and every citation still has to be pulled and read, because a citation that looks right and is not is the most expensive object in the building. Firms that use this well treat the output as a table of contents. Firms that get hurt treat the summary as the research, and the position rests on something nobody read.
Call and meeting summarization. Useful for internal continuity, weak as a record of what a client agreed to. Brief the next person on the engagement with it, but do not let it become the file's account of a client decision.
Client-facing status answers. Telling a client where their return stands is a bounded, low-risk question and a reasonable candidate, conditional on the status being accurate in a system somewhere. In most firms it is not: status lives in the preparer's head and in the last email, so the tool becomes a fast way to tell clients something wrong. That is a process problem, and fixing it pays whether or not a tool gets bought.
What consistently fails, and why the reason is structural.
Three patterns account for most of the abandoned projects in this profession, and none of them improve with a stronger model next year.
Output positioned as a tax position or a client-facing answer with no licensed reviewer attached. This never arrives labeled that way. It arrives as a pilot producing draft returns for review, where review turns out to mean re-deriving the work, or as a client chat answering substantive questions because restricting it to status felt like underusing the investment. The firm cannot delegate the judgment, so the judgment stays where it was, and the cost of the tool now sits on top of a review that got harder. When one of these stalls, what to do after a failed AI pilot is a diagnostic exercise rather than a search for somebody new to try again with.
General-purpose chat used against client data outside a sanctioned system. IRC Section 7216 and the regulations under Treasury Reg 301.7216 restrict the use and disclosure of tax return information, and the AICPA Code of Professional Conduct imposes its own duty regarding confidential client information. What matters more for a managing partner is what the behavior indicates. The person pasting client figures into a consumer tool is usually one of your stronger preparers, solving a real problem the firm has not solved for them. A policy memo ends the visible behavior and leaves the problem, and the next version happens on a personal phone where nothing is visible at all.
A tool bought to fix a process the firm has never defined. The document chase is slow, so the firm buys something to accelerate the chase. The chase was slow because no single person owned it, the request list differed by preparer, and nobody agreed what complete meant. The purchase automates the ambiguity and delivers it faster, now inside a system where it is harder to argue with. This one is easy to detect in advance: if two people in the firm describe the process differently, there is no process to automate yet, and the month spent writing it down is the higher-return project.
The compression-season constraint, which kills more of these than the technology does.
Two facts about the calendar decide more CPA firm AI outcomes than any decision about tooling. Anything that requires staff to learn a new behavior cannot be introduced between late January and April 15. And anything not adopted before the season will not be adopted during it.
The second is the one firms underestimate. Under deadline pressure people revert to the method they trust, and that is correct professional behavior rather than resistance. A preparer holding forty returns and a hard date is not going to experiment with a new intake flow. So the season does not test a half-adopted system, it discards it, and the firm reads the discarding as the tool having failed.
What follows is a cadence rather than a project plan. The window for building and training runs roughly May through November, with a smaller slot around the extension deadlines for anything narrow. The season is the proof window, the only time a firm finds out whether something holds at real volume with real client behavior. The review belongs in May, while people still remember what broke and can say so specifically. A firm that judges a system on its first season is grading it during the year it was learning.
The trap is that the decision usually gets made at the wrong point in that cycle. The firm feels the pain most acutely in March, starts looking in April, and scopes something in June, by which point the partner who felt it is thinking about something else. Momentum decays, the work drifts toward the season it was meant to serve, and lands too late to train anybody. That decay is the same mechanism behind why mid-market rollouts stall around month four. The tax calendar makes it sharper, because there is a hard date after which the window closes for a full year.
One consequence is worth acting on before anything else. Measure the current state during a season you are already running, before anything changes. A firm that never wrote down what the document chase consumed will spend the following spring arguing about whether the system helped, and that argument gets settled by whoever is most invested in the answer.
When to commission nothing at all, and what a partner can do this week.
We build custom systems for a living, so read this section with the appropriate suspicion. The firms that commission at the wrong moment are the ones who end up believing the whole category is a waste of money.
The software you already run may ship it. The tax and practice management platforms most mid-market firms own have added much of the capability described above, usually in a tier you already have or can buy this quarter. A custom version of a shipped feature costs more up front, has to be maintained, and ages against a product with a roadmap behind it. If the need is generic to the profession and the documents already live in that system, buy the tier. The build versus buy decision turns on how specific the workflow is to your firm, and the calculation worth running first is what the platform costs across your full headcount at renewal, since per-seat pricing at scale is where firms discover they picked the expensive option by accident.
Doing nothing this year is a legitimate answer. It is right if the firm is mid-migration on practice management, if the partner group is not aligned, since a build with a dissenting partner becomes a tool nobody enforces, or if the volume behind the workflow is too low for automation to return anything. And it is right if the process does not exist in written form: a year spent standardizing document naming, fixing intake, and writing down how a return moves through the office returns more than a build laid over the disorder, at a fraction of the cost.
Custom is the right instrument in a narrower set of cases than the market implies: when the workflow is specific to how your firm works, when no vendor will ship it, and when the volume behind it is real. If you are in that set, the shortlist of AI consultants working with CPA firms is worth reading for the questions it raises.
Four things a managing partner can do this week without involving a vendor. Name the question that most often interrupts your most senior tax person and count how many times it got asked last week; if the answer already existed in the firm both times, you have found the retrieval case. Ask your three heaviest preparers, with no consequence attached, which tools they already use on their own, and write down the problem each was solving rather than the tool. Count what share of last season's client email was chasing missing documents. Then put each finding in a column headed preparation or a column headed review.
Two columns and an afternoon. Firms that run it honestly come out with two candidates worth funding and one purchase they were about to make for the wrong reason. If something real lands in the preparation column and survives a full season, the signs that a firm is ready to commission a custom build will already be visible, and the first vendor conversation starts from a measured problem instead of a demonstration.