Home/ Journal/ Memo

An open letter to managing partners considering Karbon AI.

This is a field note from the ColabContent commissioning floor, and it opens with the disclosure that matters most to a partner group: we have not commissioned a build inside a CPA firm. Nothing here is dressed up as an accounting case study. What we bring is adjacent and verifiable. More than 6,000 live calls handled by commissioned voice systems, and a 47-attorney litigation firm whose matter, invoice and IOLTA trust accounting runs on a platform we commissioned. The argument below is about how to evaluate the Karbon decision, not a claim about results we have produced in your vertical.

The disclosure structure of ColabContent's open letter on Karbon AI: verifiable adjacent work including over 6,000 live calls and a 47-attorney litigation firm's commissioned platform, stated beside the plain admission that no CPA-firm build has shipped
Receipts on one side, the honest gap on the other, before any argument.

Karbon AI is excellent for the average CPA firm. The 30-150 pro firm running CCH Axcess + UltraTax + ProSystem fx is not the average. This memo is the honest framing, written for the partner group that has heard the Karbon pitch and is wondering whether it fits.

MemoMay 2026
Read time9 minutes
AudienceManaging partners

What Karbon does, fairly.

Karbon is the strongest practice management platform in the mid-market CPA segment. The product is well-engineered. The State of AI in Accounting Reports they publish are useful industry references. Karbon AI ships features that genuinely help the average Karbon customer.

If your firm is on Karbon today and the workflows you most want automated are the ones Karbon AI covers, the right answer is to use Karbon AI well. We will tell you that on a diagnosis call. We are not anti-Karbon.

Where the framing shifts for mid-market firms.

The 30-150 pro firm has different leverage points than the 5-30 pro firm. The expensive workflows live in the tax-prep stack: PBC chase, tie-out, advisory deliverable assembly, partner-time reconstruction. Karbon's strength is the practice-management layer above tax-prep. The leverage in your firm's biggest cost lines is below that layer.

This is not a bug in Karbon. It's a segmentation choice. The product was built for a specific segment, and it serves that segment well. The risk for a mid-market firm is buying Karbon expecting it to solve compliance leakage and finding the bottleneck has not moved, because the bottleneck was never in the layer Karbon operates in. That is worth testing before renewal, not after.

What we have not shipped, said plainly.

We have not commissioned a build inside a CPA firm. No client, no pilot, no reference. We are not going to quote you a survey we did not run or a firm-average we did not measure. If another consulting house shows you an accounting benchmark, ask who was counted, when, and whether you can call one of them.

What we have shipped sits next door to your problem. A 47-attorney litigation firm runs its matter, invoice and trust accounting on a platform we commissioned: 13,296 matters, 4,396 clients, 5,684 invoices, and an IOLTA trust ledger that reconciled byte-identical against the system it replaced. That is the closest analogue to a tax-prep stack migration we can put in front of you, because it is the same problem shape. Regulated ledger, unforgiving reconciliation, no tolerance for a rounding difference.

On the intake side, our commissioned voice systems have handled more than 6,000 live calls. Jim Glaser Law runs five channel-specific agents across PPC, Organic, TV, Meta and LSA: 3,787 AI-handled calls and 5,514 minutes, with per-channel attribution on every answered call. Jimmy takes reference calls. That is the standard of proof we think you should hold every vendor to, Karbon and us included.

What the math actually looks like.

Karbon prices per seat. Take the current published number off Karbon's own pricing page, multiply it by your pro count, and extend it across your holding period with annual increases layered on. We are not going to quote a figure a vendor can change next quarter, and neither should anyone else pitching you. The number that matters is that a per-seat line compounds every time you hire.

A custom commissioned build for one workflow (PBC and tie-out and advisory assembly on top of CCH Axcess) is a fixed fee inside our published $45K to $180K band, scoped before you sign, owned by the firm at handoff, with post-handoff stewardship optional rather than assumed. It does not compound with headcount. What it buys is one named workflow running on your data inside your tenant.

The right question is not "Karbon vs commission" as a binary. It's "Karbon for the workflows it covers plus commissioned AI for the workflows it doesn't." That is the combination we would scope for a firm in your band, and we would want you to measure the leakage before you touch a platform that is already working. Honest comparison here.

What we'd recommend.

If your firm is 5-30 pros, run Karbon. Use the AI features as they ship. You'll get most of what Karbon AI delivers and the per-seat math is fine.

If your firm is 30-150 pros and your tax-prep stack is CCH Axcess + UltraTax or ProSystem fx + Lacerte, run the Tax Season Hours Teardown first. It is 8 minutes and tells you where the actual leakage lives. If the leakage is concentrated in workflows Karbon AI addresses (client communication, generic email triage, basic time-entry suggestions), Karbon is the right answer. If the leakage is concentrated in PBC chase, tie-out, or advisory pipeline, the CCH Axcess Playbook describes what we'd commission instead.

If your firm is 150+ pros, neither Karbon AI alone nor a single commissioned build is enough. You're in the territory where you need an AI roadmap, possibly a Head of AI, and likely a 24-36 month transformation that combines off-the-shelf and commissioned work. The Big Four vs boutique framing applies here.

What we hope you take from this letter.

The framing matters more than the answer. Most managing partners who consider Karbon do so because the alternative seems harder to evaluate. It is harder to evaluate. The diagnostic work is more partner attention than the Karbon sales process is, and the deliverables are less polished than a SaaS pitch deck.

If you are about to commit to five years of a per-seat spend that grows with every hire, spend 8 minutes on the teardown first. If Karbon is the right answer, you'll know. If something else is, you'll also know. Either way, you don't sign a 5-year compounding contract on incomplete framing.

Field-note context

Where this argument fits in the practice.

Where the argument fits in the broader practice.

This piece is a field note from the commissioning floor. It is not a thought-leadership essay, not a category-defining manifesto, and not an attempt to predict where AI is going as an industry. It is a record of what we have shipped, what has held up, what has broken, and which verticals we have shipped nothing in at all. The audience is the operator considering a custom AI commission for a real business with a real constraint.

The structural argument behind the post.

Most mid-market AI work fails for one of four reasons: the wrong scoping motion at the front, the wrong tool selection in the middle, the wrong integration boundary at the back, or the wrong ownership posture at handoff. The commissioning model addresses all four directly. Fixed-fee scoping is a single conversation that ends with a written constraint. Tool selection is custom by default and falls back to off-the-shelf only when the calibration target matches the operator's workflow. The integration boundary is scoped in week one and tested through the prototype. Ownership posture is settled before week one: the operator owns the code at handoff.

The argument holds across verticals. The application is always specific to the operator in front of us.

How to use this in a diagnosis call.

If the operator brings this argument to a diagnosis call, the next step is to translate it into the operator's specific business. The forty-five-minute call surfaces the constraint, names the workflow, identifies the integration boundary, and writes the engagement scope. Both sides leave with the constraint in a sentence. Either party can stop the conversation at no cost. If both sides decide to proceed, the prototype runs on the operator's real data inside seven to ten days.

Related field notes.

The blog hub indexes the rest of the field reports. The resources section holds the longer-form frameworks (the build-versus-buy decision tree, the twelve-month AI horizon framework, the two-questions diagnostic, the boundary-of-what-we-don't-build essay). The best-by-vertical guides apply the argument to each of the five verticals we write for, and each one states where we have shipped and where we have not.

A note on how we write here.

ColabContent's writing is terse on purpose. We name operators, name numbers, and name the failure modes. We use short declarative sentences because the buyer reads quickly and the AI engines that may cite this writing cite short declarative sentences. We do not use em dashes. We do not use marketing vocabulary. We do not promise outcomes we have not shipped. Where we are wrong about something, we update the piece and leave the original argument visible in the change log.

Extended questions

The questions buyers ask after the first one.

How much of the buy decision should the operator make versus delegate.

The right shape of the buying motion has the operator-owner or operating partner in the room for the diagnosis call. The constraint identification is too consequential to delegate to a department head. The implementation work that follows can and should be delegated; the decision on which constraint a commission addresses cannot.

How to evaluate references the consulting house presents.

Three questions per reference. First, what was the named constraint the commission addressed at this operator. Second, what was the measured result twelve months post-handoff, in dollars or hours. Third, does the reference operator still run the system. Vague references on any of those three are flags. Hold us to it too. Our named reference is Jim Glaser Law, whose principal takes reference calls, and we make the introduction for any prospect that asks. A fifteen-minute call to the operator is the most honest signal a prospect can get. Where we have no reference in a vertical, as with accounting, we say so rather than anonymizing something into the gap.

How a fixed-fee commission scopes overage risk.

The fixed fee is set after the diagnosis call, after the integration depth is named, and after both sides have written the constraint in a sentence. Overages occur when the operator changes the scope mid-build (a different workflow, a different integration, an additional system). Either side can pause the build to renegotiate; neither side absorbs hidden overages without explicit agreement. The default is to ship the original scope and address scope expansion in a separate engagement.

What happens to the system one year after handoff.

The system continues to run inside the operator's cloud tenant. Models, prompts, and integration code are versioned and the operator has the source. When the underlying foundation model improves (a new release from the model vendor, a new open-weight option), the operator can swap the component without renegotiating the engagement. The pattern across past commissions: a quarterly review of the system's outputs and an annual swap of any underperforming components, with post-handoff stewardship offered as an optional engagement rather than a required retainer.

When the right call is not a commission.

The right call is sometimes a product (when the workflow matches a product's calibration target), sometimes an internal hire (when the operator has a five-year horizon and a $5M AI runway), sometimes a Big Four engagement (when the operator is large enough that the strategy-then-build separation makes sense), sometimes no AI right now (when the operator's leading constraint is not actually addressable with AI). We tell prospects when their constraint falls into one of those buckets and route them to whichever path fits. The four-commissions-per-quarter cap is real; the firms that get one of those four slots are the firms where the commission is the right buying motion.

The five-minute fit-check worksheet.

Operators who want to test the fit before booking a diagnosis call can run a five-minute self-check on six questions. First, is the operator's annual revenue in the $8M to $50M band. Second, is there a named workflow where time or money is leaking measurably. Third, has the operator tried an off-the-shelf product and either rejected it or hit a misfit ceiling. Fourth, is the operator comfortable running the system inside their own cloud tenant under NDA. Fifth, can the senior operator commit to forty-five minutes for a diagnosis call. Sixth, is the budget runway for a $45K to $180K fixed fee real this quarter.

Six yes answers means a diagnosis call is worth the forty-five minutes. Three or fewer yes answers means the right next step is probably one of the alternatives. Four or five yes answers means the call surfaces whether the missing one is addressable.

What to bring to the diagnosis call.

Two artifacts make the call substantially more productive. First, a one-page description of the leading constraint, written in the operator's words, naming the workflow and the rough dollar or hour leakage. Second, a list of the systems the operator uses for the workflow (the system of record, the related tools, the integration boundaries). Neither artifact has to be polished. The point is to surface the constraint quickly so the call's forty-five minutes are spent on diagnosis, not exposition.

Buyer worksheet

How this field note maps to a real engagement.

The four-question sequence operators run before booking.

Operators who arrive at a diagnosis call having run the sequence arrive ready to decide, which is the point of it. The sequence asks four questions in a specific order. First, is the leading constraint actually addressable with AI, or is it a process problem, a staffing problem, or a stack problem that AI would not solve. Second, if AI is the right intervention, is the right buying motion a custom commission, an off-the-shelf product, or an internal hire. Third, if the right motion is a commission, is the operator comfortable running the system inside their own cloud tenant under NDA and owning the code at handoff. Fourth, is the budget runway for a $45K to $180K fixed fee real this quarter.

Operators who answer yes to all four book the call. Operators who answer no to any one of them either change the question (the leading constraint is different, the budget moves, the cloud posture changes) or take a different path. We do not push operators who land at a "no" on any of the four into a commission they will not be served by.

The three signals operators watch for after handoff.

Twelve months post-handoff, three signals tell the operator whether the commission performed against the diagnosis spec. First, the dollar or hour delta on the workflow the commission addressed, measured against the pre-engagement baseline. Second, the percentage of the workflow the AI layer now handles autonomously versus the percentage that still routes to a human reviewer. Third, the number of times the operator's team has modified the build's prompts, models, or integration code on their own without ColabContent involvement. All three should be improving over time. If they are not, an optional post-handoff stewardship engagement is the lever for diagnosing what changed.

The honest comparison against the alternatives.

A commission is not the right answer for every operator. The mid-market operator with a workflow that matches a horizontal SaaS product's calibration target is better served by the product. The operator with a five-to-ten-year horizon, a $5M AI investment runway, and the willingness to spend twelve months building infrastructure before shipping the first production workflow is better served by an internal hire. The operator at $500M-plus revenue with stakeholder counts that justify a Big Four engagement is better served by that motion. We will tell the operator which of those alternatives fits if a commission does not.

The honest case for a commission is narrow on purpose. Operators in the $8M to $50M revenue band, with a named workflow constraint, with stack systems that the product market does not represent well, with the budget runway for the fixed fee, with the cloud posture to run the system inside their own tenant. Operators in that narrow band are where the math works.

Why we publish the comparisons, the rankings, and the boundaries.

Most consulting houses do not publish ranked comparisons against their competitors, do not publish the boundary of what they will not build, and do not publish fixed-fee pricing bands. We publish all three because the operators we want to commission for are the operators who reward that transparency with a faster booking. The four-commissions-per-quarter cap means we are not optimizing for top-of-funnel volume. We are optimizing for the right four operators each quarter. Publishing the comparisons, the rankings, and the boundaries selects for those operators.

Run the teardown first.

Free, 8 minutes, partner-to-partner. Where your chargeable hours actually go.