Home/ Journal/ Memo

Writing the scope for a custom AI build, and why most RFPs get bids you cannot compare.

Most AI RFPs fail before a vendor reads them, because they specify a technology instead of a decision. Asking for an AI chatbot returns five proposals priced on five different assumptions about what is being built, which is why the numbers cannot be compared. A scope that works names one workflow, states the current cost of running it in real units, and defines the outcome that counts as success before any bid arrives. Seven sections carry the document: the named workflow, the current-state baseline, the success bar, the systems it must integrate with named by product, the data the vendor will receive and in what form, the ownership and handoff terms, and the constraints that are genuinely real kept apart from preferences. Leave the technology out on purpose; naming a model or an architecture removes the expertise you are paying for and gives a weak vendor a place to hide behind your own spec. To make bids comparable, require every vendor to price the same workflow, state assumptions explicitly, and separate build cost from run cost. Three questions expose more than any reference check: what would make you decline this project, what do you need from us and when, and what happens in month seven. A formal RFP is also frequently the wrong instrument. For a first workflow under six figures, a one-page scope and two conversations beat a twelve-page document, and the test of whether any scope is ready is whether a stranger can read it and tell you what the system does on a Tuesday morning.

Somewhere between approving the money and signing the contract sits a document that nobody at a mid-market company was ever trained to write. The owner has agreed to a number. The COO knows which part of the operation hurts. Then somebody has to describe the work precisely enough that three or four firms can bid on the same thing, and the usual outcome is a set of proposals with a fourfold price spread, none of which describe the same project. That spread is not the market pricing risk. It is the document doing its job badly, and the cost of it lands on the buyer twice: once in the weeks spent trying to reconcile bids that were never reconcilable, and again later, when the winning vendor builds the version they scoped rather than the one that was needed. This memo covers what belongs in the scope, what to deliberately leave out, how to force comparability, and the cases where writing a formal RFP at all is the wrong move.

MemoJuly 2026
Read time11 minutes
AudienceOwner-CEOs, COOs, CFOs, Managing Partners

Why the bids came back incomparable, and whose fault that is.

Put "we want an AI chatbot for customer service" in front of four firms and you will get four different projects. The conversational-AI shop returns a customer-facing deployment with an intent model and a content program behind it. The integration consultancy returns a middleware effort, because it read the same sentence and saw the ticketing system, the CRM, and the knowledge base that would all have to talk to each other. The staffing firm returns two engineers for six months and no opinion about scope at all. Somebody returns a fixed-fee wrapper around an off-the-shelf product with a licence attached.

None of them is being dishonest. Each one scoped the work, because the buyer did not, and each scoped it toward the thing they already know how to sell. That is the rational response to an ambiguous brief. If the request describes a technology rather than a decision, a vendor has no choice but to supply the missing half from its own catalogue, and the catalogue is different at every firm you sent it to.

What the buyer is holding at the end of that is not a set of competing prices. It is four different products, four different sets of assumptions, and a spread that measures how much interpretation was handed away rather than how much the work is worth. Choosing the cheapest one is choosing the vendor who assumed the least work, which is the same thing as choosing the vendor who understood the least. Choosing the most expensive one is choosing the firm whose sales process was most thorough about inventing scope. Neither is a decision about capability, though both feel like one at the time, and this is a large part of why choosing an AI consultant tends to come down to whichever presentation was most reassuring.

The fix is not a longer document. It is a document that specifies a decision the business needs made, states what that decision costs today, and defines what a better version of it would look like. Given those three things, vendors can only differ on approach and price, and difference of approach between two competent firms is genuinely useful information. That is the entire purpose of asking more than one.

What actually belongs in the document, section by section.

Seven sections do the work. Most of them are short, and one of them is the whole ballgame.

One named workflow. Not a department, not a goal, not a category. "Reduce administrative burden in claims" is a goal. "Intake of a new claim from the moment the email arrives to the moment it is queued for an adjuster" is a workflow. The difference is that the second one can be watched, timed, and priced, and the first one cannot. If you cannot describe the process as a sequence of things a specific person does on a specific afternoon, you have not narrowed enough yet. Naming one workflow is also the point at which the data question becomes answerable rather than existential, since data readiness is a property of a workflow rather than of a company.

The current-state baseline, in real units. How many times a week does this run. How long does each instance take. Who does it, at roughly what loaded cost. What is the error rate, if anyone tracks it, and what does an error cost when it happens. This section is worth more than anything else in the document, and most RFPs omit it entirely. It converts the project from a matter of taste into a matter of arithmetic, it tells every bidder what the ceiling on value is so nobody wastes a month proposing a system worth more than the problem, and it gives you the before number you will need at the review nobody has scheduled yet. If you cannot fill it in, that is not a reason to leave it blank. It is the first piece of work.

The success bar, written before bids arrive. Decide now what result would justify the spend, and write it down while you still have no vendor in the room. Once proposals are in front of you, the metric quietly drifts toward whatever the most persuasive approach happens to be good at, and nobody in the meeting will notice it happening. A bar written in advance is also a bar you can hold a vendor to, which is a different relationship from one where success is negotiated after delivery.

What it must integrate with, named by system. Write the product names and the versions. Not "our CRM" but the actual platform, the actual edition, and whether you have API access on your current tier. Integration surface is the single largest driver of variance in build estimates, and a vendor who has to guess at it will either pad heavily or bid low and discover the problem in week three.

The data the vendor will get, and in what form. State what you can hand over, how it will arrive, and when. A promise of an export you have not confirmed anyone can produce is worse than admitting you do not know, because the vendor prices against your promise and then the schedule absorbs the difference.

Ownership and handoff terms. State who holds the code, the model configurations, the prompts, and the cloud accounts on the last day of the engagement. State what documentation and knowledge transfer you expect, and who can maintain the system afterward. Vendors price differently depending on whether they expect to hold the asset, and leaving this to the contract stage means discovering after selection that the low bid assumed permanent hosting on their infrastructure. The full case for settling this at scope time rather than at renewal time is in why code handoff matters.

Real constraints, separated from preferences. A compliance obligation, a union agreement about how work is assigned, a system you are contractually barred from modifying: these are real, they change what is buildable, and they belong at the front. A preference that the interface look a certain way, or that the vendor use a stack your IT contractor is comfortable with, is a preference. Mixing the two teaches vendors to treat all of your constraints as soft, and the one that mattered gets negotiated away alongside the ones that did not.

What to leave out on purpose, starting with the technology.

The instinct that ruins otherwise good scopes is the urge to demonstrate technical fluency. Somebody on the team has done the reading, and the document acquires a paragraph specifying a model family, a vector database, an agentic architecture, or a particular orchestration framework. It feels like rigor. It is the most expensive paragraph in the file.

You are buying an outcome, and the mechanism is precisely the part you are paying an expert to choose. Specifying it yourself replaces professional judgment with a guess made by whoever on your team read the most about AI in the last quarter, and it does so at the exact moment you had access to four firms who think about this full time. Worse, it hands a weak vendor a defence. If the spec named an approach and the result underperforms, the vendor delivered what was asked for, and the failure belongs to your document. Buyers who write mechanism into a scope routinely find themselves paying to build the wrong thing correctly.

Leave out the model, the architecture, the framework, and the hosting arrangement unless a real constraint forces one of them, in which case it belongs in the constraints section with the reason attached. Leave out the interface design, unless the interface is the deliverable. Leave out the implementation timeline you invented, and ask each vendor for theirs instead, because a proposed schedule is diagnostic: a firm that has done this before will tell you which phase is longest and why, and the answer is rarely the one buyers expect from how long a mid-market AI build actually takes.

What you gain by leaving these out is the disagreement. Two competent vendors proposing different approaches to the same named workflow is the most informative thing the whole process will produce, and it only happens if you left them room to differ.

Making the bids comparable, which is the only reason to ask more than one.

Three requirements do nearly all of it. First, every vendor prices the same named workflow. If one proposal covers intake and another covers intake plus triage plus reporting, you are not comparing prices, you are comparing ambitions. Vendors will want to expand scope, and the expansions are often good ideas; ask for them as separately priced options after the base bid, so the base numbers stay side by side.

Second, require assumptions in writing. Every bid rests on assumptions about data quality, access, internal availability, and how many exceptions the process really has. The cheap bid is usually cheap because it assumed something the expensive bid refused to assume, and once both sets are on the page you can judge which assumptions are actually true about your business. That is a question you can answer and a vendor cannot.

Third, separate build cost from run cost, and break run cost into hosting, model usage, monitoring, and support. A low build price attached to an open-ended monthly figure is the most common way a project ends up over budget without anyone having misled anyone. Ask for the expected monthly run cost at your stated volume and at three times that volume, since the second number is where usage-based pricing surprises people. The structural differences behind these numbers, and why two honest firms quote so differently, are laid out in how AI consulting pricing models work, and the range you should expect at this size of company is covered in what a mid-market AI engagement costs.

Then use the document itself to interview the vendors. Three questions belong in the RFP, and the answers tell you more than any reference call. Ask what would make them decline this project; a firm with real standards has decline criteria and will name them, while a firm that says nothing would make them walk away is telling you they will take anything and sort it out later. Ask what they need from you and when, by name and by week, because a vendor who has done this knows the project usually stalls on the client side and will tell you so before you sign rather than after. And ask what happens in month seven, once the build is delivered, the champion has moved on, and something in an upstream system changes. A vendor without an answer has never been present for that month, which is where a surprising share of these systems quietly stop being used, for reasons collected in why mid-market AI rollouts stall in month four.

When the RFP is the wrong instrument, and how to tell.

Everything above assumes you should be running a formal process. Frequently you should not, and treating an RFP as the responsible default is itself a way to lose a quarter.

A formal process costs weeks. It also changes who bids and how. Long documents reward firms that have a proposal team, which is a different capability from having good engineers, and the best small shop you could hire may not respond at all. Worse, a formal RFP invites vendors to write rather than think. The output becomes a polished artifact produced by people optimizing for a document, and the exchange that actually reveals whether a firm understands your business, the one where someone asks a question you had not considered and you realize your process has an exception you forgot to mention, does not happen through a procurement portal.

For a first engagement on one workflow below roughly six figures, a one-page scope and two conversations with each of two or three vendors will get you a better decision than a twelve-page RFP. Write the workflow, the baseline, the success bar, and the systems it touches on a single page. Send it. Talk twice. The first conversation tells you whether they understand the operation; the second, after they have thought about it, tells you whether they can do the work.

The RFP earns its overhead in four situations. When multiple stakeholders must genuinely agree, and the document is what forces them to converge on one definition of the problem before vendors are involved. When procurement requires a documented competitive process, which is common enough in regulated industries and firms with institutional investors that arguing about it wastes more time than complying. When you have enough genuinely comparable vendors that competitive bidding produces real information rather than the illusion of it. And when the build is large enough that a scoping mistake is expensive to unwind, which is the case that most justifies the weeks.

Outside those four, a formal RFP is procurement theater: it produces a paper trail that looks like diligence, consumes a month, and yields a decision you could have reached faster with a page and two calls. The theater is not free, and the projects that are worth doing rarely get better for waiting on it. If you are unsure which situation you are in, the honest tell is whether you are writing the document to make a decision or to justify one you have already made.

The short version, and the test of whether your scope is ready.

What a buyer can produce in an afternoon: name the one workflow. Write down how often it runs, how long it takes, who does it, and what that costs. Write the success bar before you talk to anyone. List the systems it touches, by product name. Note what data you can hand over and how it will arrive. State who owns the code at the end. Separate the constraints that are real from the ones that are preferences. Then add the three vendor questions and send it. That page will get you comparable bids more reliably than the twelve-page version, and it is short enough that you will actually finish it. If you want to see what the receiving end of that document looks like, the way we structure a build starts with the same diagnostic.

The test for whether the scope is finished has nothing to do with length. Hand it to someone who does not work in that part of the business, a colleague from another department is ideal, and ask them to tell you what the system will do on a Tuesday morning. If they can describe it, who is sitting where, what arrives, what the system does with it, and what a person does next, then a vendor can price it. If they cannot, no amount of additional detail about compliance requirements or technical preferences will fix the document, because the thing missing is the workflow itself. That question takes five minutes to run and it is the only quality check the document really needs.

Field-note context

What we notice when a scope document arrives.

The length of the document is inversely related to the clarity of the problem.

The strongest briefs we receive are short. They name a process, say how often it happens and what it costs, and stop. The weakest ones run to twenty pages of requirements language, security appendices, and a matrix of desirable features, and somewhere in the middle there is a single sentence describing what the business actually wants to happen differently. Length usually accumulates because the buyer was unsure what to specify and compensated by specifying everything, which spreads the signal so thin that vendors end up bidding on the appendices. If a document is long, the first thing worth doing is finding the one sentence that matters and checking whether it is true.

The exceptions are the project, and they are never in the RFP.

A scope describes the process as it is supposed to run. The work is in the cases where it does not: the accounts that get manual handling, the orders that skip a step because a particular customer negotiated it years ago, the approval that gets bypassed when the person is on vacation. Those exceptions are what separate a two-month build from a five-month one, and they almost never appear in the document, because the person writing it describes the official process while the person doing the job knows the real one. Asking the operator directly what percentage of cases go the standard way is one of the fastest ways to find out what the project really is, and the honest answer is often lower than management expects.

A missing baseline predicts an argument at review.

When a scope arrives with no current-state numbers, it tells us two things. The workflow has probably never been measured, which means the improvement will be argued about rather than demonstrated. And the buyer has not yet decided what would count as success, which means the definition will get set retroactively by whoever is most invested in the outcome. Both are fixable in a couple of weeks by someone internal with a stopwatch and a spreadsheet, and doing that work before the bids arrive changes the entire tone of the engagement, because everyone involved is then negotiating against the same number instead of against each other's impressions.

Extended questions

The questions operators ask about scoping a build.

What should an RFP for a custom AI build actually contain?

Seven things, and the technology is not one of them. One named workflow rather than a department or a goal, described at the level of what a person does on a Tuesday. A current-state baseline in real units: how many times a week it runs, how long each run takes, who does it, and what that costs today. A success bar you wrote before any bid arrived, so nobody talks you into a metric their approach happens to hit. The systems the build must integrate with, named by product and version rather than described by category. The data you will hand over, in what form, and when. Ownership and handoff terms covering who holds the code and the accounts on the last day. And the constraints that are genuinely real, such as a compliance obligation, a union agreement, or a platform you are contractually barred from touching, kept separate from the things that are only preferences.

Why do AI vendor bids come back so different from each other?

Because the buyer left the scope open, so each vendor scoped it themselves, and every vendor scopes toward what they already sell. A request for an AI chatbot reaches a firm that sells conversational front ends and comes back as a customer-facing deployment. The same request reaches an integration shop and comes back as a middleware project. It reaches a staffing firm and comes back as two engineers for six months. None of them is being dishonest. They are all answering the question you asked, which was a question about a technology rather than about a decision your business needs made. The price spread you are looking at is not a market signal about value. It is a measurement of how much interpretation you handed away, and it is the reason comparing those numbers side by side tells you almost nothing.

Should I specify the technology or model in my AI RFP?

No, and specifying it costs you twice. First, you are buying an outcome, and the mechanism is the part you are paying an expert to choose; naming the model or the architecture yourself removes the judgment you are hiring for and replaces it with a guess made by whoever on your team read the most about AI recently. Second, a stated mechanism gives a weak vendor somewhere to hide. If the spec says to use a particular approach and the result underperforms, the vendor delivered exactly what you asked for, and the failure is now yours on paper. Describe the decision, the inputs available, the outcome that counts as success, and the constraints that are real. Let the vendors differ on approach, then judge the approaches. The disagreement between two competent vendors about how to solve it is the most useful information the process will produce.

How do I make competing AI proposals comparable?

Force three things and comparability mostly takes care of itself. Require every vendor to price the same named workflow, so that when one bid is triple another you know it is a difference in approach rather than a difference in what each firm decided to build. Require assumptions to be stated explicitly and in writing, because the cheap bid is usually cheap on an assumption the expensive one refused to make, and once both are on the page you can see which assumptions are true about your business. And require build cost and run cost to be separated, with run cost broken into hosting, model usage, monitoring, and support, since a low build price attached to an open-ended monthly figure is the most common way a project comes in over budget without anyone having lied about anything.

Do I even need a formal RFP for a first AI project?

Often not, and running one anyway is a real cost rather than a neutral precaution. A formal process adds weeks, invites vendors to write proposals rather than think about your problem, and rewards firms with a proposal team over firms with good engineers. For a first workflow under roughly six figures, a one-page scope plus two conversations with each of two or three vendors will get you further, because the useful information comes out when someone asks a question you had not considered, and that exchange does not happen through a procurement portal. A formal RFP earns its overhead when several stakeholders must genuinely agree, when procurement requires a documented process, when you have enough comparable vendors for competitive bidding to mean something, or when the build is large enough that a scoping mistake is expensive to unwind. Outside those conditions it is usually procurement theater.

How does the AI Maturity Index help me write the scope?

It produces the one input the whole document depends on, which is a single named workflow instead of a category. The Index asks you to identify the process worth investing in first and to size who touches it and how often, and those two answers are most of a current-state baseline. It also surfaces where that workflow's inputs live and whether anyone is measuring the process today, which is the section buyers most often leave blank and most often regret leaving blank at review time. In about ten minutes and with no call, you get enough structure that the scope becomes an afternoon of writing rather than a month of internal meetings about what to ask for.

Not sure which workflow your scope should actually name?

Start with the AI Maturity Index. Ten minutes, no call, and it names the one workflow worth investing in first, sizes who touches it and how often, and gives you a specific enough target that you can write the scope yourself before anyone quotes you.