Why the bids came back incomparable, and whose fault that is.
Put "we want an AI chatbot for customer service" in front of four firms and you will get four different projects. The conversational-AI shop returns a customer-facing deployment with an intent model and a content program behind it. The integration consultancy returns a middleware effort, because it read the same sentence and saw the ticketing system, the CRM, and the knowledge base that would all have to talk to each other. The staffing firm returns two engineers for six months and no opinion about scope at all. Somebody returns a fixed-fee wrapper around an off-the-shelf product with a licence attached.
None of them is being dishonest. Each one scoped the work, because the buyer did not, and each scoped it toward the thing they already know how to sell. That is the rational response to an ambiguous brief. If the request describes a technology rather than a decision, a vendor has no choice but to supply the missing half from its own catalogue, and the catalogue is different at every firm you sent it to.
What the buyer is holding at the end of that is not a set of competing prices. It is four different products, four different sets of assumptions, and a spread that measures how much interpretation was handed away rather than how much the work is worth. Choosing the cheapest one is choosing the vendor who assumed the least work, which is the same thing as choosing the vendor who understood the least. Choosing the most expensive one is choosing the firm whose sales process was most thorough about inventing scope. Neither is a decision about capability, though both feel like one at the time, and this is a large part of why choosing an AI consultant tends to come down to whichever presentation was most reassuring.
The fix is not a longer document. It is a document that specifies a decision the business needs made, states what that decision costs today, and defines what a better version of it would look like. Given those three things, vendors can only differ on approach and price, and difference of approach between two competent firms is genuinely useful information. That is the entire purpose of asking more than one.
What actually belongs in the document, section by section.
Seven sections do the work. Most of them are short, and one of them is the whole ballgame.
One named workflow. Not a department, not a goal, not a category. "Reduce administrative burden in claims" is a goal. "Intake of a new claim from the moment the email arrives to the moment it is queued for an adjuster" is a workflow. The difference is that the second one can be watched, timed, and priced, and the first one cannot. If you cannot describe the process as a sequence of things a specific person does on a specific afternoon, you have not narrowed enough yet. Naming one workflow is also the point at which the data question becomes answerable rather than existential, since data readiness is a property of a workflow rather than of a company.
The current-state baseline, in real units. How many times a week does this run. How long does each instance take. Who does it, at roughly what loaded cost. What is the error rate, if anyone tracks it, and what does an error cost when it happens. This section is worth more than anything else in the document, and most RFPs omit it entirely. It converts the project from a matter of taste into a matter of arithmetic, it tells every bidder what the ceiling on value is so nobody wastes a month proposing a system worth more than the problem, and it gives you the before number you will need at the review nobody has scheduled yet. If you cannot fill it in, that is not a reason to leave it blank. It is the first piece of work.
The success bar, written before bids arrive. Decide now what result would justify the spend, and write it down while you still have no vendor in the room. Once proposals are in front of you, the metric quietly drifts toward whatever the most persuasive approach happens to be good at, and nobody in the meeting will notice it happening. A bar written in advance is also a bar you can hold a vendor to, which is a different relationship from one where success is negotiated after delivery.
What it must integrate with, named by system. Write the product names and the versions. Not "our CRM" but the actual platform, the actual edition, and whether you have API access on your current tier. Integration surface is the single largest driver of variance in build estimates, and a vendor who has to guess at it will either pad heavily or bid low and discover the problem in week three.
The data the vendor will get, and in what form. State what you can hand over, how it will arrive, and when. A promise of an export you have not confirmed anyone can produce is worse than admitting you do not know, because the vendor prices against your promise and then the schedule absorbs the difference.
Ownership and handoff terms. State who holds the code, the model configurations, the prompts, and the cloud accounts on the last day of the engagement. State what documentation and knowledge transfer you expect, and who can maintain the system afterward. Vendors price differently depending on whether they expect to hold the asset, and leaving this to the contract stage means discovering after selection that the low bid assumed permanent hosting on their infrastructure. The full case for settling this at scope time rather than at renewal time is in why code handoff matters.
Real constraints, separated from preferences. A compliance obligation, a union agreement about how work is assigned, a system you are contractually barred from modifying: these are real, they change what is buildable, and they belong at the front. A preference that the interface look a certain way, or that the vendor use a stack your IT contractor is comfortable with, is a preference. Mixing the two teaches vendors to treat all of your constraints as soft, and the one that mattered gets negotiated away alongside the ones that did not.
What to leave out on purpose, starting with the technology.
The instinct that ruins otherwise good scopes is the urge to demonstrate technical fluency. Somebody on the team has done the reading, and the document acquires a paragraph specifying a model family, a vector database, an agentic architecture, or a particular orchestration framework. It feels like rigor. It is the most expensive paragraph in the file.
You are buying an outcome, and the mechanism is precisely the part you are paying an expert to choose. Specifying it yourself replaces professional judgment with a guess made by whoever on your team read the most about AI in the last quarter, and it does so at the exact moment you had access to four firms who think about this full time. Worse, it hands a weak vendor a defence. If the spec named an approach and the result underperforms, the vendor delivered what was asked for, and the failure belongs to your document. Buyers who write mechanism into a scope routinely find themselves paying to build the wrong thing correctly.
Leave out the model, the architecture, the framework, and the hosting arrangement unless a real constraint forces one of them, in which case it belongs in the constraints section with the reason attached. Leave out the interface design, unless the interface is the deliverable. Leave out the implementation timeline you invented, and ask each vendor for theirs instead, because a proposed schedule is diagnostic: a firm that has done this before will tell you which phase is longest and why, and the answer is rarely the one buyers expect from how long a mid-market AI build actually takes.
What you gain by leaving these out is the disagreement. Two competent vendors proposing different approaches to the same named workflow is the most informative thing the whole process will produce, and it only happens if you left them room to differ.
Making the bids comparable, which is the only reason to ask more than one.
Three requirements do nearly all of it. First, every vendor prices the same named workflow. If one proposal covers intake and another covers intake plus triage plus reporting, you are not comparing prices, you are comparing ambitions. Vendors will want to expand scope, and the expansions are often good ideas; ask for them as separately priced options after the base bid, so the base numbers stay side by side.
Second, require assumptions in writing. Every bid rests on assumptions about data quality, access, internal availability, and how many exceptions the process really has. The cheap bid is usually cheap because it assumed something the expensive bid refused to assume, and once both sets are on the page you can judge which assumptions are actually true about your business. That is a question you can answer and a vendor cannot.
Third, separate build cost from run cost, and break run cost into hosting, model usage, monitoring, and support. A low build price attached to an open-ended monthly figure is the most common way a project ends up over budget without anyone having misled anyone. Ask for the expected monthly run cost at your stated volume and at three times that volume, since the second number is where usage-based pricing surprises people. The structural differences behind these numbers, and why two honest firms quote so differently, are laid out in how AI consulting pricing models work, and the range you should expect at this size of company is covered in what a mid-market AI engagement costs.
Then use the document itself to interview the vendors. Three questions belong in the RFP, and the answers tell you more than any reference call. Ask what would make them decline this project; a firm with real standards has decline criteria and will name them, while a firm that says nothing would make them walk away is telling you they will take anything and sort it out later. Ask what they need from you and when, by name and by week, because a vendor who has done this knows the project usually stalls on the client side and will tell you so before you sign rather than after. And ask what happens in month seven, once the build is delivered, the champion has moved on, and something in an upstream system changes. A vendor without an answer has never been present for that month, which is where a surprising share of these systems quietly stop being used, for reasons collected in why mid-market AI rollouts stall in month four.
When the RFP is the wrong instrument, and how to tell.
Everything above assumes you should be running a formal process. Frequently you should not, and treating an RFP as the responsible default is itself a way to lose a quarter.
A formal process costs weeks. It also changes who bids and how. Long documents reward firms that have a proposal team, which is a different capability from having good engineers, and the best small shop you could hire may not respond at all. Worse, a formal RFP invites vendors to write rather than think. The output becomes a polished artifact produced by people optimizing for a document, and the exchange that actually reveals whether a firm understands your business, the one where someone asks a question you had not considered and you realize your process has an exception you forgot to mention, does not happen through a procurement portal.
For a first engagement on one workflow below roughly six figures, a one-page scope and two conversations with each of two or three vendors will get you a better decision than a twelve-page RFP. Write the workflow, the baseline, the success bar, and the systems it touches on a single page. Send it. Talk twice. The first conversation tells you whether they understand the operation; the second, after they have thought about it, tells you whether they can do the work.
The RFP earns its overhead in four situations. When multiple stakeholders must genuinely agree, and the document is what forces them to converge on one definition of the problem before vendors are involved. When procurement requires a documented competitive process, which is common enough in regulated industries and firms with institutional investors that arguing about it wastes more time than complying. When you have enough genuinely comparable vendors that competitive bidding produces real information rather than the illusion of it. And when the build is large enough that a scoping mistake is expensive to unwind, which is the case that most justifies the weeks.
Outside those four, a formal RFP is procurement theater: it produces a paper trail that looks like diligence, consumes a month, and yields a decision you could have reached faster with a page and two calls. The theater is not free, and the projects that are worth doing rarely get better for waiting on it. If you are unsure which situation you are in, the honest tell is whether you are writing the document to make a decision or to justify one you have already made.
The short version, and the test of whether your scope is ready.
What a buyer can produce in an afternoon: name the one workflow. Write down how often it runs, how long it takes, who does it, and what that costs. Write the success bar before you talk to anyone. List the systems it touches, by product name. Note what data you can hand over and how it will arrive. State who owns the code at the end. Separate the constraints that are real from the ones that are preferences. Then add the three vendor questions and send it. That page will get you comparable bids more reliably than the twelve-page version, and it is short enough that you will actually finish it. If you want to see what the receiving end of that document looks like, the way we structure a build starts with the same diagnostic.
The test for whether the scope is finished has nothing to do with length. Hand it to someone who does not work in that part of the business, a colleague from another department is ideal, and ask them to tell you what the system will do on a Tuesday morning. If they can describe it, who is sitting where, what arrives, what the system does with it, and what a person does next, then a vendor can price it. If they cannot, no amount of additional detail about compliance requirements or technical preferences will fix the document, because the thing missing is the workflow itself. That question takes five minutes to run and it is the only quality check the document really needs.