The pilot that could not have failed, and could not have succeeded either.
Run one test on a pilot proposal before funding it. Ask what result would have made the company stop, the finding or the number that would have caused everyone to walk away for a year. If nobody in the room can produce one, the pilot is not an experiment, and no amount of execution converts it into one. It will run, people will form impressions, and at the end those impressions get sorted by whoever is most senior in the meeting. That is not information. It is a vote taken slowly and at expense.
It happens because pilots get scoped as demonstrations rather than as questions. Somebody wants to show the leadership team what the technology can do, which is reasonable, and a demonstration is designed to succeed. A question is designed to be answerable in either direction. The two look identical for three weeks and diverge completely at the end, because a demonstration has no failure state and cannot settle anything. It produces a next step, and the next step is usually another demonstration.
What a pilot should produce instead is one of three sentences with a name attached. We are commissioning a build of this workflow. We are buying this product for this team. We are dropping this and looking again in twelve months. Anything else is a deferral.
This matters more at $8M to $50M than it does further up. A large company runs a portfolio of pilots and budgets for most of them returning nothing. A company your size runs one at a time, and the currency it spends is not the license fee. It is six weeks of attention from the one operations person who could also have been fixing the scheduling problem. Those weeks are the scarce thing, and a pilot that cannot end in a decision spends them for nothing.
Write the decision, the decider, and the date, before you look at a single tool.
Three lines on one page, before a vendor is contacted. The first is the decision the pilot exists to make, stated as a choice between named options, not a topic. Should we use AI in service is a topic. Do we commission a build for first-response drafting, buy the module our help desk already sells, or keep the current process for another year is a decision that can be answered. Naming two live options keeps it honest; a single-option pilot is a yes-or-nothing vote, and the nothing rarely gets a hearing.
Putting build and buy on the page also front-loads a choice most companies make by accident, and it has a logic of its own, worked through in the build versus buy decision for a mid-market company. Settling it in the charter costs an hour. Discovering it after a successful pilot on the wrong instrument costs a rebuild.
The second line is the decider, and it is one name. Not the leadership team, not a steering group, not whoever happens to be in the room. That person needs two authorities that usually travel together: to spend what a yes would cost, and to stop the effort over the objection of whoever championed it. Shared ownership here reliably manufactures a deferral, because a group that cannot agree adjourns, and adjourning feels like diligence.
The third line is a date, and it does more work than the two lines above it. Put a meeting on the calendar the day the pilot starts, with the decider in it, four to eight weeks out. A pilot with no scheduled decision meeting does not end; it fades, which is a more common failure than a pilot that fails. Keep the box short enough that everyone involved can hold the whole thing in their head. Builds run on a longer clock, which is the subject of how long a mid-market AI build actually takes.
Then name who does the work during the pilot, by person rather than department, and say what comes off their plate to make room. The pilot's real cost is that person's attention, and an unstaffed pilot is the ordinary way a decent idea dies. It is the same neglect that strands a project halfway through, described in why mid-market AI rollouts stall in month four. A pilot assigned to somebody already at capacity has produced its result already; it will take three months to report it.
Pre-register the bar, and the before it gets measured against.
A success bar written after the results arrive is not a bar. It is a description of the results. It is not a character flaw; it is what happens to any threshold still negotiable when the number lands. The team that hoped for half the time and got twelve percent will find a reason twelve is encouraging. The skeptic will find a reason half is disappointing. Both readings are available because nobody committed to a number while the answer was unknown.
So write two numbers instead of one. The first is the threshold at which you commission or buy; the second is the level below which you drop it. Each needs a metric, a direction, a value, and a stated moment of measurement, and the decider agrees both before the first result exists. Two rather than one because the gap between them is the ambiguous zone, and naming it in advance keeps you from arguing later about whether a middling result counts.
The bar is a comparison, so it needs a before, and the before has to be captured on the process you have right now, while the new tool does not exist. It cannot be recovered afterward; once the new process is running, whatever you reconstruct is a memory. How to count it properly is laid out in how to measure the return on a mid-market AI engagement.
Plenty of mid-market workflows leave no trace to measure, which is inconvenient rather than disqualifying. Two weeks of manual capture usually settles it: the people doing the task note when it started, when it finished, and whether the output came back for a fix. That is the same question as whether a historical record exists at all, one of the checks in data readiness for one named workflow. If two weeks of counting is more than the organization will do, learn that before you spend anything.
Both outcomes have to be actionable, or what you are running is theater.
Before funding anything, ask the decider what happens on the Monday after a clean negative. Not what they would conclude. What they would do, in the calendar week after the meeting. If the honest answer is that we would keep looking at it, or try a different vendor, or nothing in particular would change, the pilot has no negative outcome and therefore has no outcome. Do not run it.
A real negative action is specific and slightly uncomfortable. We cancel the trial and tell the team the current process stands for twelve months. We take the initiative off the operating plan. We write down the two conditions that would make us look again and the date we check them. Each is something a person does and other people notice.
The positive side deserves the same test and fails it more often than people expect. A yes with no budget behind it and no slot in the next quarter is the same non-decision as a no with no action. If a successful pilot would have to wait for the next planning cycle anyway, run the pilot then instead. Nothing decays faster than a proven result waiting on a budget conversation.
The pilot that ends in neither action is where the lasting damage sits. Inside the company it does not read as a null result; it reads as evidence the idea did not work, and that hardens into a belief that this business is not the kind of business AI helps. Reversing it is most of the work described in what to do after a failed AI pilot. A pilot built to force one of two actions cannot leave that residue, because whichever way it lands, something happened.
The comparison you can actually run, at the volume you actually have.
The clean instrument is a randomized split, and you almost certainly cannot run one. A company producing sixty quotes a month and splitting them thirty and thirty can detect a large difference and nothing else, and the moment you split by person instead of by item you have also split by skill, tenure, and territory. Insisting on the textbook design at this size gets you a pilot that runs for a year or a comparison that looks rigorous and is not.
Three weaker designs are honest and workable. Same-person before and after uses each participant's own prior period as the control, which removes the between-people variation that ruins small splits. The paired sample takes fifty real cases already completed the old way, runs them through the new process, and compares outputs on identical inputs, which can be done before anyone changes how they work. Alternating periods, one week new and one week old, spreads out anything seasonal.
Then say out loud what your design cannot rule out. Same-person before and after is confounded by learning, by seasonality, and by the fact that people work differently when they know they are being watched. Paired samples miss how the process behaves under real time pressure. Alternating weeks assumes the weeks are comparable, and often they are not. None of that makes the design useless. It is a reason to demand a large effect rather than a small one.
That is the discipline hiding inside the volume problem. At your size only large effects are detectable, so only large effects are decidable, and a pilot aimed at a five percent improvement is a way to generate a number you will argue about. Pick the workflow where you expect the difference to be obvious to anybody looking at it. If you need statistics to see whether it worked, the workflow was the wrong choice before the tool was.
Enthusiasm is the most common false positive, and the hardest one to argue with.
The usual way a mid-market pilot produces a wrong yes is that everybody liked it. The team was engaged, the sessions went well, people said it saved them time, and the decider arrives at the meeting with a strong positive impression and no number that moved. That is a real finding about the mood in the room and nothing at all about the workflow, and it converts into a six-figure commitment more often than a flat demo kills a good idea.
The mechanics are ordinary. Participants usually volunteered, which selects for people already inclined to like it. They know they are being watched. Telling the owner that the initiative they funded is not working is socially expensive and almost never anybody's job. And novelty genuinely feels like improvement for a few weeks, which is why a four-week pilot flatters a tool and an eight-week one stops doing so.
Collect the sentiment anyway, because a tool nobody will use is worthless whatever it does. Just never let it occupy the slot where the number goes. If the bar was missed and the mood was good, the honest sentence is that the tool is pleasant and did not work. The inverse is better news than it looks: a number that moved with a team that disliked the experience is an adoption problem, which is a training question rather than a verdict.
Then there is the genuinely ambiguous result, where the number lands between your two thresholds. The instinct is to expand: more users, longer runway, see whether it improves at scale. That instinct is almost always wrong, because expansion converts a cheap unresolved question into an expensive one and adds variables to a picture that was already unclear. Run it once more instead, same box or shorter, with one thing changed. Usually the ambiguity traces to something specific, such as two of the five participants never really using it.
Four situations where a pilot is the wrong instrument, including the one nobody says out loud.
We are hired to design and build custom AI systems, and diagnosis sits at the front of that funnel, so read this section with that in mind. In four situations a pilot is worse than no pilot, and in all four the diagnosis fee is money you should not spend. The first is the workflow small enough to simply fix. If the whole thing is two people and four hours a week, the pilot costs more attention than the problem does.
The second is the tool that is cheap and reversible. A per-seat subscription you can cancel at the end of the month does not need an instrument; trying it is cheaper than studying it. Piloting it means spending three months of committee attention to avoid spending two hundred dollars. Reserve the apparatus for decisions that are expensive to unwind: a custom build, a platform migration, a change to how a team works.
The third is the pilot that exists as cover for a decision already made. Somebody has decided, and the pilot is there to make the decision look considered, or to give a dissenting party the feeling of having been consulted. It costs what a real pilot costs and returns nothing, and everyone involved learns the answer was known in advance and treats the next pilot the same way.
The fourth is the company with nobody to run it. A pilot needs one competent person with real hours for the duration, somebody who will chase the participant who stopped using it in week two and notice when the measurement stops being recorded. Without that, the pilot does not fail cleanly. It decays, produces a fog of partial data, and gets remembered as proof that the technology did not work. Solve the staffing question first, or shrink the decision until no pilot is required.
What to write down this week, before anybody gets a call.
The whole design fits on one page and takes an afternoon, and none of it involves a vendor. Write the decision as a choice between at least two named options, one of which is keeping the current process. Write the decider, one name. Write the date of the decision meeting and put it on that person's calendar now; a decision nobody can schedule is a decision nobody owns yet. Write the two numbers, the one that means yes and the one that means no, and when each gets measured.
Then do the part that catches most of the errors. Read the page to the person who would have to do the work during the pilot and ask what would come off their plate. Read it to the person who would have to act on a no and ask what they would actually do that week. If either conversation produces a shrug, you have found the defect while it is still free to fix.
Naming the workflow the page is about is the one prerequisite, and it is where most operators stall. The diagnosis we run starts there. If you would rather sort it yourself first, the AI Maturity Index gets you to a single named workflow, who touches it and how often, and whether its output is something a person reads or something written into a record. That last fact is what the bar and the comparison design both turn on.