Home/ Journal/ Memo

Security questions before an AI build, and the copies nobody counts.

The security question most buyers ask first, whether the model provider trains on company data, has the shortest answer and the smallest consequence, because it is settled in a document that takes a minute to read. The exposure that matters sits in the chain underneath it: the vendor's own staff and subcontractors, the vendor's subprocessors, the logging layer that stores prompts and outputs so engineers can debug them, and the copies of real records that accumulate in development environments and evaluation sets. So the first question is a count. From your system of record to the answer on someone's screen, how many places does a record come to rest, and who controls each one. The second question sizes how hard you press, and it turns on what the system can do rather than what it can see, because a read-only search over internal documents and a system with write access to a record of account are different risk objects that most diligence treats identically. Where the data carries an obligation you owe someone else, client financials, health information, privileged material, employee records, the question is which layer touches it and under whose contract. Prefer every question whose answer is an artifact you can hold, a subprocessor list, a retention setting, an access export, a clause, over any question a vendor can answer with reassurance. And settle deletion before you sign, with the copies named.

Commissioning a custom AI build hands a new company standing access to the records that carry your legal obligations: the client files, the payroll data, the contracts, the material you hold on someone else's behalf. The cost of getting that wrong is rarely the breach an operator pictures. It is a notification you owe a client, a confidentiality term you did not realize you were subcontracting, an insurer asking who else had access, and a build you have to unwind at the point where it was starting to work. Most of that is avoidable at the diligence stage, and almost none of it is avoided by the question buyers usually lead with.

MemoJuly 2026
Read time11 minutes
AudienceOwner-CEOs, COOs, Operations Leaders

The question everyone asks first, and the safest link in the chain.

Whether the model provider trains on your data is the question that opens almost every one of these conversations, and it deserves about four minutes of attention. The answer is written down. Model providers publish their terms for business use, and the vendor building your system either passes those terms through to you in writing or does not. Read both, restate the answer in your own contract, and the question is closed. It is a real question with a checkable answer, which is exactly what makes it the wrong place to spend the leverage you have before signing.

What it hides is that the model provider is one party in a chain of several, and it is the party with the most public scrutiny, the most standardized contract, and the most to lose from handling business data badly. The links with no published policy page are the ones specific to the firm you are about to hire: the engineers who will read your records to build against them, the contractors that firm subcontracts to when a deadline moves, the hosting and monitoring tools their system depends on, and the copies of real data that pile up in staging environments over the course of a build. Nobody publishes a page about any of that, and no buyer asks unprompted.

The general work of separating a serious partner from a reseller is a different exercise, worked through in how to choose an AI consultant for a mid-market company, and one paragraph of it says to ask where your data lives during the build and who can see it. This memo is that paragraph opened up, because for a company holding other people's records it is the part of the diligence with the longest tail and the least written about it.

Count the copies, and name who controls each one.

The exercise that reorganizes everything else takes twenty minutes and no technical knowledge. Take one record, a client file, a claim, an invoice, and trace it from your system of record to the answer that appears on somebody's screen. Write down every place it comes to rest along the way, and write the controlling party next to each one. The list is longer than the architecture diagram suggests, and the length of it is the exposure.

A typical custom build produces six or seven resting places. There is the application database the system runs on. There is a vector store holding embeddings generated from your documents, which is a copy, whatever the word embedding suggests, because the passages it points back to are your text. There is the logging and trace layer that stores the prompt, the retrieved passage, and the output, which exists because nobody can debug a wrong answer without it. There are backups behind all of those, on a rotation nobody has looked at. There are the development and staging environments, where a slice of real production data was pulled early so the team could build against reality rather than against invented rows. There is the evaluation set. And there is whatever gets pasted into the vendor's own ticketing system the first time somebody debugs a case with a colleague.

Each of those sits under a different party, a different contract, and a different retention default. If you cannot name the controlling party for a row, that row is your first question for the vendor. This is a separate question from whether the data is any good, which is the subject of data readiness for one named workflow. Usable and safe are different properties, and a company can pass one and fail the other, because the same records are being examined for two different reasons.

Size the diligence by what the system can do, not by what it can see.

Most vendor security questionnaires ask the same questions of every project, which over-taxes the harmless system and under-taxes the dangerous one. The variable that should set the depth is not the sensitivity of what the system reads. It is what the system is permitted to do, because that decides what a mistake costs and whether it can be reversed.

Three tiers cover almost everything a mid-market company commissions. A read-only system that retrieves from internal documents and answers a person can see everything and change nothing, so its worst failure is disclosure to your own staff of something they were not supposed to see, which is a real problem with a bounded shape. A system with write access to a record of account can change what the company believes to be true about a customer, a balance, a schedule, or a case, and its worst failure is a change nobody notices for a month. A system that sends anything outside the company, an email, a filing, a quote, a message to a counterparty, has a failure mode that cannot be recalled at all.

The asymmetry runs against intuition. The read-only system often touches more sensitive material than the write-capable one, since pointing a retrieval layer at the whole document store is easy and pointing a write integration at one field is deliberate. It is still the lower risk object. For the read-only tier the questions are about copies and access, and they can be settled quickly. For the write tier a different set applies: precisely which fields and records can it change, who approved that scope, is every action logged against an identifiable actor, is there a reversal path, and what threshold sends a case to a person instead. Those are operational questions rather than security ones, which is why they fall through the gaps in a standard questionnaire. The same capability axis decides a good deal about what to commission in the first place, which is the argument in the build versus buy decision for a mid-market company.

Regulated material is a contract question, not a technology question.

For client financials, health information, privileged legal material, and employee records, the question a buyer usually asks is whether it is permitted to use AI on this at all. That framing produces long conversations and no decisions. The answerable version is narrower: which layer of the stack touches this category, and under whose contract. An obligation travels with the data. If you hold records under an engagement letter, a client confidentiality term, a carrier requirement, or a professional duty, that constraint follows the record into every one of the resting places you just listed, and the relevant test is whether each party holding a copy is under an instrument that carries it.

Which makes the subprocessor list the most useful document to request, in writing, before signing. It names the parties underneath your vendor: the hosting provider, the model provider, the trace tool, whatever handles storage or search on the way through. Read it against the agreements you have already signed with your own clients, because some of those restrict disclosure to subcontractors without notice or consent, and a vendor cannot tell you whether you are in breach of a contract they have never read. This is the one place where an hour of your counsel's time is genuinely worth buying, and it is an hour on one page rather than a review of the whole build.

There is also a cheaper move that gets skipped because it feels like a retreat. Scope the regulated category out of version one. If the workflow covers the cases that carry no special obligation and routes the rest to a person, the paperwork shrinks to something an operator can finish in a week, and the build ships while the harder question gets answered properly. That belongs in the scope document rather than in a side conversation, and where it goes is covered in writing the scope for a custom AI build.

Prefer the answer you can hold, because you do not have a CISO.

The person running this diligence at an $8M to $50M company is an owner, a COO, or a finance lead who has eleven other things happening this week. There is no security team to hand it to, and hiring an assessor for a single project usually costs a visible fraction of the build. So the practical value of a question is not how penetrating it sounds. It is how well a non-specialist can judge the answer, which means the best questions are the ones whose answer is an artifact rather than a sentence.

The rewrite is mechanical once you see it. Instead of asking whether the data is encrypted, which invites a yes that tells you nothing, ask for the list of people at their company who can read your production data, by name or role, and how that list is reviewed. Instead of asking whether prompts are retained, ask what the retention is set to today in the tool that stores them and to see the setting. Instead of asking whether they are secure, ask for the subprocessor list. Instead of asking whether they will delete your data, ask which clause says so and what it names. Every one of those answers is a document, a screenshot, a configuration value, or a clause, and every one can be judged by somebody who has never configured anything.

A reassuring answer costs a vendor nothing to give, and it is not evidence of anything except that they have answered the question before. An artifact costs them either the truth or a lie in writing, and firms behave differently when that is the choice. The reaction itself is most of the finding: a vendor who works this way sends the documents over within a day, and one who does not keeps the conversation verbal and warm.

The exit, where the copies you forgot are still sitting.

What you receive when an engagement ends is the ownership question, and it is settled by the contract and confirmed at handoff, which is worked through in owning your AI and what a real code handoff contains. The security question is its mirror image, and it is almost never asked: what does the vendor no longer hold. Those are two different clauses, and a company can get the first one perfect and never think about the second.

A promise to delete your data on request covers the obvious copy and leaves the rest, because whoever performs the deletion will reasonably read it as the production database. Name the others in the agreement: backups, the vector store and the embeddings derived from your documents, logs and traces including whatever the observability tool retains on its own schedule, the evaluation set, every development and staging copy, and any records sitting in the vendor's support tooling from a debugging session. Ask what the subprocessors underneath are obliged to do and on what timeline, because a vendor cannot delete faster than the parties they depend on, and most of them have never checked.

Then ask for confirmation in writing within a stated number of days. None of this is adversarial and none of it is expensive to agree before signing; a firm that intends to comply will treat it as routine. It becomes awkward only in the other order, when you are asking a company you have just left to do unpaid work on your behalf, at the exact moment your leverage is gone.

The honest counter-case, and the diligence that is theater.

We build custom systems for a living, so weigh this accordingly and then take it seriously, because there are several common situations where everything above is a waste of a quarter. The clearest one is a genuinely low-sensitivity internal workflow. If the system reads your own marketing copy, your published documentation, or a knowledge base you would happily hand a competitor, then run the test out loud: if the entire input corpus appeared on the internet tomorrow, what actually happens. When the honest answer is that almost nothing happens, the standard questions about retention and access can be settled in an afternoon, and six weeks of review buys no protection while costing real momentum. Security review is one of the more respectable ways for a company to avoid making a decision, and it is worth noticing when that is what is happening.

The second case is the platform you already run. Your CRM, your help desk, your accounting system, or your productivity suite already holds the same records under an agreement negotiated on behalf of a customer base far larger than you, and their terms are frequently stronger than anything a small build shop will sign. If their AI tier does the job you were about to commission, the security question is already answered and you skipped it entirely. Check what is included in a plan you are paying for before you introduce a new party to the data. The cost side of that trade runs the other way over time and is worked through in the real cost of off-the-shelf AI at scale, but on this particular question the incumbent usually wins.

The third is the sharpest, and it lands on more mid-market companies than the other two combined. A company with no data classification at all, where nobody has ever written down which categories of record carry an obligation and which do not, is in no position to interrogate a vendor about handling. It is asking a stranger to be more careful with your records than you are, and the vendor's answers cannot be evaluated because there is no internal standard to evaluate them against. That work is a week of internal effort, it costs nothing but attention, and it makes every vendor conversation afterward shorter. Do that first.

What an operator can do this week, with no vendor in the room.

One page and an afternoon produce most of the value here. Write three lists. The first is actions: everything the system would be permitted to do, sorted into read, change, and send outside the company. Most operators can fill this in from the workflow description they already have in their head, and the sorting is the point, because the three columns carry completely different consequences.

The second is resting places: every location a record would come to rest, with the controlling party named beside it. Your own systems, the vendor's environment, the hosting provider, the model provider, the trace store, the vector store, the backups. Any row where you cannot name the party is a question for the vendor, and the count of those rows tells you how much of this system you currently understand. The third is obligations: for each category of input, whose promise you are keeping. A client engagement letter, an employee handbook, a carrier requirement, a regulator, or nobody. That last answer is a real answer and the most common one, which is worth knowing before you treat every input as privileged.

Then draw one line. Anything sitting in the change or send column that touches an input row carrying an obligation is where your entire diligence budget goes. Everything outside that intersection gets the short version. That single page will do more to focus a vendor conversation than any questionnaire you can download, and it is the same input the structured diagnosis we run starts from before anything gets scoped. If you have not picked the workflow yet, the AI Maturity Index gets you to the named process and the resting-places map in about ten minutes, which is where this exercise has to begin anyway.

Field-note context

What we look at when a buyer asks where the data goes.

The evaluation set is the copy nobody deletes.

To build anything that works, a team needs a set of real cases with known correct answers to test against. That set gets assembled by hand, it is expensive in effort, and it is deliberately preserved because rebuilding it is painful, which makes it the most durable copy of your records produced anywhere in the engagement. It usually lives outside the production system: in a repository, in a storage bucket, sometimes in a spreadsheet on an engineer's machine. It is also the copy that deletion clauses miss, because the clause names the database and the evaluation set is not the database. Worth asking at scope time rather than at handoff: what real records are in it, whether they can be redacted or synthesized instead, where the set lives, who can read it, and whether it is named in the deletion terms. There is a tension here worth naming, because the same evaluation set is something you want handed to you at the end, since prompts without evals cannot be safely changed. Receiving it and knowing what is in it are two separate asks, and buyers tend to make the first one and skip the second.

The logging layer is a vendor you never evaluated.

Every serious AI system keeps traces of what it did: the prompt, the passages it retrieved, the output, often the user who asked. You want that, because without it nobody can explain a wrong answer. The part that goes unexamined is that the tool storing those traces is frequently a third-party observability product with its own hosting, its own staff, and its own retention default, chosen by an engineer as infrastructure rather than by anyone as a business decision. It never comes up in the security conversation, and it holds the most concentrated copy of sensitive material in the entire system. Not a table you would have to join to make sense of, but the actual question a person asked and the actual passage from your files that answered it, in plain text, indexed for search. Three questions close it: which tool, exactly what is stored, and what the retention is set to today. A fourth is worth asking too, which is whether the tool can be configured to keep the metadata without the content, since that is often possible and rarely the default.

Ask who can read production, and ask for the list.

At a large vendor, access control is a system with reviews and an audit trail. At a ten-person shop it is often a habit: everyone with a laptop holds the keys because that is how the work gets done and nobody wrote it down. That is not automatically disqualifying, and a small team can handle your records more carefully than a large one, but it is a fact you should hold rather than assume. The question is cheap and completely answerable. Who at your company can read our production data, by name or by role. How is that list reviewed, and how often. What happens to access when somebody leaves. Ask the same about contractors and subcontracted engineers, which are common in this market and rarely mentioned unprompted, because a subcontractor is a party to your data who has no contract with you at all. A vendor who works this way produces a short list within a day. A vendor who does not will describe their policies instead, and the difference between those two responses is the finding.

Extended questions

The questions an operator asks before handing over access.

What security questions should I ask before commissioning a custom AI build?

Ask where a record comes to rest and who controls each of those places, because the exposure is the chain rather than any single vendor. Walk one record from your system of record to the answer on a screen and name every copy it makes along the way: the production database, the vector store holding embeddings, the logging or trace tool, the backups, the development and staging environments, the evaluation set, and anything pasted into the vendor's own support tooling. For each one, ask who controls it and under what contract. Then size how hard you press by what the system can do rather than by what it can see, since a read-only search and a system with write access to a record of account are different risk objects. Ask who at the vendor can read your production data, by name or role. Ask which subprocessors are involved and get the list in writing. Ask what deletion covers at the end and get the copies named. Prefer every question whose answer is a document, a configuration setting, or a clause over any question a vendor can answer with reassurance.

Does a custom AI build mean my company data is used to train someone else's model?

It is the question buyers lead with and it is the one with the shortest answer, because training and retention terms are written down in the model provider's published terms and in the vendor's agreement, and checking both takes minutes rather than weeks. Read them once, get the answer in the contract, and then move on, because that link in the chain is the most scrutinized and the most standardized one you will deal with. The exposure a mid-market buyer actually carries sits underneath: the vendor's own engineers and any subcontractors they use, the hosting and observability tools that store prompts and outputs by default, and the copies of real records that accumulate in development environments and evaluation sets over the course of a build. Those parties are specific to the firm you hired, they are not covered by anyone's published policy, and nobody publishes a page about them. Settle the training question in a minute and spend the diligence you have left on the links that are actually unexamined.

Where does company data actually go inside a custom AI system?

Further than most diagrams show. A record leaves your system of record and typically comes to rest in six or seven places: the application database the system runs on, a vector store holding embeddings derived from your documents, a logging or trace layer that keeps the prompt, the retrieved passage, and the output so engineers can debug a wrong answer, the backups behind all of those, the development and staging environments where a slice of real data was pulled to build against, the evaluation set curated to test whether output is any good, and the model provider's API in transit. Embeddings are a copy, logs are a copy, and an evaluation set is the most durable copy of the group because it is expensive to assemble and deliberately preserved. Each of those sits under a different party and a different contract, and the count is the exposure. The useful exercise before signing is to draw that path on one page and write the controlling party next to every stop on it.

What should an AI contract say about deleting our data when the engagement ends?

It should name the copies, set a deadline, and require written confirmation that it happened. A clause promising that your data will be deleted on request covers the obvious copy and leaves everything else, because the person performing the deletion will read it as the production database. Name the rest explicitly: backups and their rotation schedule, the vector store and the embeddings derived from your documents, logs and traces including whatever retention the observability tool applies on its own, the evaluation set, every development and staging copy, and any records sitting in the vendor's support or ticketing system from debugging. Ask what the subprocessors underneath are obliged to do and on what timeline, since a vendor cannot delete faster than the parties they depend on. Then ask for certification in writing within a stated number of days. All of this is routine to agree before signing and awkward to request afterward, when you are asking a firm you are leaving for a favor rather than holding them to a term.

How much security diligence does a low-risk internal AI tool need?

Very little, and treating every project the same way is how a company spends six weeks of review on a drafting tool and then rushes the system with write access. The test is what would happen if the entire input corpus appeared publicly tomorrow. For a tool running over your own marketing copy, published documentation, or public filings, the honest answer is that almost nothing happens, and the standard questions about retention and access can be settled in an afternoon. Weeks of review there buy no protection and cost real momentum, which is its own kind of loss. Scale the effort to two things instead: whether the inputs carry an obligation you owe someone else, such as client financials, health information, privileged material, or employee records, and whether the system can change a record or send something outside the company. Where the answer to both is no, run the short version. Where the answer to either is yes, that is where the entire diligence budget belongs.

How does the AI Maturity Index help sort this?

It produces the two inputs this whole assessment turns on, before a vendor is in the room. The Index asks which process is worth investing in first, who touches it and how often, and where that process gets its inputs, which is most of the resting-places map you need in order to ask anything useful about handling. It also surfaces what the process actually does at the end, whether it produces an answer for a person to read or changes something in a record, and that is the distinction that decides whether you run the short version of this diligence or the long one. For an operator without a security team, having those two answers written down converts a vague worry about data into a page of specific questions with checkable answers. Ten minutes, no call, and you walk into the vendor conversation knowing which questions are worth your leverage and which ones you can settle by reading a document.

Not sure how much diligence your build actually needs?

Start with the AI Maturity Index. Ten minutes, no call, and it names the one workflow worth investing in first, shows where that process gets its inputs, and tells you whether it produces an answer for a person to read or changes something in a record, which is what decides how hard you press a vendor.