Mufasa Labs
← BlogStrategyAugust 23, 2026

AI vendor evaluation framework

Score an AI vendor on data residency, retention, logging, subprocessors, exit, eval access, and who can see your prompts — before the pilot becomes the production path.

Vendor evaluation for AI is a list of facts you can put in a contract and then verify in a tenant. If you cannot answer where prompts live, who can read them, how long they stay, and how you leave, you have a dependency you cannot operate.

This framework is the questions we ask before a tool sits on the approved list. It pairs with AI governance as engineering: a secure AI gateway in front of model calls, a living inventory row for the vendor system, a risk tier, and a light map to NIST AI RMF. GOVERN and MAP are the contract and the inventory. MEASURE is logs and eval access. MANAGE is exit and the incident path when the vendor fails.

Use this post when a salesperson says "we do not train on your data" and you need to know what that means. The company shape is the enterprise AI governance framework.

Demand a residency story you can point to on a map

"The cloud" is not a location. You need the region or regions where prompts, completions, embeddings, uploaded files, and any derived indexes are stored and processed. You also need what happens when a support engineer or a failover event moves a copy.

Ask, in writing:

  • Which regions store prompts, completions, files, and embeddings at rest.
  • Which regions process inference. Processing is not the same as storage.
  • Whether a support or success team can pull a customer workspace from another geography.
  • What happens on failover. A quiet replica in a second country is still residency.
  • Whether you can pin a region in the contract, not only in a UI checkbox that resets on a plan change.

Yes/no:

  • The answer names countries or regions, not "global infrastructure."
  • The data classes you will send are allowed in those regions under your own policy.
  • Legal has seen the same document security saw.

If the vendor cannot pin residency, treat that as a tier constraint. Internal drafts of public copy may still be fine. Customer PII is not. Risk-tier the use case, not the logo.

Retention and deletion have to be testable

Retention is how long the vendor keeps prompts, files, and derived artifacts after you click delete, after a contract ends, and after a user leaves. "We do not train on your data" does not answer any of those. Training opt-out, product telemetry, abuse review, and backup copies are separate stores.

Get these in the order:

  • Default retention for prompts, completions, uploads, and embeddings.
  • Whether product improvement, safety review, or human labeling can see your content, and for how long.
  • Deletion: API or admin action, scope (workspace vs tenant), and the time until backups are gone.
  • Whether you can set a shorter retention than the default.
  • What remains after offboarding: invoices, logs the vendor keeps for their own compliance, copies in subprocessors.

Yes/no:

  • You can delete a known prompt and later confirm it is gone, or you have a contractual audit right that names the method.
  • Backups have a stated outer bound. "As soon as practicable" is not a bound.
  • Training and fine-tuning on your content are off unless you opt in, in the contract, not in a blog post.

If you cannot test deletion, assume the data stays. That assumption should change the tier or block the data class.

Logging must serve you, not only the vendor

You will need to reconstruct a bad answer. The vendor's internal logs are not your logs. Ask what you can export, what you can stream to your SIEM, and what the vendor's staff can see when they debug.

Minimum you should be able to pull:

  • Who called, when, which model, which workspace.
  • Whether a file was attached, and a handle for that file.
  • Enough prompt and completion text for an incident, with access limited to your admins.
  • Admin actions: who exported, who changed retention, who added an integration.

Ask the inverse, too: who at the vendor can see prompts. Support, model-quality, and trust-and-safety are different teams. Each needs a named purpose, a named access path, and a record you can request.

Yes/no:

  • You can export last week's traffic without a ticket that takes a week.
  • Vendor staff access is off by default or gated, and you can see that it happened.
  • Your own gateway can sit in front so PII redaction and cost controls run in your environment even when the vendor's UI is the place people type.

If the only log is a usage invoice, you cannot do MEASURE. Do not put high-risk data in that tenant.

Subprocessors and exit are part of the same question

Subprocessors are everyone else who might see the prompt: model providers, annotation vendors, support tools, storage, analytics. Exit is whether you can leave without losing the work you already paid to create.

On subprocessors:

  • A current list, not a page that says "may include affiliates."
  • Which subprocessors see content, versus which see only metadata or billing.
  • Notice period when the list changes, and a right to object or terminate for a new content-touching party.
  • Whether the primary vendor's "no training" promise also binds the model provider behind them.

On exit:

  • Export of prompts, files, eval sets, and conversation history in a documented format.
  • Export of any fine-tune, index, or custom model you paid to build, or a clear statement that you cannot have it.
  • A timeline to wipe the tenant after export.
  • Whether features you built on their API have a portable equivalent or are a one-way door.

Yes/no:

  • You can name the model provider if the vendor is a wrapper.
  • You have run a sample export during the pilot, not promised to run one later.
  • Legal has an exit clause that matches the technical export, including embeddings and uploaded corpora.

A vendor you cannot leave becomes an unofficial source of record for prompts.

Require eval access before the tool touches production data

If you cannot run your cases against their model or their product behavior, you are buying a demo. Eval access means you can hold a set of prompts, expected behaviors, and failure cases, and re-run them when they change the model.

What to require:

  • A way to pin or at least record the model version the product is calling.
  • Notice when that version changes, or a changelog you can subscribe to.
  • Permission to send a synthetic or redacted eval set through the same path production will use.
  • Clarity on whether their "accuracy" numbers were measured on your domain or on a public leaderboard.

Yes/no:

  • You ran your own cases in the pilot tenant.
  • A model change is a change you can see on the inventory row.
  • Failures can block a data-class expansion even if the pilot "felt fine."

The vendor does not get a bypass on the same promote gate developers use.

Decide who may see prompts, including your own people

The last question is not only the vendor. It is everyone: vendor staff, subprocessors, your admins, your managers, and any integration that syncs chats into Slack or a ticket system.

Write the audience before you turn the tool on:

  • Which roles in your company can read raw prompts.
  • Whether managers can browse a team's history by default (usually they should not).
  • Whether an integration copies prompts into a second system that has a weaker retention policy.
  • Whether the vendor uses prompts for human review, and whether you can opt out.

Yes/no:

  • The allow list of readers is shorter than the allow list of users.
  • A prompt that contains residual PII is treated as the source data class, not as "just a chat."
  • Shadow use of a personal account of the same product is out of scope for this contract. Personal accounts are a shadow AI finding, not an enterprise control.

Score residency, retention, logging, subprocessors, exit, eval access, and prompt audience. Then risk-tier the use case. Low-risk work can proceed. High-risk work waits until the answers are in the contract and the tenant.

If you want help turning those answers into gateway rules, inventory rows, and a NIST-mapped pack, talk to an engineer.

Want this working in your business?

Every post on this blog comes from systems we've actually built. Book a 30-minute call and we'll map the same playbook to your stack.