Mufasa Labs
← BlogStrategyAugust 23, 2026

Shadow AI risks

Employees are already pasting work into public chatbots. How to discover shadow AI, risk-tier it, and pave or kill — without treating every unofficial tool as a moral failure.

Shadow AI is already in the company. That is not a culture problem and it is not a character judgment. It is a traffic pattern. Someone had a deadline, the official tool was slow or missing, and a public chatbot was one tab away.

The risk is specific. Customer records, contracts, credentials, and unreleased numbers leave the tenancy. The company cannot reconstruct the prompt, cannot prove the vendor will not retain it, and cannot tell a customer or a regulator what happened. The same person may also be using a harmless grammar check on a public blog draft. Those two acts are not the same incident.

Do not start with a ban-all memo. Start with discovery. Then risk-tier. Then either pave the useful work onto a logged path or shut down the flows that cannot stay. That is the same sequence as AI governance: discover, risk-tier, implement guardrails in your environment, operationalize. The readiness gaps post names the pattern — official tools too slow, work leaving the tenancy, counted as "excitement." It is not excitement. It is a missing paved road.

Discovery first, or you will police the wrong thing

You cannot risk-tier a rumor. You need a list of tools, data, and teams. The list will be incomplete on day one. That is fine. An incomplete list you update is better than a policy that assumes only the licensed copilot exists.

Look in places that already produce evidence:

  • Browser and CASB logs for known public model domains and browser extensions.
  • Expense and procurement for "AI" vendors that never went through security.
  • SaaS admin consoles for features that quietly turned on a model (search, summarize, write).
  • Ticket and Slack history for "I asked ChatGPT" as a work method, not a joke.
  • A short anonymous survey that asks what people actually used last month, with no threat of punishment for the first round.
  • Outbound DLP or email rules that already catch large pastes, if you have them.

What you are trying to record per finding:

  • Tool name and whether it is a personal account or a company tenant.
  • Team and a rough volume (daily, weekly, one-off).
  • Data class people say they paste, and data class you can infer from the job.
  • Whether the output is used internally or sent to a customer.
  • Whether anyone already asked for an official alternative.

Yes/no for discovery:

  • You can name the top five unofficial tools by use, not by fear.
  • You have at least one finding from logs, not only from interviews.
  • Shadow use is written into the model inventory as "unsanctioned," not left in a side spreadsheet.
  • Legal and security saw the method before you ran a company-wide hunt.

Discovery that starts with a threat produces silence. Silence is the opposite of an inventory.

Risk-tier the flow, not the brand name

"ChatGPT" is not a risk tier. A marketing intern drafting social copy from a public brief is a different flow than a claims adjuster pasting a medical record. The brand can be the same. The control cannot.

Use the same three tiers you would use for sanctioned systems:

  • Low: public or internal non-sensitive input, human reads the output, nothing goes back into a system of record.
  • Medium: confidential business data, internal decision support, still a human signer.
  • High: customer PII, health, payment, credentials, or any write-back into production systems.

Yes/no for each shadow finding:

  • Data class is assigned with the same rubric as official systems.
  • "The employee said they stripped the name" is not accepted as a control for high-risk data. Residuals identify people.
  • Personal accounts of otherwise-reputable products are treated as unsanctioned for company data, even if the product has an enterprise SKU you do not own.
  • Vendor marketing about training opt-out is noted, not treated as a contract you have.

High-risk customer and PII flows cannot stay on personal ChatGPT. That is not a values statement. You have no tenant, no log, no deletion you can prove, and no way to put a human-in-the-loop gate on a write-back that already happened in someone else's cloud.

Low-risk flows can stay for a while if you can see them. "Stay" means: named, logged or at least acknowledged, time-boxed, and pointed at an official path when one exists. Killing a grammar tool used on public copy while a claims desk still pastes files into a personal account is theatre.

Pave the work that is actually useful

Most shadow use is a request for a faster tool. If the official path is a three-month review for a low-risk summary, people will keep the dirt path. The fix is to make the paved road shorter.

Paving looks like this:

  • Put an approved tool on the allow list with the data classes it may see.
  • Put company traffic through a secure AI gateway so prompts are logged, PII can be redacted, and cost is visible.
  • Give the team the same job they were already doing — draft, summarize, extract — inside the tenant.
  • Time the approval for low-risk use in days. Publish that number so people believe it.
  • Show the before-and-after: same task, official tool, no lecture.

Yes/no before you call a flow paved:

  • The user who was on the shadow tool has been shown the official one, not only emailed a policy.
  • Latency and login friction are close enough that the official tool wins on a Tuesday afternoon.
  • The inventory row flipped from unsanctioned to sanctioned, with an owner and a tier.
  • You can pull last week's prompts for that team without asking them to self-report.

If you cannot pave this month, say so, and say what is allowed in the meantime. A vacuum is how personal accounts come back.

Kill the flows that cannot be made safe

Some work should stop this week. Not because AI is bad. Because the data cannot leave.

Kill, immediately, when you find:

  • Restricted data in a personal or consumer chatbot.
  • Production credentials, connection strings, or badges pasted "to debug."
  • Customer files uploaded to a tool with no contract and no tenant.
  • Any agent or plugin that can send mail, file a ticket, or change a record from a personal account.

How to kill without a morale collapse:

  • Name the flow and the data class, not the person, in the first announcement.
  • Offer the paved alternative in the same note, or a dated plan if the alternative is not live.
  • Use technical blocks where you can (DNS, CASB, DLP, extension control) so the rule is not a suggestion.
  • Record the kill in the inventory as closed, with the date and the residual risk if people will try a new domain tomorrow.

Yes/no:

  • High-risk personal-account use has a block, not only a slide.
  • Exceptions are written, time-limited, and owned. "Just this vendor demo" without an owner is still shadow.
  • You re-check logs after the block. Displacement to a new domain is the default, not a surprise.

Killing without discovery is how you ban the tool security already knew and miss the one in a browser extension. Killing without a paved alternative is how the same paste moves to a phone.

Proportionate controls beat a single ban

Once the list exists, controls should match the tier. That is the whole point of risk-tiering.

Low-risk, allowed if logged:

  • Public or internal non-sensitive drafts on an approved tool or, temporarily, on a known personal tool you have chosen to tolerate with a sunset date.
  • Spot-check a sample of prompts. You are looking for data-class drift (someone started pasting customers), not tone.

Medium-risk, allowed only on a company tenant:

  • Confidential decks, source code that is not credentialed, internal tickets without customer attachments.
  • Gateway or equivalent logging. No personal accounts.
  • Owner named. Eval if the output feeds a decision other people will trust.

High-risk, not allowed on shadow paths:

  • PII, health, payment, unique identifiers, live customer files.
  • Write-backs. If the model can change a record, it is a production system. It belongs in the SDLC and on the inventory, not in a browser tab.

Yes/no for the control set:

  • The allow/deny list an employee can find matches what the gateway and CASB actually do.
  • Cost controls exist so a paved tool does not become an unbudgeted surprise and push people back to free personal accounts.
  • A review cadence re-opens the shadow list. New tools will appear. That is a discovery input, not a failure.

Operationalize or the list goes stale in a quarter

Shadow AI is not a project with an end date. New features land inside software you already bought. New consumer apps show up in a weekend. The operational piece is boring and required.

Put on the calendar:

  • A monthly pass over CASB/DNS/extension logs for new domains.
  • A quarterly pass over the inventory's unsanctioned and recently-paved rows.
  • A one-page refresh of the allow/deny list when a vendor changes terms.
  • A drill: can security reconstruct a suspected paste from last Tuesday?

Yes/no:

  • The shadow list has a last-reviewed date.
  • Product and procurement tell security when they turn on an AI feature, instead of security finding it in a demo.
  • Training is the two lists and the two examples (low-risk draft vs customer file), not a seminar on ethics.

You will not reach zero unofficial use. You can reach a state where high-risk flows are blocked, low-risk flows are known and logged, and the official path is the one people actually take. That is governed enough to defend.

If you do not yet know which flows you have, do not start with a 40-page policy. Score the gap, then go look at traffic. Take the 3-minute AI readiness assessment.

Want this working in your business?

Every post on this blog comes from systems we've actually built. Book a 30-minute call and we'll map the same playbook to your stack.