Method
Generic AI training tells you how to write a prompt. It does not tell you what your organization should actually approve. This is the process I use to answer that question, and the one I teach inside every engagement.
Before a team uses AI on something, I ask two questions. First, how bad would it be if AI got this wrong? Second, how sensitive is the information involved? The answer to those two questions tells you whether a use case is fine to run today, fine with a person double-checking it, worth a small careful trial, or not worth attempting yet. That is the whole method. Everything below is how I apply it in practice, and how I help your organization’s own decision-makers apply it after I am gone.
Most organizations I meet have either no AI policy or a policy written once, in the abstract, that nobody consults. Both produce the same outcome: staff use these tools anyway, privately, with no guidance about what is safe. A policy tells people what happened to be decided last quarter. A method lets them decide the case in front of them, and lets whoever approves new tools at your organization, legal, privacy, risk, or a dedicated compliance function, apply the same standard consistently.
The steps below come from building and shipping AI in situations where being wrong is expensive, and from the evaluation work that any high-stakes AI deployment actually requires.
The rubric
Every proposed use case lands in one of four situations. Each one carries a standing answer, so the decision does not get relitigated every time someone has an idea.
Use it. Drafting internal summaries from public or non-sensitive material. This is not where your attention needs to go.
Allowed, with a person reviewing the output before it leaves the building. Most valuable professional work lands here.
Only in a tool cleared for that kind of data, and only after the retention and training terms are confirmed in writing.
Sensitive data plus a costly mistake if it goes wrong. Revisit once the tool, the contract, and the evidence of accuracy all improve. Not before.
Teams leave a workshop with this rubric, branded and theirs to keep, along with a draft policy written against their own environment.
Every organization already has data categories with different rules attached. In a health system that is protected health information, internal-only material, and public material. In a firm it is privileged material, client-confidential material, and public material. I write down which sanctioned tool each category may enter, and which it may not. This becomes a one-page safe-use map specific to your environment, and it is what stops the most common failure: a capable person pasting the wrong thing into the wrong window.
How bad would it be if AI got this wrong, and how sensitive is the data involved? Every proposed use case gets scored on those two questions, which places it in one of four situations below, each with a standing answer. This takes about ninety seconds per use case once a team has the rubric, and it turns an anxious open-ended debate into a quick decision.
The question is not whether the model is right. It is how a professional confirms it quickly. For each approved use case I name the checkpoint: ask the model to point to the source it relied on, ask it to flag the parts it is least sure about, and have a person spot-check the highest-stakes items by hand. The work still gets reviewed. It just arrives faster, and everyone knows in advance who is checking what.
Before a tool is approved I want answers on data retention, whether your data trains their model, whether they will sign a business associate agreement where one is required, audit logging, access control, and what evidence they can show that the tool is accurate. A vendor who cannot answer that last one is asking you to take safety on faith.
Approved use cases go out with an owner, a guardrail level from the rubric below, a way to measure success, and a date to check back in. After that, someone has to keep owning the ongoing questions: what happens when the tool changes and starts behaving differently, how new use cases get reviewed, and how staff who are already using AI on their own get brought inside the policy instead of pushed further outside it.
I spent two decades inside healthcare before AI was the job, at Kaiser Permanente, Huron, and McKinsey’s healthcare practice. I then shipped large language model products in a regulated environment at Verily, with legal, privacy, and security sign-off.
Today, in my work on AI safety, I hold systems to a bar where more than 99 percent of outputs have to pass human review, because when AI is wrong in a clinical or legal context, the cost is real. This method is that same discipline, human review, structured risk scoring, and clear ownership, translated into terms a leadership team can actually run.
Full backgroundThe workshop program walks a leadership team through this method on their own use cases, and they leave with the rubric, a safe-use map, and a 90-day plan.