Lesson 4 · AI Governance Basics
Proportional risk tiering and human oversight
- Length
- 28 minutes across 7 sections
- You will be able to apply
- The Highest Dimension
- You will produce
- Commit Statement
- You will work
- 3 gated questions
Personalize the practice
Apply this to your environment
These details adapt the application prompts and coach questions. They do not affect your score.
The tier is the highest dimension, not the average of them.
Core question
risk tiering
Assign the tier before you see the scoring method
Required practiceEach OmniCorp use case has incomplete but decision-relevant facts. Assign the risk tier you would defend now, before reading the Highest Dimension framework.
Proportionality is the only thing that makes governance survivable. Apply the same scrutiny to a meeting summariser and a credit decision and one of two things happens: the credit decision is under-examined, or the meeting summariser is examined so heavily that people stop declaring their AI use at all. Both outcomes are worse than doing nothing.
So tiering is necessary. It is also the single most gameable control in governance, because the person assigning the tier is usually the person who benefits from a lower one.
The pathology: Rounding Down
Watch a tiering conversation and you will see the same movement every time. The system has one alarming characteristic and five unremarkable ones. The discussion spends ninety seconds on the alarming characteristic, establishes that it is not as bad as it first sounded, and then settles on a tier that reflects the five. Nobody lied. The reasoning was even quite careful. The tier is still wrong.
The counter is a scale that does not permit averaging.
A moderate tier that feels reasonable and understates the actual risk.
- Averaging (many low, one high)
- A moderate tier that feels reasonable and understates the actual risk.
- Averaging (all moderate)
- A moderate tier that happens to be correct but arrived by the wrong method.
- Highest Dimension (one high)
- A high tier driven by the dimension that sets it, with oversight designed around that dimension.
- Highest Dimension (all low)
- A low tier that is genuinely low because no single dimension is elevated.
The Highest Dimension
Definition
Score an AI system independently on six dimensions: impact if wrong, autonomy, data sensitivity, reversibility, scale, and regulatory exposure. Score each one on its own, without reference to the others. The system's tier is the highest score on any single dimension. Never the mean, never the mode, never an overall impression. One dimension at the top puts the whole system at the top, and that is the entire point of the instrument.
This is not negotiable, and it is not a heuristic to apply when convenient. A system that is unremarkable in five dimensions and severe in one is a severe system. The five comfortable dimensions are not mitigation. They are the reason nobody looked at the sixth.
When to Use It
Tier before build, before launch, and on any change to autonomy, audience, data or volume. Volume is the trigger most often skipped: a system that moves from four hundred to forty thousand decisions a month has changed dimension even though nothing about its logic has changed.
Re-tier when a system that was advisory starts being followed without question. Autonomy is behavioural, not architectural: a recommendation nobody ever overrides is a decision, whatever the design document says.
How to Apply It
- Score each dimension separately and write the score down before discussing the next one.
- Take the maximum, and say the number out loud before anyone argues about mitigations.
- Record the dimension that set the tier, because that is what oversight must be designed around.
- Apply mitigations to reduce a dimension, and re-score it, rather than adjusting the tier directly.
| Dimension | Low | Moderate | High |
|---|---|---|---|
| Impact if wrong | Rework absorbed internally | A customer is materially inconvenienced | Money, health, liberty or eligibility is affected |
| Autonomy | A person composes the final output | A person reviews before release | The system acts without review |
| Data sensitivity | Public or internal | Confidential business or personal data | Special category, children, or criminal data |
| Reversibility | Undone in minutes by the operator | Undone with effort and a conversation | Cannot be undone once acted on |
| Scale | Tens of instances a month | Thousands | Continuous, at population scale |
| Regulatory exposure | No specific regime engaged | Sector guidance applies | A statutory obligation attaches to the decision |
Worked example 1 of 3
OmniCorp Financial proposed an AI system to pre-screen hardship applications and flag those likely to qualify for a payment holiday. The design group tiered it as moderate. Marisa Delgado made them score the dimensions one at a time.
- Marisa Delgado
- Impact if wrong.
- Product lead
- High. A missed flag means someone in hardship does not get contacted.
- Marisa Delgado
- Autonomy.
- Product lead
- Moderate. An adviser reviews every flag before contact.
- Marisa Delgado
- And the ones it does not flag? Who reviews those?
- Product lead
- Nobody. They stay in the standard queue.
- Marisa Delgado
- Then on the negative path autonomy is high, and impact is high. This is a high-tier system. Say the number before we discuss what we can do about it.
The oversight that followed was designed around the dimension that set the tier: a mandatory sample review of unflagged applications, not more review of the flagged ones. The team's original proposal had concentrated all its scrutiny on the path that already had a human on it.
Why This Works
Scoring dimensions in isolation defeats averaging structurally rather than by asking people to resist it. Once the six numbers exist on paper, taking the maximum is arithmetic, and arithmetic does not care how reasonable the room finds the system. The judgment is confined to each individual dimension, where it belongs and where it can be argued about specifically.
Recording which dimension set the tier is what makes the oversight proportionate rather than merely heavy. A high tier driven by scale needs monitoring and thresholds. A high tier driven by irreversibility needs a checkpoint before the irreversible step. Designing the same oversight for both wastes effort in one case and misses the risk in the other.
Worked example 2 of 3Optional depth
Dr. Naomi Ellery tiered an ambient transcription tool at OmniCorp Health that produces draft clinical notes for clinician review. Impact moderate, autonomy moderate, reversibility moderate, scale moderate, regulatory exposure moderate. Data sensitivity: high, unavoidably, because the content is clinical. One dimension, and the system is high tier. The oversight designed around that dimension was about storage, access and deletion rather than about note quality, which the moderate dimensions already handled adequately. The tier did not make the project harder. It made it point at the right thing.
Worked example 3 of 3Optional depth
Priya Raghunathan re-tiered an OmniCorp Logistics route-optimisation system that had been low tier for two years and had performed faultlessly. Nothing about its logic had changed. Its volume had risen from two hundred to nineteen thousand route decisions a day, and dispatchers had stopped overriding it because it was almost always right. Two dimensions had moved, scale and autonomy, without a single line of code changing. The system was now high tier and had been for months, governed by a tiering decision made when it was a pilot.
Edge Cases and NuancesOptional depth
A system's tier can differ between its positive and negative paths, as at OmniCorp Financial, and the tier is the higher of the two. Mitigations legitimately reduce a dimension, but only if they are implemented and verified, so a planned mitigation reduces nothing. Composite systems must be tiered as assembled rather than component by component, because three low-tier components chained together produce autonomy nobody scored. And a system can be high tier and still proceed: the tier determines the oversight required, not permission, and treating a high tier as a refusal is what teaches people to round down in the first place.
Whether a statutory obligation attaches to the decision the system influences.
- Regulatory exposure
- Whether a statutory obligation attaches to the decision the system influences.
- Scale
- The volume of decisions or outputs, from tens a month to continuous population-level.
- Reversibility
- Whether an action can be undone in minutes, with effort, or not at all once taken.
- Data sensitivity
- The classification of data the system touches, from public to special category.
- Autonomy
- Whether a person composes, reviews, or never sees the output before it acts.
- Impact if wrong
- The consequence of an incorrect output, from internal rework to harm to health or liberty.
Knowledge check
A system scores moderate on five dimensions and high on data sensitivity alone. A colleague proposes moderate tier because the system is mostly moderate. What is the correct tier and why?
Common Failure Modes
The scale end to end
Alan Brixmoor tiered a proposed OmniCorp Public system that would rank benefit review cases by likelihood of error, so caseworkers could work the highest-value queue first. The team had proposed low tier: it makes no decisions, it only orders a list.
Scored separately: impact if wrong, high, because a case ranked low is a case worked late and the affected person is on a reduced payment while they wait. Autonomy, high, because nobody reviews the ordering. Data sensitivity, high, health and financial. Reversibility, moderate. Scale, high, every case in the directorate. Regulatory exposure, high, statutory decision timescales attach.
Five high dimensions on a system described in good faith as only ordering a list. Alan's point to the team was not that they had been careless: it was that ordering a list is a decision about who waits, and that this is invisible precisely because no decision is recorded anywhere. The oversight designed around it was queue-age monitoring with a hard escalation on any case aged beyond a threshold, regardless of rank, plus a monthly sample of the bottom of the queue. The system launched. It launched with a tier that told the truth about it.
Decision point
Jo Halvorsen at OmniCorp Studio uses AI to draft client-facing project updates. Impact if wrong is moderate, autonomy is moderate since she reads every one, data sensitivity is moderate, reversibility is moderate, scale is low at maybe thirty updates a month, and regulatory exposure is low. Her studio manager proposes moderate tier and light oversight. Then Jo mentions that two clients are public sector and their updates are appended verbatim to a statutory progress report. What is the tier, and what oversight follows?
Self-check
Mark the level that describes you today. Nothing is submitted.
| Behaviour | Ready | Developing | Not yet |
|---|---|---|---|
| Deriving the tier | |||
| Handling mitigations | |||
| Keeping the tier current |
Commit
Commit Statement
Complete every line in your own words, then sign and date it. Attach your six scores and the dimension that set the tier.
| Window | Field application |
|---|---|
| Days 1 to 7 | Score the six dimensions for one system, separately and in writing, and compare the maximum with its current tier. |
| Days 8 to 21 | Find one system whose oversight was designed around a different dimension from the one that set its tier, and redirect it. |
| Days 22 to 30 | Check the override rate on one advisory system. A falling rate is an autonomy increase, and it requires a re-tier. |
Six dimensions and a maximum will give you an honest tier for a system you can see. They will not tell you how to hold that tier when a senior sponsor disputes it, how to evidence the assessment to an external reviewer, or how to run a portfolio where two hundred systems are tiered by forty different people to the same standard. Making tiering consistent and defensible across an organisation is the capability the paid programs develop next.
Learner feedback