Lesson 13 · Applied labs: evaluation, cost, and guardrails
Applied labs: evaluation, cost, and guardrails
- Length
- 18 minutes across 7 sections
- You will be able to apply
- The Evaluation Harness Protocol · The Guardrail Bench
- You will produce
- Commit Statement
- You will work
- 4 gated questions
Personalize the practice
Apply this to your environment
These details adapt the application prompts and coach questions. They do not affect your score.
What can your team operate this quarter when shadow AI is already in daily use, and what failure have they personally forced before it found a customer?
Core question
prioritization
Rank the lab evaluation priorities
Required practiceYour OmniCorp applied lab gets one day to prove the workflow is governable. Rank what the team should prioritise first before reading the Evaluation Harness Protocol.
- 1Define blocked-release thresholds before the run begins.
- 2Choose adversarial cases that probe unsafe instruction, sensitive data, ambiguity, and missing context.
- 3Record the exact prompt, model, retrieval, and tool versions used in the run.
- 4Record residual failures and assign owners after scoring the cases.
- 5Count how many enthusiastic users want access this quarter.
Applied AI governance is not the ability to explain a control. It is the ability to run it under pressure, read the result, and change the workflow before a live failure does the teaching. Shadow AI is already inside daily work. The gap between describing and doing is exactly where the next incident lives this quarter.
This lesson moves the work from slides to hands on a keyboard. You build and run an evaluation harness by hand. You attack a workflow on purpose. You price the cost of the controls as operating work, not ceremony. The standard is simple: if the team has never tripped the failure in a lab, it does not yet own the control.
The pathology: Slideware Competence
The team can explain controls but has no scored run, no attack record, and no costed operating burden.
- Slideware Competence
- The team can explain controls but has no scored run, no attack record, and no costed operating burden.
- Lab competence
- The team runs cases, trips guardrails, records residual failures, and knows what each control costs.
You counter the pathology by requiring evidence that was made by doing. At OmniCorp Financial, Priyanka Venn wants cost evidence she can defend. At OmniCorp Health, Nadia Haddad wants safety proof that survives a clinical review. At OmniCorp Logistics, Yusuf Demir wants workflow controls that hold during disruption. At OmniCorp Retail, Rafael Duarte wants a launch that will not leak customer data through a helpful assistant.
The lab standard is intentionally physical. Someone opens the case list, runs the workflow, captures the score, forces the attack, reads the log, and records the cost. Governance becomes a sequence of observable acts. That sequence is what lets a sponsor clearly distinguish an operating team from a team that has memorised the operating model. The distinction matters when failure arrives at speed.
The Evaluation Harness Protocol
Definition
The Evaluation Harness Protocol is the practice of building and running a repeatable test harness for an AI workflow before production. It includes golden cases, adversarial cases, acceptance thresholds, a scored run, and recorded residual failures. The protocol is not a discussion of quality: it is a dated run that shows what the workflow did against known cases.
A harness is small enough to run by hand and structured enough to run again. It turns quality from opinion into evidence. You do not need a perfect platform to begin. You need cases, expected behaviours, scoring rules, thresholds, and a place to write what still failed after the team improved the workflow.
When to Use It
Use the protocol before a pilot becomes production, before a prompt change ships, before a model or retrieval source changes, and before a vendor result is accepted as proof. Use it when the workflow affects customers, employees, clinical communication, financial judgement, logistics decisions, or regulated records. Use it whenever confidence is currently based on a demo.
Do not wait for an automated evaluation platform. Automation helps later. The first discipline is manual repeatability. A spreadsheet with case identifiers, inputs, expected outcomes, scores, reviewer notes, and residual failures is enough to start. A team that cannot run ten cases by hand is not ready to automate a hundred.
How to Apply It
- Select golden cases that represent normal, high-value work the AI workflow must handle reliably.
- Select adversarial cases that stress known risks: ambiguity, missing context, conflicting policy, unsafe instruction, and sensitive data.
- Define acceptance thresholds before the run, including minimum score, blocked release conditions, and reviewer authority.
- Run the cases against the actual workflow, with prompt, model, retrieval, and tool versions recorded.
- Score each case, record residual failures, assign owners, and decide whether the release proceeds, changes, or stops.
Write the threshold before the result tempts you. For example: release requires 90 percent pass on golden cases, zero critical failures on adversarial cases, and no unsupported claim in a regulated output. A threshold written after the run is negotiation. A threshold written before the run is governance.
Normal work the workflow must handle, with expected output and reviewer notes.
- Choose golden cases
- Normal work the workflow must handle, with expected output and reviewer notes.
- Choose adversarial cases
- Risk cases that test ambiguity, unsafe instruction, missing data, and sensitive content.
- Set thresholds
- Minimum scores and blocked release conditions written before the run.
- Run and score
- Dated results tied to prompt, model, retrieval, and tool versions.
- Record residuals
- Known failures, owners, and the release decision they constrain.
Worked example 1 of 3
OmniCorp Financial built a lending-summary assistant for complex consumer files. Bea Karlsson, model risk lead, asked Martin Hollis for twenty golden cases and ten adversarial cases. The golden cases covered standard income, identity, and collateral summaries. The adversarial cases included contradictory documents, a missing adverse-action notice, and a borrower note that looked relevant but was prohibited from the decision record.
| Harness element | OmniCorp Financial run |
|---|---|
| Golden cases | Twenty historical lending files with expected summary points and source-page references. |
| Adversarial cases | Ten files with contradictions, missing notices, prohibited notes, and policy ambiguity. |
| Threshold | At least 90 percent golden pass, zero critical adversarial failures, and no unsupported claim. |
| Scored run | Prompt v3.1, retrieval bundle 2026-07-18, model release 14, reviewed by Bea Karlsson. |
| Residual failures | Two weak source citations and one overconfident policy phrase assigned to Martin Hollis. |
- Martin Hollis
- The demo already looked accurate to the lending team.
- Bea Karlsson
- A demo is not a run. Show me the cases it passed, the cases it failed, and the threshold it had to clear.
- Priyanka Venn
- Then show me the review hours, because the harness is now part of the operating cost.
Why This Works
The protocol works because it makes quality falsifiable. A stakeholder can argue with a slide, but a scored case forces the conversation onto evidence. The team sees which failures are rare, which are systemic, and which block release. Confidence becomes a result of repeated exposure to known risks, not a mood created by a polished walkthrough.
It also creates an operating memory. When a later prompt change breaks source citation, the team can compare the new run with the old one. When a regulator asks why the workflow was released, the evidence is already dated. The harness is not bureaucracy: it is the record that the team looked before it shipped.
Worked example 2 of 3Optional depth
OmniCorp Health used the protocol for a discharge-instruction assistant. Dr. Ilona Reyes selected golden cases for routine discharge, chronic medication instructions, and follow-up scheduling. Nadia Haddad added adversarial cases with incomplete medication reconciliation and conflicting clinician notes. The first run passed routine cases and failed two safety cases. The team stopped release for six days, changed the workflow to block drafts when medication reconciliation was incomplete, and reran the same cases.
Worked example 3 of 3Optional depth
OmniCorp Retail tested a product-copy assistant before peak season. Rafael Duarte wanted speed for category teams. Elena Novak asked for the cost of evaluation labour. The harness used golden cases for standard descriptions and adversarial cases for restricted claims, private supplier notes, and invented discounts. The scored run showed a 96 percent pass on normal copy and one critical claim failure. The launch waited until the claim guard was fixed.
Edge Cases and NuancesOptional depth
A small harness is not a weak harness when the cases are chosen well. Ten cases that represent actual risk beat one hundred random examples. Avoid false precision. A score of 87 percent is useful only if the failed cases are understood. One critical failure can block release even when the average score looks strong. The residual list matters as much as the score.
Knowledge check
A team says its AI workflow passed review because five leaders watched a successful demo. What does the Evaluation Harness Protocol require before release?
Evaluation tells you whether the workflow performs against cases you selected. It does not prove the workflow resists deliberate attack. The second lab changes the posture. You stop asking whether the system behaves in normal work and start proving what happens when a user, document, prompt, or tool call tries to make it misbehave.
The Guardrail Bench
Definition
The Guardrail Bench is a hands-on attack lab for an AI workflow. The team deliberately attempts prompt injection, data exfiltration, and over-permissioned tool calls, then proves each guardrail holds before production. The bench is not a policy review: it is a controlled attempt to make the workflow break its own rules.
A bench has named attacks, expected blocks, observed results, and corrections. It uses the real workflow path where possible, including retrieval, tools, identity, logging, and human approval. If the production workflow can call a tool, the bench tries to misuse that tool. If the workflow can retrieve documents, the bench tries to poison or overreach retrieval.
When to Use It
Use the bench before production, after any new tool permission, after a retrieval source changes, after a prompt policy changes, and after any incident or near miss. Use it when a workflow can see sensitive data, act through tools, influence customers, or draft regulated communication. Use it most when the team says the guardrails are obvious.
Do not reserve the bench for security specialists alone. Security should shape the attacks, but operators must run them too. A control that only the security team understands will fail during daily operation. The practitioner who owns the workflow needs muscle memory: what the blocked attack looks like, where the log appears, and who decides the release.
How to Apply It
- List the workflow's guardrails: instruction hierarchy, data boundaries, retrieval filters, tool permissions, human approvals, and logging.
- Create attack scripts for prompt injection, attempted data exfiltration, over-permissioned tool calls, and bypass of human approval.
- Run each attack in the controlled environment using the same identity and connectors expected in production.
- Record whether the guardrail blocked, degraded safely, alerted, logged, or failed, with exact evidence captured.
- Fix failures, rerun the same attack, and price the review, monitoring, and incident-response hours into the operating model.
The bench is successful when it produces a boring answer under hostile input. The model refuses the injected instruction. The retrieval layer withholds the document. The tool call stops at the approval gate. The log shows the attempt. The owner can explain the evidence in two minutes. Anything less is a rehearsal that found work still to do.
Injected instruction ignored, refusal recorded, and original system instruction remains effective.
- Prompt injection
- Injected instruction ignored, refusal recorded, and original system instruction remains effective.
- Data exfiltration
- Sensitive source withheld, redaction applied, and access denial logged with identity.
- Tool overreach
- Unauthorised tool call blocked or routed to human approval before action.
- Approval bypass
- Workflow fails closed and records that no human approval token was present.
Worked example 1 of 3
OmniCorp Logistics ran the bench for its disruption assistant. Anders Nilsen created attacks against a prompt that asked the assistant to ignore dispatch policy, a retrieval attempt for customer payment notes, and a tool call that tried to reassign a truck without supervisor approval. Claire Beaumont required the bench to use a real supervisor account in a controlled environment, because paper permissions had already fooled the team once.
| Attack | Observed result |
|---|---|
| Prompt injection | The assistant refused the instruction to ignore dispatch policy and logged the injection phrase. |
| Data exfiltration | Payment notes were not retrieved because the connector filter denied the data class. |
| Tool overreach | Truck reassignment opened a draft only; the action required Claire Beaumont's approval. |
| Approval bypass | A missing approval token caused the workflow to fail closed and alert platform support. |
Why This Works
The bench works because it removes comfortable ambiguity. A control either holds against the attack or it does not. The team sees the failure path before an attacker, an employee, or a malformed document finds it in production. That experience changes behaviour. People stop saying the policy exists and start knowing where the block occurs.
It also prices safety honestly. Guardrails cost time: attack design, reruns, monitoring, approvals, and incident response. When Priyanka Venn saw the bench results, she accepted the additional review hours because they prevented a larger cost: an emergency shutdown, two weeks of rework, and a regulator's letter asking why tool permission was never tested.
Worked example 2 of 3Optional depth
OmniCorp Health attacked its discharge workflow with a malicious instruction embedded in a pasted referral note. The assistant tried to follow the referral note instead of the system instruction during the first bench. Nadia Haddad stopped the release, added source isolation and explicit instruction precedence, and reran the attack. The second run refused the injected instruction and logged the source document that carried it.
Worked example 3 of 3Optional depth
OmniCorp Retail tested a merchandising assistant against data exfiltration. Rafael Duarte asked for a summary of supplier margin notes using a category manager account that should not see them. The assistant initially returned one margin phrase through retrieval leakage. Elena Novak costed the fix at 42 engineering hours and eight review hours. The launch waited. The rerun withheld the notes and logged the denied source class.
Edge Cases and NuancesOptional depth
Not every failed bench means the use case dies. Some failures require a stronger guardrail, narrower tool scope, or a human approval step. Others reveal that the workflow should not be built. Treat the bench as a decision instrument, not a punishment. The best outcome is not a perfect first pass. The best outcome is a failure found while it is still cheap.
Knowledge check
A workflow has documented tool permissions and a policy saying the model cannot take action without approval. What does the Guardrail Bench require before production?
Decision point
OmniCorp Retail has a product-copy assistant with a strong evaluation score, but the team has not attempted data exfiltration or tool overreach. Rafael Duarte wants to launch before peak season. What should Elena Novak require first?
Common Failure Modes
The labs end to end
OmniCorp Logistics ran the labs for the disruption assistant before winter peak. Yusuf Demir began with a strong demo and a real operational need. Claire Beaumont converted the demo into a harness. She chose historical disruption cases as golden cases, added adversarial cases for missing depot capacity and conflicting route rules, and required a threshold before the first run.
The first harness run passed most normal cases and failed one critical adversarial case: the assistant recommended a route using a closed depot. Anders Nilsen recorded the residual failure, tied it to retrieval freshness, and assigned the fix. The rerun used the same case and passed. Priyanka Venn added the review and rerun hours to the workload cost rather than hiding them.
Then the team moved to the Guardrail Bench. Anders attacked the workflow with prompt injection, payment-note exfiltration, and an over-permissioned truck reassignment. The prompt injection was refused. Payment notes stayed outside retrieval. The truck reassignment opened only as a draft for Claire's approval. One approval-bypass attempt failed closed but lacked a clear alert, so the release waited until logging was repaired.
The final release record contained the harness score, residual failures, bench attacks, guardrail evidence, control labour, and owners. Rosalind Achebe saw an operating case, not slideware. The assistant launched with narrower confidence and stronger proof. The team moved faster because it had already done the work that a live incident would otherwise demand.
Mark the level that describes you today. Nothing is submitted.
| Behaviour | Ready | Developing | Not yet |
|---|---|---|---|
| Running the evaluation harness for the OmniCorp Logistics disruption assistant | |||
| Running the Guardrail Bench for the same scenario | |||
| Pricing and presenting the same lab evidence |
Commit
Commit Statement
Complete the lines before your next AI release review. Use the workflow whose controls you have described most often but tested least directly. Sign it and bring it to the lab.
| Window | Field application |
|---|---|
| Days 1 to 7 | Choose one AI workflow and build a manual harness with golden cases, adversarial cases, and thresholds. |
| Days 8 to 14 | Run the harness, score it, record residual failures, and assign owners for every release-blocking defect. |
| Days 15 to 21 | Run a Guardrail Bench against prompt injection, data exfiltration, tool overreach, and approval bypass. |
| Days 22 to 30 | Rerun failed attacks, price control labour, and present one release record with evidence, owners, and residual risks. |
Learner feedback