TeleosynLearning Free access

Lesson 4 · AI Foundations

Verification, hallucinations, and accountability

Length
28 minutes across 7 sections
You will be able to apply
The Evidence Contract
You will produce
Commit Statement
You will work
2 gated questions
Reading

Which artifact settles this claim, and who opens it?

Core question

Generative systems invent facts, sources, quotations, statistics and legal citations, and present all of them in the register they use for correct answers. This is not a defect that will be patched out. It is the direct consequence of a system optimised to produce probable text rather than true statements, and probable text includes the citation that ought to exist.

Because the model will not mark its own inventions, marking them is your work. And the person who acts on an AI-generated claim owns that claim exactly as if they had written it themselves. There is no version of this where the tool is accountable.

Next: The pathology: Format Credibility

The pathology: Format Credibility

A fabricated citation looks like a citation. It has an author, a year, a plausible title and a journal that exists. A fabricated statistic carries a unit and a comparison period. A fabricated quotation has quotation marks around it. The signals people use to assess credibility at speed are all signals of form, and form is precisely what a language model reproduces perfectly.

PrincipleFormat Credibility is the reflex of accepting a claim because it is shaped like a verified one. It is the specific mechanism by which fabrications survive review: the reviewer is not lazy, they are pattern-matching on the wrong feature. A correctly formatted reference is evidence that the model knows what a reference looks like. It is evidence of nothing else.

The defence is not to distrust everything, which is unaffordable, but to decide in advance which claims require settling and what would settle them. Verification is a plan, not an attitude.

Next: The Evidence Contract

The Evidence Contract

Definition

An Evidence Contract is written before generation and states, for each class of claim the output will contain, the artifact that settles it and the person who opens that artifact. Three columns: claim class, settling artifact, opener. A claim class with no settling artifact does not get published. A settling artifact with no named opener does not get opened.

The contract is deliberately about classes rather than claims. You cannot enumerate the claims before the output exists. You can always enumerate their kinds.

When to Use It

Write a contract for any output that will be acted on, sent, published, or filed. That is a lower bar than it sounds: a summary that changes what a colleague does is being acted on. The contract takes four minutes to write and is reusable for every future run of the same task.

Rewrite it when the output starts carrying a new class of claim. A drafting task that begins producing numbers has acquired a class it did not have last month.

How to Apply It

  1. List the claim classes the output will contain: names, figures, dates, quotations, citations, obligations, causal explanations.
  2. For each class, name the one artifact that settles it, and prefer the original over any summary of it.
  3. Name the person who opens that artifact, and make sure it is not the person who generated the output where the stakes justify separation.
  4. Delete any claim whose class has no settling artifact, rather than softening its wording.
Interactive modelEvidence Contract workflowflow · 4 elements
01
List claim classes

Enumerate the kinds of claims the output will contain.

The contract is written before generation so verification is a plan rather than an afterthought.
List claim classes
Enumerate the kinds of claims the output will contain.
Name settling artifacts
For each class, identify the original that would prove or disprove it.
Assign openers
Name the person who opens each artifact, separate from the generator.
Delete unsettleable claims
Remove any class with no settling artifact rather than softening it.
Claim classSettling artifactOpened by
Figures and percentagesThe source report, with the arithmetic redoneThe analyst, not the drafter
Quotations and commitmentsThe recording or the signed documentThe reviewer
Citations and referencesThe cited document itself, openedWhoever will be asked about it
Obligations and entitlementsThe contract, policy or regulation textThe accountable owner
Causal explanationsNo artifact settles theseDelete or attribute as an untested hypothesis

Worked example 1 of 3

Alan Brixmoor at OmniCorp Public was preparing a briefing on take-up of a benefits programme. The AI-drafted summary included the sentence that take-up had risen to sixty-two percent, up eighteen percent year on year. It read like every other sentence in the document.

Alan Brixmoor
What class is that sentence?
Policy analyst
Two classes. A figure and a comparison.
Alan Brixmoor
So which artifact settles it, and who opens it?
Policy analyst
The quarterly take-up extract. I can open it now.
Alan Brixmoor
Then open it and redo the arithmetic. Do not check whether the number appears. Recalculate it.
Policy analyst
Take-up is sixty-two percent. The prior year was fifty-seven. That is a five point rise, not eighteen percent.

The figure was real. The comparison was invented, and it was invented in exactly the form a real comparison takes. Had the analyst checked only that sixty-two percent appeared in the source, the briefing would have gone to a minister with a fabricated growth rate in it.

Why This Works

Splitting claims into classes defeats Format Credibility because it moves the decision away from the sentence. You are no longer asking whether this looks right. You are asking which column it belongs to, which is a question your judgment cannot be talked out of by good prose.

Naming the opener defeats the other half. Verification that belongs to everyone belongs to no one, and the person who generated an output is the worst possible reviewer of it: they already know what it was supposed to say, so they read what they intended rather than what is there.

Worked example 2 of 3Optional depth

Marisa Delgado at OmniCorp Financial asked for a summary of regulatory obligations affecting a servicing change. The draft cited three guidance documents. Two existed. The third had a real issuing body, a plausible reference number and a title that described precisely the thing Marisa was hoping to find. Her contract required that citations be opened by the person who would be asked about them, which was her. She opened them. Two took a minute; the third took four, because she kept assuming she was searching badly.

Worked example 3 of 3Optional depth

Dr. Naomi Ellery at OmniCorp Health received an AI-drafted literature summary containing a causal explanation: that a change in discharge protocol had reduced readmissions. The protocol change was real and the readmission drop was real. Nothing in the source supported the link between them. Her contract classes causal explanations as unsettleable, so the sentence was rewritten as two observations and one explicitly labelled hypothesis. Nothing was lost except a claim nobody could defend.

Edge Cases and NuancesOptional depth

Some artifacts are themselves AI-generated, which makes them worthless as settling evidence; trace to a human-authored or system-of-record original or treat the class as unsettleable. Some claims are true but unverifiable within your access, and the honest handling is to attribute rather than assert. And some contracts fail not because the artifact was wrong but because the opener was nominal: a named reviewer who is copied on an email has not opened anything. If you cannot say when the artifact was last opened, the contract is decorative.

Interactive modelVerification strengthladder · 3 elements
01
Presence checking

Confirms that a token appears in the source, but not the relationship around it.

Each rung represents a stronger form of verification, from weakest to strongest.
Presence checking
Confirms that a token appears in the source, but not the relationship around it.
Value confirmation
Opens the artifact and confirms the figure matches.
Claim settling
Recomputes the relationship, including comparisons and causal links.

Knowledge check

An AI-drafted briefing states that programme take-up rose to sixty-two percent, up eighteen percent year on year. An analyst confirms that sixty-two percent appears in the source extract. Is the claim verified?

Answer first, then check.
Next: Common Failure Modes

Common Failure Modes

Failure modePresence checking. What it looks like in the moment: you search the source document for the number, find it, and move on. Your eyes go to the figure and stop there. You have confirmed that the number exists somewhere in the source, and you have confirmed nothing about the sentence it appears in. The cost when this happens: fabricated relationships between real values survive review intact, and they are the most damaging class of error because every component is verifiable. The correction: settle the claim, not the token. If the sentence contains a comparison, recompute the comparison.
Failure modeSelf-review. What it looks like in the moment: the person who ran the prompt also checks the output, usually within a minute of generating it, often in the same window. They scroll rather than open anything. They are not cutting corners: they simply know what the text was meant to say, and they read that. The cost when this happens: the specific errors that survive are the ones consistent with the author's expectation, which are also the ones a reader is least likely to question. The correction: for any claim class where being wrong costs somebody else, the opener is not the generator. This is not negotiable.
Next: The contract end to end

The contract end to end

Trevor Okafor at OmniCorp Retail was asked to produce a quarterly supplier performance summary from four systems, and wanted AI to assemble the narrative. He wrote the contract before he wrote the prompt.

Four classes. Delivery figures, settled by the logistics extract, opened by the logistics analyst. Quality rates, settled by the returns ledger, opened by the quality lead. Contractual obligations, settled by the supplier agreement, opened by Trevor. Causal explanations of why a supplier slipped, settled by nothing, therefore labelled as hypotheses or removed.

The first draft contained nine claims across the four classes. Seven settled cleanly. One delivery figure was correct but attributed to the wrong quarter. One paragraph explained a quality decline by a change of manufacturing site, which was true and had nothing to do with the decline; it was rewritten as a hypothesis and flagged for the supplier review meeting. Total verification time: nineteen minutes, of which fourteen were spent on the class Trevor had least expected to be wrong. He kept the contract and reused it every quarter.

Decision point

Priya Raghunathan at OmniCorp Logistics receives an AI-drafted incident narrative for a customer. It states that the delay was caused by a port closure, cites a port authority advisory with a reference number, and gives the closure window. Priya recognises the port and remembers the closure. She has ten minutes before the customer call. What does she verify?

Confidence before seeing the analysis
Commit, calibrate, and name contrary evidence first.
Next: Self-check

Self-check

Knowledge checkTake the last AI-assisted document you sent. List its claim classes. For each, name the artifact that settles it and the person who opened it, with the date. Any row you cannot complete is a claim you published on the strength of its formatting.

Mark the level that describes you today. Nothing is submitted.

BehaviourReadyDevelopingNot yet
Classifying claims
Settling rather than presence checking
Separating generator from opener
Next: Commit

Commit

Commit Statement

Complete every line in your own words, then sign and date it. Keep your first Evidence Contract with it, so the commitment has something concrete attached.

WindowField application
Days 1 to 7Write an Evidence Contract for one recurring AI-assisted output. Three columns, no more than six rows.
Days 8 to 21Run the contract on every instance of that output and record which class fails most often.
Days 22 to 30Move the opener for your highest-stakes claim class to someone who did not generate the output, and keep it there.

An Evidence Contract makes your own verification honest and repeatable. It does not create the independent assurance function, the evidence retention, or the audit trail that lets an organisation prove to a regulator that verification actually happened on a given day. Building that provable layer is the capability the paid programs develop next.

DisclaimerGeneral guidance only. All organisations and people named in this lesson are fictional. Regulated organisations should confirm requirements with a qualified professional before relying on this material.
Required practice must be complete.