AI FOUNDATIONS · COURSE 05
THE AI INSTITUTE
Source-led executive learning · reviewed August 2026
What support would justify relying on the output?
A system can perform well in a demonstration and fail in the work that matters. Reliance requires support built from representative tasks, explicit quality criteria, material failure analysis and a boundary that survives contact with reality.
Your decision
By the end of this module, you will decide whether an evaluation pack is sufficient for a bounded use—and design the smallest additional test that would change a reliance decision.
1. Data fitness is always fitness for a purpose
“We have lots of data” is not a support statement. Data is useful when it is relevant to the task, representative of the intended operating context, sufficiently accurate and complete for the decision consequence, current enough for the claim, and permitted for the proposed use. Each dimension can fail independently.
Relevance asks whether the data contains the information needed for the task. Representativeness asks whether important groups, channels, time periods and difficult cases appear in proportions or deliberate stress samples suitable for the decision. Quality asks whether labels, fields, provenance and missingness can support evaluation. Permission asks whether the organisation may collect, use, disclose and retain the information in this system and jurisdiction. Currency asks whether changes in policy, products, language or behaviour have made the support stale.
A dataset can be accurate but irrelevant, large but unrepresentative, current but unauthorised, or permitted but too incomplete for a consequential decision. Record these as separate findings. The NIST AI Risk Management Framework emphasises that risk and trustworthiness must be considered in context and across the lifecycle [C005-S02]. Australian privacy guidance adds organisation-specific duties and due diligence questions when commercially available AI products process personal information [C005-S05]. This course is not legal advice; obtain qualified advice for your situation.
2. Define the reliance decision before the test
An evaluation is not “see how good it is”. Write the decision it will inform. Will the output be used only to suggest a draft? Will it route work, prioritise human attention, recommend an entitlement or trigger an action? Who is affected if it is wrong? Can a person detect and correct the error before consequence? The same model result can be acceptable for brainstorming and unacceptable for an automatic payment decision.
Turn the decision into criteria. For a policy-answer assistant, criteria might include whether the response addresses the question, whether each material statement is supported by a displayed approved source, whether the source is current, whether the answer expresses uncertainty when sources conflict, and whether escalation occurs for excluded topics. Do not collapse these into one vague score.
Set the verification threshold before seeing results. Pre-agreed thresholds reduce the temptation to move the goalposts after an attractive demonstration. Include a stop condition for a failure whose consequence cannot be tolerated, even if the average is high.
3. Build a representative and deliberately difficult task set
Random sampling from historical work may reproduce the common cases but miss exactly the situations leaders need to understand. Begin with the intended population, then stratify by meaningful dimensions: channel, document type, customer group, language, region, product, complexity, time period and consequence. Add deliberate challenge cases such as conflicting sources, missing information, novel requests, unusual formats and adversarial or ambiguous wording.
Keep evaluation examples separate from examples used to configure prompts, retrieval or rules. Otherwise the system may be tuned to the test. Record the origin, inclusion logic and limitations of each evaluation segment. Documentation practices such as datasheets for datasets and model cards provide useful questions about motivation, composition, uses, evaluation and limitations [C005-S09; C005-S10]. They do not replace an evaluation designed for your own workflow.
Sample size should follow the decision and expected variation, not a universal magic number. Ten hand-picked questions can reveal a defect but cannot establish reliability across a diverse policy corpus. A large undifferentiated sample can also mislead if severe but rare categories disappear in the average. Seek statistical advice for high-consequence estimates and uncertainty intervals.
4. Measure failure, not only success
An overall accuracy or pass rate can hide asymmetry. A missed urgent complaint may matter more than an unnecessary escalation. A false positive in a low-consequence search result differs from a false allegation about an employee. Define material error types and review them separately.
For generative outputs, useful measures may include citation presence, citation support, completeness against an approved checklist, unsupported material claims, correct abstention, consistency and reviewer effort. For extraction or classification, use class-specific precision and recall, severe-miss counts, calibration or confidence behaviour, and exception-routing performance as appropriate. Plain-language decision records are more valuable than technical metrics no owner understands.
Segment every material result. If the average pass rate is 90 per cent but one high-consequence category is 55 per cent, the operating claim must reflect that category. The outcome may be restrict, revise or stop—not necessarily reject the system entirely.
5. Grounding and retrieval help, but do not close the support question
A retrieval component can supply approved documents to a generative system and display links. This can improve relevance and verification, but it creates new evaluation questions. Does retrieval cover the whole intended corpus? Are documents current and permissioned? Does the selected passage actually support the answer? Does the system ignore a conflicting or more authoritative source? Does it abstain when nothing sufficient is retrieved?
Separate source presence from source support. A response can display a real link that does not justify its conclusion. Open the source, inspect the cited passage and compare the answer’s strength with the support. Where currency matters, include a content owner and expiry or review process.
6. Worked case: policy-answer assistant
Composite teaching case
An organisation wants an internal assistant to answer travel and leave questions using approved policy documents. The first evaluation contains ten questions written by the project team. Nine answers are judged “helpful”. The project reports 90 per cent accuracy.
The support review finds that all questions concern standard travel, none concern leave eligibility, conflicting documents, regional variations or recently changed rules. “Helpful” has no rubric. Reviewers do not open citations. One answer confidently combines two valid passages into a conclusion neither supports.
The team rebuilds the test around the intended use. It maps policy topics and consequence, samples common questions, and deliberately adds ambiguity, conflicts, missing coverage, current changes and excluded employee-relations issues. Two authorised reviewers independently score answer relevance, citation support, completeness, uncertainty and escalation. Materially unsupported entitlement advice is a stop condition. Results are segmented by topic and difficulty. The initial operating boundary is a draft answer with source display for reviewer confirmation; the system cannot determine individual eligibility.
7. Contrast case: ten selected questions
Misconception
A product owner chooses ten questions that were already used while configuring the assistant. The answers look good, so the team declares the system accurate. No baseline, independent review, failure taxonomy or excluded-use test is recorded.
The exercise demonstrates that the configured system can answer those familiar questions plausibly. It does not support reliance across the policy workload. Repair it by separating configuration and evaluation sets, mapping the intended population, adding realistic and difficult cases, defining criteria and making the reliance threshold explicit.
8. Practice: audit support sufficiency
Produce support
Use the audit template on a fictional or de-identified use. Record the intended use and consequence, available data, missing segments, permissions/currency check, representative tasks, quality criteria, material failures, baseline, reviewer process, threshold, stop condition, owner and monitoring trigger.
Completion support: five representative test cases, one deliberately difficult case, one segmented measure, one material stop condition and one clear statement of what the support would still not prove.
Download editable support sufficiency audit (CSV)9. Decision summary
Support is sufficient only for a named use within a named boundary. Start with consequence, inspect data fitness, build realistic and difficult tasks, define failure-sensitive measures and pre-agree the gate. Report uncertainty and uncovered segments as results—not as embarrassing footnotes.
Transfer prompt
Ask the owner of one proposed AI use: “Which result would make us restrict or stop?” Use de-identified examples. Do not upload organisational datasets, policy files or personal information to Moodle or the Learning Partner.