AI FOUNDATIONS · COURSE 05

THE AI INSTITUTE

Source-led executive learning · reviewed August 2026

What is the smallest test that will change a real decision?

An experiment is not a small rollout. It is a bounded support system: a hypothesis, comparator, representative work, measures, controls and a decision gate agreed before results are known.

Your decision

By the end of this module, you will write a 30-day experiment charter that can support a scale, revise or stop decision without pretending the optional 30-day run has already occurred.

1. Name the decision, not only the ambition

“Explore AI for productivity” cannot be tested. A useful experiment begins with a decision the owner must make at the end: whether to expand a bounded writing-assistant practice to another proposal team, revise the source-verification workflow, or stop because quality and confidentiality controls do not hold.

Write a hypothesis that links a specific intervention to observable outcomes and the assumptions connecting them. For example: “For eligible internal proposal drafts, an approved source-linked assistant plus a review rubric will reduce median first-draft time without increasing critical source errors. If quality, adoption and capacity-redirection gates pass after four weeks, the proposal director will decide whether to extend the test.” This is falsifiable. “AI will transform proposals” is not.

The smallest useful test targets the uncertainty most likely to change that decision. If output quality is unknown, evaluate representative tasks before workflow integration. If adoption and review behaviour are uncertain, run a contained workflow experiment. Do not expose customers or employees merely to make a test feel real.

Experiment learning loop moving through decision and hypothesis, baseline and design, bounded run with controls, support review, and scale revise or stop, then recording observed result, interpretation and next decision.
Experiment learning loop. Controls surround the run; support changes a named decision. A full text alternative follows the lesson.

2. Establish the baseline and comparator

A baseline describes the current method using the same population and measures proposed for the intervention. Record volume, time period, quality, exceptions and relevant segments. Historical data may be adequate for a low-consequence learning test if definitions are stable; otherwise collect a short prospective baseline.

A comparator helps separate intervention effects from demand, seasonality, team composition or another change. It may be the same participants alternating eligible tasks, a similar team continuing the current method, or a before-and-after design with explicit limitations. Do not withhold a necessary service or protection to create a control group. For consequential settings, seek methodological and specialist review.

Predefine inclusion and exclusion. “All proposals” hides variation. “Internal, English-language proposals using three approved service families; exclude regulated advice and client-confidential attachments” is operational. Record why exclusions exist and what support would be required to expand them.

3. Select representative participants and tasks

Choose participants who reflect the intended user roles, experience and workflow conditions. A sponsor and two enthusiasts cannot establish ordinary adoption. Participation should be informed; users need to know the purpose, permitted data, monitoring, how results affect them and how to report a concern. A voice reflection may capture judgement but cannot grade accent, speech style or disability; provide a text alternative.

Sample common tasks, meaningful segments and deliberately difficult exceptions. Keep configuration examples separate from evaluation. If the system changes during the experiment, record the version and decide whether support remains comparable. The Australian Government Guidance for AI Adoption and NIST AI RMF materials support lifecycle, testing, monitoring and accountability practices [C005-S04; C005-S13; C005-S02]. Apply them proportionately and verify current versions.

4. Measure value, quality, adoption and risk together

One headline metric cannot govern the decision. Select a small balanced set. Value-path measures connect the task change to the workflow outcome, such as usable capacity redirected or end-to-end cycle time. Quality measures use an observable rubric and material-error taxonomy. Adoption measures show eligible use, persistence, override and review behaviour. Risk measures capture prohibited-data events, severe errors, affected-person concerns, security events and boundary breaches.

Define the unit, population, calculation, source, owner and threshold for every measure. Median time is often more informative than an average when a few difficult tasks dominate. Segment material results. Record reviewer effort and downstream rework so apparent task savings are not double counted.

Do not optimise a metric that undermines the actual outcome. A target for drafts per hour can encourage unnecessary drafts. A target for low override may suppress appropriate challenge. Pair speed with quality and give reviewers permission to stop.

5. Put controls inside the experiment

The charter incorporates the safe-use boundary from Module 4: approved environment, permitted information, excluded decisions, verification, affected-person protection where relevant, escalation triggers, containment and owners. Controls are part of what the experiment evaluates. If the required review effort is impractical, that is a legitimate result.

Assign one accountable experiment owner, a workflow owner, a support reviewer and relevant privacy, security, people or legal contacts. Separate vendor support from decision ownership. Agree who can pause the experiment immediately and who may approve resumption.

6. Pre-agree scale, revise and stop

Write the gate before the run. A scale decision requires all material thresholds and controls to pass for the defined boundary; it does not authorise enterprise rollout. A revise decision names the failed assumption, changed design and support needed before another gate. A stop decision can follow an uncontained severe failure, disproportional review burden, lack of usable value or inability to operate within the approved boundary.

A mixed result is normal. The writing assistant may reduce draft time and improve consistency while source verification fails for one service family. The correct decision may be to restrict the supported corpus, redesign citations and retest—not average the results into a pass.

7. Worked case: four-week proposal experiment

Composite teaching case

Eight proposal specialists volunteer across two teams. Eligible work is internal first drafts for three approved service families. The assistant runs in a managed environment and may use de-identified discovery notes plus approved service content. It cannot process client-confidential attachments or generate regulated advice.

The current-method baseline covers 40 recent eligible proposals. The four-week run targets 40 comparable drafts. Measures are median draft and review time, accepted-first-draft rubric result, critical source error count, eligible-use adoption, overrides, prohibited-data events and usable capacity redirected to qualified discovery. Reviewers open every citation. A critical unsupported service claim triggers immediate pause.

Scale means extend only to one additional trained team if there are no critical source errors, quality meets the threshold, total draft-plus-review time improves and the manager verifies usable capacity redirection. Revise means restrict a failing service family or change retrieval/review and repeat the affected support. Stop means an uncontained confidentiality event, persistent critical errors or no usable outcome after review effort.

8. Contrast case: enterprise rollout called a pilot

Failure pattern

Every employee receives a licence. No baseline, eligible-task definition, participant communication, approved data boundary, version record, quality rubric or stop authority exists. Success is measured by logins and anecdotes. After three months, leaders call the activity a pilot and ask whether it worked.

This is an uncontrolled rollout. Narrow the decision, stop unsupported uses, establish boundaries and support, and create a contained experiment. Historical login data cannot reconstruct missing quality and consequence support.

9. Communicate support without spin

The learning update separates observed result, interpretation, uncertainty and next decision. “Median total handling time fell in the eligible sample” is observed. “The change appears related to the new workflow” is interpretation. “Demand mix and volunteer participants limit transfer” is uncertainty. “Restrict one service family and repeat the citation test” is a decision.

Report negative and mixed findings. A stopped experiment can be valuable if it prevents unsafe scale or identifies an enabling dependency. Preserve the decision record, version, support and owner so later teams do not repeat the same uncertainty.

10. Practice: write the charter

Signature artefact

Complete the charter and learning-update tool using a fictional or de-identified opportunity. Review it against the published rubric. Ask the Learning Partner to identify missing support, but do not ask it to write your assessed submission.

Completion support: a named decision and owner; falsifiable hypothesis; baseline/comparator; participants/tasks; balanced measures; boundary and controls; and pre-agreed scale/revise/stop rules.

Download experiment charter (CSV) Download learning update (CSV)

11. Decision summary

A disciplined experiment converts uncertainty into a proportionate decision. Keep the test bounded, measurement balanced, controls real and gate pre-agreed. Planning the 30-day run completes this course artefact; running it remains an optional transfer activity.

Transfer prompt

Send the charter to the accountable owner for a go/no-go review. Do not upload confidential data, real incidents or identifiable participant information to Moodle or the Learning Partner.

Review the annotated exemplar

Educational material, not legal advice. Do not enter confidential, personal or commercially sensitive information into an unapproved AI service.