Module 1 · Foundation
What can you credibly claim this system can do?
A good AI decision begins with a precise claim. “It understands our customers” is not precise. “It groups these complaints into themes that a service manager can verify” can be tested.
Decide whether an AI capability claim is defensible now, needs a test in your context, or is misleading—and locate the accountable human decision around it.
1. Start with the whole system, not the magic word
At business level, AI is useful shorthand for systems that infer how to produce outputs—such as a prediction, recommendation, classification, summary or generated response—from inputs and objectives. This shifts the conversation away from whether a product “thinks” and towards observable behaviour.
A model is only one component. It transforms inputs into outputs. A product wraps models with an interface, instructions, retrieval, permissions and monitoring. A workflow connects the product to people, data and hand-offs. An organisational decision is the consequential choice: whether to contact a customer, approve a payment, change a policy or allocate work.
A model may classify a complaint accurately in a test set. That does not prove a live triage system will route complaints reliably, because the interface, input quality and exception rules matter. Even a reliable triage system does not prove customer outcomes will improve, because staffing, response authority and follow-through sit in the workflow. Keep the claim at the layer actually supported by support.
2. Predictive and generative outputs require different questions
A predictive or classification system estimates a label, score or likely outcome: for example, grouping a message by topic or estimating demand. A generative system produces new text, images, audio or code by continuing patterns learned from data and context. Both can be useful; neither arrives with truth, relevance or permission guaranteed.
For a classifier, ask how the categories were defined, whether evaluation examples represent the live population, which errors are most consequential, and what happens when confidence is low. For a generative system, ask whether the response needs sources, how fabricated or unsupported content will be detected, what information may enter the service, and who verifies the output before it affects anyone. “The answer sounds fluent” is never an evaluation criterion.
Outputs vary when inputs, context or system settings change. This does not make the system unusable. It means the operating design must tolerate variation and evaluation must sample realistic tasks—including awkward exceptions—not merely repeat a polished demo.
3. A capability claim has four parts
- Task: observable work with a verb and a bounded object—for example, classify incoming messages into approved service categories.
- Support: the test that supports reliance—performance on representative messages, segmented by channel and category, with review of material errors.
- Boundary: the context outside which the claim does not apply—approved English-language channels, current category definitions, no automatic customer response.
- Accountable decision: the person who accepts, overrides or acts on the output—the service operations manager owns routing policy; agents can correct a route.
Claims without these parts invite capability inflation. “The assistant understands policy” confuses fluent output with verified coverage. “The tool will save 30 per cent” jumps from model activity to realised value. “Human in the loop” names neither a decision nor the proof that person receives.
4. Worked case: customer feedback triage
A national service organisation receives feedback through email, web forms and call-centre notes. The team is considering a classifier to assign one of twelve approved themes and a generative model to draft a weekly synthesis. A service analyst checks low-confidence classifications and samples routine ones. The customer experience manager decides which recurring issues enter the improvement backlog.
Weak claim: “The platform understands our customers and tells us what to fix.”
Decision-quality claim: “For current English-language feedback in three approved channels, the system can propose one of twelve themes and draft a source-linked synthesis. We will test theme accuracy and missed high-severity complaints on a representative sample. Analysts verify exceptions and source links; the customer experience manager decides priorities.”
The revised claim does not promise customer improvement. It separates classification, generation and management judgement. It also makes a support request possible: sample design, category-level error rates, severity misses, source-link coverage and analyst override rate.
| Input | System output | Support needed | Human decision |
|---|---|---|---|
| Feedback text from approved channels | Proposed theme and source-linked synthesis | Representative accuracy; severe-issue recall; source-link validity; exceptions | Analyst corrects routing; manager prioritises service changes |
5. Contrast case: the impressive demonstration
A vendor loads six carefully chosen reviews and produces an elegant summary. The sponsor concludes that the product “understands our entire customer base”. No one tests multilingual messages, call notes, duplicated complaints, sarcasm, novel issues or source traceability.
The demonstration establishes that the configured system produced plausible outputs for six examples. It does not establish representative performance, operational fit or value. The right response is a bounded test: agree the categories, construct a representative and deliberately difficult evaluation set, blind the review, define material error types and set an verification threshold before live use.
Interactive decision scenario
You have one week before the steering meeting
A vendor’s six-example demo impressed the sponsor. You need to recommend the next step. What do you do?
6. Practice: diagnose and rewrite
Sort each statement into defensible now, test in our context, or misleading. Choose two important claims and rewrite them with task, support, boundary and accountable decision.
- “The chatbot answers staff policy questions accurately.”
- “The classifier assigns an approved category to 92 of 100 held-out test messages.”
- “The copilot eliminates administrative work.”
- “The draft remains subject to the authorised officer’s source check and approval.”
- “Retrieval means the answer cannot hallucinate.”
- “For these three document types, the extractor proposes five fields for clerk verification.”
Completion support: two rewritten claims, one named evaluation method, one operating boundary and one accountable decision.
7. Decision summary
Do not ask whether “AI works” in the abstract. Ask what component performs which task, what support supports reliance in the intended context, where the boundary sits, and who remains accountable for the consequential decision. A credible leader can say “we do not know yet” and turn that uncertainty into a proportionate test.
Rewrite one live claim used in your organisation. Use a de-identified description. Do not paste customer, employee or confidential material into Moodle or an AI assistant.