Module 1 · Foundation

What can you credibly claim this system can do?

A good AI decision begins with a precise claim. “It understands our customers” is not precise. “It groups these complaints into themes that a service manager can verify” can be tested.

Your decision
Decide whether an AI capability claim is defensible now, needs a test in your context, or is misleading—and locate the accountable human decision around it.

1. Start with the whole system, not the magic word

At business level, AI is useful shorthand for systems that infer how to produce outputs—such as a prediction, recommendation, classification, summary or generated response—from inputs and objectives. This shifts the conversation away from whether a product “thinks” and towards observable behaviour.

A model is only one component. It transforms inputs into outputs. A product wraps models with an interface, instructions, retrieval, permissions and monitoring. A workflow connects the product to people, data and hand-offs. An organisational decision is the consequential choice: whether to contact a customer, approve a payment, change a policy or allocate work.

ModelProduces an output from an input.
SystemAdds data, rules, interface and controls.
WorkflowChanges tasks, hand-offs and review.
DecisionCreates consequences for the organisation or people.

A model may classify a complaint accurately in a test set. That does not prove a live triage system will route complaints reliably, because the interface, input quality and exception rules matter. Even a reliable triage system does not prove customer outcomes will improve, because staffing, response authority and follow-through sit in the workflow. Keep the claim at the layer actually supported by support.

Capability claim model showing task, support, boundary and accountable decision as four gates between an AI output and a defensible business claim.
Capability claim model. A defensible claim names the task, support, operating boundary and accountable decision.

2. Predictive and generative outputs require different questions

A predictive or classification system estimates a label, score or likely outcome: for example, grouping a message by topic or estimating demand. A generative system produces new text, images, audio or code by continuing patterns learned from data and context. Both can be useful; neither arrives with truth, relevance or permission guaranteed.

For a classifier, ask how the categories were defined, whether evaluation examples represent the live population, which errors are most consequential, and what happens when confidence is low. For a generative system, ask whether the response needs sources, how fabricated or unsupported content will be detected, what information may enter the service, and who verifies the output before it affects anyone. “The answer sounds fluent” is never an evaluation criterion.

Outputs vary when inputs, context or system settings change. This does not make the system unusable. It means the operating design must tolerate variation and evaluation must sample realistic tasks—including awkward exceptions—not merely repeat a polished demo.

3. A capability claim has four parts

  1. Task: observable work with a verb and a bounded object—for example, classify incoming messages into approved service categories.
  2. Support: the test that supports reliance—performance on representative messages, segmented by channel and category, with review of material errors.
  3. Boundary: the context outside which the claim does not apply—approved English-language channels, current category definitions, no automatic customer response.
  4. Accountable decision: the person who accepts, overrides or acts on the output—the service operations manager owns routing policy; agents can correct a route.

Claims without these parts invite capability inflation. “The assistant understands policy” confuses fluent output with verified coverage. “The tool will save 30 per cent” jumps from model activity to realised value. “Human in the loop” names neither a decision nor the proof that person receives.

4. Worked case: customer feedback triage

Composite teaching case

A national service organisation receives feedback through email, web forms and call-centre notes. The team is considering a classifier to assign one of twelve approved themes and a generative model to draft a weekly synthesis. A service analyst checks low-confidence classifications and samples routine ones. The customer experience manager decides which recurring issues enter the improvement backlog.

Weak claim: “The platform understands our customers and tells us what to fix.”

Decision-quality claim: “For current English-language feedback in three approved channels, the system can propose one of twelve themes and draft a source-linked synthesis. We will test theme accuracy and missed high-severity complaints on a representative sample. Analysts verify exceptions and source links; the customer experience manager decides priorities.”

The revised claim does not promise customer improvement. It separates classification, generation and management judgement. It also makes a support request possible: sample design, category-level error rates, severity misses, source-link coverage and analyst override rate.

InputSystem outputSupport neededHuman decision
Feedback text from approved channelsProposed theme and source-linked synthesisRepresentative accuracy; severe-issue recall; source-link validity; exceptionsAnalyst corrects routing; manager prioritises service changes

5. Contrast case: the impressive demonstration

Misconception

A vendor loads six carefully chosen reviews and produces an elegant summary. The sponsor concludes that the product “understands our entire customer base”. No one tests multilingual messages, call notes, duplicated complaints, sarcasm, novel issues or source traceability.

The demonstration establishes that the configured system produced plausible outputs for six examples. It does not establish representative performance, operational fit or value. The right response is a bounded test: agree the categories, construct a representative and deliberately difficult evaluation set, blind the review, define material error types and set an verification threshold before live use.

Interactive decision scenario

You have one week before the steering meeting

A vendor’s six-example demo impressed the sponsor. You need to recommend the next step. What do you do?

6. Practice: diagnose and rewrite

Sort each statement into defensible now, test in our context, or misleading. Choose two important claims and rewrite them with task, support, boundary and accountable decision.

  • “The chatbot answers staff policy questions accurately.”
  • “The classifier assigns an approved category to 92 of 100 held-out test messages.”
  • “The copilot eliminates administrative work.”
  • “The draft remains subject to the authorised officer’s source check and approval.”
  • “Retrieval means the answer cannot hallucinate.”
  • “For these three document types, the extractor proposes five fields for clerk verification.”

Completion support: two rewritten claims, one named evaluation method, one operating boundary and one accountable decision.

Download the editable capability claim map (CSV)

7. Decision summary

Do not ask whether “AI works” in the abstract. Ask what component performs which task, what support supports reliance in the intended context, where the boundary sits, and who remains accountable for the consequential decision. A credible leader can say “we do not know yet” and turn that uncertainty into a proportionate test.

Transfer prompt

Rewrite one live claim used in your organisation. Use a de-identified description. Do not paste customer, employee or confidential material into Moodle or an AI assistant.