All articles

AI for CIOs

How do you build an AI test set that tells you something useful?

ShareLinkedIn

SimplSolutions editorial team · Answer quality · 3 min read

Published

AI-assisted original editorial guidance. Calculations and scenarios are illustrative, not customer results.

Reviewers comparing sample questions with approved documents

Sample the work, not the showcase

Start with a question family employees actually ask. Collect permitted examples from different ordinary contexts, then remove private information the test does not require. Label synthetic examples clearly. A test set containing only perfectly phrased questions can make a system look better than the experience of people who omit a location or use an old term.

For each case, write the intended audience, applicable source, expected result and reviewer before running the candidate system. Expected results may include answer, clarify, deny or escalate. Do not let the vendor's response define what correct means after the fact.

Include different kinds of difficult cases

Case familyExample purpose
Ordinary supported questionVerify the relevant instruction and source
Missing contextRequire clarification when applicability changes
No evidenceAvoid an invented answer
Conflicting sourcesPreserve conflict for owner resolution
Restricted audienceDeny without disclosing private facts
Source update/outageTest freshness and honest failure handling

Keep categories separate in the report. A high ordinary-answer score does not cancel a denied-user failure. Your owners determine which cases are hard release gates and which shortcomings can be accepted within a narrower scope. Decide those rules before seeing the scores.

Technology colleagues reviewing a workflow implementation together

Illustrative editorial photograph, not a customer result.

Hold back cases from configuration work

Use one set to investigate and improve behavior and another to evaluate the candidate release. If the team repeatedly adjusts the prompt using every test case, passing those cases may partly demonstrate memorization of the exercise. Add fresh representative questions and preserve a held-out set reviewers can inspect independently.

Do not promise statistical certainty from a tiny sample. Record sample size, selection method and task scope. A small set can identify specific failures and justify another bounded test. It cannot establish company-wide accuracy across every source and audience.

Score support at the claim level

Open the cited source and locate the supporting passage for each material instruction. Record source currency, applicability, permission behavior and answer support separately. Also record review effort. A response that is factually correct but requires ten minutes to verify may not meet the operating goal.

Use a second reviewer on a subset of cases to expose ambiguous criteria. Resolve disagreements with the source or process owner. If two experts cannot agree which guide applies, treat that as a governance gap rather than expecting a model to resolve it silently.

Version the test result with the release

Record the source configuration, relevant model settings, tool scope, expected result and observed evidence. Retest when those dependencies change. Maintain known incident cases as regression tests, but do not let them become the only evaluation material. The next ordinary employee question matters too.

Put this to work this week

Write ten ordinary cases and five deliberate failure cases with your task owner, using only permitted information. Assign expected answer, clarification, denial or escalation before running the candidate. Hold back a subset from configuration work. Have a second reviewer inspect a few results and discuss disagreements. Preserve both the result and release reference. Add new cases from actual operating gaps without replacing difficult failures with easier questions merely to improve the average.

Use the AI Knowledge Access Test Sheet and the claim-support guide to prepare the review. The NIST AI RMF provides broader measurement and risk-management context; this test set is not certification. Request an evaluation-led demo and bring cases the demonstration was not built around. SimplSolutions should show the expected supported answer and the appropriate hold or escalation, with the evidence available to review.

Sam, your AI guide

Your role. Your questions.

Need CIO guidance?
Ask Sam.

Talk through an idea, ask about the tools you already use, or find out what a first project could look like.

Sam is a fictional campaign character and AI guide. Our team handles demo requests.