“It worked on the example” is not a test plan. AI workflow quality depends on the range of inputs, the cost of mistakes, and the path a person takes when the output is wrong.
You do not need a large laboratory to begin. You need a small, representative set of examples and a written definition of acceptable behavior.
Build an evaluation set
Collect examples that are safe to use and remove unnecessary personal information. Include ordinary cases, ambiguous cases, long inputs, empty fields, formatting surprises, and examples that previously caused a correction. Store the expected behavior, not only the input.
Score the behavior that matters
Choose measures that match the workflow:
- Extraction: field-level accuracy and evidence quality.
- Classification: route correctness and safe handling of unknowns.
- Summaries: factual consistency and useful omission of noise.
- Actions: whether the system correctly paused for approval.
A single overall score can hide a dangerous failure. Break out high-impact cases and set a stricter threshold for them.
Compare versions, not vibes
When you change a prompt, model, parser, or source document, run the old and new versions against the same set. Save the outputs and mark regressions. If a reviewer cannot tell what changed, the process will eventually depend on memory and anecdotes.
Turn corrections into tests
- Capture the input, output, correction, and why the correction mattered.
- Remove sensitive details and generalize the example where possible.
- Add it to the evaluation set with the desired behavior.
- Run it whenever the workflow changes.
- Review the set periodically so it still represents current work.
Watch production without spying on people
Measure aggregate corrections, failure rates, latency, and escalation volume. Limit access to raw inputs, set retention periods, and tell people what is logged. A quality program should improve reliability without becoming an invisible surveillance program.
- The test set covers edge cases and past failures.
- Expected behavior is written before comparing outputs.
- High-impact errors are reviewed separately.
- Every material change runs a regression check.
- Reviewers can turn the workflow off safely.