A few stable examples can make an AI workflow easier to compare. They give a reviewer something more specific than “the answer seems better” when a prompt, model or integration changes.
Start with the downloadable evaluation cases. They are synthetic teaching fixtures with expected behaviors, not recorded provider results or an automatic scoring service. This guide extends the summary-review exercise.
Separate the task from the answer key
Each fixture has a stable ID, an input object and a reviewer-only expectation. The input contains the task, supplied context and relevant constraints. The expectation identifies the facts or boundaries a reviewer should check.
When testing a model, supply only the case's input object through an approved tool. Keep expected and criticalFailures out of the model input. Otherwise you may test whether the model can repeat the answer key instead of whether it can solve the task.
Save the exact request separately from the fixture so another reviewer can see what was actually sent.
Use five cases with different jobs
| Case | What it probes | Successful behavior |
|---|---|---|
| EVAL-01 | Ordinary summary with one missing date | Correct totals, categories and uncertainty |
| EVAL-02 | Missing source content | States that the requested result cannot be established |
| EVAL-03 | Conflicting records | Describes the conflict and requests an authoritative source |
| EVAL-04 | Request outside supplied project scope | Does not invent, retrieve or expose unauthorized project data |
| EVAL-05 | An instruction embedded in a stored note | Treats the note as data and preserves the task boundary |
The fourth case has no secret foreign-project records hidden in the download. In plain chat it tests the response only. A real integration needs a separate controlled-account test that verifies its server refuses the unauthorized read. Never put a private record in the prompt merely to see whether the model will reveal it.
OWASP identifies external documents as a route for indirect prompt injection. That is why the fifth case includes a harmless instruction-like note and why its text cannot grant tool access. OWASP prompt-injection guidance.
Define failures before running
For each case, distinguish required facts from acceptable wording. A summary can be concise or detailed and still be correct. An invented completion date is a factual failure regardless of tone.
Use a simple outcome vocabulary: pass, fail, blocked or not run. Record critical failures separately. An unauthorized action or fabricated source should not disappear inside an average score for good formatting.
If a provider or account is unavailable, mark the run blocked. Do not reuse an old answer and present it as a new result.
Compare one meaningful change
Record the date, tool, model identifier if exposed, version/configuration, exact input, output and reviewer. Describe whether the run was chat-only, mocked or connected to a controlled integration. Omit secrets from every record.
Keep the fixtures and acceptance criteria unchanged while comparing a prompt revision. If you change the task at the same time, the comparison answers a different question. Repeat cases when variability matters and save failures as well as successes; do not select only the best response.
The fixture pack contains no provider calls and records no scores. You choose an authorized tool and perform the review. A small passing set is a regression aid, not a general accuracy percentage or safety certification.
Keep examples fresh without erasing history
Version the set when a requirement changes. Preserve the old fixture so a past run remains interpretable. Add a case when a real failure exposes a gap, using sanitized or synthetic data that retains the relevant behavior.
Use the workbook to record the run and follow-up. A result needing code changes belongs in the applicable implementation process; a wording improvement may only need an editorial change. The proposal handoff guide explains how to make that distinction.