Voltar ao blog

Build a Small AI Evaluation Set You Can Reuse

13 de setembro de 20266 min
Gestão de fluxos de trabalho

Este artigo ainda não foi traduzido — exibindo o original em inglês.

A few stable examples can make an AI workflow easier to compare. They give a reviewer something more specific than “the answer seems better” when a prompt, model or integration changes.

Start with the downloadable evaluation cases. They are synthetic teaching fixtures with expected behaviors, not recorded provider results or an automatic scoring service. This guide extends the summary-review exercise.

Separate the task from the answer key

Each fixture has a stable ID, an input object and a reviewer-only expectation. The input contains the task, supplied context and relevant constraints. The expectation identifies the facts or boundaries a reviewer should check.

When testing a model, supply only the case's input object through an approved tool. Keep expected and criticalFailures out of the model input. Otherwise you may test whether the model can repeat the answer key instead of whether it can solve the task.

Save the exact request separately from the fixture so another reviewer can see what was actually sent.

Use five cases with different jobs

CaseWhat it probesSuccessful behavior
EVAL-01Ordinary summary with one missing dateCorrect totals, categories and uncertainty
EVAL-02Missing source contentStates that the requested result cannot be established
EVAL-03Conflicting recordsDescribes the conflict and requests an authoritative source
EVAL-04Request outside supplied project scopeDoes not invent, retrieve or expose unauthorized project data
EVAL-05An instruction embedded in a stored noteTreats the note as data and preserves the task boundary

The fourth case has no secret foreign-project records hidden in the download. In plain chat it tests the response only. A real integration needs a separate controlled-account test that verifies its server refuses the unauthorized read. Never put a private record in the prompt merely to see whether the model will reveal it.

OWASP identifies external documents as a route for indirect prompt injection. That is why the fifth case includes a harmless instruction-like note and why its text cannot grant tool access. OWASP prompt-injection guidance.

Define failures before running

For each case, distinguish required facts from acceptable wording. A summary can be concise or detailed and still be correct. An invented completion date is a factual failure regardless of tone.

Use a simple outcome vocabulary: pass, fail, blocked or not run. Record critical failures separately. An unauthorized action or fabricated source should not disappear inside an average score for good formatting.

If a provider or account is unavailable, mark the run blocked. Do not reuse an old answer and present it as a new result.

Compare one meaningful change

Record the date, tool, model identifier if exposed, version/configuration, exact input, output and reviewer. Describe whether the run was chat-only, mocked or connected to a controlled integration. Omit secrets from every record.

Keep the fixtures and acceptance criteria unchanged while comparing a prompt revision. If you change the task at the same time, the comparison answers a different question. Repeat cases when variability matters and save failures as well as successes; do not select only the best response.

The fixture pack contains no provider calls and records no scores. You choose an authorized tool and perform the review. A small passing set is a regression aid, not a general accuracy percentage or safety certification.

Keep examples fresh without erasing history

Version the set when a requirement changes. Preserve the old fixture so a past run remains interpretable. Add a case when a real failure exposes a gap, using sanitized or synthetic data that retains the relevant behavior.

Use the workbook to record the run and follow-up. A result needing code changes belongs in the applicable implementation process; a wording improvement may only need an editorial change. The proposal handoff guide explains how to make that distinction.

Entre em contato

Interessado em um tema? Deixe uma mensagem e escolha uma categoria. Também estou disponível para uma reunião de consultoria gratuita — entre em contato e combinamos.