Thirty real inputs before the first prompt
Collect thirty real inputs and decide what a correct output is for each. Write the scorer. Get a number for a trivial prompt. Then write the real prompt and treat it as code that has to pass.
evals/tickets.jsonl one case per line: input, expected, source
evals/run.ts prints accuracy and the failuresEvery correction from production goes into the set. The prompt is disposable. The set is the asset.
ai-engineeringtesting
Longer version: the post this came from.