Few-shot examples come from the eval set, not from imagination
Examples you wrote are clean and typical. Examples from production are messy and are the cases the model gets wrong. Pick the three from the eval set's failures that a good example would have fixed.
input: <a real ticket, trimmed>
output: billingRefresh them when the failures change. An example that no longer matches a failure mode is just tokens.
ai-engineeringllm