Prompts are code. Treat them that way.
Small wording changes can move accuracy a lot. Version prompts, test them on a fixed set, and let real code do the maths.
If you have shipped an LLM feature, you have probably seen this: someone changes one sentence in a prompt, and a week later a different part of the product behaves strangely. Prompts look like text, but they behave like code. They deserve the same care.
Version and test them
Keep prompts in git, next to the code that uses them, with a version number. Build a fixed test set of real inputs and expected outputs. Every prompt change runs against that set before it ships, and you compare the score with the current version. No score, no merge.
Climb the ladder only as far as you need
Start with a plain, direct instruction. Modern models follow clear instructions well. If that is not enough, add a few examples. Then ask for step by step reasoning. Then sample several answers and take the one that passes a check. Last, add tools and retrieval. Each step costs more tokens and more time, so stop where your test set says you are good enough.
Examples are powerful, and biased
The examples you pick and the order you put them in change the result. Models tend to repeat the most common label in the examples, and they lean toward the last one they saw. Pick examples that look like real inputs, balance the labels, and test with the order shuffled. If results jump around, your prompt is fragile.
Ask for structure
When code reads the output, ask for a schema and validate it. I define the output as a typed model, for example with Pydantic, and reject anything that does not fit. A clear error early is much better than a broken record three systems later.
Let code do the maths
Models are good with language and unreliable with arithmetic. If a task needs counting, sums or dates, let the model write or call code and use the result. It is faster to debug and it is right every time.
What to track per prompt version
- Score on the fixed test set
- How much the score moves when you shuffle examples or tweak formatting
- Token cost and latency
- Rate of outputs that fail the schema