Prompts are code. Treat them that way.

AI EngineeringFrom shipping LLM features to production2 min read
prompt v1prompt v2prompt v3Test setfixed casesv162%v278%v391%

Small wording changes can move accuracy a lot. Version prompts, test them on a fixed set, and let real code do the maths.

If you have shipped an LLM feature, you have probably seen this: someone changes one sentence in a prompt, and a week later a different part of the product behaves strangely. Prompts look like text, but they behave like code. They deserve the same care.

Version and test them

Keep prompts in git, next to the code that uses them, with a version number. Build a fixed test set of real inputs and expected outputs. Every prompt change runs against that set before it ships, and you compare the score with the current version. No score, no merge.

Climb the ladder only as far as you need

Plain instructionAdd examplesAsk for stepsSample + voteTools + retrievalmore accuracy, more cost and latency
Each step usually adds accuracy, and always adds cost and latency.

Start with a plain, direct instruction. Modern models follow clear instructions well. If that is not enough, add a few examples. Then ask for step by step reasoning. Then sample several answers and take the one that passes a check. Last, add tools and retrieval. Each step costs more tokens and more time, so stop where your test set says you are good enough.

Examples are powerful, and biased

The examples you pick and the order you put them in change the result. Models tend to repeat the most common label in the examples, and they lean toward the last one they saw. Pick examples that look like real inputs, balance the labels, and test with the order shuffled. If results jump around, your prompt is fragile.

Ask for structure

When code reads the output, ask for a schema and validate it. I define the output as a typed model, for example with Pydantic, and reject anything that does not fit. A clear error early is much better than a broken record three systems later.

Let code do the maths

Models are good with language and unreliable with arithmetic. If a task needs counting, sums or dates, let the model write or call code and use the result. It is faster to debug and it is right every time.

What to track per prompt version

  • Score on the fixed test set
  • How much the score moves when you shuffle examples or tweak formatting
  • Token cost and latency
  • Rate of outputs that fail the schema
Faizan Khan

Faizan Khan is an AI and data engineer in Berlin. Working on something like this? Book a 30 minute call or email hello@faizankhan.me.