The harness matters more than the model

AI EngineeringFrom building AI agents for clients3 min read
ModelToolsEvalsContextMemory

Most of the quality in an AI agent comes from the code around the model. How I think about tools, context, memory and evaluation.

When an agent gives a bad answer, the first instinct is to blame the model and try a bigger one. Sometimes that helps. Most of the time the problem sits somewhere else: in the code around the model. People now call that code the harness, and in my work it is where most of the engineering happens.

What the harness is

Think of the model as the brain. The harness is the body and the workspace. It decides which tools the model can use, what goes into its context, where it keeps notes, how it hands work to helpers, and how you check the result.

Modelthe brainToolsContext playbookFiles as memorySub-agentsEvaluatorread onlyRouterfast or reasoning
The parts of a harness around a model.

Two teams can use the exact same model and get very different results, only because their harness is different.

The parts I care about most

Tools with clear contracts

Every tool gets a typed input, a typed output and a clear error. I write tools so they can be tested without the model at all. If a tool is flaky, no prompt will save you.

Context as a playbook

It is tempting to keep adding rules to the system prompt every time something breaks. After a few months the prompt is long, contradictory and expensive. I keep a short, curated list of lessons instead, and I remove old entries as often as I add new ones.

Files as memory

Long tasks do not fit in a context window. Logs, intermediate results and error traces go into files or a database, and the agent reads what it needs. As a bonus, a crashed run can continue where it stopped.

Sub-agents with limits

Independent subtasks can run in parallel helpers. Each one gets a timeout, a budget and a clear return format. Without limits, parallel agents become parallel bills.

A router

Not every request needs a slow reasoning model. Simple questions go to a fast model, hard multi step problems go to a reasoning model with a capped budget. This one decision often cuts cost more than any other change.

An evaluator the agent cannot touch

The thing that scores the agent must be separate and read only. If the agent can influence its own grade, it will learn to please the grader instead of solving the task.

How the harness gets better

The most useful habit I know is boring: read the failures.

Run tasksLog every runGroup failuresSmall fixBeats eval set?yes: ship it. no: throw it away
A simple improvement loop for agents.

Log every run. Once a week, group the failures into patterns and fix the most common one. Then treat the fix like a pull request: it only ships if it beats the current version on a fixed set of test tasks. That is continuous integration, applied to agents.

Research teams now run this loop automatically, with one model proposing changes and another testing them. For most companies the manual version is already a big step up, and it keeps a human in charge of what ships.

What to take away

  • Before you change the model, check the tools, the context and the evaluation.
  • Keep prompts short and curated. Put long state in files.
  • Route easy work to fast models and save reasoning models for hard work.
  • Every change must beat a fixed test set before it ships.
Faizan Khan

Faizan Khan is an AI and data engineer in Berlin. Working on something like this? Book a 30 minute call or email hello@faizankhan.me.