How to Debug AI Agents

When an agent lets you down, experiment by tweaking one of its inputs at a time: harness, model, runtime, or purpose.

Does this ever happen to you? You give an agent a task and it does some of it, but doesn’t quite do everything you were expecting it to do.

I want to share my playbook for debugging agents. By the end of this post, you’ll understand which inputs you can control and how to reason about which one is to blame when your agent doesn’t seem to do what you want it to.

We are generally all building on the same foundational capabilities from the big AI providers, but the fun part—and the opportunity—is tuning agents to do our bidding.

There are four inputs we generally have in our control:

Agent inputs

1. Harness

2. Model

3. Runtime

4. Purpose

If you want to see meaningful agent improvement, you first need to know which input isn’t working and needs to be adjusted. Let’s step into these inputs one by one, and then I’ll explore examples of how I’ve debugged poor output.

Harness

Harness

A harness provides the agent’s capabilities and orchestration.

It’s like a supersuit with cool tools that the hero can reach for to fight crime. You can pick up a standard harness like Claude Code or Codex, or build your own to understand the inner workings of agents. The more control you want, the more configurable your harness may need to be—or you might even start sewing your own spandex.

Model

Model

The model is intelligence.

The model is the superhero that dons the suit. The interface of a model is strictly input and output: a beautiful functional interface of inputs and a generated response. The harness and model lines get blurred easily because the harness operates largely through the intelligence of the model.

Harness and model swimlane handoffs Harness and Model participants exchange task instructions, branch from a tool call decision into a generated response or tool operation, then restart the cycle with new context. Task + instructions + tools Yes · Execute tool Execute tool Add tool response to context No · Return response to Harness Generated response Harness Model Tool call?

This is how the harness/model loop works:

  1. The harness sends the model the task, instructions, and available tools so the model can generate a response.
  2. The model answers one core question: do I have everything I need to respond, or should I use one of these tools from the harness first?
  3. If it needs a tool, the model specifies the tool and its input arguments for the harness to execute.
  4. The harness executes the tool, adds the result to the context, and starts the loop over. If the model has what it needs, it sends the generated response back.

The model is strictly input and output. The harness does the orchestration; the model is the intelligence behind it, deciding whether it needs more information or has a response ready.

Runtime

Runtime

Runtime is where the agent is running and the permissions under which it operates.

Superheroes have a city in which they enact justice, and there is usually some limitation in the world that the hero has: They don’t use guns, they melt from kryptonite, or they can’t wear capes.

You have probably felt the difference in an agent’s runtime when comparing a web-based chat agent that mostly answers questions with a more powerful agent that can operate programs on your computer, save and delete files, and more because it has far more permissions and capabilities at its disposal. You could use the same model and harness, but if they run in a constrained environment, you are giving that hero kryptonite and limiting its potential.

Purpose

Purpose

Purpose tells the agent what outcome to pursue.

Every compelling hero has a tragic backstory that informs their mission to fight crime. Typically, the more compelling and relatable the story of the hero, the more we get attached to our hero’s world and particular cut of spandex.

The agent’s input prompt is its purpose in its simplest form. More advanced agents use supporting instructions and context—such as AGENTS.md, Obsidian notes, and conversation history—that supplement the input prompt without explicitly including it in the input, giving the agent a more comprehensive purpose.

Agent reliability often improves when the agent receives relevant context, clear constraints, and a useful definition of done. More context is not automatically better. Irrelevant information, contradictory instructions, and an unclear definition of done can all lead to poor outcomes.

Purpose and harness context loop An input prompt enters the Harness lane. The harness appends its configured context before sending the prompt and appended context to the Model lane. Input prompt Append configured context • System prompt • Skills • MCP tools • AGENTS.md Prompt + appended context Harness Model

Purpose + Harness

Purpose is how you specifically inform an agent of its purpose. The harness is the technical context in which you do so: system prompts, orchestration, and hidden tools that you likely don’t control beyond which harness you use or build.

You could give Hermes and Claude Code the same purpose and see different behavior because of the differences between their harnesses. The model receives that assembled context and uses it to determine what to say or do next.

Agent Debugging Steps

Now that we have broken down the inputs of an agent, let’s return to the original problem: my agents aren’t operating as I expect. How do we troubleshoot them?

These are the debugging steps that have worked for me:

Describe the failure

Ask why the agent didn't do the thing you expected. Be specific in your expectation.

Identify the most likely input at fault

Form a hypothesis and compare it against the agent's diagnosis.

Change one input

Be a good scientist: change one input at a time as your hypothesis.

Run the task again

Compare the new result against the previous result and note the differences.

Agent Debugging Examples

A skill or tool isn't being used consistently

The harness makes skills and tools visible, discoverable, and available to the model .

The model still needs to recognize when they are relevant and choose to load or call them without an explicit instruction every time. A more intelligent reasoning model may be needed.

Competing skills can also be the problem: if you have 14 writing skills, their descriptions may compete for the same tokens, so the skill you expect is not automatically selected.

Check the inputs

Harness
Model
Runtime
Purpose

The agent didn't open a PR, and I have to ask it manually every time

Most likely a purpose problem: AGENTS.md or the input prompt could have gotten you that outcome.

It could also be a runtime constraint where the agent is missing permissions or shell access to open a PR.

Check the inputs

Harness
Model
Runtime
Purpose

The agent solved the request but didn't use an existing convention

This is a common problem: agents want to finish quickly, but don't produce uniform, maintainable work as requirements change.

Purpose is where I'd start: the agent produced a working result, but did not recognize or apply an existing convention. If the convention is known, it should be easy for humans and agents to discover in the codebase or in documented guidance; otherwise the purpose is underspecified.

If convention expectations are well documented, the model is the next most likely culprit. More reasoning and more capable models can lead to better solutions at both the macro and micro levels.

Check the inputs

Harness
Model
Runtime
Purpose

Agent is burning through my tokens

I think we've all been bitten by this one. I'd start with the purpose , especially any lifecycle expectations you have set for every agent in files like `AGENTS.md`. I was bitten once because I demanded Playwright screenshots be used and inspected by every agent before sending work back to me. A token diagnostic prompt pointed me toward optimizing screenshot size and frequency. A picture may be worth a thousand words, but it can also be worth a thousand tokens.

If you haven't added a token-heavy process to the purpose , look at the harness : it determines how much input is sent to the model. Some harnesses use more tokens than others.

Check the inputs

Harness
Model
Runtime
Purpose

The agent didn't fix the bug

Purpose is where I'd start: are you defining the bug and expected behavior well enough?

Runtime is next if the agent lacks the permissions or capabilities to reproduce the bug and iterate until it is fixed.

Check the inputs

Harness
Model
Runtime
Purpose

With these examples, you’ll notice I bounced between inputs in all of them. Agents are multi-causal: changing different inputs could still lead to the same expected outcome.

It is a different way of thinking from a purely deterministic engineering approach, to be sure. Mapping agent behavior to these four inputs is how I’ve seen the best success at forming effective hypotheses for improving agents.

I hope this helps. If you have questions, share them in the comments and I’ll do what I can to help. Happy debugging—I’m looking forward to seeing how you whip your agents into shape.

Let's learn together

Struggling to get reliable results from agents?