Mert Sahin

The harness is the product

· 4 min read · Harness · Agents

TL;DR

  • In April 2026, Anthropic traced a drop in Claude Code quality to an effort default, a caching bug and one line of system prompt. The API was not affected.
  • The harness is everything around the model: prompts, tools, permissions, skills, hooks and defaults. It decides what users get.
  • Keep it in git, pin what the vendor can change, run evals on every change and roll out slowly.

What happened

  • Jun–Oct 2025: Claude Code adds hooks and subagents, then plugins and Agent Skills that load only when a task needs them.
  • Sep 2025: Factory's Droid tops Terminal-Bench. Droid on Sonnet beats every other agent running on Opus.
  • Feb 2026: Mitchell Hashimoto names harness engineering: when the agent makes a mistake, add an AGENTS.md line or a real tool so it never makes it again.
  • Mar 2026: Anthropic writes that "every component in a harness encodes an assumption about what the model can't do on its own."
  • Apr 2, 2026: Birgitta Böckeler sums it up as "Agent = Model + Harness", with guides that steer before the agent acts and sensors that check after.
  • Apr 23, 2026: Anthropic's postmortem. On March 4 the default effort went from high to medium. On March 26 a change meant to clear old thinking from sessions idle for over an hour kept clearing it every turn. On April 16 a system-prompt line capped text between tool calls at 25 words. The fixes: "a broad suite of per-model evals for every system prompt change", soak periods and gradual rollouts.
  • Sep–Oct 2026: ant apply manages agents, skills and memory stores as files with a committed lockfile. Claude Code mods can rewrite prompts and gate tool calls, and are not sandboxed.

Take this: a harness release checklist

  1. One home. AGENTS.md, settings, permissions, hooks, skills and the MCP server list live in one repo and change through pull requests.
  2. Pin what the vendor can change. Set model and effort explicitly. The April regression started with a default.
  3. Eval every change. Run a fixed task set from your own codebase on main and on the branch; read some transcripts. Trigger it on every pull request touching AGENTS.md, .claude/ or .mcp.json.
  4. Roll out slowly. Pilot group, soak period, one-commit revert.
  5. Record the version. Store the harness commit with every session log and eval result, so a regression can be bisected.
  6. Keep the agent out of its own harness. In CVE-2025-53773, a prompt injection made Copilot switch off its own approval prompts.
  7. Review extensions like dependencies. Plugins and mods run code on your machine.
  8. Delete scaffolding when a new model no longer needs it, then rerun the evals.

A checked-in .claude/settings.json that pins the model and effort and keeps the agent out of its config:

{
  "model": "claude-opus-5-5",
  "effortLevel": "high",
  "permissions": {
    "allow": ["Bash(npm run lint)", "Bash(npm run test *)"],
    "deny": ["Read(./.env)", "Edit(/.claude/**)", "Edit(/AGENTS.md)", "Bash(git push *)"]
  },
  "hooks": {
    "PostToolUse": [
      { "matcher": "Edit|Write", "hooks": [{ "type": "command", "command": "npm run lint" }] }
    ]
  }
}

My take

Most of what I build is harness. The MCP servers that expose SAP ABAP development to agents have read-only profiles and per-action confirmation gates. The Eclipse plugin that brings a coding agent to SAP developers runs on internally hosted models, and its tool allow/deny lists decide what the agent can touch. Changes pass CI gates with automated tests and an LLM-as-judge eval. None of that is the model. The April postmortem is the case for giving a one-line config change the same gates as a code change.