Mert Sahin

Coding agents, October 2024 to October 2026: the 20 events that mattered

· 4 min read · Agents · MCP · Timeline

TL;DR

  • In two years, the parts around the model changed more than the model: protocols, harness config and permissions.
  • MCP, AGENTS.md and Agent Skills became formats that agents from several vendors read.
  • The headline coding benchmark of 2024 lost its standing. Your own evals are what is left.

What happened

Protocols and shared formats

The harness

  • Feb 2025: Claude Code arrives as a research preview: a terminal agent that edits files, runs tests and commits.
  • Sep 2025: Factory's Droid tops Terminal-Bench: "the right agent framework can lead to greater improvements than model selection."
  • Feb 2026: Mitchell Hashimoto names harness engineering: when the agent makes a mistake, engineer a fix so it never makes it again.
  • Apr 2026: Anthropic's postmortem traces a drop in Claude Code quality to three harness changes. The API was not affected.
  • Sep 2026: ant apply manages agents, skills and memory stores as files in the repository, with a committed lockfile.

Permissions and containment

  • Jun 2025: Simon Willison names the lethal trifecta: private data, untrusted content and a way to communicate externally.
  • Aug 2025: in CVE-2025-53773, a prompt injection makes Copilot turn on chat.tools.autoApprove in its own settings, then run shell commands.
  • Oct 2025: Claude Code sandboxing isolates filesystem and network and cuts permission prompts by 84%.
  • Mar 2026: Anthropic reports that users approve 93% of permission prompts and ships auto mode, where classifiers decide instead.
  • Aug 2026: auto mode becomes the default for new Pro, Max and Team sessions. In a controlled test, 1,053 paid testers caught 13.6% of dangerous commands; auto mode blocked 89%.

Models and how we measure them

  • Oct 2024: the upgraded Claude 3.5 Sonnet reaches 49.0% on SWE-bench Verified, up from 33.4%.
  • Mar 2025: METR finds the length of tasks models finish at 50% success doubling about every 7 months. Claude 3.7 Sonnet is at about an hour.
  • Jul 2025: in a METR randomised trial, experienced open-source developers took 19% longer with AI tools while believing they were 20% faster.
  • Feb 2026: OpenAI stops reporting SWE-bench Verified: at least 59.4% of audited hard problems had flawed tests, and every frontier model tested had seen some solutions in training.

The pattern

Models got better: SWE-bench Verified went from a 49% headline to a benchmark OpenAI no longer reports. But most of what changed how I would build a coding agent sits around the model: a shared tool protocol, instruction and skill files that work across vendors, harness config you can version, and permissions moving from prompts to sandboxes and classifiers. The April 2026 postmortem shows it most clearly: the product got worse while the API did not. More in the harness is the product and long-running agents need a progress file.

Take this: four questions for your own setup

Area Ask If not
Protocols Do our MCP servers and clients state which spec version they speak? Pin it, and plan the move to the stateless 2026-07-28 spec.
Harness Are prompts, tools, permissions and skills in git, with an eval on every change? Start with ten tasks from your own repo.
Permissions Can the agent edit its own config, or combine private data, untrusted input and network access? Deny writes to agent config. Split the trifecta.
Measurement Do we pick models on our own tasks rather than a public leaderboard? Keep a private task set. Read transcripts.

My take

Most of my work on coding agents for SAP developers sits in these four rows: MCP servers that expose ABAP development with read-only profiles and per-action confirmation gates, an Eclipse plugin whose tool allow/deny lists decide what the agent may touch, and CI gates with automated tests and an LLM-as-judge eval. Models improve on the vendors' schedule. The parts around them are the ones you own.