HarnessTax: How Much Does the Harness Matter for Coding Agents?
-
Study Scope & Setup:
- Evaluated 21 model–harness combinations using 7 models across 3 agent harnesses (Claude Code, Codex CLI, and Pi) on two benchmarks: SWE-bench Lite and Terminal-Bench 2.0.
- Controlled variables by testing on identical sampled tasks, using native high-effort settings, capping attempts at 100 turns, and normalizing direct API token pricing.
-
The "Harness Tax" Phenomenon:
- Changing the harness has minimal effect on task resolution rate (within ±2% on SWE-bench Lite and ~±5% on Terminal-Bench 2.0), but dramatically alters token expenditure (diverging by up to 2×–5× for identical success rates).
- This overhead is driven by initial context bloat, verbose system prompts, and heavy tool schemas injected into the context window before work begins.
-
Minimalist Harnesses Remain Highly Competitive:
- Pi, a lightweight open-source harness, proved competitive with or superior to heavy proprietary harnesses in both resolution rate and token efficiency.
- For example, running Claude Fable 5 resolved ~97%–98% of tasks across both Claude Code and Pi, yet Claude Code cost approximately twice as much ($1.33 vs. $0.67 per attempt).
-
First-Party Harness Mismatch:
- Models do not necessarily perform best inside their creator's native harness.
- In 9 out of 12 head-to-head comparisons across Anthropic and OpenAI models, third-party or alternative harnesses delivered higher success rates than the model provider's default CLI.
-
Core Practical Takeaway:
- Accepting a coding agent's default harness incurs an unnoticed "harness tax" in the form of inflated API bills or premature subscription rate limits.
- Agent benchmarking should decouple models from scaffolds, treating harness complexity and tooling architecture as separate empirical engineering trade-offs.