For coding agents that work while you're away

Agents that pick up tomorrow where they stopped today

One command installs rules, a docs kit, checks and session hooks into your repository. Your agents stop losing context, stop writing into each other's zones, and stop overpaying just to start a session.

  • Measured on live machines, not synthetic tasks
  • Cost law fit R² 0.96–1.00
  • Checked against 50 billion tokens of real sessions
  • Built for Claude Code

The problem

A coding agent is great at one task with a human beside it. It breaks when nobody is.

Several agents, many sessions, and no human holding the whole project in their head at every moment — that's where it falls apart. What we keep seeing on live projects:

The smart zone is gone before work starts

Start-up reading eats more than half of the smart zone — the stretch of context where a model still reasons well. The task gets what's left.

Docs rot

Agentic development moves so fast that docs written for humans go stale almost as soon as they're written. A person can usually work out which copy is right. For an agent, a stale doc is fatal: it believes what it reads.

Dead weight in every session

Left to keep docs on their own, agents do it clumsily: closed issues and useless notes pile up, and every session drags them along for the life of the project.

Agents wander into each other's zones

Splitting a big project into smaller subprojects should help. Instead, agents get confused and work outside their own zones — and the chaos only grows.

Context gets lost

After auto-compaction, docs the agent had read leave a trace of “I remember reading it” — and not a single value.

Long sessions burn budget quadratically

Every turn resends the whole context, and the context grows every turn. So the input bill grows with the square of the session: b·n + g·n²/2. Twice the turns — roughly four times the bill.

Agents can't see their own spend

Eager to please, agents try to report what they've spent — but they have no tools for it and don't know where to look. So they make the wrong calls and misread what they find.

Who it's for

Built for agents that run without you

Yes

  • Solo developers and small teams who run agents in parallel and for long stretches: several sessions, several zones (backend, frontend, mobile), work that resumes the next day or the next week.
  • Anyone who wants to hand an agent the whole loop — take a task, do it, verify, commit — without a human at every step. Soon, 2.0
  • Anyone paying for tokens who wants to see where they go.

No

  • An “assistant in the IDE”, where you set the task, know the layout and say “run the tests”. You already supply the context; an outside study found a context file adds +4% success for up to +19% cost in that mode.
  • One agent in one repository. Splitting into zones is pure overhead there.
If you sit next to your agent, you don't need this. If your agents work while you're away, they don't work without it.

How it works

Four principles

  1. 01

    Ownership by location

    A zone is the folder an agent was started from; its docs live right there. It writes only to its own kit — a problem it finds next door becomes an entry in the neighbour's issues file. Agreements go stale; locations don't.

  2. 02

    Files split by how they change, not by topic

    Map, open issues, decisions, archive, contract with neighbours — each with its own rule and its own token budget.

  3. 03

    Forgetting is built in

    A closed issue moves to the archive whole and at once; links to it keep working. Session start cost depends on the number of open issues, not on the age of the project. There is no session log on purpose: 92% of its entries duplicated git.

  4. 04

    A script checks, not the agent

    Budgets, format, links — verified by machine at the end of every session. An agent can't testify about itself, so accounting is built on transcripts, not on the agent's reports.

What's inside

Everything a zone needs, installed with one command

Available

Rules

CLAUDE.md, root + zone: zone ownership, what to read and when, working rules, what to update after work.

Available

Docs kit per zone

map, issues, decisions, archive, howto — each with a token budget.

Available

One-command install

Deploys the kit into any project, filling in zone names.

Available

Session hooks

Inject the zone map at start, log reads, daily summary, context-window meter with thresholds, a “where we left off” snapshot for the next session.

Available

Setup skill

The agent asks for zones and prefixes and rolls out the standard — including migrating existing docs.

Available

Task intake skill

Turns a request into a task: items with a checkable “done when”, questions for the human.

Available

The standard

Every rule with its reason and its measurement.

Partial

Docs checker

Machine check of the docs, run by a hook at session end. Ships checking the rules budget; our internal build runs thirteen checks.

Soon — 2.0

Unattended loop

The agent works through task items on its own, with safeguards and a gate that asks the human.

Soon — 2.0

Delivery & updates

Archive + installer; update the harness on a live project without touching your own files.

Soon — 2.0

Telemetry

A metrics catalogue and one collector over transcripts: who spends what, who makes mistakes.

How we differ

The genre is taken. The job isn't.

There are popular rule collections with huge GitHub followings. All of them are rulebooks and packaging, and none of them targets agents working unattended.

What you get Rulebooks Harness
Zone ownership by location, no writing into another zone's kit —✓
One address for a neighbour's contract; copying its spec is forbidden — a copy goes stale with no diff to warn you —✓
Archive of closed entries with links that keep working —✓
Token budget for every file —✓
“When → what to read” table: required reading split from read-on-demand —✓
Machine check of the docs by a hook at session end —✓
One-command install —✓
Numbers from live machines, not synthetic tasks —✓

They sell rules. We measure whether they work.

Honest numbers

Nobody else in the genre publishes one

Some collections ship telemetry on by default — and still list the effect as “not published”. Ours come from transcripts on disk (Claude Code, September 2026), never from what the agent says about itself.

The cost law

input = b·n + g·n²/2

Cumulative input is base × turns plus the area under a growing context. Context grows linearly per turn (R² 0.96–1.00); the formula matched ten runs within 1.5% and held across 50 billion tokens of real sessions.

What the law gives

T* = 5·√(b/g)

The optimal moment to start a fresh session: 21–43 turns, 60–98k tokens of context. Restarting there, the law predicts up to 72% less input on long runs.

+68k

tokens per extra 1k tokens in the root CLAUDE.md, over a 708-turn session — for every agent. 16% of that session's bill was resending unchanged text.

500+

counted runs, across three models.

48×

input billed vs. final context on one 708-turn session.

Outside context

ETH Zurich, February 2026: a context file generated by an agent lowers success by 0.5–2% and costs 20% more; one written by a human adds +4% for up to +19%. A tool named in the context file gets called 50× more often. Our reading, stated openly: in “assistant” mode, docs barely pay off. We build for “tool” mode — the one that study didn't see.

Why it's a business

Rules are free. Knowing whether they work isn't.

Rulebooks are free by design

The genre leader gives away 25 skills and earns on courses. We tested “collect rules, sell a bundle” — and dropped it.

“Maintained rules” won't sell either

A subscription to maintained rules was dropped too: it competes with a free catalogue and can't prove its value.

Measurement is the reason to pay

The open spot: show users their own number — what they spend, where they overpay, whether the rules help in their project. Telemetry from the wild can't do that — there's no control group. An A/B inside the repository can. Technically hard — which is why the spot is empty.

A narrower, worse-served market

Narrower than “everyone with a coding agent”, and worse served: agents working unattended, in parallel, across several zones.

FAQ

Fair questions

“There's a study showing AGENTS.md doesn't help.”

There is, and we agree — for “fix one bug in a mature repo with a human beside you”. It didn't measure the second session, the neighbour you can break, or work with nobody watching. That's exactly what the harness is for.

“Isn't this just markdown files?”

The rules are. The difference is the machinery: one-command install, budgets, a machine check, hooks that hand the agent what it needs and count what it spends.

“What does it cost in tokens?”

The rules layer stays within a budget, and a script enforces it. The harness is an investment: a bare agent is always cheaper on a single task; the harness pays off over the long run.

“Only Claude Code?”

For now, yes.

Want to see it on your repository?

Write to us — we'll answer and can open the code for your review.

hello@claude-harness.tech

Smart zone — the part of the context window where a model still reasons well. Past it the model doesn't run out of room; it gets worse: follows instructions less reliably, loses details, makes more mistakes. Practitioners put the edge at roughly 40% of the window, about 80–100k tokens on today's models. The term is coding-agent practitioners' slang; the edge is a rule of thumb, not a measured constant.