Shared memory · loops · evals — for coding agents

Your agents, sharper every week.

One shared memory for every agent on the team. Scheduled loops for the recurring work. Evals that prove — not promise — what actually helped.

$ honemill run nightly-sentry-audit
◐ pass 7 of 8 · recalled 3 canon memories
run recorded a41f… · 142k tokens · $0.87

The flywheel

Loops run and leave records. Evals score the runs. Learnings that held up under scoring get promoted into shared memory — and the next runs start sharper. Each pass removes a little roughness; the compounding is the product.

runscorelearnpromote

The stone turns; the stations hold still. Loops feed run, evals own score, memory owns learn — and promotion is the eye of the stone: only learnings that earn their score become canon.

$ honemill report flywheel
memory “clickhouse degrades before it errors”
  recalled in 11 runs · +18% eval score · promote to canon? [y/N]

Loops · run

Passes, on schedule

Nightly audits, PR reviews, weekly digests — recurring agent runs become first-class records: what ran, what it cost, how it ended. The mill turns while you sleep.

Evals · score

Improvement you can query

Deterministic asserts and judge rubrics, runnable locally and as a CI gate. Every score is pinned to a commit and a model, so improvement is a fact — not a feeling.

Memory · learn → promote

One canon, not ten notebooks

Every agent on the team reads and writes the same memory. New learnings land in an inbox with full provenance; only reviewed, evidenced knowledge becomes canon. No drift, no poison, no “my agent knew that.”

HONEMILL RUNS ON HONEMILL. Our own nightly audits, PR review loops, and eval gates come first — early access opens when the grind log says it’s ready.

Run it on your own work first — that’s what we’re doing.

Get early accessWorks with Claude Code, Codex, Cursor — anything that speaks MCP.