Case study · Public, one-person agentic system
A governed agentic pipeline you can audit end to end
A collection of 160+ interactive reference pages, built and maintained by one person plus AI agents, held to one written spec and shipped through a human merge gate. The pipeline is the thing worth looking at: it makes the next page cheap and consistent, and every commit is browsable in public.
Why this one is different
The other case studies on this site describe work inside companies, where the strongest numbers are internal and I ask you to take them at the altitude I can share. This one is public end to end: a system run by one person plus AI agents, with every claim below checkable on the live site and in the git history.
Prompting a model into one good-looking page takes an afternoon. Doing it 160 times, to the same accuracy and accessibility bar, with the rules written down and enforced, is an engineering problem. The pipeline is the artifact here. Any individual cheatsheet is output.
The pipeline
A new reference moves through a fixed six-step sequence. The first four steps are the agent’s; the last two are the gate. No page reaches the main branch without clearing all six.
- Research. Verify version-sensitive facts against primary sources before writing, so a hallucination gets caught before it reaches a page.
- Outline to three depths. Fundamentals, working knowledge, and edge or advanced material, laid out before any prose.
- Fill to a density floor. Populate every section to a floor of roughly twenty substantive entries, with no hollow stubs.
- Self-verify. Run the page against the testing checklist for coverage, accuracy, platform, and accessibility, and fix the gaps it finds.
- Browser-verify. Serve locally, confirm a clean console, and prove the pinned CDN assets actually loaded under Subresource Integrity.
- Review and commit. A human reviews the diff, then it ships as one commit bundling the page and its social preview image.
The quality lives in the spec. A lucky prompt gets caught at step four.
The spec as acceptance criteria
The center of gravity is a single checked-in file, AGENTS.md. Its rules are treated as binding acceptance criteria: a page that violates them is not done, regardless of how polished it looks. The spec enforces a coverage contract (three depths, no hollow sections), an atomic-entry rule (every entry needs a one-line definition and at least one concrete example with realistic values, never foo and bar), a density floor, an accuracy gate (“verify, do not recall”), visible freshness dating on volatile facts, and a platform baseline (pinned Bootstrap with SRI, native disclosure widgets, light-dark() theming, WCAG 2.2 AA).
Governance, defined as checkable things
“Governed” here means four specific mechanisms that bound what the agent can ship and keep the result accountable afterward.
- The gate is a person. A testing checklist plus a human reviewing the diff sit between generated output and the main branch. The model proposes; a person reviews and holds the merge. The checklist writes a report and the person writes the verdict.
- A public, immutable audit trail. Every page ships as a version-controlled commit with a descriptive message. The change-history page renders the entire git log, commits and diffs and per-file revisions, read-only, straight from the repository.
- Supply-chain integrity by default. Every CDN link and script carries a sha384 Subresource Integrity hash computed from the real bytes. A tampered or swapped asset fails the check and never runs.
- Disjoint ownership. When work fans out to parallel workers, each owns exactly one page, its HTML and its preview image, and nothing else. The shared files that could collide are reserved for the supervising agent. Blast radius is bounded by who is allowed to touch what.
This is the same invariant as the regulated-lender and enterprise-scale work: the agent drafts against a written standard, and a named human holds the merge.
Delegation by cost, and a closed loop
The pipeline spends frontier tokens on judgment and commodity tokens on construction. A frontier model does the work that needs taste and verification: picking topics against live traffic, authoring the spec, and reviewing results. Implementation of each page is delegated to a cheaper model, and in batches to parallel workers with disjoint ownership.
The loop closes with measurement. A nightly job pulls per-page view counts from Cloudflare analytics into a popularity file that ranks the gallery and powers a public traffic dashboard. That signal feeds back into the first step: what people actually read informs what gets built next. Ideation, spec, build, deploy, measure, and back to ideation.
What the agent wrote, and what I approved
The AI has wide latitude to write. It does not have latitude to ship. The commit trail keeps those two apart on its own: the diff a person reviewed is the diff that shipped, and both are public.
Limits
This runs at personal scale. There is no production headcount behind it and no on-call rotation, and the governance model is the part worth copying. Comprehensiveness is enforced by the spec and the checklist, with no independent audit behind either, so the process and the coverage floor are the claim, and any individual page is arguable on its own topic. The commit count grows over time; the figure here is the live count as of mid-2026, and the history page carries the current one.
Where it fits
The Antech case study shows agentic engineering carrying load at enterprise scale, with the strongest numbers internal. The regulated-lender case study shows AI as a governed product in a regulated domain. This one shows the same governance discipline at personal scale, with the spec, the merge gate, the audit trail, and the results all sitting in public. The general pattern, platform before agents, is written up in the Governing Agentic AI field guide, which is itself one page in this collection.
How this site is built, the engineering exhibit ↗ Email me about this Where I fit best ← All case studies