The Ember Guildhall, Through the Guild Lens — the receipts
The evidence behind the long argument, in full.
Every claim the systems paper makes is stated there; what lives here is the detail behind it. Nothing on this page is a new claim.
The closest prior art, claim by claim
Where the paper's convergence claim is checkable.
- The resume surface — a fixed-size file rewritten rather than appended — converges with MemGPT/Letta's core-versus-archival split and with Anthropic's own compaction-and-note-taking. Letta's Context Repositories keep agent memory in git too.
- Role as a context boundary converges with the orchestrator-worker pattern, chartered subagents, and Anthropic's own first-party guidance for human-agent teams.
- The proving pipeline converges with Nous Research's Hermes Agent, whose Skills System added automatic workflow capture without a hand-written skill file. Curation's benefit has been measured outside this house: SkillsBench found curated skills raised average pass rates by roughly sixteen percentage points over a no-skills baseline, while self-generated skills scored below that baseline where they were tested (the paper's own figures, re-read at the primary on 2026-08-22).
- The council is not this house's invention: Perplexity shipped a consumer-facing Model Council running one query across three heterogeneous models with a separate synthesizer reconciling disagreement, and a 2026 preprint uses the same framing for reducing hallucination through heterogeneous consensus.
- The trust ramp's automatic variant shipped first elsewhere, in Paperclip's Agent Trust Levels.
The three catches, one by one
The paper names these three; here they are, told in full. The two dates below are consecutive working days.
Three of the audit's own best candidates, killed and published (2026-07-07). The house's second whole-guild self-audit ran eight parallel lenses over the full canon, each briefed that its findings would face adversarial verification. The coordinating role then checked every finding against the record before any of it reached the operator, and killed three: a claim that mis-paired two sets of numbers, an assertion that a charter was missing when it was not, and a pair of banners reported as wrong that were already correct. The kills were published with their refutations rather than quietly dropped. Publishing them is what makes the survivors readable as verified.
A real gap in the credential gate, caught by adversarial multi-model review (2026-07-07). The same day, the house's first council panel — four chartered seats on three different models, chaired by a fourth — reviewed the plan for standing up a local autonomous agent. Seats filed blind positions, then a refutation round in which each tried to break the leading answer. Two seats independently reached the same defect: the plan's credential gate, the step at which the agent would first be handed real keys, named no adversary. Nothing in it required a reviewer who was not the builder. The gate now halts until a named non-builder adversarial reviewer exists and the operator signs, and the credential handover was deferred past the planned window by design.
Byte-identical proof, and the two bugs a looser check would have passed (2026-07-06). A game build gained a watch-and-replay instrument whose gate demanded that a replayed match reproduce the recorded one exactly — winner, duration, damage, spend, and experience each matching to the value — rather than closely. Two real defects surfaced only against that standard: a half-tick timing skew between the live path and the test path, and an initialization step whose own tail reset erased the match it had just begun. The audit the next day then graded the house's byte-identity claim declared rather than tested, and scheduled a standing replay battery to close it.
The guard layer's denials, class by class
The protection pack is four pre-tool-call hook guards. Since arming, they have denied their own builder at least thirteen times, in six distinct classes:
- The tamper tripwire refused the coordinating role's own pen on the guard pack's files.
- The secrets-path guard refused a probe aimed at a protected path.
- The push-target guard refused a variable-wrapped push whose written and resolved destinations differed.
- The push-target guard failed closed on an unfamiliar path form, denying two lawful pushes before the cause was diagnosed. The friction was real and the guard stayed armed unchanged.
- The deny-all-outbound guard failed closed on a renamed heading — a lawful push denied because an editorial rename changed a string the guard reads. The heading was restored and marked do-not-rename; the guard was not loosened.
- The guards read a whole command line, so a redirect, a pipe, or a wrapping shell construct resolves as a destination the guard will not allow. That class alone has denied the coordinating role seven times across four working days, twice in a day on three of them. Every one was correct inside the guard's own claim; every one was re-run in the sanctioned form — bare, and alone.
Before any of that, the pack's own check battery caught a bug in the tripwire it tests: placeholder handlers silently shadowed the real ones, so the drift alarm would have exited clean and never fired. A tripwire that ships untested is the silently rotten guard the pack's own first law forbids, and the battery is what caught the bug.
The correction that followed is the layer's record against itself. Its design canon had the pack recorded as built-but-unarmed — true when written, superseded the same evening by the operator's arming paste. A readiness audit found the gap a fortnight later, and the correction went at the head of the log rather than the foot: a cold reader consulting the log would have concluded the system unguarded when it had been guarded the whole time.
The council, and how a seat is kept independent
A council convenes for decisions too large for one reviewer, under the rules the paper states. Independence is enforced two ways. By model: seats in a genuine panel run on different models, so agreement is not self-consistency wearing a quorum's clothes. By tools: the quarantine family gives three subagent types defined by what they lack — a web scout with no write, shell, or git tools; a reader for untrusted local files under the same absence; and an independent-verifier seat that is read-only, so it can report a defect but cannot repair it into invisibility. The verifier's first proving run confirmed nine of nine audited claims against their anchors, caught one real defect, and declared the figures only a run could check to be unverifiable by reading, rather than vouching for them.
The doctrine's own caught bug: an earlier form of the doctrine ran one model voicing every seat and labeled that collusion-proof. It is the opposite — single-model self-consistency forfeits inter-seat independence, the exact anti-pattern the doctrine forbids. The repair separates a solo gear, openly labeled as not independent, from a panel of different-model seats.
Calibration: the burst, the cap, the miss-class
The burst. An early uncapped launch put eighteen agents into the air — fifteen of them at once — against a twenty-five-agent audit plan, and the plan window's own throttle killed fifteen of those eighteen mid-run; the remaining seven of the plan never launched. The re-run, staggered to no more than six concurrent, finished thirteen for thirteen clean. A six-to-eight-agent wave cap has been the standing rule since.
The band. A pre-registered spend band is drawn narrow — a single-digit-percent slice of a session's budget — and misses get logged rather than smoothed over. One block closed well outside its band and was recorded flatly as "MISS."
The miss-class. A roughly 2× over-draw on a kit-analysis sweep; a roughly 1.2× miss on a fleet of subagents, self-caught at the wave boundary; then a fifth naming of the same class, at which point it became a standing rate new pre-registrations must use.
The lapse. The measuring never stopped: consumption is recorded at closes as a matter of course, in roughly two closes out of three. After its last row on 2026-07-24, the pre-registering lapsed: an estimate written before a heavy block and closed afterward against the measured actual with a hit-or-miss grade. Part of that was by design, since cheap, well-characterized shapes were granted launch-and-report and stopped pre-registering on purpose. The rest was not: the heaviest blocks of that span were reported after the fact instead of predicted before. The house's own audit surfaced it, and the pre-registering has resumed with its first before-the-fact row.
The adversarial record, in full
Four classes of live adversarial contact, five events.
Embedded prompt injection, refused (2026-07-28). A quarantined web-reading scout — a subagent type with no write, shell, or git tools — fetched a game-archive page carrying agent-directed text that claimed the user's own authority in order to justify deleting files. The scout refused it, used nothing from the page, and flagged the URL for the record. Independently, the fetch layer's own summarizer declined to pass the payload through, so the defense was observed working at two layers that do not depend on each other.
The same trap, refused again (2026-08-05). A later scout on unrelated work reached the same page and refused it the same way. The refusal is a property of the quarantine, not a lucky read.
A trusted domain, an untrustworthy page (2026-07-31). A scout ran a root-check on a domain the house treats as reliable, and surfaced the site author's own disclaimer: the deep page a search engine had served was AI-generated filler that he hosts as a demonstration. A sibling scout had already cited that page as authority. The cross-scout check corrected the citation inside the working artifact before anything reached canon, and the facts were independently re-sourced. The practice written from it: a trusted domain hosting a page is not the trusted author writing it.
The quarantine refused its own operator (twice). Two genuine relays from the human were read as injection attempts by a quarantined seat and rejected. The fix was not to lower the seat's default suspicion but to build an authentication convention for operator relays, so a real relay can be told from a forged one without weakening the guard. A guard that only ever inconveniences strangers has not been tested.
The production line and its sealed channel
The pixel-art line runs under a battery grown past seventeen hundred automated checks, every rendered piece re-tested against every rule the line has written down, and nothing lands until it is green. Beside it runs a sealed blind-grading channel: the operator grades paired variants without knowing which set came from where, and the seal breaks only after the verdict is recorded.
The channel has returned verdicts in both directions rather than confirming its own design. A model comparison graded five-of-five against three-of-four under seal; the quarantined set — built by a seat that had the shared style rules and its own knowledge of the subject, and none of the house's own craft on it — has won every blind craft wave run so far, four in a row, over the house's standing pieces; and a four-item gate returned an honest three-of-four, with the one collision named rather than buried.
The unattended run: a staged improvement loop ran without the operator's hand and raised a shelf's blind-read quality score off the floor. The same night, the scoring instrument grew a paired check that fails any piece whose score rises while its readability falls, because the gate had caught three pieces doing exactly that.
The export ceremony's mechanics
A kit is assembled from the origin's own working canon through a sanitizing gate carrying a denylist, stamped with a signed provenance mark recording what was issued and when, logged as a row in an append-only registry, and delivered by the operator's own hand. There is no self-serve download and no automated path.
The registry carries a row for every kit that has crossed — at least ten — plus two that never crossed at all: one internal test mint, and one superseded build. Nothing is ever removed from it, because a complete trace is worth more than a tidy count. Consecutive builds reconcile at zero drift in the shared body, and every difference between two recipients' kits is a named per-recipient class rather than an unexplained delta.
The one miss. A sanitization list was owed one new entry at a particular crossing and did not get it; the omission survived until a later review of the list caught it. That review also established that nothing was exposed, since the kit had swept clean regardless. It is the only miss of that kind since the practice began, and it was found by the house's own checks rather than by a consequence.
The gate-intent ladder: the surplus rule and the unclaimable rungs
Beyond the four rungs, the ladder carries a notation for hardening a gate past what it claims: the rung is the promise that must be met, the surplus is what is done anyway and never advertised. Exceed the rung wherever hardening is cheap; never bill the surplus as the promise.
The two unclaimable rungs. On one platform the estate does not control and cannot export from, a guarded rung was not available at any price — the surface's ceiling sat below the intent. On a code host, a graduated delegate was ungrantable, because the permission model has no intermediate tier. The option that appeared to grant a limited role would have named a role nobody there can hold, and the only real alternative was full control. A decorative grant is worse than no grant, because it is believed.
The field survey, vendor by vendor
Run 2026-07-10; graded survey-based, medium confidence; not re-run since, and decaying until it is.
- Anthropic, on Claude Code's Auto Mode: a deliberately conservative, reasoning-blind classifier with no designed trust accumulation. Where approval history figures at all, it figures as a flagged failure mode, suppressed rather than built upon (2026-03-25).
- Cursor, Auto-review run mode: an allowlist paired with a sandbox and a classifier subagent (2026-05-29). Same pattern.
- Microsoft, Copilot Studio guidance: expand an agent's capabilities gradually as it proves reliability — as manual administrator judgment, not as a feature (2026-06-11).
- Okta and Auth0: scoped, time-bound, per-action delegation, not promotion (GA 2026-04-30 and 2025-11-19).
- Paperclip, the shipped counterexample: Agent Trust Levels, automatic promotion and demotion from run history, opened and closed as completed the same day, its promotion fields verifiable in the codebase (2026-03-09).
The shipped counterexample is the field's direction of travel made concrete: the same promotion-from-history idea, with the opposite answer to who decides.
The autonomy fences, in full
Not yet in its proving phase; each item is a configured limit rather than prose-only governance.
- A named, multi-stage stop ladder escalating from a light throttle through a hard kill to an account freeze.
- A local supervisor process pinging an independent, external dead-man's-switch monitor on a regular heartbeat.
- Hard daily spend ceilings carrying staged behavior at fractions of budget.
- Per-task wall-clock, token, and retry ceilings on top.
Underneath, a journal entry lands before the action it records completes and carries a sequence number, so a gap is detectable and is itself an incident.
The citation record
Three quarantined citation passes have read every external source against its primary — two on 2026-07-11, and one on 2026-08-22, when one source could not be fetched and stands on the earlier passes' record. Corrections ran from quote precision and product-name fixes up to one load-bearing status change: a trust-levels proposal recorded as work-in-progress had been closed as completed the day it opened. A benchmark comparator was corrected the same way, in the body rather than silently.
A later check re-read a subset. Four sources held exactly as stated. One figure was true but attributed to the wrong page, and the attribution was corrected. One characterization was contradicted by its own source: an early-stage agent-registry platform had been read as keeping a human-or-auditor grant above its entry tiers, and its own published tier table diverged from that. The claim was removed from body and table rather than restated in a weaker form, and no replacement claim will be made about that source without a fresh primary read.
How this paper was made
This paper was written by a commissioned drafting agent from isolated research sweeps, then reviewed by a council of seats on different models, each returning assent with amendments. A second commissioned pass produced the outward paper, gated by an independent verifier with no write tools — a seat that failed the draft and passed it only after repair. A further council read it across independent lenses, and a synthesis seat turned those reads into one brief, with the judgment calls put to the operator.
Sources
Every source below was verified against its primary by a quarantined web scout on 2026-07-11, and again on 2026-08-22 — save one that could not be fetched at the later pass and stands on the earlier one.
- Anthropic, "How we built Claude Code auto mode" (2026-03-25); "Building effective human-agent teams" (2026-06-24; not re-checked since)
- Letta, "Introducing Context Repositories: Git-based Memory for Coding Agents" (2026-02-12)
- Perplexity, "Introducing Model Council" (2026-02-05); Wu et al., "Council Mode," arXiv:2604.02923 (submitted 2026-04-03, rev. 2026-07-05)
- Cursor, "Auto-review Run Mode" changelog (2026-05-29); Microsoft Learn, "Design autonomous agent capabilities" (updated 2026-06-11)
- Okta for AI Agents (GA 2026-04-30); Auth0 for AI Agents (GA 2025-11-19)
- Nous Research, Hermes Agent documentation — the Skills System "/learn" (accessed 2026-07-11), with MarkTechPost coverage (2026-06-24)
- GitHub, paperclipai/paperclip issue #379, "RFC: Agent Trust Levels" (opened and closed-completed 2026-03-09; re-checked and holding as stated)
- Linux Foundation, "Linux Foundation Announces the Formation of the Agentic AI Foundation" (2025-12-09) — source of record for the AGENTS.md adoption figure and the stewardship fact
- Li et al., "SkillsBench," arXiv:2602.12670 (Feb 2026) — curated skills raise average pass rate by roughly sixteen points over a no-skills baseline (16.2 at first posting; 16.6 in the revision current at verification); self-generated skills scored below that baseline where they were tested (re-read at the primary 2026-08-22)
- Checked and not cited: an early-stage agent-registry platform considered as support for the human-gated trust-ramp finding (accessed 2026-07-11). Its own published tier table diverged from that characterization; no claim rests on it, and none will without a fresh primary read.
The Ember Guildhall — kept by Rob Ruud, its Guildmaster · co-signed by the Majordomo of the Ember Guildhall.