## Summary
This PR is Phase 4 of 5 in [AERIE-2174 — Port Aerie–Sindri document field reconciliation](https://linear.app/builder-team/issue/AERIE-2174). It reconstructs Aerie's reconciliation lifecycle, settlement and scheduled automation from current main, tracked by [AERIE-2455 — Reconstruct the Phase 4 reconciliation lifecycle from current main](https://linear.app/builder-team/issue/AERIE-2455) and implementing the approved [AERIE-2196 lifecycle slice](https://linear.app/builder-team/issue/AERIE-2196).
It adds pinned Workflow Instance start, bounded polling and recovery, complete-run inspection, Aerie-owned final-output settlement, scheduled verified commit, and monitored cron ownership.
Production effect: controlled rollout. The cron entries become scheduled when deployed, but Workflow start and Site writes remain behind the existing environment, registration, active-Site and arbitrary-cardinality rollout gates. Merging this PR does not publish assets, bind credentials, enroll Sites or start an E2E run.
---
## Why
Phase 3 established the sole verified Site-write boundary but deliberately left it without a production caller. This slice adds the fenced lifecycle that can claim due work, start the pinned Sindri instance exactly once, inspect a completed run, persist its validated proposal and schedule that existing verified commit. It prevents duplicate starts, stale or late settlement, unbounded retries and stuck leased work before Phase 5 adds Site-facing presentation.
---
## Business Value
- Reconciliation work can progress through one bounded, recoverable lifecycle without agents writing Site fields directly.
- Stable idempotency, leases and retry caps prevent duplicate Sindri runs and stuck executions.
- Aerie retains final authority over settlement, policy, provenance, audit and every Site write.
- Operators can see the three lifecycle schedules and hourly discovery through existing monitoring surfaces.
---
## How does it work
1. reconciliation/coordinator.ts claims the oldest due execution, rechecks registration and rollout authority, prepares protected read access and starts the pinned Workflow Instance with stable aerie-reconciliation:v1:${executionRef} idempotency.
2. reconciliation/sindriRuntime.ts exposes exactly three private operations over the existing authenticated Sindri transport: start the pinned instance, read run status and inspect the complete run.
3. reconciliation/coordinatorPollRecovery.ts polls leased executions, caps attempts at three, rejects stale or late results and recovers one eligible expired start from the bounded 16-row scan.
4. reconciliation/finalOutputSettlement.ts validates the complete inspection, persists the canonical proposal and citations atomically, and schedules the existing internal commitVerified({ executionId }) boundary only after successful settlement and write-gate checks.
5. crons.ts and cronRegistry.ts register and monitor the three one-minute lifecycle owners and the hourly REBL3 discovery owner while preserving every existing schedule.
---
## Scope
### Included in this phase
- Pinned Workflow Instance start, status polling and complete-run inspection.
- One-winner claims, lease recovery, three-attempt caps and late-result fencing.
- Final-output settlement before the sole scheduled verified commit.
- Existing active-Site and arbitrary-cardinality rollout checks at lifecycle boundaries.
- Three lifecycle cron owners plus the Phase 1 hourly discovery registration.
- The deletion-reduced AERIE-2270 lifecycle test surface.
- AERIE-2477 repair of retry classification, bounded uncertain-start recovery, freshness, resolver availability, field-owned citation associations and marked-recovery rollout fencing, with seven new named focused regressions.
- Exact final PR diff paths (Phase 4 plus AERIE-2477):
chat/convex/_generated/api.d.tschat/convex/automations/cronRegistry.ts
chat/convex/automations/monitoring.test.ts
chat/convex/crons.ts
chat/convex/reconciliation/admin.test.ts
chat/convex/reconciliation/admin.ts
chat/convex/reconciliation/commit.test.ts
chat/convex/reconciliation/commit.ts
chat/convex/reconciliation/coordinator.ts
chat/convex/reconciliation/coordinatorPollRecovery.test.ts
chat/convex/reconciliation/coordinatorPollRecovery.ts
chat/convex/reconciliation/coordinatorStart.test.ts
chat/convex/reconciliation/finalOutputSettlement.test.ts
chat/convex/reconciliation/finalOutputSettlement.ts
chat/convex/reconciliation/operator.test.ts
chat/convex/reconciliation/propertyAcquisitionFieldPolicy.ts
chat/convex/reconciliation/readiness.test.ts
chat/convex/reconciliation/reads.ts
chat/convex/reconciliation/sindriRuntime.test.ts
chat/convex/reconciliation/sindriRuntime.ts
chat/convex/reconciliation/validator.test.ts
chat/convex/reconciliation/validator.ts
chat/convex/sindri/client.test.ts
chat/convex/sindri/client.ts
scripts/check-monitoring-cron-coverage.test.mjs
### Deliberately excluded for later phases
- Site-facing evidence, provenance and decision-lineage UI — Phase 5 / AERIE-2197.
- Reconciliation-specific Forge inspector presentation — the generic inspector remains unchanged.
- Asset publication, materialisation, credential binding, Site enrollment and rollout configuration.
- New live execution during review repair — the controlled Austin/Roswell E2E passed before Mercy review; no further deployment, activation or shared-data mutation is part of AERIE-2477.
- Obsolete generic sindri/runs.ts, workflows.ts and workflows.test.ts modules.
- Specifications, inventories, reports, handoffs, logs and temporary agent artifacts.
---
## Test plan
### Automated validation
- focused lifecycle and Sindri client tests — 47/47 passed (pnpm --dir chat exec vitest run --project edge convex/reconciliation/readiness.test.ts convex/reconciliation/operator.test.ts convex/reconciliation/coordinatorStart.test.ts convex/reconciliation/coordinatorPollRecovery.test.ts convex/reconciliation/finalOutputSettlement.test.ts convex/reconciliation/sindriRuntime.test.ts convex/sindri/client.test.ts --maxWorkers=1)
- complete reconciliation suite — 97/97 passed (pnpm --dir chat exec vitest run --project edge convex/reconciliation --maxWorkers=1)
- monitoring suite — 92/92 passed (pnpm --dir chat exec vitest run --project edge convex/automations/monitoring.test.ts --maxWorkers=1)
- cron ownership — 3/3 passed (node --test scripts/check-monitoring-cron-coverage.test.mjs)
- root suite — 152/152 passed (pnpm test:root)
- Chat and Convex typecheck — passed (pnpm --dir chat typecheck)
- architecture boundaries, Convex paths, read bounds and test architecture — passed (pnpm lint:boundaries, pnpm lint:convex-paths, pnpm lint:read-bounds, pnpm lint:test-architecture)
- exact-path Biome — passed
- git diff --check — passed
- pre-repair preservation evidence — four lifecycle modules matched the accepted source byte-for-byte; all nine AERIE-2270 test hashes matched; the current Sindri 2026-09-20 contract is preserved
- generated declaration — only the authorised eight module-registration lines changed and full typecheck passes. convex codegen was attempted but could not run without a bound CONVEX_DEPLOYMENT; no deployment selector or credential was copied because binding one was outside the reconstruction safety boundary
- original Phase 4 diff scope — 17 authorised paths before review repair; the final 25-path inventory includes the 14-path AERIE-2477 repair
- AERIE-2477 focused repair tests — 71/71 passed across seven files, with seven named focused regressions (pnpm --dir chat exec vitest run --project edge convex/reconciliation/coordinatorPollRecovery.test.ts convex/reconciliation/coordinatorStart.test.ts convex/reconciliation/finalOutputSettlement.test.ts convex/reconciliation/commit.test.ts convex/reconciliation/admin.test.ts convex/reconciliation/reads.test.ts convex/reconciliation/validator.test.ts --maxWorkers=1)
- AERIE-2477 complete reconciliation suite — 104/104 passed across 15 files (pnpm --dir chat exec vitest run --project edge convex/reconciliation --maxWorkers=1)
- AERIE-2494 focused target/source/recovery regressions — 3/3 passed after 3/3 expected red failures; both full start/recovery files 23/23 passed (pnpm --dir chat exec vitest run --project edge convex/reconciliation/coordinatorStart.test.ts convex/reconciliation/coordinatorPollRecovery.test.ts --maxWorkers=1)
- AERIE-2494 complete reconciliation suite — 107/107 passed across 15 files; foundation/discovery 20/20, monitoring 92/92, cron ownership 3/3 and root 152/152 passed. Chat/Convex typechecks, architecture/path/read-bound/test-architecture checks, five-path Biome and git diff --check passed; two independent read-only reviewers returned PASS on the exact diff
- AERIE-2477 Chat/Convex typechecks, architecture boundaries, Convex paths, read bounds, test architecture, exact 14-path Biome and git diff --check — passed; two independent read-only reviewers returned PASS on the exact cumulative diff
### E2E and Ready-head continuity
The controlled reconciliation E2E passed at authored Phase 4 head 543a44b91502a7f3f14063ab4fb0165b14bac832; no product-code repair was required. Before marking the PR Ready, current main (68924d9d10e45ab31ec44fe19e9b648146bc2591) was merged into the branch as 79ea86b0026fcfeab3be8a6d351c696723bb88c7. The intervening main commit changes only three Forge/Sindri list UI files already present in the PR base. All 17 authorised Phase 4 path blobs are byte-identical between the E2E-tested authored head and the Ready head, and the pre-repair PR diff against current main remained exactly those 17 paths. The first hosted CI and Mercy review ran against that Ready head; the repair push requires new-head CI and Mercy review.
### Branch update after the second repair
After AERIE-2494 was committed as f89ccb58f4a2d40cfa5f8d180fb10bfe42d202b4, current main (1a8f4580af84d7090fb1f56e16b93c2b20b6a533) was merged via GitHub's Update branch as 680f0b3ac40da34c4f98a53c1b68ac45fe9e43bd. The merge commit has exactly the repair commit and current main as parents; all 25 PR-owned path blobs are unchanged. The full 25-path binary PR diff is byte-identical across the old and new bases (SHA-256 a3baa26b061531133cb0c678cfa24f0f588471348eb87b3edcc03ff0596eca48). Upstream main changed other surfaces, including Rhodes Worker reconciliation integration; the earlier controlled E2E was not rerun. A second GitHub Update branch merged current main (ff7809c7f310899777cc1d73e721ab2b16e9d0af) into that branch as bd83ff6a85ca7d471aa5b10089c5d58361cbcf06 before restoring Ready. The AERIE-2494 repair commit remains an ancestor and the 25-path PR diff remains byte-identical at SHA-256 a3baa26b061531133cb0c678cfa24f0f588471348eb87b3edcc03ff0596eca48. Required CI and review must evaluate the current Ready head, not either prior merge head or the old reviewed repair head.
### Time for Implementation
About 4 to 6 engineer-weeks without AI assistance, including source archaeology, reconstruction, lifecycle concurrency review, reduced-test preservation, integration validation and controlled E2E.
### Manual QC
E2E-Prepper ran the controlled Phase 4 lifecycle against the personal development environment using pinned Workflow Instance wfi_Q7R8sSOZLbs and contract 2026-09-20. The only manual product entrypoint was rebl3Discovery/orchestrator:runHourlyDiscovery; discovery-page, poll, settlement and commit functions were never invoked manually.
- Austin: execution reconciliation-receipt-6608a976-92c2-456b-8e0e-0b83a9dcb73c, Sindri run run_2e5AeDmwAFn, completed / updated / commit; 2 fenced sources, 7 citations, 12 decisions, 3 updated fields, 3 provenance rows, revoked read grant and 1 write audit.
- Roswell: execution reconciliation-receipt-02f473dc-c391-432d-b88e-00e60bc202a0, Sindri run run_sJmIe6ESgci, completed / updated / commit; 2 fenced sources, 7 citations, 12 decisions, 4 updated fields, 4 provenance rows, revoked read grant and 1 write audit.
- Combined durability: 2 accepted proposals, 14 citations matched to fenced source tuples, 24 field-history rows, 7 current provenance rows, 2 revoked grants, 2 write-audit rows and 0 nonterminal executions.
- Runtime quality: 6/6 Agent activations completed, 0 runner failures, with the Quality Bar, three attempts and 60-second timeout preserved.
- Cleanup: observe/commit allowlists emptied, start and automatic-write kill switches restored, personal Worker binding restored, temporary Worker/processes/files removed and the exact PR worktree left clean.
No production action or upstream REBL3, Rhodes or Due Diligence write occurred. The final bounded mode-600 evidence artifact contains no secrets, source content or provider payloads.
Evidence-hygiene disclosure: an initial read-only convex data workflowRuns attempt returned malformed/truncated JSON, and the local parse exception echoed part of one historical source quote into the isolated E2E agent tool transcript. That failed attempt produced no artifact, retained log or credential exposure. Workflow-run table inspection was removed from the verifier, temporary runner/tail logs were deleted, and the final evidence was generated only from sanitized Aerie metadata. This transcript-only disclosure does not change the green product result.
## Review repairs and contract clarifications
### First review
Review [5310603172](https://github.com/AI-Builder-Team/Aerie/pull/1492#pullrequestreview-5310603172) was classified under [AERIE-2477](https://linear.app/builder-team/issue/AERIE-2477/preserve-retryability-citation-ownership-and-rollout-fences-in-phase-4) at exact reviewed head 79ea86b0026fcfeab3be8a6d351c696723bb88c7.
Six blocking findings are accepted for a bounded repair. The start-path finding contains two independently reachable failures, so the repair owns seven focused durable contracts:
- unexpected status or inspection exceptions must use the existing bounded retry path rather than terminal status_response_mismatch;
- transient read-preparation failures must leave a claimed execution recoverable rather than record permanent bad input;
- action exceptions and repeated unrecognized or malformed start responses must preserve stable idempotency, advance the bounded recovery state and reach an explicit durable outcome rather than remain starting;
- only an explicit freshness mismatch may terminalize as evidence_not_current; unexpected operational failures must remain recovery failures;
- citation resolver unavailability must remain retryable rather than become terminal citation_invalid;
- intentional cross-array citation reuse must persist every field-owned association needed by verified history and provenance;
- marked recovery must recheck the current Site rollout and effective mode before retrying.
Two blocking claims are rejected:
- reconciliation has one global pinned Workflow registration shared by arbitrary-cardinality Sites. The operator refuses a second row, referenced rows cannot be deleted, and the registry deliberately enforces the singleton. Multiple retained registrations require out-of-contract database corruption; they are not a supported multi-Site state;
- the exact 16-row expired-lease scan is an approved bound. Running and verifying rows are independently claimed and drained by the one-minute poll owner, so they can delay but do not permanently starve a later starting row under the scheduled lifecycle.
The five review-deferred items remain outside this repair: expired/revoked grant cleanup, renewal between marked retry attempts, successful pending/running poll counting, runtime-classification test coverage and the related grant-expiry variants. They have no demonstrated incorrect durable outcome on the current head.
The repair retains the existing idempotency key so a possible remote partial success is rediscovered without creating a duplicate run. The initial claim advances durable attemptCount before the external call; an uncertain response retains its lease and the repaired recovery advances bounded pollAttemptCount, then terminalizes explicitly at the cap. Structured permanent failures retain sindri_start_permanent, distinct from retry exhaustion. The validator retains separate unowned and field-owned citation associations through settlement, verified commit, field history and provenance. The persisted association bound is 440 = 200 unique source citations + 12 allowed fields × 20 per-field citations; the wire bound remains 200 unique citations. Revival of a marked revoked grant is freshness-fenced before access becomes readable.
The repair preserves Aerie write authority, the generic Sindri boundary, exactly three private runtime operations, the global pinned registration, the exact 16-row scan, stable idempotency, three attempts, cron ownership and the E2E-tested happy path. No schema, migration, generated contract, deployment, activation, credential, new E2E, Site enrollment, Phase 5 or upstream-write surface changed.
Before repair, exact-head validation was green: focused lifecycle/client 47/47, reconciliation 97/97, monitoring 92/92, cron ownership 3/3, root 152/152, Chat/Convex typecheck and architecture/static checks. The controlled Austin/Roswell E2E remains green. On the exact pre-commit repair diff 8c1945bbe89fd3d269dfbecf49f07e1e83b345821fc143089924ce5e84b66214 against reviewed head 79ea86b0026fcfeab3be8a6d351c696723bb88c7, seven named focused red → green regressions pass within the 71/71 focused tests; reconciliation is 104/104 across 15 files. Chat/Convex typechecks, architecture/path/read-bound/test-architecture checks, Biome on all 14 repair paths and git diff --check pass; two independent read-only reviewers returned PASS. The reviewed repair changes only 14 files; the full PR scope after push comprises the 25 paths listed above.
### Second review
Review [5317780426](https://github.com/AI-Builder-Team/Aerie/pull/1492#pullrequestreview-5317780426) covers exact head 678122289153476365a20ba0ad4176105f1e7c27. [AERIE-2494](https://linear.app/builder-team/issue/AERIE-2494/preserve-stale-evidence-reason-during-reconciliation-start-preparation) accepts one proven durable blocker: prepareReconciliationStartInput collapsed a typed target-revision or source-facts freshness mismatch into unavailable. Initial start then terminalized a changed target/source as read_access_unavailable; the same misclassification was reachable when a target changed between separate recovery-preparation and read transactions. The five-path repair preserves a private typed stale result, recording terminal stale with the exact target_revision_stale or source_facts_stale reason and revoking the matching grant. Genuine inaccessible input still records read_access_unavailable; unexpected failures remain recoverable. Registration, lease, grant binding and hash checks remain in place.
Two reported blockers are rejected on their actual paths:
- An ordinarily expired or already revoked sole grant does not prevent terminal cleanup. registry.ts validGrant validates hash, safe timestamps, expiresAt > createdAt and an optional revocation timestamp; it does not require expiresAt > now or an unrevoked row. markTerminal therefore patches the execution and revokeIfLive only revokes a still-live grant. Duplicate or malformed inventory is a structural error, not expiry or ordinary revocation.
- The observe-mode poll fixture is not excluded by a commit-list entry. Its registration has global activationMode: observe; rolloutPolicy.ts admits membership in either pilot list, then global observe bounds its effective mode to observe, matching the seeded execution and the recovery guard. Changing its list would not repair a failing production/test path.
The runtime 2xx cast is checked by the start, status and inspection caller guards before a durable success. The review itself defers its classification robustness question, observe-success settlement coverage, discovery cron owner-table coverage and transport-classification tests. Four suggestions/nits remain hardening only. No such deferred or rejected item is bundled into this repair, and no previously approved global-registration, 16-row scan, citation, three-attempt or idempotency contract is changed.
Three new named regressions failed on the old behavior and pass on the repair: target freshness, source freshness and the separate recovery transaction gap. The exact pre-commit five-path binary diff is 3fc473456900a05bff027bbc144f14eeec878e3032de221b4776189fa7a64bfb, committed as f89ccb58f4a2d40cfa5f8d180fb10bfe42d202b4; complete reconciliation passes 107/107, retained foundation/discovery 20/20, monitoring 92/92, cron ownership 3/3, root 152/152, and focused start/recovery 23/23. Typechecks, architecture/static checks, Biome and whitespace checks pass; two independent read-only checks passed the same diff. The PR still has exactly the 25 paths listed under Scope. No deployment, activation, credentials, E2E rerun, migration, Site enrollment, external traffic or upstream write was part of either review repair.
### Third review: invalid-state claims and unchanged boundaries
Review [5319650419](https://github.com/AI-Builder-Team/Aerie/pull/1492#pullrequestreview-5319650419) covers Ready head bd83ff6a85ca7d471aa5b10089c5d58361cbcf06. All eight non-Mercy required checks succeeded on this head (the remaining code-smith check was skipped). The reviewer withdrew its earlier global-registration, 16-row scan, ordinary expired-grant and observe-fixture objections. Its four remaining blocking claims rely on persistent inventory corruption or on a typed-error consumer that does not exist in the reviewed production paths; no supported mutation can produce their premise:
- Queued input revision (coordinator.ts:370–372). readiness.ts:createReceipt inserts the queued execution, its sole input revision and captured sources, then patches inputRevisionId in one Convex mutation. A failed insertion/patch rolls that entire mutation back. There is no production deletion or second insertion of input revisions. loadInput returns null for a missing/duplicated row only if a database invariant has already been broken outside this supported writer. The proposed terminalization of an arbitrary corrupted row is not a repair of a reachable partial-persistence state.
- Receipt unavailable sentinel (reads.ts:464–473). The public-facing loader translates ReceiptUnavailable into a clean userError. Actual read/start preparation calls the unchecked loader inside prepareReadAccess and classifies ReceiptUnavailable explicitly. Other callers do not branch on that sentinel: commit catches any loader rejection as its existing failure, settlement catches it as a receipt error, recovered-start preparation catches it as recovery-required, and validator HTTP returns a clean denied/error boundary. No typed classifier is bypassed by the translation. The third-review request would alter internal error semantics without a proven wrong outcome.
- Settlement exception (finalOutputSettlement.ts:728–732). The action uses a structured rejected result for malformed/missing final output and resolver-unavailable for transient citation lookup. Its settlement mutation explicitly handles absent/expired grants, changed registration and changed receipt freshness. The cited missing input/receipt rows, duplicate registration and malformed grant inventory are out-of-contract persisted corruption: the receipt writer is atomic, registration is global and singleton, and the read-grant writer checks the indexed sole-grant inventory. The generic catch intentionally avoids labeling an unexpected mutation/platform exception as a business failure; no supported durable-invalid path was shown to be collapsed.
- Poll terminal grant inventory (coordinatorPollRecovery.ts:506–523). grantInventory accepts an ordinarily expired or previously revoked sole grant; it rejects duplicate or structurally malformed rows. reads.ts:prepareReadAccess first reads the bounded indexed inventory and either renews its one matching row or inserts a first grant; the path does not manufacture a duplicate. Transactional validation before terminalization protects the grant/lifecycle invariant; skipping it to force a terminal patch would turn corruption into apparently valid evidence. The recovered-start inventory guard has the same provenance boundary.
The five missing-test findings and nine hardening/nit findings are retained for human consideration, not treated as proof of a broken shipped production path. This review made no code, deployment, activation, credential, E2E or upstream-write change. The PR remains at the unchanged 25-path diff SHA-256 a3baa26b061531133cb0c678cfa24f0f588471348eb87b3edcc03ff0596eca48; we will not push an empty or unrelated commit merely to retrigger review.
### Fourth review and final combined repair
Review [5320089243](https://github.com/AI-Builder-Team/Aerie/pull/1492#pullrequestreview-5320089243) covers bd83ff6a85ca7d471aa5b10089c5d58361cbcf06. [AERIE-2498](https://linear.app/builder-team/issue/AERIE-2498/revoke-a-revived-marked-recovery-grant-when-read-secret-disappears) accepts its one proven blocker. A marked recovery can atomically revive its previously revoked grant, then the action can be interrupted. When the retained lease expires and the read secret is absent on the next attempt, recovery previously terminalized the execution as read_access_unavailable without revoking the still-live grant. The focused repair validates the sole grant's structure and execution/Site/target binding even without a secret-derived hash, rejects unbound inventory before any write, and revokes a valid bound grant in the same mutation that terminalizes the execution. A supplied expected hash remains checked; no empty-secret hash is invented.
The legacy-citation compatibility finding assumes that field-blind Phase 4 settlement rows were deployed and left nonterminal before this still-unmerged Phase 4 writer. The current production base has no Phase 4 settlement/commit producer. The controlled personal-development Austin/Roswell E2E produced completed, committed executions and zero nonterminal rows. Inferring field ownership from a hypothetical old unowned citation row would weaken verified provenance, so no speculative migration or compatibility fallback is included. The review withdrew the previously rebutted registration, sentinel, queued-input, grant-inventory and observe-fixture blockers. Missing coverage and 15 suggestions/nits remain nonblocking hardening/polish, not this security repair.
The repair is exactly two paths, with one named regression that fails on the reviewed code (terminal execution retains an unrevoked grant) and passes after the repair, including an unbound-grant negative check. Its pre-commit binary diff SHA-256 is aaa59a3cb04353181ce2139d5064ccc9e8487d71784cb2db123e2c104769a3d5; two independent read-only reviewers passed that exact diff. The complete reconciliation suite passes 108/108 across 15 files; retained monitoring/Sindri tests pass 107/107, cron owners 3/3 and root 152/152. Chat/Convex typecheck, architecture/path/read-bound/test-architecture checks, exact-path Biome and whitespace checks pass.
New main (fe883884f58b01e31c3b33d385b3b39a43e65f84) introduced one content conflict in scripts/check-monitoring-cron-coverage.test.mjs: this PR's three reconciliation owner assertions and main's two capacity owner assertions are both preserved, with the combined registration count 42. The integrated and committed merge tree is 92d3e82cbf0207a64b21767e3c0c65f539b01360. On that combined tree, reconciliation 108/108, retained monitoring/Sindri 107/107, cron owners 3/3, root 152/152, Chat/Convex typecheck, architecture/static checks, three-path Biome and git diff --check pass. The final PR diff against this main remains exactly 25 paths, SHA-256 94b6f224a52e829d06a6e1f6f3d0cf5542f9bfd0fc014ca868f1d9f5ab6b7132. One merge commit 71817e6c1bc2660610820bc0548c42527d7987dc carries both the two-path repair and conflict resolution; it was pushed once to the PR branch. Required hosted CI and automated review must evaluate this exact new head.
The earlier controlled Austin/Roswell E2E passed before the review repairs. Its live scenario has not been rerun after repair or the later main updates; the current-head automated suites and hosted CI remain the post-repair gates. This change does not deploy, activate, bind credentials, enroll Sites, change schema, run live E2E, mutate shared data or perform upstream writeback. Aerie alone owns Site writes, Sindri remains generic, and the one global registration, 16-row scan, three private runtime operations, stable idempotency, three attempts and citation bounds stay unchanged.
### Fifth review: inspection lease and nonblocking findings
Review [5320866146](https://github.com/AI-Builder-Team/Aerie/pull/1492#pullrequestreview-5320866146) is for exact head 71817e6c1bc2660610820bc0548c42527d7987dc, where all nine non-Mercy hosted checks completed green or intentionally skipped. [AERIE-2503](https://linear.app/builder-team/issue/AERIE-2503/fence-in-flight-inspection-past-the-poll-lease-without-duplicate) accepts the in-flight inspection lease problem. claimOnePoll claims a 120-second lease; a completed status changes the row to verifying and nextPollAt=now, retaining that original lease. pollDue first invokes status GET with no shorter timeout; that call can itself cross the initial lease and be reclaimed before the completed transition. If status completes inside the initial lease, pollDue invokes inspection with no sub-lease timeout. If inspection crosses the remaining lease, another poll can reclaim the same due row and inspect again. The original valid result loses its token fence, and repeatedly slow responses can prevent settlement. AERIE-2503 repairs both in-flight calls: only the private reconciliation status and inspect GETs get a cancellable 90-second operation budget starting before capability/link preflight; completed → verifying renews the same token and sets its lease/next-poll due time to the transition time plus 120 seconds. Authorization remains before link/network. The validated five-path fix is now in the pushed merge commit e4e3acba642c0b4fafe22833a692d2b95cbe156a. The existing one-registration/16-row/start-attempt/idempotency/citation and Aerie-only-write boundaries remain authoritative.
The purported late-start stale-grant leak at coordinator.ts:611 assumes that the starting row has failureCode=registration_control_changed and a live grant under the original lease. That state has no supported producer. The only writer of that marker on a starting row is registry.ts:fenceExecutions, which in the same transaction revokes the indexed sole grant. coordinatorStart.test.ts exercises an operator transition during a deferred Sindri response and asserts the grant is already revoked before recordStartSuccess records the late run ID, and remains revoked afterward. Recovery can revive a marked grant, but claimExpiredStart first replaces the lease token: the original recordStartSuccess immediately returns recovery_required on its old-token check, while recovered-start recorders own the new token and revoke on stale terminalization. A change only to environment rollout config does not write the marker; recordStartSuccess takes its !registration/unmarked recovery_required branch instead of the cited terminal patch. Thus this direct terminal patch cannot leave a newly live grant in the supported path; do not add a redundant write to address an impossible marker/live-grant state.
The pending/running poll reschedule does not advance the failure counter, intentionally preserving the actual provider running state. The review itself defers this finding: a legitimate long-running provider run is not an exhausted failed attempt, and no incorrect durable result is demonstrated. The cron owner, runtime classification, negative capability/link, status/inspection mapping, settlement observe and operator authorization requests are missing-test or hardening proposals, not shown shipped defects. The 13 suggestions/nits likewise remain outside this narrow repair. No deploy, activation, credentials, live E2E, shared-data mutation, or upstream write occurs during review repair.
The first uncommitted repair candidate was rejected at the parent pre-push gate despite two PASS reviews: placing lazy first-start Sindri user-link provisioning inside the shared provider-error classifier would have turned a missing-config or transport exception into sindri_start_permanent or a consumed attempt, rather than preserving startDue's recovery_required state. A fresh gated corrective task restored the original boundary: an outer timer try/finally covers the two bounded GET operations, while a narrow inner provider fetch/body catch leaves capability and lazy link preflight errors outside classification. One direct sindriRuntime.test.ts regression is red on the rejected candidate and green after correction; no start timeout or generic Sindri behavior change is introduced.
The final five-path pre-commit corrective diff SHA-256 is 8042a3f36ca1b7dfc5b28c71f68e1cddeba73d24619e85add98734833fa72c6f. Three focused red→green contracts pass; the complete reconciliation suite is 111/111 across 15 files and the two independent read-only reviews passed the same diff. Against latest main a0c078111f973630c819fa843719b12144796583, the conflict-free integrated tree is 82f94ce8683f78931d9a51d79582b4a395f85099, and its PR diff is 25 paths, binary SHA-256 e8f9af8e8e68ff2ca9f44d59e4b7c5844b0258d71fc03d3a8cbc1be4d490f899. On that combined tree, reconciliation 111/111, retained monitoring/Sindri 107/107, cron composition 3/3, root 152/152, Chat/Convex typecheck, architecture/Convex-path/read-bound/test-architecture, five-path Biome and whitespace checks pass. Required hosted CI and automated review must assess the pushed exact head e4e3acba642c0b4fafe22833a692d2b95cbe156a. The Austin/Roswell E2E has not been rerun since the earlier controlled run, and this review repair does not deploy, activate, enroll Sites, bind credentials, mutate shared data or write upstream.
### Sixth review: settlement error flow at current head
Review [5321871235](https://github.com/AI-Builder-Team/Aerie/pull/1492#pullrequestreview-5321871235) evaluated exact pushed head e4e3acba642c0b4fafe22833a692d2b95cbe156a; all nine non-Mercy checks completed successfully or were intentionally skipped. Its free-text claim that a malformed inspection payload or citation resolver exception escapes before settlement does not match the supported producer-to-consumer flow. inspectReconciliationFinalOutputsV1 snapshots and bounds the provider's parsed JSON, returning accepted_output_missing or accepted_output_invalid rather than throwing for malformed/oversized output. validateReconciliationFinalOutputV1 catches structural parsing errors as rejected, catches each citation resolver rejection as resolver_unavailable, and catches final canonicalization/digest errors as rejected. settleReconciliationFinalOutputV1 builds a structured SettlementArgs for each of those results and calls the settlement mutation; the rejected branch terminalizes with a failure outcome, while resolver_unavailable intentionally returns recovery_required without labeling transient evidence unavailability as an invalid proposal. The claimed unhandled malformed-payload/resolver path has no reachable supported producer.
On resolver_unavailable, the execution remains verifying under its finite lease. This is an explicit retry state, not a silently lost terminal result: after the lease expires, claimOnePoll can reclaim the due verifying row and re-read the run/inspection; finalOutputSettlement.test.ts verifies that resolver unavailability leaves no partial settlement writes. A persistently unavailable dependency remains unavailable, but manufacturing a permanent business failure or fabricated citation ownership for it would change the approved evidence boundary. Unexpected platform/database failures outside these classified responses are not proof of malformed provider data, and generic exception-to-terminal conversion would mask their cause.
The request for a cap on successful pending/running polls remains deferred in the review metadata and would misreport a truly running provider run. The preflight timer's inability to interrupt an arbitrary stalled capability/link await is also marked deferred; normal polling follows a successful start, which created an active service link, and no production link-deactivation writer exists. Redaction regex, observe-path, cron owner, runtime classification, capability and operator findings are missing coverage or hardening, not demonstrated shipped defects. The 13 suggestions remain nonblocking. This round makes no code, deployment, activation, credential, E2E, shared-data or upstream-write change and does not push an empty commit. Its COMMENTED/auto-approve-withheld verdict is not treated as an approval under this PR's exact-head merge gate; no manual tagged review is requested.
### Seventh review: marked recovery order at current head
Review [5323918288](https://github.com/AI-Builder-Team/Aerie/pull/1492#pullrequestreview-5323918288) covers exact current PR head 7ed459f202f72a48cb11f5846974bc95c29ae281 after main was merged. The 25-path PR diff against current main is byte-identical to the prior validated PR diff, SHA-256 e8f9af8e8e68ff2ca9f44d59e4b7c5844b0258d71fc03d3a8cbc1be4d490f899; all nine non-Mercy hosted checks on this head are green or intentionally skipped. The review withdrew the earlier settlement-exception objection.
The single high-severity marked-recovery finding at coordinatorPollRecovery.ts:1418 reverses the actual order. recordRecoveredStartFailure loads the sole indexed grant, checks its structural/execution/Site/target/hash binding, and returns recovery_required if mismatched. It then checks !registration.active || grant.revokedAt !== undefined before its revocation patch. When a bound grant is live and registration is active, it patches revokedAt once, then follows the retryable branch to schedule leaseExpiresAt and return retryable, or the permanent/exhausted branch to patch a failed outcome and return failed. There is no post-patch revoked-state guard capable of turning either result into recovery_required. The review's own machine-readable finding marks this premise deferred and states that the described post-revocation invalidation is not present. No code or test change is warranted for an impossible ordering. The 19 suggestions/nits remain nonblocking coverage or hardening.
This evidence update makes no commit, push, deploy, activation, credential, E2E, shared-data or upstream write. The current verdict is again COMMENTED because auto-approve was withheld for a human merge. It is not an exact-head approval under our merge gate, and no further manual tagged review is issued from this session.
### Eighth review: interrupted unmarked recovery (repair pushed)
Review [5324040853](https://github.com/AI-Builder-Team/Aerie/pull/1492#pullrequestreview-5324040853) covers reviewed head 7ed459f202f72a48cb11f5846974bc95c29ae281. [AERIE-2541](https://linear.app/builder-team/issue/AERIE-2541/revoke-an-unmarked-recovery-grant-when-preparation-is-interrupted) accepts one reachable blocker: for an unmarked expired-start recovery, prepareRecoveredStart(phase: "prepare") accepts an exact bound, unrevoked but expired grant and extends its TTL. If the following prepareReadRef rejects, recordRecoveredStartFailure(kind: "preparation_interrupted") returned recovery_required without revoking that now-live grant because the execution was unmarked. This left evidence-read authority live after work was interrupted. An unmarked *already revoked* grant cannot pass the preparation guard; the proven path is an expired, unrevoked grant.
Stage 1 produced and received explicit parent SEND approval before completion on frozen diagnosis SHA-256 29d4e4d482b221cdc68fede6d3127ad78257e1eb1f0018443847de1d55ec46f0; Stage 2 had a fresh owner. The accepted exact two-path repair, pre-commit SHA-256 46a42971d2f64b31512046efc0aac43343116b9e477fdd81ef426e6df13b41a7, removes only the unmarked early bypass and revokes the sole currently live, matching execution/Site/target/hash grant after a fenced preparation interruption. It does not patch execution state, attempts, lease, scheduling or failure classification. The same recorder also covers a later failed confirmation fence. One real-seam regression was red at the old head (revokedAt undefined), green after repair. Owner validation: recovery file 15/15, reconciliation 112/112 across 15 files, Chat/Convex typecheck, architecture/path/read-bound/test-architecture checks, two-path Biome and whitespace PASS; two independent read-only reviewers PASS on the exact diff SHA. Parent independently applied that byte-identical patch to a clean worktree at the reviewed head; parent reconciliation 112/112, architecture/path/read-bound/test-architecture and whitespace PASS. Parent typecheck in the new worktree was interrupted by its execution timeout; owner typecheck on the identical patch passed. No code or review artifact outside the two exact paths is authorized.
The reported triggerRevision bound mismatch is not a reachable queued pilot revision. The only production receipt writer (enqueueAfterPromotionHandler → createReceipt in readiness.ts) mints triggerRevision or coalescedThroughRevision from derivePromotionTriggerRevision: the fixed 28-code-point reconciliation-promotion:v1: prefix plus 64-character SHA-256 hex digest, exactly 92 code points, below both 128- and 256-code-point guards. The 129–256-character scenario requires a non-production row writer. The review's structured metadata explicitly defers this finding. We preserve the shared persisted contract; no speculative cross-phase bound migration is in this repair.
The reported nested inspection identity mismatch is possible only if Sindri's single authenticated /v1/runs/{runId}/inspect response violates its own producer contract: controlPlaneReads.ts:inspectRunHandler obtains runStatus and get_workflow_snapshot using the same supplied runId, then projects that snapshot's run through toPublicRun. Aerie's getRunStatus and inspectCompleteRun use its fenced run ID; validInspection checks runStatus.run.id, workflow-instance ID and completed status before settlement, and settleFinalOutput rechecks the claimed run ID against the execution and proposal. No supported route combines another run's snapshot with this run's status; additional nested-ID validation would be defensive hardening against an internally inconsistent Sindri implementation, not a demonstrated cross-run producer path.
The other 19 suggestions are nonblocking coverage/hardening. Boundaries remain: Aerie owns Site writes, Sindri is generic, exactly three private operations, global pinned registration, exact 16-row recovery scan, three start attempts, stable start idempotency, field-owned citations and no live activation or upstream writes. All nine non-Mercy checks at the reviewed head are green or intentionally skipped. The description and untagged classification comment were updated before the single repair commit/push. Parent committed 498b800ed6ca2580e40b2251f4d2af0f2148887a on reviewed head 7ed459f... with only those two files and pushed it once; remote PR head matches and the integration worktree is clean. The PR diff against current main remains 25 paths, binary SHA-256 f7bcb97e309d5bba5a4c31b46ab847185bb4ffcb08d35821a4ebe7f9b6e9a7f6. Fresh hosted CI and Mercy-Watcher review of this exact head are pending; no manual trigger was issued. No deployment, activation, credential binding, Site enrollment, live E2E, shared-data mutation, upstream write, manual Mercy trigger or empty push occurred. Merge still requires exact pushed-head approval and green required CI.