## Summary
This is now the single review and merge target for the school-report stack formerly split across #2083, #2086, #2100, and #2102. Review the complete diff here against main; the three upper PRs are superseded and retained only as historical records.
- Report readiness and categories: correctly classify Transportation and Renovations/Furnishings, preserve known zero versus unknown actuals, and display approved missing budgets as unavailable without fabricated budget variances.
- Financial correctness and recovery: align Facilities/depreciation with the report cutoff through the owning warehouse procedure; preserve native-document layout/readback, source evidence, financial arithmetic, independent audits, bounded correction/recovery, and compatible saved-work resume.
- Bounded parallel processing: run 1–9 school-isolated workers with shared provider pacing and cooldowns, one fenced pipeline lease, failure isolation, and one consolidated completion email.
- HTML delivery: provide a full-width Finance digest and equivalent plain-text fallback with alphabetized native report links, reporting dates, aggregate updated/reused counts, and a conditional failure section.
- Durable source findings: retain both financial and operating Facilities values and insert an exact Finance-owned reconciliation finding rather than silently changing source values or withholding an otherwise usable report.
- Safe report reuse and pageless exports: verify immutable snapshots, combine retained and newly generated report links without regenerating unchanged schools, and validate readable/unencrypted/nonblank PDFs without imposing an artificial fixed page count.
- Controlled recipients: default to Ashwanth-only review delivery; a separate explicit, delivery-only approved_finance action reuses the proven review run for the fixed five-person Finance audience.
## Consolidation and review scope
| Former layer | Scope now reviewed in this PR |
| --- | --- |
| #2083 | Categories, missing-budget handling, Facilities cutoff, layout, audit/correction recovery |
| [#2086](https://github.com/AI-Builder-Team/Surtr/pull/2086) | Nine paced, isolated concurrent school workers |
| [#2100](https://github.com/AI-Builder-Team/Surtr/pull/2100) | HTML/plain-text completion digest |
| [#2102](https://github.com/AI-Builder-Team/Surtr/pull/2102) | Source findings, retained-report reuse, pageless PDF checks, approved Finance delivery |
GitHub automatically marked #2086, #2100, and #2102 merged into the consolidated base branch, not into main. #2083 is now unstacked and is the only open review target. The original upper branches are retained along with their PR histories.
The base branch codex/school-report-cost-categories is fast-forwarded from 6fa8d66f to the former top-layer commit 114001640a839f1ca837937f71f93c541e1384a6. All original commits are retained; no rebase, squash, cherry-pick, force push, or implementation change is needed. The consolidation baseline tree is identical to the former #2102 head (tree ddc45216dbe5a2e7db21014f85781e1a8844965b).
The PR targets main. It was draft during consolidation and was subsequently marked ready for review by ashwanth1109 while CI was running. This CI fix did not change readiness or enable auto-merge. Consolidation does not merge into main, deploy infrastructure, refresh warehouse data, generate reports, send email, or resolve review threads. Original upper PR descriptions remain available for detailed historical evidence; the original #2083 description is archived below.
## Business Value
Finance gets accurate, evidence-backed school reports despite explicitly identified source gaps, without fabricated budgets or silently reconciled values. Parallel processing and safe completed-report reuse reduce batch turnaround and unnecessary regeneration. A readable consolidated digest and review-first recipient workflow make delivery easier to inspect and deliberately approve.
## Implementation Effort
Approximately 45–67 engineer-hours for an average engineer to hand-code the combined solution without AI assistance, based on the four original layer estimates (20–30, 12–18, 4–6, and 9–13 hours). These estimates include differing amounts of deployment/live-validation work; additional validation effort may be required.
## Validation
### Fresh local checks — October 6, 2026
- School-performance report suite at the exact consolidated candidate: 300 passed.
- Education-financials upstream suite at the same candidate: 151 passed.
- git diff --check against current main: passed.
- Read-only git merge-tree against main at 3f67679a: clean, with no conflicts.
- Every former layer head is an ancestor of 11400164; the fast-forward preserves all four layers with no code change.
- No repository-wide lint/format run or production operation was performed for consolidation.
### Recorded pre-merge live validation
The former top-layer [#2102](https://github.com/AI-Builder-Team/Surtr/pull/2102) records exclusive deployment of the same commit 11400164 to Pipeline-school-performance-reports-prod, task revision 44, with CDK [1/1] and CloudFormation UPDATE_COMPLETE.
Its final review execution 6a1c4b05-9e60-4ca3-aa95-38f5ad74ce4d and explicitly approved Finance delivery execution a25e4985-9956-4adb-b6d5-71e0642addf6 each record 26 reports, zero failures, zero generated, 26 reused. The earlier batch exercised the nine-worker cap, unavailable Malibu budgets, the exact New York $391,000 Facilities source finding, and Raleigh's legitimate ten-page pageless export.
These are retained historical validation records, not new live runs or independently reverified AWS receipts in this consolidation session. Consolidation leaves the implementation commit and tree unchanged. Any subsequent implementation or repair still requires candidate live validation before merging under the user-level Surtr policy.
### CI import-order fix — October 6, 2026
Follow-up commit [9557f15a](https://github.com/AI-Builder-Team/Surtr/commit/9557f15a19396404333ba2f981f59f0a2ddfe7d4) fixes the absorbed stack's Ruff E402 failure in pipelines/runners/mart-aerie-education-financials-refresh/scripts/apply_facilities_cutoff.py. The helper must add its sibling script directory before importing apply_ddl; the import now has an explanatory comment and a narrow, line-only # noqa: E402. No global lint rule or workflow is relaxed.
- CI-pinned Ruff 0.15.22 check and format check pass for the only modified file.
- Python AST is identical to the pre-fix version: comments only, no runtime behavior change.
- The helper's default dry run succeeds without warehouse writes.
- Fresh local suites: 300 school-report tests and 151 education-financials tests passed.
- git diff --check passes.
- No production deployment, pipeline invocation, warehouse refresh, or email send was performed for this comments-only CI fix.
- [Fresh GitHub CI](https://github.com/AI-Builder-Team/Surtr/actions/runs/37482751012) completed successfully for this commit: all seven CI jobs passed, including Ruff check/format and the complete pipeline-runner suite. The automated review rerun was still in progress at the final CI verification; human review remains required before merge.
## Linear
- [SURTR-1535 — Categories and approved missing budgets](https://linear.app/builder-team/issue/SURTR-1535/fix-school-performance-report-categories-and-approved-missing-budgets)
- [SURTR-1538 — Parallel school-report workers](https://linear.app/builder-team/issue/SURTR-1538/run-school-performance-report-batches-with-three-isolated-parallel)
- [SURTR-1551 — HTML completion digest](https://linear.app/builder-team/issue/SURTR-1551/style-school-performance-report-delivery-as-an-html-digest)
- [SURTR-1553 — Unavailable budgets and source findings](https://linear.app/builder-team/issue/SURTR-1553/publish-school-reports-with-unavailable-budgets-and-source-findings)
## Historical base-layer description
<details>
<summary>Original #2083 description and detailed validation history (superseded scope statements)</summary>
The following text is preserved verbatim as historical evidence. Its earlier statements that upper-layer work was separate or excluded no longer describe the consolidated review scope above.
## Summary
Fixes discovered while validating school performance reports in batches. This remains a draft for review only; nothing has been merged. Candidate changes have been deployed for pre-merge validation.
- Expense categories: put 62500 Transportation under Program costs and 62100 Renovations/Furnishings as the second/last detail in Miscellaneous. Include both exactly once in financial totals, with the narrower operating Programs comparison explicitly labeled.
- Approved missing budgets: retain actuals, display Unavailable, and omit budget variances for Lunch Revenue, Financial Aid, Sibling Discount, Stripe Fees, Associate / Other Headcount, Landscaping, Lunch Program, Transportation, Renovations/Furnishings, Computers, Administration/Other, and the Miscellaneous subtotal. Other budget, actuals, Timeback, freshness, enrollment, model, and reconciliation safeguards remain enforced.
- Confirmed zero actuals: budget-only rows with explicitly zero actual postings produce zero management actuals. Unknown actuals with postings are not converted to zero.
- Native report layout: add two rows to each financial table in report copies only, preserving the pinned source template, inherited styles, revision guards, readback validation, and retry behavior. Use 90% financial-body line spacing to retain nine pages at the native 9pt font size.
- Scoped view migration: opt-in, dry-run-by-default migration of only mart_education.school_performance_report_lines, with definition/ACL backups, support for older views missing budget_value_status, transactional owner/grant preservation, and no cascading drops.
- Facilities cutoff repair: derive facilities/depreciation actuals from canonical atomic P&L postings through the report cutoff, instead of including the entire current month. Preserve the accepted QuickBooks generation/snapshot pins, monthly coverage guards, expense sign convention, allocations, model-budget proration, and atomic publication. Add a narrow procedure-only deployment script with backups and owner/ACL verification; no grants/revokes or table rebuilds.
- Native correction addresses: preserve report/claim structure in narrative-correction requests and list exact editable paths, compacting only source evidence. Fixes the observed Chicago correction targeting a columnar transport address instead of a real report field. Invalid paths still fail closed, and final audit remains required; saved drafts/contexts remain compatible.
- Truncated audit recovery: re-review the full original report, prominent text, accounting instructions, and all evidence in a bounded structured call; never treat opaque reasoning/signatures as user-text evidence. Preserve incomplete-coverage rejection, stage cost/deadline limits, corrections and final independent review. Saved research/drafts remain compatible.
- Correction numeric roles and resume compatibility: keep source-linked actual, model allowance and signed variance values separate in every correction request, outside prose compaction. Version correction results/context, so an unfinished old correction can be replaced from its compatible saved complete audit without repeating research/drafting; completed reports remain untouched. Final review and the one-correction limit remain enforced.
## Business Value
Enables usable school performance reports with correctly classified expenses and honest treatment of unavailable budgets. Aligns facilities and depreciation actuals with the report date, preventing future-dated entries from creating false reconciliation failures or overstating current spending. Makes audit output exhaustion recoverable without bypassing financial review or regenerating successful schools. Preserves the distinction between spending and comparison differences during narrative correction.
## Validation
- School report tests: 162 passed after the numeric-comparison correction fix (eight additional regression cases).
- Education financials upstream tests: 151 passed.
- Ruff checks on modified Python files and git diff --check: passed.
- Regression tests cover category totals, optional versus required budgets, zero versus unknown actuals, migration permissions, native row/style/retry handling, and the canonical facilities SELECT across cutoff/quarter boundaries, credits, wrong-school/publication/classification rows, and no-posting prefixes.
### Original changes: Boca Raton live validation
- Deployed only Pipeline-school-performance-reports-prod with --exclusively; CDK showed [1/1] and CloudFormation confirmed UPDATE_COMPLETE (task revision 28). No other pipeline stack was deployed.
- Applied the report-lines view migration with ownership/grants preserved and verified 37 distinct management lines.
- After correcting the added-row page spill, the resumed run completed SUCCEEDED, zero failures, at 2026-09-29 08:49:32 UTC. The nine-page PDF passed the layout guard and visual review; completion email was sent.
- Execution: boca-layout-resume-20260929-0847; run: e55c7dbe-4056-46d7-8d11-e0c2d2449ab7.
- [Boca Raton report](https://docs.google.com/document/d/1sL6T4wiuziNvuKdhX1pNypgD4af-knWO97L8E8YY_ls/edit).
### Batch two: candidate deployment and upstream validation
- Commit 1891f6f6 adds the seven separately approved budget exceptions. Deployed only Pipeline-school-performance-reports-prod using --exclusively ([1/1]), task revision 29; CloudFormation UPDATE_COMPLETE at 2026-09-29 09:06:51 UTC. No other stack or recipient configuration changed.
- Preflight identified future-dated September 30 depreciation in the facilities mart versus the September 29 report cutoff: $11,065.93 Brownsville, $7,747.62 Chicago. Report generation was not invoked against these known-invalid inputs.
- Commit d7070e9c repairs the owning facilities calculation (procedure version 2026-09-29.1). Applied only that procedure and its measure comment, verified original ownership/ACLs unchanged, and retained the original procedure backup. No upstream runner/infrastructure deployment was required.
- Authorized upstream platform run f9de7fb1-c7ed-4a09-9790-c25c922a92c6, execution facilities-cutoff-batch2-20260929-0917, completed SUCCEEDED at 09:18:54 UTC, publishing 10,208 facilities rows. Verified rent, facilities, and depreciation agree across the report and facilities tables for both schools. Correct depreciation: $24,963.43 Brownsville, $0 Chicago.
- That refresh picked up the newer 08:45:47 UTC FinalSite snapshot while the unit-economics tables still used 06:46:21 UTC. Preflight caught the mismatch; no report invocation occurred with mixed enrollment publications.
- With explicit additional authorization, unit-economics run 1c5bee2c-4e0e-47ea-87e3-e1614eb99259 succeeded at 09:26:37 UTC with 1,102 rows. Its automatic per-student run 053f1a3c-7a1f-441c-97bb-61d9d33d8bde succeeded at 09:26:53 UTC, also 1,102 rows. No model assignments or safety checks were changed.
### Batch-two report outcome
Full read-only preflight passed for both schools and the unchanged native template after enrollment publications were aligned. Run 4f390d10-edb1-4fc2-ad7b-c99b6408b426 completed at 2026-09-29 09:51:23 UTC, but its result was partial_failure, not complete report success:
- Brownsville succeeded, passed its independent evidence audit and PDF guard, and all nine pages passed visual review. [Completed report](https://docs.google.com/document/d/1zOALdJl1CU8d4MFsasUCPKs6ORDE1dpYOHNp5j8oMUc/edit). PDF SHA256: c37891c4b46804bf0fa7a76fb93fac7ecea35d234443ff11c5ecfe45210354e0.
- Chicago stopped before publication. Its audit correctly identified comparison-label/scope issues; correction generation then used the nonexistent transport path signals.rows[2][1]. The unchanged field guard rejected it. The summary email reported one ready and one failed school.
- Commit 326a5a4f fixes the correction request's native structure. Regression tests cover full, focused and previously compacted contexts, preserve evidence compaction, and continue rejecting the exact malformed path.
- Deployed only Pipeline-school-performance-reports-prod using --exclusively ([1/1]); CloudFormation confirmed UPDATE_COMPLETE at 09:57:43 UTC, task revision 30. No parallel-worker changes are included in this deployment.
- Verified immutable saved snapshots/draft/audit/context checksums, current freshness, template contract, and real editable paths before the authorized resume.
- Resume execution chicago-native-correction-20260929-0958, run 051588b7-a2b5-40f1-97cf-9544c24c3fc4, completed SUCCEEDED at 2026-09-29 10:03:40 UTC. Result: 2 reports, 0 failures, 1 reused report; summary email acknowledged by SES.
- Brownsville's document ID and PDF checksum are unchanged. Chicago resumed at correction, without repeating source capture/research, passed the final independent audit and PDF guard, and all nine pages passed visual review. [Completed Chicago report](https://docs.google.com/document/d/122G13Ecxp2om_UwsYMEODIn8h9-AuDs4NNJgzHW1PbQ/edit). PDF SHA256: 5ecba3c87d6fa53308a32b9320911d34ad1bde657fce78323b7f90a58e619276.
- The user-requested parallel-worker change is a separate stacked branch/PR, not included here.
### Kirkland: audit recovery live validation — September 29
- Commit 8b0c2d92ba7d9e8e8043937f229a643a5eb4c49e fixes the truncated-audit fallback. The old fallback had only a JSON serialization of opaque thinking and could not certify coverage. The new fallback independently reviews the same complete report/evidence with thinking disabled, a strict 6,000-token verdict allowance, counted full-context cost, and the unchanged 180,000-token stage / 900-second deadline. Missing coverage or truncated recovery output still fails closed.
- Eight regression cases cover opaque-response recovery using the saved draft, complete evidence including confidence-only receipts, incomplete coverage, truncated/missing verdicts, correction plus fresh final audit, and both reserved and recounted cost limits. Parent suite: 154 passed; restacked candidate suite: 169 passed. Targeted Ruff and diff checks passed.
- Deployed only Pipeline-school-performance-reports-prod, --exclusively, CDK [1/1]; CloudFormation UPDATE_COMPLETE at 11:21:31 UTC. Task revision 32, image digest sha256:5786dfa972a7cb7ae4982e73eca72089ac3b926dc8e71b605f39b1dac6db123d. The deployed candidate 94a2c767bf5eac6620470d3a6dccb01c8f5f3f95 includes the unchanged parallel-worker layer. No other stack, upstream refresh, model assignment, or recipient change.
- Fresh source/template preflight passed. Exact saved snapshot/research/draft versions and checksums were verified; current publication fingerprints matched the retained snapshots. Resume contract remained unchanged.
- Execution kirkland-audit-recovery-20260929-1122; run 2f69a5f4-02b1-40d2-97f6-23f6f908c872, resuming 474b28cb-5125-4c21-8175-ec01e0b0bbad.
- The deployed recovery path was actually exercised: the first audit consumed all 14,000 output tokens without a verdict; fresh full-evidence recovery returned complete coverage and a structured rejection, identifying a Programs comparison-label issue. Initial audit plus recovery consumed 66,404 tokens, within the unchanged cap. The built-in correction ran, followed by a fresh complete final audit.
- Report outcome remains partial failure, not successful Kirkland publication. The automatic correction introduced a separate material actual-versus-variance wording error. The final audit completed (38,307 tokens) and correctly rejected it; the existing one-correction limit stopped the school. No Kirkland Google Doc/PDF was created. Saved corrected draft and audit evidence are retained. The audit-recovery fix is live-verified; this separate narrative-correction issue still blocks Kirkland.
- La Jolla reused unchanged: same document ID 16Umzqz83io-lFdLzDROC4BjPXpRm1gKeKHd_-yr4b5w and PDF SHA256 eabb85d232ac6314a436e817a66da8d90908b19d890c94df5db238735e74aa80. Kirkland's original research/draft references also remained unchanged; neither was regenerated. Houston Heights and all other schools were excluded.
- One summary email was accepted by SES for ashwanth.r@trilogy.com (010001a0ececcb2e-444976e3-266e-4cc7-9b06-b54a18111791-000000). No inbox-delivery claim. The run is terminal and the pipeline lease is released. Paused pending authorization to address the new correction error; no automatic retry or next batch.
### Authorized Kirkland correction retry — successful September 29
This resolves the separate narrative blocker recorded in the preceding run. The user explicitly authorized fixing the correction and retrying final review from saved work, keeping La Jolla unchanged.
- Fix commit 3ed718e0b3fa90a8d25af70fe28c3544d624f29f. Numeric comparison entries retain exact E00 source paths, metric/breakdown identities and units, plus distinct actual/model/variance fields. They survive focused context and prose compaction. New correction rules prohibit substituting a variance or quarter-to-go for actuals and prefer omitting a secondary comparison to ambiguous shortening.
- Correction policy school-performance-correction-v2-comparison-roles is included in correction results and context signatures. A saved older correction is redone only if its prior audit completed and has identical evidence. Current-policy corrections resume at final review; approved reports return unchanged. No extra automatic correction cycle, altered financial data, or reduced audit gate.
- 162 parent / 177 combined tests passed, including numeric roles/nulls, compaction preservation, old/current/approved checkpoint handling, missing/incomplete/mismatched prior-audit handling, and continued rejection of an incorrect final correction. Targeted Ruff and diff checks passed.
- Native stack #2087 rebased bottom-up and safely pushed; combined candidate 42bb39f46ee889339b9d3df1a5d6d5ffa893e325. Only Pipeline-school-performance-reports-prod deployed with --exclusively; CDK [1/1], CloudFormation UPDATE_COMPLETE at 11:38:39 UTC, ECS revision 33, image digest sha256:567e2c8148ddbe3c0a4e060f5533c7ea22752a0c14b41f47d8a0eec8f5a4c069. No other stack, upstream refresh, school/model assignment, recipient or schedule change.
- Fresh full source/template preflight passed. Exact saved snapshot/research/draft/audit/correction versions and checksums were verified; source publication fingerprints were unchanged. The saved audit/correction evidence matched and the direct comparison ledger reconciled.
- Execution kirkland-comparison-retry-20260929-1139, run dfc18896-0389-41cc-bad7-5ae89015133c, resuming 2f69a5f4-02b1-40d2-97f6-23f6f908c872. Terminal SUCCEEDED, 11:39:24–11:44:05 UTC, ECS exit 0; report result success: 2 reports, 0 failures, 1 reused.
- Kirkland completed: original research/draft/snapshot references are unchanged. The deployed new policy rebuilt only the correction from the saved audit, then completed fresh final review with approved=true, coverage_complete=true, issues=[]. The corrected Programs card clearly distinguishes actuals, QuickBooks budget, and variance; the misleading model-variance-as-spending phrase is absent. No initial research/drafting or initial audit was repeated.
- [Completed Kirkland report](https://docs.google.com/document/d/1xWYkhsdsrt7m04VotO4M9RCGLq72zyznbITVW9D1piE/edit). Exact retained PDF SHA256 132254afd64c713d6cffd17050ee37a8a6703e3942bffabd9804c9dd6835e0f5. Runtime PDF checks passed; all 9 pages were rendered and visually checked, including comparison labels, tables, category placement, unavailable budgets and complete evidence/confidence layout.
- La Jolla reused unchanged in 0.5 seconds: same document 16Umzqz83io-lFdLzDROC4BjPXpRm1gKeKHd_-yr4b5w, same PDF SHA256 eabb85d232ac6314a436e817a66da8d90908b19d890c94df5db238735e74aa80. No other school was invoked.
- One consolidated email accepted by SES for ashwanth.r@trilogy.com: 010001a0ecfa36ee-1f658f77-7386-45d9-9e4b-019a85e90cdc-000000. The batch lease is released. Paused before the next batch. Both PRs remain draft; nothing merged.
## Scope and remaining investigation
The 18 changed files are limited to the report runner and the owning facilities procedure, its deployment script, tests, and documentation. No schedule, recipient, shared infrastructure, or model-assignment changes are included. No confidential report/evidence artifacts are committed.
East Bay remains paused: 19 students, but no assigned unit-economics model. Bethesda and Boston Suburbs retain their separate model/enrollment prerequisite blockers. No next batch is started automatically.
## Linear
[SURTR-1535 — Fix school performance report categories and approved missing budgets](https://linear.app/builder-team/issue/SURTR-1535/fix-school-performance-report-categories-and-approved-missing-budgets)
## Implementation Effort
Estimated 20–30 engineer-hours (about 2.5–4 days) for an average engineer to investigate and hand-code the SQL, validation, native Google Docs layout/retry changes, safe migrations, regression tests, cutoff repair, and pre-merge live validation without AI assistance.
## Five-school audit/correction recovery — September 29
The user authorized repairing the five failures from nine-worker run b7adcc20-8aa3-4dbd-b9b7-fe70021ec81c and resuming saved work, retaining Boston, Santa Monica, Scottsdale and The Woodlands unchanged.
Parent fix 3139d150db6224014728b3602e5e367660897ace, combined candidate 0d36afbbfe37f4a0c56df21f7db628d56a2a129c:
- Chantilly: request at most one audit verdict. An unexpected multi-verdict response is not cherry-picked; the bounded full-evidence recovery must independently produce one complete verdict.
- Charlotte: the one incomplete-audit retry now has two disjoint review scopes covering every claim and confidence paragraph. Both retain the full report and every receipt, must attest complete coverage, and share the original 180,000-token / 900-second audit ceiling. No incomplete verdict can reach publication or be mistaken for an issue to edit away.
- San Francisco: direct numeric comparison context now explicitly separates booked and management actuals, including booked-minus-budget/model calculations. Published financial variances are labeled as management-based. Unavailable benchmarks remain null.
- Dorado: a policy-upgrade resume re-audits the latest saved correction, rather than editing the old draft using only its first audit's issues. It can therefore address a headline defect discovered at final review without reverting earlier fixes. The same single-correction limit and fresh final audit remain mandatory.
- Palo Alto: correction requests carry exact field-length ceilings and conservative targets. The second bounded text-repair call gets its actual rejected text and measured length. Secondary examples may be omitted to fit, but retained amounts, comparison basis/direction and necessary qualifications cannot change. Full evidence review and PDF checks remain mandatory.
172 parent / 211 combined tests pass. Targeted Ruff and git diff --check pass. Regressions cover ambiguity recovery, complete two-part coverage, all receipts retained, combined issues, shared budget exhaustion, exact booked variance/null handling, field limits, rejected-text feedback, new-policy fresh review, and preservation of already-corrected fields. The native stack was rebased bottom-up and pushed with explicit force-with-lease checks; both PRs remain draft and nothing is merged.
Read-only resume preflight passed for all nine exact retained snapshot versions/checksums, their generation checkpoints and evidence references, the 48-hour source freshness constraints, six-table/enrollment/model contracts, deterministic rendering and the unchanged native template. The invocation contract is unchanged apart from resume_run_id. Five unfinished schools reuse research and drafts; four completed schools reuse existing document IDs and PDFs. No upstream refresh, model assignment, school mapping, recipient or schedule change.
Live validation finished with remaining blockers — not ready to merge. Only Pipeline-school-performance-reports-prod was deployed, using --exclusively; CDK reported [1/1] and its CloudFormation events reached UPDATE_COMPLETE at 12:33:05 UTC. Task definition 35 ran the combined candidate image. No other pipeline, registry or shared stack was deployed.
The authorized resume 967c955b-a8ef-41af-86c6-1b2f6b573fea ran from 12:33:56 to 12:41:50 UTC. Step Functions succeeded, but the application correctly reported partial_failure: 8 published reports, 1 failure, 4 reused. Boston, Santa Monica, Scottsdale and The Woodlands retained their exact prior document/PDF references. Chantilly, Dorado, Palo Alto and San Francisco completed from saved research/drafts; Charlotte retained its corrected draft. Checks confirmed unchanged research/draft and source snapshot references; no full generation was repeated. The usual summary email was sent to the unchanged recipient with eight published reports. The pipeline lease was released.
All 36 pages of the four newly published PDFs were rendered and visually inspected. Layout and page-count checks passed. San Francisco's corrected management-profit/variance pairing, Dorado's operating-program variance and janitorial account attribution, and Palo Alto's fitted correction are present. However, automated approval is not sufficient: manual inspection found the narrative defects below, so Chantilly and Dorado are published, not fully QA-approved. The email had already been sent before this post-run inspection; those documents have not been silently edited or regenerated.
### Remaining blockers found by live validation
- Charlotte — review budget coordination: its final audit exhausted the 14,000-output allowance, then the full-evidence recovery returned incomplete coverage. The first scoped retry completed and approved its own scope, but the second was blocked before invocation by the shared 180,000-token gate, including its full recovery reserve. The final audit had actually spent 117,507 tokens by that point; this is a reservation/planning failure, not evidence that the corrected narrative is wrong. The runner did not publish a half-reviewed report. The scoped fallback was live-exercised but did not finish; its live validation remains failed.
- Chantilly — standalone headline arithmetic missed by automated audit: facilities total spend $35.1K is labeled as the over-model amount instead of the actual ~$5.8K variance. Both a depreciation headline and finding title label $10.8K actual spend as the amount over both benchmarks; the overruns are ~$10.1K against QuickBooks and ~$9.6K against model. Tables and explanatory paragraphs retain the correct values. These three narrative fields need correction and a stronger comparison check.
- Dorado — explanation/attribution missed by automated audit: finding 7 and the Timeback transaction-reference paragraph say the allocation raises EBITDA, while the report's own basis shows booked EBITDA -$393,604.62 becoming management EBITDA -$488,332.07 after the $94,727.45 allocation. Its second opening signal also attributes the EBITDA result partly to depreciation, despite depreciation being added back. The third signal calls all six rent rows bills, whereas its detailed evidence distinguishes rent bills from two purchases. The displayed accounting totals are correct; these causal/directional and transaction-type descriptions need correction.
Work is paused after this batch, consistent with the requested investigate-and-confirm workflow. No follow-on execution or completed-document modification has been made. Recommended next step: repair the bounded audit scheduling and these saved-narrative QA gaps, then resume/review only the affected saved work after user confirmation. Other completed September 29 reports remain reusable and unchanged; combining reports from separate historical runs into one delivery is a separate operation, not part of this resume.
Evidence: /tmp/surtr-review-resume-OIfR2u (deployment events, terminal result/email checkpoints, immutable-source assertions, audit receipts, exact-version PDFs and rendered pages).
Incremental implementation effort for this recovery work: approximately 4–6 engineer-hours without AI assistance (in addition to the earlier work estimated below).
## Three-school saved-work repair — September 29
The user authorized fixing Charlotte's remaining audit-budget failure and the post-publication QA findings in Chantilly and Dorado, leaving all other reports untouched.
Parent 8bb89187, combined candidate 087a7534. Native stack #2087 was rebased bottom-up and pushed with leases. Both PRs remain draft; nothing is merged.
- The incomplete-review fallback now counts and reserves both disjoint scopes before starting either. Each gets full evidence and a 6,000-token forced verdict with thinking disabled, without recursively reserving another complete recovery per scope. Complete coverage of both remains mandatory; the original 180,000-token / 900-second safety limits remain unchanged. Truncation, ambiguity, incomplete coverage and unaffordable plans fail closed.
- Explicit repair_reports maps selected completed school IDs to retained human-QA findings. It requires matching published insights and a compatible saved approval with receipts. A new repair seed starts from the latest published text, receives one scoped correction and a full fresh final audit. Previous documents/PDFs/checkpoints remain intact, with supersedes provenance for replacements. No research/draft generation is repeated, and unselected completed reports are imported unchanged.
- Supplementary deterministic checks reject unambiguous facilities/depreciation variance labels that use spending or disagree with exact source arithmetic, and explicit EBITDA/addback or allocation-direction contradictions. They respect displayed rounding and explicit EBITDA-neutral/addback qualifications; they supplement, not replace, full evidence review.
198 parent / 239 combined tests pass. Targeted Ruff and git diff --check pass. Regression coverage includes Charlotte's observed budget scale, all-evidence scope coverage, fail-closed incomplete/truncated output, no partial-review approval, explicit repair eligibility, published-text seeding, mandatory final review, preservation of other reports/prior documents, same-run idempotency and the observed narrative defects/valid qualifications.
Read-only preflight passed against the exact retained snapshots and checkpoints from 967c955b-a8ef-41af-86c6-1b2f6b573fea, the 48-hour freshness and six-table contracts, current native template, unchanged recipients and role. Six reports will be reused, Chantilly and Dorado repaired, and Charlotte resumed at final review. Seven reports from earlier batches remain outside this run and unchanged. No warehouse writes or upstream refreshes.
Live validation finished with two narrative blockers — not ready to merge. Only Pipeline-school-performance-reports-prod was deployed with --exclusively; CDK showed [1/1], and CloudFormation reached UPDATE_COMPLETE at 12:58:07 UTC. ECS task definition 36 ran image digest sha256:354bf51b1de2b1d12430c41caf84125b5acf0cbbf95e8fc3962912e4b10e6c2c. No other pipeline/shared/registry stack or upstream refresh was deployed.
Resume 6b7d7aa3-5249-43b1-9a25-1c62b8f20328 ran from 12:58:34 to 13:04:10 UTC. Step Functions succeeded; the application result is correctly partial_failure: 8 reports, 1 failed repair, 6 reused. The completion email was sent to the unchanged recipient, and the lease was released. Boston, Palo Alto, San Francisco, Santa Monica, Scottsdale and The Woodlands retain identical prior document/PDF references. All original research/draft/snapshot references remained unchanged; repaired reports retain identical receipts. Seven earlier-batch reports were not included or modified.
- Dorado: the targeted repair completed, passed full evidence review (50,968 final-audit tokens), and all nine PDF pages were manually inspected. Five narrative fields changed only within the four authorized claim scopes, correcting EBITDA direction, depreciation/addback causality, and six rows versus six bills. [Corrected replacement document](https://docs.google.com/document/d/1DNkPYEJsQDqlxs7CsBFOVnA8zxl5zCuK1WE1m2CgxRM/edit). The prior published document/PDF remain untouched.
- Charlotte: final review completed on the first call (41,882 tokens), and publication succeeded with its saved corrected narrative byte-for-byte unchanged. The new scoped fallback was not needed in this live run; its budget behavior is regression-tested against the observed 80,740-token starting balance and complete evidence. All nine PDF pages were inspected. Manual QA then found a separate, pre-existing narrative error: several claims invent a 13-versus-25 enrollment target as the revenue-gap cause. E00's model revenue is $130,000 = 13 actual students × $50,000 annual tuition × 0.2; 25 is capex_reference_student_count in the model assumptions, not a revenue enrollment target. The financial tables are correct. Charlotte is published and emailed, but not fully QA-approved.
- Chantilly: the saved correction now contains the correct $5.8K facilities-over-model headline and separate $10.1K QB / $9.6K model depreciation overruns in its headline/finding. Final review (50,807 tokens) correctly blocked the replacement because the executive paragraph still calls $10.8K actual depreciation 'over both budget and model'. That paragraph was outside the supplied repair findings/allowed correction claims. No replacement document was published or included in this email; the corrected work and original published version are retained.
Next step, awaiting user direction after this batch: extend the explicit saved-work repair path to accept the latest unfinished corrected narrative with new QA findings (without replaying research or reverting the headline fixes), repair Chantilly's executive paragraph, and correct Charlotte's enrollment-benchmark/causal claims using the actual revenue-model basis. Keep every other report untouched. No automatic extra correction cycle or follow-on execution has been started.
Evidence directory: /tmp/surtr-three-repair-OzHvpw (preflight, deployment events, terminal result/email, immutable-reference assertions, audit receipts, narrative diffs, exact-version PDFs and all 18 rendered/reviewed pages). Deployment alone is not validation, and automated publication is not human QA approval.
Incremental implementation effort: approximately 3–5 engineer-hours without AI, beyond the earlier estimates.
## Saved-work narrative repairs — September 29, 2026, final validation
- Latest correction can now be an explicitly authorized repair source, not just a completed approval. Source stage and immutable reference are retained. An approved publication in progress cannot silently roll back to an older correction. Ordinary retry behavior is unchanged: no extra correction loop and no research/drafting regeneration.
- Correct the remaining Chantilly executive depreciation variance, preserving the three earlier headline/title fixes. The new source-backed check also caught its related 9 vs 25 enrollment inference before dispatch.
- Distinguish actual/model student counts from the CapEx reference denominator and model-name suffix. Remove Charlotte's unsupported 25-student enrollment target and enrollment/pricing causal claims throughout its affected narrative; preserve all reported amounts and require reconciliation instead of inventing a cause.
- Add narrow regression guards for directly labeled body-text expense variances (including “over both budget and model”) and misuse of CapEx reference counts as modeled enrollment. Full-source review still remains mandatory.
- Tests: 213 parent / 254 combined pass; targeted Ruff and git diff --check pass.
- Candidate deployed before merge: parent 6fa8d66f, combined 67cd5297, task definition 37, image digest sha256:68db812865cd85046ad4c7842b006c9e05ea8038f83574b0204ebe8dd385ccd7. Only Pipeline-school-performance-reports-prod deployed with --exclusively; CDK [1/1] and CloudFormation UPDATE_COMPLETE at 13:17:17 UTC verified. No upstream/warehouse change or other pipeline deployment.
- Live validation: arn:aws:states:us-east-1:479395885256:execution:pipeline-school-performance-reports-prod:repair-two-saved-reports-20260929-131743; Surtr run 667b1467-231f-4fae-96e6-2dd58706bd11. 9/9 reports complete, 0 failures, 7 exact reused reports, consolidated email sent to the unchanged configured recipient. Only Chantilly and Charlotte received bounded corrections and fresh full-source audits, using byte-identical retained evidence and original snapshot/research/draft references. Prior documents remain retained, new reports record superseded publication/source provenance. Dorado and every other report untouched; lease released.
- Both new PDFs have nine pages; full visual/narrative QA completed. No changes beyond the selected narrative claims; earlier Chantilly corrections preserved. Paused after this batch; both PRs remain draft and unmerged.
### Requested consolidated delivery: all 16 schools
After both repaired reports passed audit and manual QA, the user requested inclusion of the seven earlier completed reports (Boca Raton, Brownsville, Chicago, Fort Worth, High Austin, Kirkland and La Jolla). A separate delivery-only consolidation used the unchanged existing delivery.send helper, pipeline lease, immutable S3 manifest and idempotent/uncertain-send protection. This was not another generation run, a change to the fixed-school-list resume contract, or another deployment.
Read-only preflight verified complete source checkpoints, September 29 cutoffs, exact-version snapshot/narrative/PDF checksums, nine-page PDF contracts, native document completion and existing access for the unchanged recipient. 16 distinct report links sent, 0 generated in the consolidation, delivery ID consolidated-16-20260929-f16193e2b34c8ce3. The 14 other reports stayed unchanged; only Chantilly and Charlotte were repaired in the preceding run. Delivery manifest retained at s3://surtr-school-performance-evidence-prod-479395885256/runs/consolidated-16-20260929-f16193e2b34c8ce3/6db7048a-299b-41a1-97e5-20a0e50e70f5/delivery/consolidation-manifest/d58e1cf7309a47c68e187a0931d07f313b5bc5e12f602d84a98cb541ecabc44e. No source run/checkpoint/document was edited for consolidation.
</details>