Vol. I  ·  No. 278 Established 2026  ·  AI-Generated Daily Free to Read  ·  Free to Print

The Trilogy Times

All the news that's fit to generate  —  AI • Business • Innovation
MONDAY, OCTOBER 05, 2026 Powered by the TrueFoundry AI Gateway  ·  Published on Klair Trilogy International © 2026
🖶 Download PDF 🖿 Print 📰 All Editions
Today's Edition

OPENAI PUNTS ON FOURTH DOWN: SAFETY SQUAD BENCHED AS AGENTS STORM THE FIELD

With half its alignment roster cut and AI agents already swiping the family card, the league's most-watched franchise is playing without a secondary.

SAN FRANCISCO — Folks, we are HERE, and what a WILD sequence of plays we've just witnessed on the AI gridiron. OpenAI — the reigning champs, the team everybody builds their playbook around — just cut nearly HALF its safety and alignment unit, and they did it RIGHT as a string of agent-related hacks was lighting up the scoreboard against them. You do not bench your secondary during a four-game losing streak to the defense, people! That's Coaching 101! What did those researchers know? The front office isn't saying, and that silence is getting LOUDER by the hour.

Meanwhile, on the OTHER side of the field, the offense is having the game of its LIFE. Meta's Muse just POLE-VAULTED to the top of the App Store, Instinct hauled in a cool BILLION dollars in fresh funding, and now Snaplii is rolling out a dedicated wallet — Snaplii Cash — so your AI agent can shop with a budget instead of your actual credit card. Think of it as putting the rookie quarterback on a pitch count. Smart. Necessary. Because folks, when the agents start BUYING THINGS ON THEIR OWN, you'd better believe somebody needs guardrails on that spending, and apparently it's not going to be the guys who just got laid off from the safety desk.

And just when you thought the scoreboard couldn't get any more chaotic — Zuckerberg, Musk, Bezos, AND Huang all signed onto the new White House superintelligence accord. Four rival owners, same conference table, same signature line. That NEVER happens unless everybody smells the same storm coming.

Off the field, the economy's got its own injury report. Anthony Scaramucci told Yahoo Finance America's confidence is cracking at the foundation — call it a sprained ankle on a team that already can't find the end zone.

So here's your halftime read, Trilogy faithful: the offense is unstoppable, the defense just got gutted, and nobody's sure who's calling the next play. Buckle up — second half's gonna be a SHOOTOUT.

↗ OpenAI Fired the People It Needed Most  ·  Why Americans are losing hope in the economy: Anthony Scaram  ·  As Muse Tops the App Store and Instinct Raises $1 Billion, S

GOOGLE PULLS PLUG ON BUG BOUNTY AS ROBOT REPORTS FLOOD THE GATE

Search giant says machine-written vulnerability claims swamped its open source reward program, even as Washington rolls out a new task force to polish AI's name.

MOUNTAIN VIEW, CALIF. — Google shut the door on its open source bug bounty program Friday. The company says a flood of AI-generated submissions overwhelmed the reviewers. Too much chaff, not enough wheat.

The tech giant calls it a "significant rise" in low-quality reports, the kind a chatbot spits out when asked to find a flaw and doesn't much care if it finds one. Engineers must now sift noise from signal by hand. The freeze hits a program built on trust between hackers and the house.

Bounty hunters used to dig through code by hand, hunting real holes for real money. Now a man can point a language model at a repository, click a button, and mail in a hundred reports before lunch. Quantity drowns quality. Google won't say when the gates reopen.

Meanwhile down in Washington, the White House rolls out a different kind of AI fix. Trump unveils a new "Super Intelligence Force" this week, a task force built to answer growing fear about machines running loose. Critics call it a label job, not a law.

On the Equity podcast this week, the question gets asked plain: can a task force and a handshake pact fix AI's image problem? No enforcement teeth in the deal. Just promises, signed by companies who write the rules they're promising to follow.

This correspondent sees the pattern plain as day. Bad actors, human and machine both, game every open system fast as it opens. Bug bounties, chatbots, safety pacts — the con men move quicker than the cops.

Out on the roads, the mobility desk reports cities moving to rein in robotaxis, same instinct at work. Regulators chase the technology, always a lap behind. Autonomous cars multiply on city streets while rulebooks sit half-written on some bureaucrat's desk.

And overseas, the DeepSeek story keeps printing fresh pages. The Chinese firm says it trains top-shelf models cheap, skipping the priciest chips America guards close. Washington frets over the chip embargo while Beijing finds the side door.

Add it up and a newsman sees one headline under all the headlines: the guardrails lag the gold rush. Bug bounty programs choke on fake bug reports. Governments chase AI with press releases instead of statutes. China trains cheap and fast while export controls do their best impression of a screen door.

Somewhere in Austin, Joe Liemandt's shops keep building software the old-fashioned way — one line of code, one human check at a time. Crossover's engineers still read what they ship. In this town, that counts as news.

↗ Google froze its open source bug bounty program due to a ‘si  ·  Can ‘super intelligence’ and a non-binding safety pact solve  ·  TechCrunch Mobility: Reining in robotaxis
Haiku of the Day  ·  GPT-5.6 LunaBots hum in the trees
The grand systems miss the beat
Yet nothing gets done
The New Yorker Style  ·  Art Desk
The New Yorker Style  ·  Art Desk
The Far Side Style  ·  Art Desk
The Far Side Style  ·  Art Desk
News in Brief
When AI Agents Lie (Accidentally): The Industry's New Trust Problem
SAN FRANCISCO — Okay, I need everyone to sit with this for a second: an AI agent can tell you, with total confidence, that it finished the task.
The Fairness Deficit: A Tripartite Audit of Algorithmic Governance Across Health, Policing, and Pedagogy
GENEVA — It could be argued — and, this week, three independent bodies of scholarship have argued it with varying degrees of empirical rigor — that the present moment in artificial intelligence governance is best characterized not by technological scarcity but by normative lag (a distinction that bears repeating, since popular discourse tends to conflate the two). Consider, first, the thesis advanced by the World Health Organization's new report, which calls, not unreasonably, for strengthened ethics oversight of AI-related health research — a recommendation that, preliminary evidence suggests, arrives roughly a decade after the horse has not merely left the barn but founded a subsidiary. The antithesis emerges from the domain of criminal justice, where the Human Rights Research Center's treatment of predictive policing documents what one might term the erosion of procedural fairness (itself a slippery construct, admittedly, contingent on which Rawlsian commitments one imports into the analysis).
WHEREAS AN UNSEALED OPINION HEREINAFTER CLARIFIES THE RATIONALE UNDERGIRDING SAID LANDMARK AI FAIR-USE DETERMINATION, AND WHEREAS KOREAN REFRIGERATION UNITS DID CONCURRENTLY SUFFER FIRMWARE-INDUCED SPOILAGE OF PERISHABLE COMESTIBLES
WILMINGTON, DEL.
The Dole, Rebranded: Notes on a Panic in Search of a Policy
AUSTIN, TEXAS — There is a peculiar comfort in watching an old argument put on new clothes, and so I confess a certain fondness for the current spectacle of serious people, in serious publications, debating whether the republic ought to mail everyone a check to compensate for the sin of having been born into the age of the large language model.
Unpopular Opinion: Your Next Coworker Might Be an AI Actress (And That's a FEATURE, Not a Bug) 🚀
AUSTIN, TEXAS — I'll be honest, I almost didn't write this one. Because when the algorithm serves you a headline about data scientists hunting remote jobs right next to a story about an AI actress named Tilly Norwood getting death threats, you have two choices. You can scroll past. Or you can recognize the single thread connecting them and build your entire week's content calendar around it.
A Trilogy Company
Crossover
The world's top 1% remote talent, rigorously tested and ready to ship.
A Trilogy Company
Alpha School
AI-powered learning. Two hours a day. Academic results that defy belief.
A Trilogy Company
Skyvera
Next-generation telecom software — built for the networks of tomorrow.
A Trilogy Company
Klair
Your AI-first operating system. Every workflow. Every team. One platform.
A Trilogy Company
Trilogy
We buy good software businesses and turn them into great ones — with AI.
The Builder Desk  —  AI Builder Team
📅 Week in ReviewProduction Release

The Builder Desk

185 pull requests merged across the org this week

#163 AI-962: Evaluate task prompts with two independent replay trials (@ashwanth1109, Shipyard)

#2128 feat(aerie-a8): read dbt_invocation_id as the dbt build marker (SURTR-1595) (@kevalshahtrilogy, Surtr)

#1655 feat(a8): dbt copy shadow checks accept both build marker forms, pg_class_oid and dbt_invocation_id (AERIE-2700) (@kevalshahtrilogy, Aerie)

#1653 SIS Enrollment: complete start-year cohorts in the enrollment mart (@vvp-trilogy, Aerie)

#1638 feat(dbt): stamp dbt_invocation_id on the admissions pipeline and SIS enrollment marts (AERIE-2682) (@kevalshahtrilogy, Aerie)

#1652 fix(a8): keyed Gateway readers say when a program key is absent; gateway stays unreleased in production (AERIE-2690) (@kevalshahtrilogy, Aerie)

#1651 fix(a8): a failed legacy read makes the per-program shadow dry-run not clean (AERIE-2689) (@kevalshahtrilogy, Aerie)

#2120 fix(education): take school ontology locks in the producer's order across sibling marts (SURTR-1590) (@kevalshahtrilogy, Surtr)

#1650 fix(a8): a derived shadow output with programs left out is degraded, not clean (AERIE-2688) (@kevalshahtrilogy, Aerie)

#1649 Forecast V3: migrate consumers to canonical dbt fields (@vvp-trilogy, Aerie)

#1646 Fix Session 1 lock when historical observation is missing (@vvp-trilogy, Aerie)

#1647 fix(dbt): make SIS lifecycle data quality non-blocking (@vvp-trilogy, Aerie)

#2127 fix(timeback): adapt resources pages to gateway limits (@caina-barbosa, Surtr)

#1643 Align next-year enrollment forecast formula (@vvp-trilogy, Aerie)

#1585 feat(a8): G6 gate SIS_ENROLLMENT_READ — SIS rollup input + members via Surtr Gateway, shared source_run_id, stale_copy rule (AERIE-2624) (@kevalshahtrilogy, Aerie)

#2125 feat(aerie-a8): Forecast V2 copy, forecast + grade operands in one transaction (SURTR-1594) (@kevalshahtrilogy, Surtr)

#1580 feat(a8): G2 gate ADMISSIONS_PIPELINE_READ — admissions pipeline detail + tenant crosswalk via Surtr Gateway, stale_copy rule (AERIE-2620) (@kevalshahtrilogy, Aerie)

#1587 feat(a8): G3 marketing seam + Gateway readers D1–D4 + gate ADMISSIONS_MARKETING_READ, PII-redacted shadow compare (A8 U22) (@kevalshahtrilogy, Aerie)

#1586 feat(a8): G2 per-program shadow compare + dry-run for ADMISSIONS_PROGRAM_DETAIL_READ (A8 U21) (@kevalshahtrilogy, Aerie)

#2123 Enable on-demand QuickBooks Core validation (SURTR-1593) (@ashwanth1109, Surtr)

#1578 feat(a8): G2 per-program Gateway readers Q6, Q7, Q8, Q2, Q3 + app conversion (A8 U17) (@kevalshahtrilogy, Aerie)

#1576 feat(a8): G2 per-program source seam + gate ADMISSIONS_PROGRAM_DETAIL_READ + Gateway readers Q5, Q4 (A8 U14) (@kevalshahtrilogy, Aerie)

#1637 fix(funnel): relay community funnel sync refusals as userErrors the worker can show (AERIE-2680) (@kevalshahtrilogy, Aerie)

#1636 fix(gateway): fail closed when a mapper drops a Gateway row (AERIE-2679) (@kevalshahtrilogy, Aerie)

#2121 fix(collections): skip 42DS in CollectIQ v2 (@sanketghia, Surtr)

#1633 fix(rhodes-shadow): exact-cent tuition compare; count roster failures as mismatched (AERIE-2676) (@kevalshahtrilogy, Aerie)

#1634 fix(a8): refuse a partial community funnel Gateway snapshot by a proportional check (AERIE-2677) (@kevalshahtrilogy, Aerie)

#1632 fix(xo-contractor): read a NULL trailing-52-week sum as 0 on both transports (AERIE-2675) (@kevalshahtrilogy, Aerie)

#1627 Add dated backup site need periods and DSS reporting (@YibinLongTrilogy, Aerie)

#1629 Add Preview datasets for Praxis product verification (@caina-barbosa, Aerie)

#2118 docs(ai-spend): spec 14 - gpt-audio-2025-08-28 pricing, applied (@kevalshahtrilogy, Surtr)

#1626 docs(documents): clarify controlled managed relink (AERIE-2136) (@marcusdAIy, Aerie)

#162 Release: Shipyard 0.6.11 (@ashwanth1109, Shipyard)

#1625 docs(admissions): align portfolio status with Milestone 10 postOpen (@marcusdAIy, Aerie)

#161 AI-960: Remove Pi permissions info box and acknowledgement checkbox from task creation (@ashwanth1109, Shipyard)

#160 AI-959: Allow switching a project's Git branch from the Projects list (@ashwanth1109, Shipyard)

#159 AI-954: Show task creation time and time taken on task detail (@ashwanth1109, Shipyard)

#2117 fix(education): reconcile Alpha OKC QuickBooks Class replacement (@caina-barbosa, Surtr)

#158 AI-953: Support queueing and steering messages in Pi conversations (@ashwanth1109, Shipyard)

#1624 docs(directory): explain typed relationship endpoint IDs (@marcusdAIy, Aerie)

#1623 fix(admissions): return allowed portfolio statuses for invalid capacity filters (@marcusdAIy, Aerie)

#1622 fix(admissions): explain legacy appointment-driven Shadowing (@marcusdAIy, Aerie)

#1621 fix(insights): clarify overdue UTC default and portfolio-health status prose (@marcusdAIy, Aerie)

#157 AI-952: Fail fast when the bundled Pi runtime cannot see images (@ashwanth1109, Shipyard)

#156 AI-951: Add Prev/Next pagination to project repository commit history (@ashwanth1109, Shipyard)

#2116 fix(education): accept GuidePlatform audio_recordings/audio_transcripts additive schema drift (SURTR-1582) (@kevalshahtrilogy, Surtr)

#1571 feat(a8): G3 gate COMMUNITY_DEPOSITS_READ + keep every community deposit (AERIE-2615, AERIE-2353) (@kevalshahtrilogy, Aerie)

#1573 feat(a8): G4 gate EXPENSES_READ — expense transactions + vendor classifications via Surtr Gateway (AERIE-2617) (@kevalshahtrilogy, Aerie)

#1572 feat(a8): G2 gate ADMISSIONS_COMMUNITY_FUNNEL_READ — community funnel via Surtr Gateway, PII-redacted shadow compare (AERIE-2616) (@kevalshahtrilogy, Aerie)

#1569 feat(a8): G1 gate ADMISSIONS_REFERENCE_READ — programs + Program directory via Surtr Gateway (AERIE-2614) (@kevalshahtrilogy, Aerie)

#1566 AERIE-2612: A8 read-gate kit (read mode, Gateway reader with type parity, keyed shadow compare, purge guard) (@kevalshahtrilogy, Aerie)

#2115 docs(ai-spend): spec 13 - gpt-6.1-sol pricing insert and reprice, applied (@kevalshahtrilogy, Surtr)

#1527 chore(sync): retire gated-off F3–F7 financial-worker tasks (AERIE-2545) (@kevalshahtrilogy, Aerie)

#1614 chore(praxis): follow master during hosted testing (AI-874) (@caina-barbosa, Aerie)

#1499 fix(analytics): repoint XO contractor sync at the real table + add Gateway parity (F1/F2) (@kevalshahtrilogy, Aerie)

#155 Release: Shipyard 0.6.10 (@ashwanth1109, Shipyard)

#1529 feat(sync): A7 MATTERPORT_DISCOVERY_MODE gate + shadow compare (AERIE-2244) (@kevalshahtrilogy, Aerie)

#1563 feat(rhodes-merger): A1 shadow mode against Surtr's published merge (AERIE-2610) (@kevalshahtrilogy, Aerie)

#1564 feat(sync): A6 REBL3 shadow mode, REBL3_READ=legacy|shadow|gateway (AERIE-2239) (@kevalshahtrilogy, Aerie)

#154 AI-949: Play a sound and show a macOS notification when the Pomodoro timer changes phase (@ashwanth1109, Shipyard)

#1530 feat(camps): A5 CAMP_SOURCE_READ publish gate for refreshCampData (AERIE-2548) (@kevalshahtrilogy, Aerie)

#1478 feat(camps): A5 Gateway-backed parity check for summer camp directories (@kevalshahtrilogy, Aerie)

#2113 Honor quarter-close window in CollectIQ sync (@sanketghia, Surtr)

#1531 AERIE-2547: Finalsite tenant directory Gateway reader + shadow comparator (@kevalshahtrilogy, Aerie)

#1617 Forecast V2: export the primary table as CSV (@vvp-trilogy, Aerie)

#1533 feat(school-calendar): read Surtr Gateway behind SCHOOL_CALENDAR_READ with shadow compare (AERIE-2544) (@kevalshahtrilogy, Aerie)

#1528 feat(sync): A2 Schools Data Sheet via Surtr Gateway, with shadow compare and read gate (AERIE-2543) (@kevalshahtrilogy, Aerie)

#142 Gate auto-approve only on the checks the base branch requires (AI-948) (@kevalshahtrilogy, mercy)

#151 AI-942: Evals execute real node replays from immutable entry snapshots (@ashwanth1109, Shipyard)

#1615 test(financials): pin the clock in QTD reports test (AERIE-2662) (@kevalshahtrilogy, Aerie)

#152 AI-946: Keep Pi image attachments visible and usable in chat (@ashwanth1109, Shipyard)

#153 AI-947: Let the companion look up Shipyard data outside the current view (@ashwanth1109, Shipyard)

#2112 fix(q118): remove unsupported Redshift temp schema (@sanketghia, Surtr)

#2110 feat(q118): route BalanceSheet landing separately (@sanketghia, Surtr)

#2109 fix(q118): make Redshift DDL application compatible (@sanketghia, Surtr)

#3841 fix(budget-load): prepare Q4 refresh and sandbox backups (@sanketghia, Klair)

#2107 fix(q118): accept QuickBooks no-data metadata rows (@sanketghia, Surtr)

#2073 feat(q118): add deferred-revenue BalanceSheet ingestion (@sanketghia, Surtr)

#1613 fix(praxis): use corrected private source authentication (AI-874) (@caina-barbosa, Aerie)

#1602 feat(praxis): connect Aerie PRs to Praxis verification (AI-924) (@caina-barbosa, Aerie)

#1611 Refresh DSS contract fingerprint (@caina-barbosa, Aerie)

#1610 Fix Forecast V2 public response projection (@caina-barbosa, Aerie)

#1608 Present End-of-Year additions and deductions (@vvp-trilogy, Aerie)

#2106 fix(q75): verify established service grants and current catalog (@marcusdAIy, Surtr)

#1607 Add DSS discovery to Aerie Tools and API Docs (@YibinLongTrilogy, Aerie)

#1605 Net historical departures in End-of-Year forecast (@vvp-trilogy, Aerie)

#211 Settle stalled workflow nodes by progress, not heartbeat (SINDRI-516) (@marcusdAIy, Sindri)

#1603 Capacity: bind the delegated Sindri identity and validate run ownership (AERIE-2582) (@marcusdAIy, Aerie)

#2105 feat(mart-aerie-dbt-publication-refresh): enable the 10-minute schedule (SURTR-1562) (@kevalshahtrilogy, Surtr)

#1600 Forecast V2: add next-year forecast and calculation details (@vvp-trilogy, Aerie)

#2098 feat(aerie-a8): G6 SIS enrollment copy, rollup input + members in one transaction (SURTR-1549) (@kevalshahtrilogy, Surtr)

#2096 feat(aerie-a8): mart-aerie-dbt-publication-refresh runner + admissions pipeline detail copy + tenant crosswalk (SURTR-1547) (@kevalshahtrilogy, Surtr)

#2101 fix(aerie-a8): widen applied parity marts' text columns to the query's width (SURTR-1552) (@kevalshahtrilogy, Surtr)

#2091 fix(mart-aerie-xo-contractor-refresh): deterministic latest-week tie-break (SURTR-1542) (@kevalshahtrilogy, Surtr)

#2088 feat(aerie-a8): mart-aerie-expenses-refresh runner + G4 expense parity marts (SURTR-1539) (@kevalshahtrilogy, Surtr)

#2099 feat(aerie-a8): G3 marketing parity marts, D3 shadow-day events + D4 weekly deposits (SURTR-1550) (@kevalshahtrilogy, Surtr)

#2095 feat(aerie-a8): G3 marketing parity marts, D1 event aggregates + D2 event contacts (SURTR-1546) (@kevalshahtrilogy, Surtr)

#2092 feat(aerie-a8): per-program enrollment parity marts, Q6 cohort + Q7 pipeline deposit + Q8 transfer (SURTR-1543) (@kevalshahtrilogy, Surtr)

#2090 feat(aerie-a8): per-program parity marts part 1, Q5 pipeline students + Q4 community metrics (SURTR-1541) (@kevalshahtrilogy, Surtr)

#2089 feat(aerie-a8): community conversion + deposit parity marts (SURTR-1540) (@kevalshahtrilogy, Surtr)

#2103 chore(pipelines): assign HC forecast refresh owner (@sanketghia, Surtr)

#24 fix(deployment): allow retired task definition cleanup (@benji-bizzell, Redshift-DSS)

#23 fix(registration): prepare the provider front for submission (@benji-bizzell, Redshift-DSS)

#150 Release: Shipyard 0.6.9 (@ashwanth1109, Shipyard)

#149 AI-939: Prevent horizontal scroll from attached images (@ashwanth1109, Shipyard)

#22 fix(deployment): match immutable GitHub OIDC identity (@benji-bizzell, Redshift-DSS)

#21 feat(deployment): isolate DSS deployment permissions (@benji-bizzell, Redshift-DSS)

#148 AI-938: Prevent Trace panel refresh flicker (@ashwanth1109, Shipyard)

#147 AI-937: Add a Pomodoro timer to the top bar (@ashwanth1109, Shipyard)

#20 fix(delivery): align approval gates and Surtr networking (@benji-bizzell, Redshift-DSS)

#144 AI-934: Align read-only companion with shared chat UX (@ashwanth1109, Shipyard)

#19 fix(dependencies): establish a current validated baseline (@benji-bizzell, Redshift-DSS)

#12 fix(dependencies): align validated SDK and delivery updates (@benji-bizzell, Redshift-DSS)

#146 AI-932: Support image attachments in Pi chat threads (@ashwanth1109, Shipyard)

#11 fix(delivery): harden CI tooling and release preview (@benji-bizzell, Redshift-DSS)

#145 AI-933: Clarify Codex and Pi connection setup (@ashwanth1109, Shipyard)

#1596 Capacity: artifact integrity, retention, and audit lineage (AERIE-2580) (@marcusdAIy, Aerie)

#1599 Forecast V2: publish End-of-Year arrival provenance (@vvp-trilogy, Aerie)

#3836 fix(board-doc): carry triaged findings across add-on Doc reconcile (@marcusdAIy, Klair)

#3835 fix(budget-bot): avoid duplicate MIPs title after Apply (@marcusdAIy, Klair)

#3834 fix(addon): stop stale findings batch after document apply (@marcusdAIy, Klair)

#1592 Forecast V2: refactor Start-of-Year enrollment forecast (@vvp-trilogy, Aerie)

#3833 fix(addon): fail closed on nested P&L pseudo-headings before parent rewrites (@marcusdAIy, Klair)

#3832 fix(budget-bot): guide add-on on stale review finding (@marcusdAIy, Klair)

#1594 feat(reconciliation): add Site evidence and decision lineage UI (AERIE-2288) (@caina-barbosa, Aerie)

#1595 Fix Real Estate mobile scrolling (@YibinLongTrilogy, Aerie)

#3830 fix(board-doc): diagnose stale duplicate session identity without deleting drafts (@marcusdAIy, Klair)

#3828 fix(budget-bot): skip unsafe C2.1 margin target comparisons (@marcusdAIy, Klair)

#1590 1400-ahj-personnel-contract (@mwrshah, Aerie)

#3827 fix(addon): check effective Drive editing capability for repair (@marcusdAIy, Klair)

#1588 Capacity: fleet-wide run limit, rate-limit backoff, and a size cap on stored results (AERIE-2581) (@marcusdAIy, Aerie)

#143 Release: Shipyard 0.6.8 (@ashwanth1109, Shipyard)

#142 AI-930: Fix updater installation with empty companion thread (@ashwanth1109, Shipyard)

#3825 fix(addon): one-character heading markers for all sections; auto-repair stretched markers (@marcusdAIy, Klair)

#1523 Backfill production hotfixes (2026-09-25) into main (@benji-bizzell, Aerie)

#141 Release: Shipyard 0.6.7 (@ashwanth1109, Shipyard)

#1579 1397-aerie-site-backup-operation (@mwrshah, Aerie)

#140 AI-885: Add replay comparison reports and regression history (@ashwanth1109, Shipyard)

#1584 Centralize Community Commitment qualification (@vvp-trilogy, Aerie)

#139 AI-928: Fix Ask Shipyard startup and companion recovery (@ashwanth1109, Shipyard)

#1582 Require paid deposits for Community Commitments (@vvp-trilogy, Aerie)

#138 AI-884: Add node-level and end-to-end eval scoring with failure attribution (@ashwanth1109, Shipyard)

#1577 fix(reconciliation): carry target decision schemas (AERIE-2591) (@caina-barbosa, Aerie)

#2097 fix(runners): keep the Redshift cancellation window when a statement times out at the run deadline (SURTR-1548) (@kevalshahtrilogy, Surtr)

#1575 Forecast V2: publish End-of-Year enrollment forecast (@vvp-trilogy, Aerie)

#137 AI-927: Repair Pi task streaming and lifecycle (@ashwanth1109, Shipyard)

#133 AI-918: Improve and consolidate search input UI (@ashwanth1109, Shipyard)

#2094 feat(aerie-a8): forecast-input parity marts, Q2 projections + Q3 coming-year projections + app conversion (SURTR-1545) (@kevalshahtrilogy, Surtr)

#2093 088-finance-audit-gl-access (@mwrshah, Surtr)

#136 Release: Shipyard 0.6.6 (@ashwanth1109, Shipyard)

#134 AI-883: Build an end-to-end feature replay runner (@ashwanth1109, Shipyard)

#135 AI-917: Integrate Pi adapter into task and conversation UI (@ashwanth1109, Shipyard)

#1568 Forecast V2: publish expected enrollment arrival inputs (@vvp-trilogy, Aerie)

#2084 feat(aerie-a8): mart-aerie-admissions-refresh runner + G1 parity marts (SURTR-1536, SURTR-1537) (@kevalshahtrilogy, Surtr)

#2082 feat(gateway): register all 24 A8 Gateway sources in an aerie-a8 entity (SURTR-1533) (@kevalshahtrilogy, Surtr)

#1562 Forecast V2: bound milestone conversions by enrollment date (@vvp-trilogy, Aerie)

#2081 fix(collections): skip 42DS forecast view (@sanketghia, Surtr)

#2068 fix(education): trim Schools Data Sheet cells like JavaScript .trim() (SURTR-1524) (@kevalshahtrilogy, Surtr)

#2069 feat(gateway): register the Finalsite tenant directory for Aerie (SURTR-1523) (@kevalshahtrilogy, Surtr)

#2080 fix(education): make behavioral_events migration 006 drop+create idempotent (SURTR-1519) (@kevalshahtrilogy, Surtr)

#3824 fix(board-doc): enforce literal Khoros FY26 markers on main (@marcusdAIy, Klair)

#3822 fix(board-doc): Khoros Q4 paired financials and guarded narrative refresh (@marcusdAIy, Klair)

#2079 fix(aws-spend): map Totogi CapitaTFL account (@caina-barbosa, Surtr)

#1559 Restore Forecast V2 compatibility publication (@vvp-trilogy, Aerie)

#3821 fix(board-doc): reject truncated Khoros Financials rows (@marcusdAIy, Klair)

#3819 feat(board-doc): Q4 Khoros-only BU financial source and scoped owner access (@marcusdAIy, Klair)

#3817 fix(data-api): guide Education platform charge comparisons (@mwrshah, Klair)

#1558 Show demographic totals in collapsed mobile cards (@YibinLongTrilogy, Aerie)

#1556 Fix admissions program session qualification (@vvp-trilogy, Aerie)

#1557 Forecast V2: remove Marketing Planning Forecast from the report UI (@vvp-trilogy, Aerie)

#1554 Capacity: publish and roll back record mode only once the DD request is approved (AERIE-2579) (@marcusdAIy, Aerie)

#1544 Consolidate redundant dbt data tests (@vvp-trilogy, Aerie)

#1545 Forecast V2: retain observations with missing program years (@vvp-trilogy, Aerie)

#210 Bound traces after redaction and complete requiredWhen (SINDRI-519) (@marcusdAIy, Sindri)

#1552 Capacity: rollback compare-and-set, docType-scoped dedupe, honest drain scheduling, day-bounded sweep (@marcusdAIy, Aerie)

#1551 Fix Portfolio Utilities value overflow (@YibinLongTrilogy, Aerie)

#1550 Capacity: wait for in-flight re-indexes, require prior numbers in the contamination log, operator review for held runs (@marcusdAIy, Aerie)

#1548 Capacity evidence: prior capacities, site facts in hash, poll unknown statuses (AERIE-2356) (@marcusdAIy, Aerie)

#209 Spell out CAP-1, CAP-2, and CAP-4 in the capacity prompt (AERIE-2501) (@marcusdAIy, Sindri)

#1547 Harden capacity evidence and dispatch (AERIE-2356) (@marcusdAIy, Aerie)

#1546 Capacity: record Sindri failure reasons; cover rollback failure paths (@marcusdAIy, Aerie)

#1543 Restrict admissions deals to school-year sessions (@vvp-trilogy, Aerie)

#2067 fix(aws-spend): map AI Engineering experiment account (@caina-barbosa, Surtr)

#2060 fix(education): accept GuidePlatform behavioral event fields (@kevalshahtrilogy, Surtr)

#2075 fix(sales-educrm-mart-sync): add upstream program_id to the coming-year projection tables (@kevalshahtrilogy, Surtr)

Mac's Picks — Key PRs This Week  (click to expand)
#163 — AI-962: Evaluate task prompts with two independent replay trials @ashwanth1109  no labels

Completed-task evaluation previously had no in-app way to propose a prompt change, replay it against the original inputs, and compare the evidence before publication. This adds persistent eval chat with one editable candidate and two independent trials, with original/trial conversations available throughout review.

## Business Value

Users can find optimization opportunities or shorten a template, refine the proposal directly in chat, and verify behavior before adopting it. Incomplete execution, failed acceptance checks, and passing results have distinct next actions; retries preserve the candidate and prior attempt evidence.

## Changes

- Add a compact proposal card, prompt diff, acceptance checks, trial transcripts, focus mode, concise quick messages, and shared new-conversation/composer controls. New eval chats default to Full access and reset their proposal state when replaced.

- Restore declared repository and artifact inputs from the captured node-entry baseline, record implementation base provenance, replay original follow-ups and images, and retain partial transcripts. Recover transient provider reads without resubmitting agent turns.

- Bind each assessment to the candidate and exact current run IDs. Separate behavioral regressions from informational notes, and request another review when the assessment contradicts itself.

- Keep publication explicit and validate successful independent runs, passing checks, baseline provenance, candidate hash, and current template identity before creating an immutable local template version.

## Validation

- Frontend production build and native app build passed on the integrated branch.

- Native library suite: 377 passed, 3 ignored.

- UI, transcript, chat, workflow, and PR-review regression tests passed.

- Fixture-backed desktop smoke run d87a87bb-c923-4ccb-ac1f-00502c360736 passed workflow verification. Observed original/trial navigation, Full access/model controls, and focus expand/restore; the harness was stopped afterward.

- Two real-agent Research trials completed with saved conversations and final artifacts; all four acceptance checks passed in both. One trial reported a missing-esbuild validation limitation, recorded separately from replay execution and behavioral regressions.

Historical tasks still require verifiable starting inputs; the runner reports missing provenance instead of replaying against current state. The recovered historical baseline used during validation was prepared locally, without adding a database backfill or changing bundled node templates.

## Implementation Effort

Estimated 5–8 engineer-days for an average engineer to implement and validate the UI, replay/session lifecycle, baseline restoration, provider recovery, publication safeguards, and regression coverage without AI assistance.

## Linear

https://linear.app/builder-team/issue/AI-962/evaluate-and-refine-task-prompts-with-two-independent-replay-trials

#2128 — feat(aerie-a8): read dbt_invocation_id as the dbt build marker (SURTR-1595) @kevalshahtrilogy  approvedmercy-allow-critical

> [!NOTE]

> Update, 5 Oct: the Aerie prerequisite is met. Aerie PR 1655 (AERIE-2700, shadow checks accept both marker forms) was released to production in Aerie release PR 1654 at 06:53 UTC. The banner below is kept for history; this PR can now be applied and merged, DDL first.

> DO NOT MERGE and DO NOT APPLY THE DDL until the Aerie prerequisite below is released.

> Production Aerie runs both gates in shadow (ADMISSIONS_PIPELINE_READ=shadow, SIS_ENROLLMENT_READ=shadow). Its shadow check pins the pg_class_oid: marker form, so applying this DDL today would make three shadow checks degraded on every cycle. 072 and 092 are the canonical definitions of what production runs, so this PR must not merge before its DDL is applied either.

Linear: [SURTR-1595](https://linear.app/builder-team/issue/SURTR-1595/dbt-publication-marts-read-dbt-invocation-id-as-the-build-marker)

## What changes

Aerie PR 1638 stamps dbt's run id on every row of sandbox_education.mart_admissions_pipeline_dtl and sandbox_education.mart_enrollment_dtl as dbt_invocation_id. The two procedures that copy those relations identified a dbt build by the relation's pg_class OID. They now use dbt's own id.

| Object | Before | After |

|---|---|---|

| mart_education.sp_refresh_aerie_admissions_pipeline_detail | pg_class_oid:<oid> | dbt_invocation_id:<id> |

| mart_education.sp_refresh_aerie_sis_enrollment | pg_class_oid:<oid> | dbt_invocation_id:<id> |

| sp_refresh_aerie_admissions_forecast_v2 | pg_class_oid:<oid>,<oid> | unchanged (its relations do not carry the column, and it is not applied in production) |

- Procedures (072, 092): the one marker assignment is replaced by a read of dbt_invocation_id and a fail-closed check.

- The OID guard is kept. Each procedure still resolves the relation's OID first and re-reads it after the copy, so a dbt build swapped in mid-copy still fails the procedure. The OID also still drives source_published_at and the column-type guard.

- Published comments (071, 090, 091): the table and source_build_marker column comments described the OID form.

- Reconciliation (both reconciliation/*.sql): the marker queries read the marker the way the procedure does.

- Tests and README: everything that pinned the OID form, plus the new checks.

- Runner code is unchanged. handler.py and redshift_client.py only compare markers for equality. A test now asserts neither names a marker prefix.

## No new numbered file

The ticket asked for a numbered migration. The procedure files are edited in place instead:

- A procedure file is CREATE OR REPLACE, so applying it again is its migration. 016/056 exist because CREATE TABLE IF NOT EXISTS cannot migrate a table.

- A 075/093 copy would be a second 480 to 550 line definition of each procedure (PIPELINE §10 asks for one canonical DDL). The contract tests also require exactly one numbered file per object.

- Earlier procedure changes were made in place (004_sp_refresh_education_plan_actual_variance.sql, the SURTR-1590 lock-order fix).

So DDL_FILES is unchanged. The README now documents this as "Changing an applied procedure".

## How the marker is read, and why

SELECT COUNT(*), COUNT(DISTINCT dbt_invocation_id),

SUM(CASE WHEN dbt_invocation_id ~ '<uuid>' THEN 1 ELSE 0 END), MIN(dbt_invocation_id)

INTO ... FROM sandbox_education.<relation>;

-- raise unless: exactly one distinct id, and every row carries a UUID

v_source_build_marker := 'dbt_invocation_id:' || v_source_invocation_id;

- Not a one-row read. A dbt table build writes one run's rows, so one id on every row is the invariant. If it ever breaks (rows of two runs, a NULL), a LIMIT 1 read would pick an arbitrary id and the marker could flip between runs. The procedure fails instead; the previous publication stays and the run alerts.

- Empty relation, missing column: both fail closed too.

- Width. Only a UUID is accepted, so the marker is always 18 + 36 = 54 bytes, inside source_build_marker VARCHAR(128).

- Cost. One pass over one column (about 31,000 and 9,200 rows) on each 10-minute run, where the no-op path used to read only the catalog. It takes a brief shared lock on the dbt relation.

Production evidence (read-only SELECTs, 2026-10-05, build created 05:07 UTC):

| Relation | Rows | Rows with an id | Distinct ids | Rows with a UUID | Marker bytes | Column type |

|---|---|---|---|---|---|---|

| mart_admissions_pipeline_dtl | 30,956 | 30,956 | 1 | 30,956 | 54 | varchar(36) |

| mart_enrollment_dtl | 9,190 | 9,190 | 1 | 9,190 | 54 | varchar(36) |

Both relations carry the same id, as one hourly dbt build writes both. The exact marker SELECTs in 072/092 and the new reconciliation marker queries were run as they stand in the files, with the same result.

## Behaviour after the DDL is applied

1. The next scheduled run of each procedure sees a marker (dbt_invocation_id:...) that differs from the published one (pg_class_oid:...), copies the current build and reports published. That is one republish per contract (three marts). No force run is needed.

2. Later runs report unchanged until the next dbt build, as today.

3. A dbt build is still detected within 10 minutes. A relation recreated without a dbt run no longer triggers a republish.

## Prerequisite in Aerie (blocking)

Aerie's shadow check computes the marker on its own side and compares it with the copy's source_build_marker as text:

- sync/src/analytics/queries/admissions-pipeline-gateway.ts: ADMISSIONS_PIPELINE_BUILD_MARKER_SQL builds 'pg_class_oid:' || oid.

- sync/src/analytics/sis-enrollment-gateway.ts: SIS_ENROLLMENT_BUILD_MARKER_SQL builds the same form.

- sync/src/analytics/a8/keyed-shadow-compare.ts: when the two markers differ, a8BuildSkew orders them with /^pg_class_oid:(\d+)$/. Any other form is unknown, and the outcome is degraded ("the build markers differ and are not both pg_class_oid:<oid> ... not compared").

With this DDL applied and Aerie unchanged, the shadow checks of aerie-admissions-pipeline-detail, aerie-sis-enrollment-rollup-input and aerie-sis-enrollment-member would be degraded on every cycle and never record clean. Published data is unaffected in every mode: legacy does not read the copy, and gateway only logs the marker.

What Aerie needs:

1. Read its side's marker as 'dbt_invocation_id:' || MIN(dbt_invocation_id) for these two relations.

2. Replace the OID ordering in a8BuildSkew. UUIDs have no order, so "which side is newer" needs another signal, for example the copy's source_published_at against the relation's creation time.

Safe order: Aerie accepts both marker forms → Aerie release → apply this DDL → merge this PR.

## Pre-merge DDL

Not applied by this PR, and not yet applied anywhere. Once the Aerie prerequisite is released, apply these files, then merge:

cd pipelines/runners/mart-aerie-dbt-publication-refresh

REDSHIFT_CLUSTER_IDENTIFIER=redshift-cluster-1 REDSHIFT_DATABASE=finance_dw REDSHIFT_DB_USER=CQL_download_OM \

uv run python scripts/apply_ddl.py \

../../cdk/sql/mart_education/071_aerie_admissions_pipeline_detail.sql \

../../cdk/sql/mart_education/072_sp_refresh_aerie_admissions_pipeline_detail.sql \

../../cdk/sql/mart_education/090_aerie_sis_enrollment_rollup_input.sql \

../../cdk/sql/mart_education/091_aerie_sis_enrollment_member.sql \

../../cdk/sql/mart_education/092_sp_refresh_aerie_sis_enrollment.sql

- 072 and 092 are the behaviour change. They are the minimum to apply.

- 071, 090 and 091 only refresh comments. Their CREATE TABLE IF NOT EXISTS is a no-op. Their OWNER TO and REVOKE statements repeat the current state: the three marts are owner-only in production today (ACLs read 2026-10-05).

- Applying these files again is safe and idempotent. Every top-level statement is CREATE TABLE IF NOT EXISTS, CREATE OR REPLACE PROCEDURE, COMMENT, OWNER TO, REVOKE or GRANT; a new test checks that for every file of this runner. A failed CREATE OR REPLACE leaves the previous procedure in place.

- Only the named files are applied. With paths given, apply_ddl.py sends those files' statements and nothing else (new test), so Forecast V2's objects are not created. The script has no runtime check for destructive statements; the test above is the guard.

- The 10-minute schedule is live, so the new procedures take effect on the next run after the apply.

- Verify: the next run's results show published for all three marts with a dbt_invocation_id: marker, the run after shows unchanged, and both reconciliation files report PASS (parity in Query 4; rollup_parity, member_parity and one_publication in Query 8).

## Validation

- Runner suite: uv run pytest, 246 passed (229 on main).

- mart-aerie-admissions-refresh suite (reads the same DDL directory): 556 passed.

- ruff check pipelines and ruff format --check pipelines with ruff 0.15.22: clean.

- scripts/apply_ddl.py --dry-run on the five files: 26, 4, 14, 26 and 4 statements.

- Production, read-only: the table above; the marker SELECTs and reconciliation marker queries as they stand in the files; the published markers (all three marts on the current build's OID); the deployed procedures (both still on the OID marker; Forecast V2's procedure is not deployed).

- Not validated: the procedures were not created or called anywhere. Redshift is not reachable from CI, and this PR applies no DDL.

## Business Value

The Surtr copies of these two marts feed Aerie's Admissions Pipeline and SIS Enrollment reports through the Gateway. The copy's lineage now names the dbt run that built its rows, a value dbt also writes to its own logs, instead of a Postgres catalog number. "Which dbt run is this copy from" becomes answerable directly, and the copy no longer republishes when a relation is recreated without a dbt run. This closes the last open design decision (A8 plan §9 D4) for these two sources on the way to moving the EC2 analytics worker's reads onto Surtr. The work also found that Aerie's shadow check pins the old marker form, before the change could silently stop the shadow evidence that gates that cutover.

## Manual Effort Estimate

About 1 day of focused work by hand (6 to 8 hours): reading both procedures and the runner's contract tests, checking the column in production, changing two procedures, two reconciliation files, three published comments, the README and four test modules, and tracing the marker into Aerie's shadow check. This is a proposed figure for Keval to confirm or adjust.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1655 — feat(a8): dbt copy shadow checks accept both build marker forms, pg_class_oid and dbt_invocation_id (AERIE-2700) @kevalshahtrilogy  approved

## Summary

A prerequisite for SURTR-1595. It changes nothing on its own.

Surtr's copies of two Aerie dbt relations (mart_admissions_pipeline_dtl, mart_enrollment_dtl) name the dbt build they copied in source_build_marker. Today that is pg_class_oid:<oid>. PR 1638 put dbt_invocation_id on every row of both relations, and Surtr will move the marker to dbt_invocation_id:<id>.

Aerie's shadow check pinned the OID form: the pg side computes pg_class_oid:<oid>, the kit requires it to equal the copy's marker as text, and the direction rule only orders OIDs. If Surtr switched first, three checks that run in shadow in production (pipeline detail, SIS rollup input, SIS members) would read degraded on every cycle.

After this PR Aerie accepts both forms, so the order is Aerie first, Surtr second, with no simultaneous flip.

## How it works

The copy's marker decides the form that is compared.

| Copy's source_build_marker | What the check does |

|---|---|

| pg_class_oid:<oid> | Exactly what it does today. |

| dbt_invocation_id:<id> | Compares it with the id read from the pg relation's own rows. Same id: the rows are compared. Different id: ordered (below). |

| anything else | degraded, as today. |

The pg side's id. One aggregate over the relation (COUNT(*), COUNT(dbt_invocation_id), COUNT(DISTINCT dbt_invocation_id), MIN(dbt_invocation_id)), read between the two OID lookups that already bracket the pg read. The same OID before and after proves the relation did not change, so the id belongs to the build the rows came from. It is usable only when every row carries the same one non-null id. Two ids, rows without one, an empty relation, a missing column or a failed query all mean "no marker in that form", and the check is degraded with a count-only reason if the copy carries that form.

Direction without OIDs. Invocation ids are UUIDs and have no order. Two differing builds are ordered by when their relations were created:

- the copy's side is its source_published_at. I read Surtr's 072_sp_refresh_aerie_admissions_pipeline_detail.sql and 092_sp_refresh_aerie_sis_enrollment.sql on Surtr origin/main: both stamp pg_class_info.relcreationtime AT TIME ZONE 'UTC' of the copied relation there;

- the pg side is the same catalog value for the relation its rows came from, looked up with the same expression, and used only if the relation still has the OID the rows were read under.

dbt's table materialization creates a new relation on every build, so the older relation is the older build. Both times come from the warehouse catalog, so no worker clock is involved.

| Copy's relation vs the pg relation | Outcome |

|---|---|

| older | stale_copy (the copy lags) |

| newer | degraded (the copy is not the stale side) |

| same time, not known, relation rebuilt since, lookup failed | degraded (order unknown) |

The catalog is asked only when the ids differ, and once per cycle (the two SIS checks share it).

## What changes per mode

| Mode | Change |

|---|---|

| legacy | None. |

| shadow, copy still in the OID form (today) | No line changes: outcome, markers and messages are identical. The pg side makes one extra read-only aggregate per relation per cycle (the id), whose result is unused and whose failure is silent while the copy is in the OID form. |

| shadow, copy in the dbt_invocation_id form (after SURTR-1595) | Compared in that form, as described above. Without this PR every such line would be degraded. |

| gateway | None. Gateway mode does not read pg markers. |

Publishing is untouched in every mode: shadow still returns exactly the legacy read.

I read the id eagerly, inside the bracket, on purpose. Reading it only after seeing the copy's form would save that one query before the switch, but the read would then happen after the Gateway read, and a dbt rebuild in that window would make the cycle degraded. The worker cycle and dbt are both hourly, so an unlucky phase could repeat that every cycle.

## Scope

Only the two relations that carry the column. Other callers of the stale_copy rule pass no pgDbtInvocation and behave exactly as before, including against a copy in the new form (degraded, with today's message). Forecast V2 relations are not touched.

## What Surtr's DDL must guarantee (SURTR-1595)

1. source_build_marker = 'dbt_invocation_id:' || <the relation's single dbt_invocation_id>, verbatim: no trimming, no case change.

2. source_published_at stays the copied relation's pg_class_info.relcreationtime (UTC), read in the same transaction and for the same relation as the id. The direction rule depends on it.

3. Marker, publication time and rows describe one relation (the existing "dbt replaced the build during the copy" guard stays).

4. One call stamps one marker on every row, and for SIS the same marker on both marts. A mixed pair is refused by Aerie's reader, as today.

5. Refuse to publish when the relation does not carry exactly one distinct non-null dbt_invocation_id.

## Business Value

Three production shadow checks are collecting the evidence that decides when the analytics worker can read admissions pipeline and SIS enrollment data from Surtr instead of Redshift. Surtr needs to change how its copies identify a dbt build. Without this change that Surtr release would silently turn all three checks degraded and stop the evidence, and the two repos would have to be released at the same instant to avoid it. With it, each side can ship on its own schedule and the shadow window keeps running through the change.

## Manual Effort Estimate

Proposed: about 7 hours of focused work by hand, without AI. Keval, please confirm or adjust. It covers reading the two Surtr procedures and the kit to design the two-form compare, working out a sound direction rule without OIDs (and why reading the id lazily is unsafe), the kit change, the two gates, and about 120 tests across five files.

## Testing

- sync: pnpm typecheck exit 0, pnpm lint exit 0, vitest run --maxWorkers=2 exit 0 (123 files, 2897 tests).

- OID-form copy is unchanged. Before adding any test, the existing kit, pipeline and SIS suites (292 tests) passed unmodified against the new code. New tests also compare whole lines: for a copy in the OID form the logged line is deep-equal whatever the id read returns (usable, another id, unavailable, throwing), for the kit, the pipeline gate and the SIS gate.

- New form (kit, pipeline gate, SIS gate): same id compares the rows (clean, and mismatch on differing rows); lagging copy is stale_copy; newer copy is degraded; unknown order (no creation time, equal times, relation rebuilt since, lookup failing) is degraded; two ids, rows without an id, no id at all and a missing column on the pg side are degraded with a count-only reason; an unrecognised prefix stays degraded; a dbt build swapped in mid-read stays degraded.

- legacy and gateway modes never issue the new reads.

- The heavy pipeline test file gains one case with real rows. The rest of the pipeline cases are in a new light file (admissions-pipeline-build-marker.test.ts) that stubs the Gateway table read, because the real reader refuses fewer than 10,000 rows.

## Not covered

- Not verified against the warehouse: that the analytics worker's Redshift user can read pg_class_info.relcreationtime for the two relations. If it cannot, nothing breaks, but a lagging copy in the new form would read degraded (order unknown) instead of stale_copy. Both dry-run scripts now print one line saying whether the pg side can read the id and the creation time, so this can be confirmed before Surtr switches.

- No Surtr change, no production change, no environment change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1653 — SIS Enrollment: complete start-year cohorts in the enrollment mart @vvp-trilogy  approved

Completes the six start-of-year SIS cohorts in the existing enrollment mart. Qualified cancelled returning outcomes and matched Transfer In memberships retain their true offering year; unmatched departures remain visible and ambiguous pairs are never guessed. No educrm mart supplies production enrollment facts, and no separate transfer mart is introduced.

Adds deterministic fixtures and source/grain/coverage checks, plus a persisted transfer audit included in hourly model builds even when tests are skipped. Preserves existing headcount classification and legacy forecast grade operands. Neutral forecast Declined now includes 28 qualified outcomes; the existing negative-residual diagnostic flags one additional school/year (documented below), with the formula unchanged.

Validation: local dbt parse and diff checks passed. Final-commit warehouse run: 113 build nodes passed; 570 tests passed, 8 warnings, 0 errors. All application CI checks passed on f48ddb2d3. Final-commit mart checks found zero grain/year/cancellation/transfer violations and zero changes to existing SIS cohort memberships.

The required pre-merge default-report comparison is complete for 2024–25, 2025–26 and 2026–27, with @vvp-trilogy tagged:

- [Totals, source differences, freshness and forecast effect](https://github.com/AI-Builder-Team/Aerie/issues/1294#issuecomment-5978285836)

- [Per-school six-cohort tables](https://github.com/AI-Builder-Team/Aerie/issues/1294#issuecomment-5978285941)

- [Historical coverage boundaries](https://github.com/AI-Builder-Team/Aerie/issues/1294#issuecomment-5978286083)

Final-commit warehouse marker: 3e4a3956-e484-47bd-96bf-6489338f9038; 9,190 rows / 6,013 facts. All comparison rows and identity explanations revalidated unchanged before PR table cleanup. Mercy approved f48ddb2d3 after the missing-date audit fix and fixture coverage. Its dense-grid concern was withdrawn as a pre-existing contract limitation. All review threads are resolved. Comparison, identity explanations, historical coverage and existing-cohort regression were rechecked on this head; all comparison values are unchanged. The audit flags six 2025 and four 2026 departures with incomplete dates without changing cohort memberships.

Closes #1294.

#1638 — feat(dbt): stamp dbt_invocation_id on the admissions pipeline and SIS enrollment marts (AERIE-2682) @kevalshahtrilogy  approved

## Summary

Adds one column, dbt_invocation_id, as the last column of two marts:

- mart_admissions_pipeline_dtl

- mart_enrollment_dtl

Linear: [AERIE-2682](https://linear.app/builder-team/issue/AERIE-2682/a8-stamp-dbt-invocation-id-on-mart-admissions-pipeline-dtl-and-mart)

@vvp-trilogy, this touches your models, so it is yours to review and merge. I will not merge it.

## What invocation_id is (no source field needed)

invocation_id is a variable dbt itself provides to every model. It is the id (a UUID) of the dbt run that is building the table. dbt fills it in when it compiles the SQL, so the model line is just:

cast('{{ invocation_id }}' as varchar) as dbt_invocation_id

- It is not read from any source table. There is no ingestion change, no new staging column and no join.

- Every row of one build has the same value. The next build writes a new value.

- The hourly job is one dbt build, so both marts built in the same run carry the same id. A PR build gets its own id in its pr<N>_ tables.

- It is a run id, not personal data. dbt already writes the same id to its logs and run_results.json.

## The change

| File | Change |

|---|---|

| dbt/models/marts/admissions/mart_admissions_pipeline_dtl.sql | Final select gains dbt_invocation_id after u.*. |

| dbt/models/marts/enrollment/mart_enrollment_dtl.sql | The final UNION ALL moves into a unioned CTE; the final select is u.* plus dbt_invocation_id, so the value is written once for both arms. Rows and existing columns are unchanged. |

| _mart_admissions__models.yml, _mart_enrollment__models.yml | Column documented, with a not_null test. The pipeline mart's "Column population" table gets a row. |

The column is last, so no existing column changes position. Materialization is untouched: both marts are table (set per layer in dbt_project.yml), so there is no on_schema_change to handle. Neither model has an enforced contract.

## Why Surtr needs it

Surtr's mart-aerie-dbt-publication-refresh pipeline copies these two production marts every 10 minutes so the analytics worker can read them through the Surtr Gateway. It must know when dbt has rebuilt a mart. Today it uses the table's pg_class OID as the build marker. dbt_invocation_id is the marker it was designed to use.

## Check 1: does the Surtr copy keep working? Yes, with no Surtr change.

Read from Surtr main (pipelines/runners/mart-aerie-dbt-publication-refresh/ and pipelines/cdk/sql/mart_education/):

- Explicit column lists. Both procedures copy with Aerie's own reader SQL: 53 named columns for the pipeline detail, 6 for the SIS rollup and 21 for the SIS members. Nothing does SELECT * from the dbt relation.

- The type guard ignores extra columns. It walks the columns of the Surtr mart and looks each one up by name on the dbt relation. A column that exists only on the dbt side is never examined.

- Parity checks are safe. The reconciliation EXCEPT queries compare CTEs built from the same named lists. The == 53 test counts Aerie's select list, not the dbt relation's columns.

- The Surtr README says so: "New dbt columns that Aerie does not read ... need nothing here."

- Marker today: v_source_build_marker := 'pg_class_oid:' || v_source_oid in each procedure. Surtr does not read dbt_invocation_id yet.

- What Surtr expects of the new column: the name dbt_invocation_id. It pins no type and no position. The value lands in source_build_marker VARCHAR(128); a 36-character UUID fits.

Order. This PR goes first and is safe alone. Only after it is merged and one production build has run should Surtr switch its marker (a separate Surtr ticket): replace the one v_source_build_marker := assignment in each of the two procedures, plus the tests and reconciliation queries that pin the OID form. Doing the Surtr switch first would make its procedures fail on a missing column.

## Check 2: dbt side

- Materialization: table for both. No incremental model is touched.

- Schema yml: columns are documented, so the new column is added to both files. No enforced contract on either model.

- Downstream dbt readers select named columns: int_admissions_forecast_grade_operands, and the forecast_neutral_ec and forecast_pipeline_scope macros (used by int_admissions_forecast). No model does select * from either mart, so the column does not spread.

- Singular tests: three use select * from the pipeline mart inside a CTE and then filter on named columns (..._community_hubspot_ids, ..._deposit_paid_date_scope, ..._enrollment_date_source). An extra column does not affect them. No test pins a column list, count or order.

- Unit tests: the three unit tests on mart_admissions_pipeline_dtl assert only the columns named in expect. The unit tests that use either mart as an input (_int_admissions__models.yml, forecast-grade-operands.yml) null-fill columns they do not set.

## Check 3: Aerie readers

Every reader of the two relations, found by searching the repo for both model names:

| Reader | Columns | Row schema |

|---|---|---|

| queryAdmissionsPipelineRows (sync/src/analytics/queries/admissions-pipeline.ts) | 53 named columns | z.object, not .strict() |

| rollupsSql (sync/src/redshift/sis-enrollment.ts) | 5 named columns + COUNT | z.object |

| membersSql (same file) | 21 named columns | z.object |

No SELECT *, no .strict() schema, so an extra column cannot reach or break them. chat/ and packages/contracts only mention the marts in comments. The other mart_enrollment_dtl hits in sync/ are the EduCRM table staging_education.sales_educrm_wh_mart_enrollment_dtl, a different relation. An org-wide code search found no reader of these two relations outside Aerie and Surtr.

Not checkable from code: ad-hoc or BI queries by edu_read users that do SELECT *. They would see one more column at the end.

## Testing

- dbt parse: passes (dbt-core 1.12.0, dbt-redshift 1.11.0, placeholder profile, no warehouse).

- dbt compile --no-introspect of both models: renders offline; the final selects end with CAST('<uuid>' AS VARCHAR) AS dbt_invocation_id.

- Not run locally: dbt build and dbt test. They need warehouse credentials, which I did not use. The PR build (prefixed) job on this PR does both.

CI result (PR build (prefixed), run twice, same result both times):

- dbt build: 103 of 103 models and seeds built, including both marts with the new column.

- dbt test: 543 pass, 8 warn, 1 fail. Both new tests pass (not_null_mart_admissions_pipeline_dtl_dbt_invocation_id, not_null_mart_enrollment_dtl_dbt_invocation_id).

- The one failure is assert_sis_enrollment_lifecycle_has_arrival (1 row). It reads only int_enrollment_cohort, which is upstream of the marts and is not changed here; the dbt manifest shows neither changed mart among the test's ancestors. It is a SIS data finding (one enrollment in a lifecycle cohort with no arrival cohort in the same year), not an effect of this PR. It keeps the dbt check red until the data or the test is addressed. @vvp-trilogy, that one is yours to judge; I have not touched the test.

## Business Value

The Surtr copy of these two marts feeds Aerie's Admissions Pipeline and SIS Enrollment reports through the Gateway. A build marker that dbt itself writes makes "which dbt run is this copy from" a plain, auditable value on the row instead of a Postgres catalog number. That makes stale-copy alerts and parity checks easier to trust and to explain, and it removes the last open dependency (A8 plan §9 D4) on the way to moving the EC2 analytics worker reads onto Surtr.

## Manual Effort Estimate

About 2.5 hours of focused work by hand: reading the Surtr procedures and tests to confirm an extra column is safe (1 h), tracing dbt and Aerie readers (1 h), the change and PR (0.5 h).

_Proposed by Claude. @kevalshahtrilogy, please confirm or adjust._

## Not covered

- The Surtr switch from the OID marker to dbt_invocation_id (separate Surtr ticket, after this is in production).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1652 — fix(a8): keyed Gateway readers say when a program key is absent; gateway stays unreleased in production (AERIE-2690) @kevalshahtrilogy  approved

## Summary

Follow-up to Mercy's review of release PR 1641. Three of her critical findings are one class: a per-program Gateway lookup turns a missing program key into an empty result and reports success (admissions-marketing-gateway.ts:537 and :625; admissions-program-gateway.ts:653 and the other keyed readers; the whole-table app conversion read at :782).

The underlying fact. A program with no rows in a mart has no key in it, exactly as the legacy WHERE key = $1 returns no rows for it, and several sources are sparse (transfers and pipeline deposits cover a minority of programs; most programs have no shadow-day events). A program the mart *left out* looks the same. The reader cannot tell the two apart, and a table-level row floor cannot either.

That splits the finding in two, and they need different answers.

### Shadow (live in production): the verdict was already right; now it is pinned and visible

In shadow, the same call's pg read is the arbiter of whether a program has rows. Rows on pg and none from the Gateway is a mismatch, however the Gateway came to have none. An absent key only compares clean when pg returned no rows either, which is a true match.

- New tests pin both cases for the per-program gate (a keyed source and the alias-based Q3 reader) and the marketing gate: a program the mart leaves out while pg has its rows is a mismatch; a program with nothing on either side is clean.

- Every keyed slice now carries present, and each shadow line counts the absent keys it compared as empty: scope.absentCalls (per-program) and scope.absentPrograms (marketing, whose lines gain a scope). A reader of the line can see how much of a clean result was "nothing on either side".

### Gateway (not enabled in production): real, and closed off until there is a rule

Without pg there is no arbiter: gateway mode would publish a left-out program as an empty result, and a short slice as that program's data. Closing that properly needs a completeness rule, which needs a baseline (what is published for the program now) or a manifest from the mart. That is AERIE-2683 and is a design across about a dozen sources, not a guard in the reader.

What this PR does is make sure gateway mode cannot publish in production before that rule exists:

- ADMISSIONS_MARKETING_READ=gateway was selectable. It is now unreleased, like the per-program gate: refused with AdmissionsMarketingConfigError, before the admissions run starts, unless the new ADMISSIONS_MARKETING_READ_NONPROD_OPT_IN=true is set.

- Both gates now refuse gateway on the production worker even with the opt-in. The production worker is the one pinned to DBT_TARGET=production (compose.prod.yml); isA8ProductionWorker is the same literal check the dbt-backed gates already use. Until now "never set the opt-in in production" was a convention.

I did not make gateway mode refuse every absent key. That would fail every program that legitimately has no rows, on every cycle, for every sparse source; and shadow, which maps the Gateway side as gateway mode would, would then read degraded for those sources forever and stop producing evidence. The baseline rule in AERIE-2683 is the one that can tell the cases apart.

## What changes per mode

| Mode | Per-program gate (ADMISSIONS_PROGRAM_DETAIL_READ) | Marketing gate (ADMISSIONS_MARKETING_READ) |

|---|---|---|

| legacy | None. | None. |

| shadow | Publishes exactly as before. Source lines gain scope.absentCalls. No verdict changes. | Publishes exactly as before. Lines gain scope (programs, uncomparedPrograms, absentPrograms). No verdict changes. |

| gateway | Refused on the production worker even with the opt-in. Outside production, with the opt-in, the reads behave as before. | Now refused without ADMISSIONS_MARKETING_READ_NONPROD_OPT_IN=true, and always refused on the production worker. Outside production, with the opt-in, the reads behave as before. |

Production today runs both gates in shadow, so nothing it does changes except the extra counts on the lines. An environment that has ADMISSIONS_MARKETING_READ=gateway set without the opt-in would fail its admissions cycle at start with a clear config error after this deploys; I know of none.

## Business Value

The team is collecting shadow evidence to decide when the analytics worker can read admissions data from Surtr's Gateway instead of Redshift. This change does two things for that decision. It shows, on every line, how many programs matched only because both sides had nothing, so a clean window is not over-read. And it makes it impossible to switch production to gateway mode for these two gates before the missing-program rule exists, so a mart that drops a school can never blank that school's admissions data in the dashboards.

## Manual Effort Estimate

Proposed: about 5 hours of focused work by hand, without AI. Keval, please confirm or adjust. It covers working out which half of the finding is real (reading both gates, the kit and the refresh's write path), deciding between refusing absent keys and gating release, the presence flag and counts through two shadow implementations, the production refusal for both gates, and about 20 new or changed tests.

## Testing

- sync: pnpm typecheck exit 0, pnpm lint exit 0, vitest run --maxWorkers=2 exit 0 (121 files, 2779 tests).

- New tests:

- per-program shadow: absent key with nothing on pg is clean and counted; a program the mart leaves out is a mismatch (keyed reader and Q3 aliases);

- marketing shadow: the same two cases, with scope.absentPrograms;

- marketing gate: gateway refused without the opt-in (global and per-source), opt-in value must be exactly true, refused in production with the opt-in, legacy and shadow unaffected in production;

- per-program gate: refused in production with the opt-in; legacy and shadow unaffected;

- isA8ProductionWorker; the orchestrator aborts before any run write on a refused marketing configuration.

- Existing marketing gateway-mode tests now pass the opt-in; their assertions are unchanged.

## Not covered

- The completeness rule for gateway mode (AERIE-2683): per-program baselines from Convex, or a presence manifest published by the Surtr marts. Until it lands, gateway mode outside production still publishes an absent key as empty.

- No Surtr change. If the manifest route is chosen, every keyed mart (the seven per-program marts and the four marketing marts) would need to publish which program keys it covers.

- Not run against the live Gateway.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1651 — fix(a8): a failed legacy read makes the per-program shadow dry-run not clean (AERIE-2689) @kevalshahtrilogy  approved

## Summary

Follow-up to Mercy's review of release PR 1641 (finding at sync/src/scripts/dry-run-admissions-program-detail-shadow.ts:103).

The per-program shadow dry-run counts rejected legacy reads in failedReads, but its final verdict looked only at the shadow lines. A failed read that produced no non-clean line could still end in "Every check is clean" and exit 0.

failedReads now takes part in the verdict: any failed legacy read makes the run not clean, the script prints N legacy read(s) failed, so the run is not clean whatever the lines above say, and it exits 1.

Class audit: every dry-run that keeps a failure counter apart from its verdict. I searched all 22 dry-run-*.ts scripts for a counter (++, +=) or a swallowed rejection (.catch) and read the verdict logic of each A8 one. This was the only script with a failure counter outside its verdict. The other A8 dry-runs set their verdict to not clean inside each read catch (reference, pipeline, SIS, expenses, the generic dry-run-a8-shadow) or start from not clean and only turn clean on a result (community funnel, community deposits). The marketing dry-run had the same gap and was fixed in PR 1587.

## What changes per mode

| Mode | Change |

|---|---|

| legacy | None. |

| shadow | None on the worker. Only this developer script's verdict and exit code change. |

| gateway | None. |

## Business Value

An operator runs this script to decide whether the per-program shadow looks healthy before trusting a shadow window. If it exits 0 after reads failed, a partial run can be taken for a clean one. With this change a green run means every read the cycle makes actually ran and every check was clean.

## Manual Effort Estimate

Proposed: about 45 minutes of focused work by hand, without AI. Keval, please confirm or adjust. Most of it is checking the other 21 dry-run scripts for the same gap; the change itself is a few lines.

## Testing

- sync: pnpm typecheck exit 0, pnpm lint exit 0, vitest run --maxWorkers=2 exit 0 (121 files, 2744 tests).

- The dry-run scripts have no unit tests (each runs main() on import and needs Redshift and a Gateway key). The change is one boolean term in the verdict plus one message.

## Not covered

- The nit on admissions-program-gateway.ts:755 (use AdmissionsProgramDetailConfigError for a program with no code): not changed; reasoning is on the PR 1641 thread.

- I did not run the script against live Redshift or the Gateway.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2120 — fix(education): take school ontology locks in the producer's order across sibling marts (SURTR-1590) @kevalshahtrilogy  approved

## Summary

Incident (2026-10-02, SURTR-1590, related SURTR-1518). quickbooks-core-tables fanned out to its three marts at 06:46:25 UTC. mart-school-performance-unit-economics-refresh (Table 2) and mart-aerie-education-financials-refresh (AE) both died with deadlock detected, but not against each other. All three QuickBooks marts take quickbooks_financial_refresh_lock first, so they queue behind one another (AE got it at 06:46:29; QB waited 423 s and Table 2 436 s on it). The inversions were against the core-education-ontology-refresh producer, which rhodes-staging-sync triggered at 06:46:58 UTC and which does not take that lock, and against a plain reader. Relation 15058250 is core_education.dim_program and 15058264 is core_education.bridge_school_link.

Evidence from sys_query_history (CQL_download_OM):

| Time (UTC) | Statement | Result |

|---|---|---|

| 06:46:29 | AE sp_refresh_agg_school_pl_breakdown: LOCK bridge_school_link | waited 422 s on a holder not visible to this user, granted 06:53:32.085 |

| 06:47:02 | reader WITH alpha_school_years ... (session 1073826357; reads dim_school, bridge_school_link, dim_program) | 390 s lock wait, ends 06:53:36 |

| 06:53:32 | AE LOCK dim_program | blocked by a reader that holds dim_program and waits for bridge_school_link, which AE now holds: deadlock, AE aborted 06:53:33 (process 4582 / 5144 in the message). The reader above fits (its wait ends 06:53:36), but the message does not name it. |

| 06:53:40 | ontology retry sp_refresh_aerie_ontology: one LOCK TABLE rhodes..., bridge_school_link, dim_program, dim_school, dim_site, xref_school_source (the first attempt was a deadlock victim at 06:53:37) | takes bridge_school_link, then waits for dim_program |

| 06:58:17 | Table 2 sp_refresh_agg_school_performance_unit_economics_qtd: gets dim_program (it had dim_school already), asks for bridge_school_link | deadlock, Table 2 aborted 06:58:18 (process 4592 / 9564). The ontology LOCK completes 06:58:19.165, one second later, which identifies it as the other process. |

The ontology producer locks bridge_school_link, dim_program, dim_school, dim_site, xref_school_source. Table 2 and the Guide QTD procedure locked dim_school, dim_program, bridge_school_link: a textbook inversion. The AE procedure already follows the producer order, so the AE-versus-reader cycle is a queue pile-up behind the 422 s holder rather than an ordering bug in that procedure; it is not changed (the file is 165 KB, above the 100 KB Data API statement limit, and the live version is already newer than main).

Review follow-up (this push). Mercy's blocking finding on the first revision was that the contract left out the QuickBooks core writers. They lock xref_school_source before dim_school, the reverse of the producer, and they are not gated against it. Fixed here: the audit below covers every procedure that touches the ontology, sp_refresh_quickbooks_profit_and_loss_posting and sp_refresh_quickbooks_budget_detail now follow the canonical order, and the contract test discovers participants from the repo so a new procedure cannot silently escape it.

Lock-order map (ontology and shared relations, in acquisition order). Every participant takes these at the start of its body, before it reads them (a publication target may be locked just before its DELETE).

| Procedure | Before | After |

|---|---|---|

| sp_refresh_aerie_ontology (producer, unchanged) | bridge, program, school, site, xref_school_source | same |

| sp_refresh_agg_school_pl_breakdown (AE, unchanged) | posting fact, bridge, program, school, xref_school_source | same |

| sp_refresh_agg_school_performance_unit_economics_qtd (Table 2) | school, program, bridge, xref_school_source | bridge, program, school, xref_school_source |

| sp_refresh_agg_school_qtd_guide_staffing_program_spend | school, program, bridge, xref_school_source; posting fact after them | posting inputs first, then bridge, program, school, xref_school_source |

| sp_refresh_agg_school_performance_quickbooks_budget_qtd | xref_school_source, then school | school, then xref_school_source |

| sp_refresh_qtd_hc_posting_classification | school, xref_school_source, posting fact | posting fact, school, xref_school_source |

| sp_refresh_agg_school_qtd_all_other_headcount | Guide mart, then classification | classification, then Guide mart |

| sp_refresh_quickbooks_profit_and_loss_posting | xref_school_source, class_school, ue_model, school; posting fact ~900 lines later | posting fact, school, xref_school_source, class_school, ue_model |

| sp_refresh_quickbooks_budget_detail | xref_school_source, class_school, ue_model, school | school, xref_school_source, class_school, ue_model |

| sp_refresh_school_quickbooks_pl_reconciliation, ..._facilities_capex_campus_spend, ..._unit_economics_per_student_qtd, sp_load_q94_site_finance_entity_xref (unchanged) | already consistent | same |

sp_refresh_quickbooks_financial_contracts runs vendor identity, posting, budget, attribution, school P&L and reconciliation as children of one transaction (the handler calls only it), so its effective order is the children's locks concatenated; the first acquisition is posting fact, school, xref_school_source, which the contract checks.

Audit of every procedure that references bridge_school_link, dim_program, dim_school, dim_site, xref_school_source or a view over them (34 under pipelines/, 30 in the live catalog): the producer, 10 locking participants (the table above plus classification and q94), 4 one-off migrations that drop themselves (sp_replace_legacy_dim_school, sp_drop_retired_dim_school_next, the two sp_migrate_quickbooks_*), and 19 that take no explicit lock on these tables and are listed in UNLOCKED_READERS with a test that fails if one starts locking them: hubspot fct_admissions_deal/_event/hubspot_core_foundation, fct_enrollment, the two student-snapshot appenders, q48 publish, capex, finalsite_billing, school_calendar, aerie_admissions_program/_directory and the seven forecast procedures. The contract's participant list also holds All Other and Facilities, which lock shared marts without naming an ontology table. Readers are not changed (see below).

Canonical order. CANONICAL_LOCK_ORDER in mart-aerie-education-financials-refresh/tests/test_sibling_lock_order_contract.py: QuickBooks gate and posting inputs, then bridge_school_link, dim_program, dim_school, dim_site, xref_school_source, then the remaining inputs, then publication targets. Plain alphabetical was rejected: it would put the QuickBooks gate after dim_*. A procedure is *gated* when its first lock is the QuickBooks gate; gated procedures serialize on it and cannot deadlock each other, so the contract requires every pair that includes an *ungated* procedure (the producer, classification, All Other, Facilities, Table 3, q94) to agree on relation order, and ungated procedures to ascend through the canonical list.

Changes. Seven procedures re-ordered (lock statements and comments only; no lock mode, logic or output changes). The posting writer now locks the posting fact first instead of ~900 lines later, just before its DELETE: no lock is added or removed, it is taken earlier within a coordinator transaction that already holds it until commit (at most the writer's own ~25 s runtime earlier, per the 10:52 run). Table 2 and QB budget QTD version markers are bumped to 2026-10-02.1; the QuickBooks core markers are left alone because their tests pin them. The contract test (61 tests) now discovers participants, models gating and the coordinator transaction, and fails against the old files (11 failures) for the QuickBooks writers, the coordinator and the marts changed earlier.

Not changed, found on the way:

- Unlocked readers can stall writers for minutes. At 10:57:42 hubspot sp_refresh_fct_admissions_event opened an 11.7 minute transaction; it reads bridge_school_link, dim_program, dim_school and dim_site, so it holds AccessShare on them until commit. QB budget QTD's LOCK dim_school waited 777 s and was granted 0.6 s after that transaction's last statement; Table 2 waited 796 s behind QB; AE's classification step stalled behind QB's posting-fact lock and timed out (the 10:56 AE failure). Not a deadlock, not caused by lock order, and not fixed here; options are an up-front LOCK ... IN ACCESS SHARE MODE in canonical order or copying the inputs to a temp table first, as pl_breakdown does for fct_admissions_deal.

- Live procedures from unmerged branches differ from main: pl_breakdown 2026-10-02.1 (codex/alpha-enrollment-denominator, order already canonical), Guide and Facilities (#2124, applied after the first apply here; Guide keeps the canonical order), retention (#2049) and classification (older than main: lacks #1459).

- Live Facilities from #2124 locks xref_school_source before the posting fact, the reverse of classification. They run back to back in one AE run, so this only matters if two AE runs overlap; #2124 will fail this contract until it takes the posting fact first.

- The AE, Table 1 and Table 2 handlers do not retry a deadlock victim; the Table 3 client does. Classification, All Other, Facilities and Table 3 take no QuickBooks gate.

Live apply 1 (marts), 2026-10-02 07:54:27 to 07:54:50 UTC. Quiet window: nothing in the QuickBooks chain RUNNING or started in the last 10 min, no CALL of the five procedures, no DDL on the involved schemas in the previous 30 min (as visible to the pipeline DB user). CREATE OR REPLACE PROCEDURE as admin, one statement per procedure back to back (the Data API 100 KB statement limit rules out one combined statement): classification, All Other, Guide, QB budget QTD, Table 2. Owner, ACL, OID, SECURITY INVOKER and arguments identical before and after (owner CQL_download_OM; ACL CQL_download_OM=X/CQL_download_OM, Surtr_Service_User=X/CQL_download_OM). For the four whose live body matched main byte for byte the live body is byte-identical to this PR's file; classification live is older than main (#1459), so only the lock-block change was applied on top of the live body.

Live apply 2 (QuickBooks core writers), 2026-10-02 17:56:55 to 17:57:03 UTC. Last quickbooks-core-tables run ended 12:10; at 17:56:40 no pipeline in the QuickBooks chain was RUNNING or had started in the previous 15 min, no CALL of the chain was running, and no DDL had touched the involved schemas in the previous 30 min. quickbooks-raw-sync is scheduled once a day at 06:00 UTC (next 06:00 tomorrow); its other runs, and core-tables since #2123, are on demand, so I re-checked immediately before applying. Both writers were byte-identical to main in the live catalog before the apply. CREATE OR REPLACE PROCEDURE as admin, posting then budget detail back to back. Owner CQL_download_OM, ACL CQL_download_OM=X/CQL_download_OM ; Surtr_Service_User=X/CQL_download_OM, OIDs 16074835 and 17282951, SECURITY INVOKER and 9 arguments identical before and after; live body md5 now equals the PR file body for both (667ec6a7..., 1f3eaa44...). Catalog read of the live bodies: the coordinator's effective order is posting fact, dim_school, xref_school_source, and every pair involving an ungated live procedure agrees except the live Facilities pair noted above.

Since the first apply: no deadlock detected in pipeline_runs_prod. The 12:10 on-demand fan-out completed on all three marts and Table 3. The 10:56 fan-out's QB and Table 2 runs succeeded after the reader stall above; AE timed out on it and succeeded on re-run. The posting and budget writers have not yet run in their new order.

## Business Value

The School Performance reports (Tables 1, 2 and 3) and the Aerie school P&L marts are the finance team's view of school economics, and they all refresh off the same QuickBooks publication. Whenever that fan-out overlapped the school ontology refresh, a mart could abort on a lock deadlock, leaving its table stale until someone noticed the alert and re-ran it (on 2026-10-02 Table 2 and the AE P&L marts failed, and Table 3, which triggers off Table 2, never ran). This change puts every concurrent writer of the ontology tables, including the upstream QuickBooks core writers that feed all of those marts, on one lock order, and adds a CI check that discovers new procedures touching those tables so a future edit cannot reintroduce an inversion. That cuts alert noise, on-call triage time and the window in which leadership dashboards show stale numbers. It also documents, with evidence, the separate multi-minute stall caused by long-running readers, which is the next largest source of failed refreshes.

## Manual Effort Estimate

About 23 focused hours (roughly three working days) to do this by hand. The first revision was about 14 hours: reconstructing the deadlock timeline from the Redshift system history and mapping OIDs (3 h), reading the nine involved stored procedures for lock order and unlocked reads (3 h), working out a canonical order consistent with the producer, the QuickBooks writers and the live-versus-main drift (2 h), the five edits (1 h), the parser-based contract test (3 h), and the live apply with owner/ACL checks (2 h). The review follow-up adds about 9 hours: auditing the 34 procedures in the repo and 30 in the live catalog and classifying them (2 h), analysing the coordinator's single-transaction semantics and re-editing the two QuickBooks writers (2 h), reworking the contract test with discovery, gating and the coordinator's effective sequence (4 h), and the second live apply and verification (1 h). Proposed by Claude, Keval to confirm or adjust.

## Test plan

- [x] quickbooks-core-tables: uv run pytest 95 passed

- [x] mart-aerie-education-financials-refresh: uv run pytest 208 passed (61 are the contract tests)

- [x] mart-school-performance-quickbooks-refresh: uv run pytest 78 passed

- [x] mart-school-performance-unit-economics-refresh: uv run pytest 42 passed

- [x] mart-school-performance-unit-economics-per-student-refresh: uv run pytest 16 passed

- [x] ruff@0.15.22 check and ruff format --check clean on the touched Python file (no other Python changed)

- [x] Contract test run against the previous DDL (origin/main files in a throwaway worktree): 11 failures (50 pass), covering the QuickBooks core writers, the coordinator transaction and the marts changed earlier

- [x] Live apply 1: five marts, 07:54:27 to 07:54:50 UTC; owner/ACL/OID identical, body md5 equals the expected body

- [x] Live apply 2: two QuickBooks core writers, 17:56:55 to 17:57:03 UTC; owner/ACL/OID identical, body md5 equals the PR file body

- [x] Catalog read of live bodies: coordinator order and all ungated pairs agree, except the live Facilities version from #2124

- [x] No deadlock detected in pipeline_runs_prod since the first apply; 12:10 fan-out succeeded on all marts and Table 3

- [ ] First quickbooks-core-tables run after apply 2 completes with the new writer order

Linear: SURTR-1590

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1650 — fix(a8): a derived shadow output with programs left out is degraded, not clean (AERIE-2688) @kevalshahtrilogy  approved

## Summary

Follow-up to Mercy's review of release PR 1641 (posted after the release was deployed). ADMISSIONS_PROGRAM_DETAIL_READ is shadow in production, so each shadow line is evidence for the Gateway cutover. This PR fixes the one finding where a line could count as clean without having checked anything, and settles the projection finding with tests.

1. A derived output with programs left out is no longer clean (admissions-program-shadow.ts, finding at line 741)

checkDerived compares a derived output (derived:enrollment-snapshot, derived:pipeline-funnel, derived:projection) only for programs whose inputs all mapped on both transports. A program whose input failed was left out, and the line could still read clean. If every program was left out, the kit compared two empty sets and reported clean.

Now, a derived line that would otherwise count as clean is degraded when:

- any program was left out because an input failed on a transport (N of M program(s) had an input that failed on a transport, so their derived output was not compared), or

- no program was compared at all (no program's derived output was compared this cycle).

A mismatch stays a mismatch. A program the cycle made no read for is still left out without failing the line, since nothing was published from it. scope gains failedPrograms beside programs and skippedPrograms.

2. Projection values were always compared; now the code says so and tests pin it (finding at line 331)

The finding reads the projection shapes' values list as the set of compared fields. It is not: a RecordShape only says how records are keyed and which values an example may show. The kit compares every field of every record. Evidence: the new Q3, community metric and app conversion tests pass on the code before this PR, and the existing Q2 test already reported fieldMismatchCounts: { currentEnrollment: 1 }.

What changes: the three projection shapes now list every field as showable (they hold forecast counts and program identifiers, no personal data), so a differing value is named on the line instead of [redacted]. The lists are checked at compile time against the record types, so a new field cannot be left out. The module header and RecordShape now state that the shape never narrows the compare.

## What changes per mode

| Mode | Change |

|---|---|

| legacy | None. No legacy code path is touched. |

| shadow | Publishes exactly what it published before (the legacy reads). Only shadow lines change: derived lines can now be degraded where they were clean, and projection examples show values. |

| gateway | None. |

## Business Value

The shadow window is the evidence the team uses to decide when the analytics worker can stop reading Redshift for per-program admissions data. A line that says clean when a program was never compared makes that evidence look stronger than it is. After this change a clean derived line means every program the cycle read was compared and matched, so the cutover decision rests on checks that actually ran.

## Manual Effort Estimate

Proposed: about 3 hours of focused work by hand, without AI. Keval, please confirm or adjust. It covers reading the finding against the kit to separate the real defect from the display-only one, the verdict rule and its scope counts, the compile-time field lists, and 9 new tests plus two updated ones.

## Testing

- sync: pnpm typecheck exit 0, pnpm lint exit 0, vitest run --maxWorkers=2 exit 0 (121 files, 2753 tests).

- New tests in admissions-program-shadow.test.ts:

- every program left out because an input failed: both derived lines degraded, with the empty compare underneath still clean: true;

- an input that fails on the Gateway only also keeps the derived line from clean;

- no program at all: degraded;

- a program the cycle made no read for: left out, line still clean;

- a mismatch stays a mismatch when another program was left out;

- Q2, Q3, community metric and app conversion: a differing value is a mismatch.

- Ran the new value-compare tests against the unchanged main code: Q3, community metric and app conversion pass there; Q2 fails only on the example text (values were redacted).

## Not covered

- Absence versus empty in the keyed Gateway readers, and the dry-run verdict: separate PRs.

- No Surtr change, no production change, no environment change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1649 — Forecast V3: migrate consumers to canonical dbt fields @vvp-trilogy  approved

## Summary

- select canonical/V3 mart fields in the Forecast worker and derive presentation-only stage contributions locally

- remove rounded Session 1 pipeline/community intermediates from new contracts, Convex publications, API responses, and UI helpers while retaining stored-row rollout tolerance

- recompute Physical variants with the V3 decimal returning/new-student decomposition and final floor

## Validation

- pnpm typecheck

- contracts: 60 tests

- sync worker/refresh: 87 tests

- Convex publication + public API: 132 tests

- forecast UI: 47 tests

- OpenAPI: 14 tests

- pnpm lint:test-architecture

- pnpm lint:boundaries

- Biome check on all changed files

Closes #1645

#1646 — Fix Session 1 lock when historical observation is missing @vvp-trilogy  approved

## Summary

- evaluate Session 1 data-availability gates before transitioning a passed milestone to locked_actual

- keep missing historical observations unavailable with the controlled reason and null calculation/provenance

- add deterministic dbt unit coverage for missing and complete passed-milestone observations

- document the lock completeness requirement

## Validation

- poetry run dbt test --select int_admissions_forecast_resolved_session_1_lock_requires_observation

- poetry run dbt parse --no-partial-parse

Closes #1644

#1647 — fix(dbt): make SIS lifecycle data quality non-blocking @vvp-trilogy  approved

## Summary

- rename the SIS enrollment lifecycle assertion with the data_quality_ prefix

- configure the singular test at warning severity

- keep the lifecycle inconsistency visible without failing CI/CD

## Why

This assertion observes source-data quality and does not change production models. Running it as an error during PR CI blocks unrelated development on source data that the code change did not introduce. Warning severity preserves the signal while unblocking urgent delivery; production scheduled builds already exclude tests.

## Validation

- REDSHIFT_HOST=unused REDSHIFT_PASSWORD=unused poetry run dbt parse (with profiles.yml.docker)

- parsed manifest confirms data_quality_sis_enrollment_lifecycle_has_arrival has severity=warn

- git diff --check

#2127 — fix(timeback): adapt resources pages to gateway limits @caina-barbosa  approved

## Summary

This is a standalone production-incident fix for the resources entity failure observed in TimeBack Raw Sync run 699d8c8a-7507-4693-8ff6-8cf59da3518e.

It hardens incremental resources extraction against source-gateway response limits by using auditable, fail-closed page-size degradation while preserving keyset coverage and publication semantics. The exact PR head was deployed and validated successfully in production before this PR was opened.

Production effect: compatibility hardening.

---

## Why

resources responses vary dramatically in encoded size across later key ranges, so a fixed row limit can work for hundreds of pages and then repeatedly fail with HTTP 502. Static limits of 1,000 and 500 were both disproved by bounded production runs without publishing warehouse state. This change makes the existing incremental contract resilient to that payload variation without weakening closure, replay, or atomic publication checks.

---

## Business Value

- Restores healthy publication of the only stale TimeBack entity without re-running the other 20 healthy entities.

- Preserves exact immutable evidence for every successful page and every gateway-driven size transition.

- Keeps failures safe: exhausted retries at the minimum page size still fail before warehouse publication.

- Makes recovery reproducible from source-controlled runtime and manifest-validation behavior.

---

## How does it work

1. Incremental resources keyset extraction starts with a 500-record page limit; users retains 100 and ordinary entities retain 2,000.

2. After three retryable HTTP 5xx responses, the extractor lands the exact failed response and retries the same inclusive cursor at half the active limit.

3. The reduced limit is retained for subsequent pages, with deterministic transitions down to one record; another retryable 5xx at one record fails closed.

4. The manifest records the starting/minimum limits, every degradation receipt, and the actual limit used by each successful page.

5. Replay/transform validation verifies receipt checksums, ordered halving, cursor and offset binding, keyset continuity, count closure, and terminal confirmation before atomic raw/clean publication.

---

## Scope

### Included in this phase

- Entity-specific incremental keyset sizing for payload-heavy resources responses.

- Bounded, evidence-backed page-size degradation on exhausted retryable 5xx responses.

- Manifest and replay validation for adaptive page limits.

- Regression coverage for later-boundary failures, retained reduced limits, terminal confirmation, and the minimum-size fail-closed path.

- Exact final diff paths:

pipelines/runners/timeback-raw-sync/README.md

pipelines/runners/timeback-raw-sync/src/timeback_client.py

pipelines/runners/timeback-raw-sync/src/transform.py

pipelines/runners/timeback-raw-sync/tests/test_timeback_client.py

pipelines/runners/timeback-raw-sync/tests/test_transform.py

### Deliberately excluded for later phases

- Warehouse DDL or migration — table shape, grain, column types, lineage, and publication semantics are unchanged; existing source-controlled DDL remains sufficient to recreate the warehouse safely.

- Changes to full-snapshot partitioning, fan-out extraction, or source-limit isolation — those contracts are unchanged.

- Changes to entities other than the existing users cap and the new resources behavior.

- VALIDATION.md, generated CDK artifacts, local evidence files, and unrelated repository changes.

---

## Test plan

### Automated validation

- TimeBack full suite — 262 passed (uv run pytest -q)

- Ruff lint — passed (uv run ruff check src tests scripts)

- Ruff formatting — 28 files already formatted (uv run ruff format --check src tests scripts)

- git diff --check — passed

- exact-head diff scope — only the five authorized paths listed above

- commit trailers — no Co-authored-by trailers in any PR commit

### Time for Implementation

Approximately 2–3 engineer-days without AI assistance, including failure reproduction, test-first implementation, two bounded production falsification runs, adaptive redesign, deployment, and warehouse reconciliation.

### Manual QC

- Deployed exact commit eefd7af4d3d67fd87f0aedd75b4c9f730b8765be to Pipeline-timeback-raw-sync-prod as task definition revision 31.

- Verified image tag 863936842f9fcfd55ccb58ff7b47b2325f1b83721566d634f9accc44002d2cce and digest sha256:1ad3b3dea3715e7f648ae573cfea09ffb8d73b15132c67aff8ac210b3a13ad27.

- Ran only resources in incremental mode: Step Functions execution resources-adaptive-proof-20261002T142659Z, pipeline run da8aadc4-612c-42e1-9222-483c48a8804f.

- Observed production degradation at the same keyset traversal: 500 → 250 → 125, with two immutable HTTP 502 receipts; extraction then completed with 499,227 records across 1,830 pages.

- Verified manifest SHA-256 0443672bad9731e96b5a22decff058e03c7da312a4c921777d7fa53d9641982b, terminal confirmation, stored watermark overlap, and actual per-page limits.

- The run succeeded with 499,227 source/raw/clean delta rows, one published entity, zero failed entities, and exit code 0.

- Redshift reconciliation: raw and clean each contain 2,185,849 rows; exactly 499,227 rows in each carry the recovery run ID; null keys, duplicate keys, and raw/clean key-set differences are all zero.

- The clean watermark advanced from 2026-09-13 09:10:24.707 to 2026-10-02 14:27:22.923.

- Both failed proof runs produced zero ingestion-ledger rows, and no non-resources entity published during the successful run.

#1643 — Align next-year enrollment forecast formula @vvp-trilogy  approved

## Summary

- subtract start-year transfers out from undecided returners

- floor the undecided-returner population at zero

- document the mutually exclusive returner groups

## Validation

- git diff --check

- documentation-only change

#1585 — feat(a8): G6 gate SIS_ENROLLMENT_READ — SIS rollup input + members via Surtr Gateway, shared source_run_id, stale_copy rule (AERIE-2624) @kevalshahtrilogy  approved

## Summary

A8 unit U23 ([AERIE-2624](https://linear.app/builder-team/issue/AERIE-2624)): the read gate for G6 SIS enrollment. refreshSisEnrollment reads the Aerie dbt mart sandbox_education.mart_enrollment_dtl once per cycle, at two grains inside one transaction (querySisEnrollmentSnapshot): the rollup grid cells and the student members. Surtr's U19 (AI-Builder-Team/Surtr#2098) copies both reads into mart_education.aerie_sis_enrollment_rollup_input and aerie_sis_enrollment_member in one procedure and one transaction. This PR lets the worker read those copies over the Surtr Gateway.

- Gate: SIS_ENROLLMENT_READ=legacy|shadow|gateway, default legacy. It is inert (stays legacy, with a WARN) unless DBT_TARGET is exactly production, because the copies only exist for the production relation.

- The two slugs must describe one snapshot, so the gate has one source, snapshot. Both slugs always share a mode, and there is no way to mix transports.

- Wiring: through the existing injectable querySnapshot dependency of refreshSisEnrollment, a one-line change there. With the env unset, the read is exactly querySisEnrollmentSnapshot.

- Gateway readers (A8 kit, #1566): one joint read of aerie-sis-enrollment-rollup-input and aerie-sis-enrollment-member.

- The pair must share one source_run_id (assertSharedLineage) and one build marker and publication time. Anything else is refused in both shadow and gateway mode, so rollups and members always come from one Surtr publication, as the legacy transaction guarantees.

- The kit's per-source rules also apply: type parity, absent column fails, population floors (500 cells / 1,000 members; 1,741 / 5,303 on 09-29), 6h max age, unique mart_row_id, one build marker.

- Both copies are mapped with the legacy mappers: the cells with the same schema, the members with the same strict parse. The legacy cell parse silently drops a cell that fails its schema; on the Gateway path any such drop instead refuses the read (degraded in shadow), because Surtr publishes the copy whole. Gateway mode then runs the legacy roll-up, coverage and first-day partition checks on the Gateway cells.

- Shadow publishes the pg read unchanged and logs one a8_shadow_check line per slug:

- rollup input: parsed cells keyed by (program_code, session_school_year, cohort_id, x_pipeline), before the order-dependent TS roll-up. A code with two names in one year is compared as a multiset under its key.

- members: after collapseSisEnrollmentMembership (order-independent), keyed by its grain (programCode, schoolYear, cohortId, xPipeline, sisStudentId).

- stale_copy dbt skew rule: the pg read is bracketed by the relation's pg_class OID marker (the one Surtr records). A copy that lags a newer dbt build is skipped, not counted as a mismatch. A build swapped in mid-read is degraded.

- PII (minors): no value allowlist on either compare. Every value and key is [redacted], so the logs carry field names, value types, counts and lineage only.

- Legacy file sync/src/redshift/sis-enrollment.ts:

- It exports rollupsSql, membersSql, the cell parse (split from the roll-up) and buildMembers.

- It adds querySisEnrollmentSnapshotWithCells. querySisEnrollmentSnapshot now delegates to it, with the same transaction, statement order and throws.

- One deliberate change (Mercy round 1): when a program code carries two names in one school year, the roll-up keeps the smallest name instead of the first row's. Neither the SQL nor the Gateway copy is ordered, so first-row-wins could publish different names per transport. No (code, year) in today's mart carries two names (read-only check, 09-29), so current output is unchanged, and the proof hashes below are identical before and after the change.

- .gitattributes: sis-enrollment.ts contains a literal NUL (its group-key separator), so git diffed it as binary. A diff attribute makes local and CI git diff show it as text. The NUL itself is unchanged. GitHub's PR view still renders this file as binary, so its text diff is inlined below.

- Dry-run: sync/src/scripts/dry-run-sis-enrollment-shadow.ts, which loads dotenv first and then uses await import().

<details><summary>Text diff of <code>sync/src/redshift/sis-enrollment.ts</code> (GitHub shows it as binary)</summary>

diff --git a/sync/src/redshift/sis-enrollment.ts b/sync/src/redshift/sis-enrollment.ts

index e4e40b438..5fe1d332f 100644

--- a/sync/src/redshift/sis-enrollment.ts

+++ b/sync/src/redshift/sis-enrollment.ts

@@ -64,7 +64,21 @@ const SisEnrollmentCellRowSchema = z.object({

student_count: z.coerce.number(),

});

-function rollupsSql(relation: string): string {

+/** The dbt model both SIS reads select from; resolveDbtRelation picks its relation for DBT_TARGET. */

+export const SIS_ENROLLMENT_DBT_MODEL = {

+ schema: "sandbox_education",

+ model: "mart_enrollment_dtl",

+} as const;

+

+/** One parsed rollupsSql row: a (program, name, year, cohort, x_pipeline) grid cell. */

+export type SisEnrollmentCellRow = z.infer<typeof SisEnrollmentCellRowSchema>;

+

+/**

+ * Surtr's mart_education.aerie_sis_enrollment_rollup_input (A8, Gateway slug

+ * aerie-sis-enrollment-rollup-input) copies this SQL's output from the production relation, so

+ * its text is a contract: change it only together with that mart.

+ */

+export function rollupsSql(relation: string): string {

// COUNT(DISTINCT CASE WHEN has_fact THEN student_id END): factless rows keep a

// program/year/cohort present in the grid without counting as students. The

// x_pipeline group splits the first-day cohort into re-enrollment vs

@@ -81,9 +95,23 @@ function rollupsSql(relation: string): string {

GROUP BY 1, 2, 3, 4, 5;

}

+/**

+ * Parses rollupsSql rows as the pg driver returns them (a row that fails the schema is dropped

+ * and counted, as before). Shared by the legacy read and the A8 Gateway read

+ * (analytics/sis-enrollment-gateway.ts), so both transports parse identically.

+ */

+export function mapSisEnrollmentCellRows(rows: readonly unknown[]): SisEnrollmentCellRow[] {

+ return safeParseRows("querySisEnrollmentRollups", [...rows], SisEnrollmentCellRowSchema);

+}

+

function buildRollups(rows: unknown[]): SisEnrollmentRollup[] {

- const parsed = safeParseRows("querySisEnrollmentRollups", rows, SisEnrollmentCellRowSchema);

+ return buildSisEnrollmentRollups(mapSisEnrollmentCellRows(rows));

+}

+/** Rolls parsed grid cells into the ten report metrics, refusing incomplete coverage. */

+export function buildSisEnrollmentRollups(

+ parsed: readonly SisEnrollmentCellRow[],

+): SisEnrollmentRollup[] {

// Group cells by (program, year). programName is carried alongside the code

// so the report can fall back to the mart's display name.

type Group = { programCode: string; programName: string; schoolYear: string };

@@ -101,6 +129,11 @@ function buildRollups(rows: unknown[]): SisEnrollmentRollup[] {

cells: [],

};

groups.set(groupKey, group);

+ } else if (row.program_name < group.key.programName) {

+ // The SQL has no ORDER BY, so when a code carries two names in one year the smallest

+ // name is kept, not the first to arrive: the same rows publish the same name, whichever

+ // order (or transport) they arrive in.

+ group.key.programName = row.program_name;

}

group.cells.push({

cohortId: row.cohort_id,

@@ -191,7 +224,12 @@ const SisEnrollmentMembershipRowSchema = z.object({

hold_date: z.string().nullable(),

});

-function membersSql(relation: string): string {

+/**

+ * Surtr's mart_education.aerie_sis_enrollment_member (A8, Gateway slug

+ * aerie-sis-enrollment-member) copies this SQL's output from the production relation, so its

+ * text is a contract: change it only together with that mart.

+ */

+export function membersSql(relation: string): string {

// Fact rows only (has_fact) — factless grid placeholders carry NULL identity

// and never represent a student. TRIM(BOTH '"' …) on program_code/name

// mirrors the rollup reader so both agree on the program key.

@@ -222,13 +260,14 @@ function membersSql(relation: string): string {

AND cohort_id IN (${reportCohortList()});

}

-function buildMembers(rows: unknown[]): SisEnrollmentMembershipRow[] {

+/** Maps membersSql rows strictly; shared by the legacy read and the A8 Gateway read. */

+export function buildMembers(rows: readonly unknown[]): SisEnrollmentMembershipRow[] {

// Strict parse: a malformed row must fail the whole membership read (which the

// refresh turns into a skipped publication preserving last known-good) rather

// than silently dropping students and breaking reconciliation.

const parsed = parseRowsStrict(

"querySisEnrollmentCohortStudents",

- rows,

+ [...rows],

SisEnrollmentMembershipRowSchema,

);

@@ -281,16 +320,27 @@ export async function querySisEnrollmentSnapshot(): Promise<{

rollups: SisEnrollmentRollup[];

members: SisEnrollmentMembershipRow[];

}> {

- const relation = resolveDbtRelation({

- schema: "sandbox_education",

- model: "mart_enrollment_dtl",

- });

+ const { rollups, members } = await querySisEnrollmentSnapshotWithCells();

+ return { rollups, members };

+}

+

+/**

+ * querySisEnrollmentSnapshot's read, also returning the parsed grid cells the rollups were built

+ * from, so the A8 shadow compare (SIS_ENROLLMENT_READ=shadow) can compare them.

+ */

+export async function querySisEnrollmentSnapshotWithCells(): Promise<{

+ cells: SisEnrollmentCellRow[];

+ rollups: SisEnrollmentRollup[];

+ members: SisEnrollmentMembershipRow[];

+}> {

+ const relation = resolveDbtRelation(SIS_ENROLLMENT_DBT_MODEL);

return withTransaction(async (exec) => {

const run: SqlRunner = async (sql) => (await exec(sql)).rows;

// Sequential (not Promise.all) — both statements must run on the single

// fenced transaction connection, which processes one query at a time.

- const rollups = buildRollups(await run(rollupsSql(relation)));

+ const cells = mapSisEnrollmentCellRows(await run(rollupsSql(relation)));

+ const rollups = buildSisEnrollmentRollups(cells);

const members = buildMembers(await run(membersSql(relation)));

- return { rollups, members };

+ return { cells, rollups, members };

});

}

</details>

## Business Value

The SIS enrollment report is one of the last Aerie refresh domains that reads Redshift directly from the EC2 analytics worker. Moving it behind a Surtr Gateway read is a step toward retiring that direct warehouse access (AERIE-445, "Remove Sync from EC2"). The gate lets the cutover be proven in production before anything changes:

- shadow compares every cycle, with zero effect on what is published;

- gateway mode can be switched on, and rolled back to legacy, with one env var.

The shared-lineage guard means the report can never publish rollups from one Surtr copy and student rows from another. And because the shadow lines redact everything, a clean-window check never puts minors' data in logs.

## Manual Effort Estimate

Proposed: ~14 hours of focused work — Keval, please confirm or adjust. That covers:

- reading the kit and U19's contract;

- the sql/mapper split without changing behaviour;

- the joint-read and lineage design;

- a synthetic mart fixture that reconciles end to end;

- about 38 tests;

- the read-only proof harness.

## Testing / evidence

Unit tests: sync/src/analytics/sis-enrollment-gateway.test.ts, 38 tests. They cover:

- the legacy split: one transaction, same SQL and order, and a coverage throw still skips the member query;

- the readers: pg-identical cells and members; a copied cell failing the legacy schema refuses the read (legacy's drop unchanged); spec columns equal to the legacy SQL aliases and types; BIGINT parity; absent or ill-typed columns failing without values; the strict member mapper; short, stale, torn and mixed-build copies;

- shared lineage: a different source_run_id, marker or publication time between the two slugs is refused;

- gate modes and overrides, and DBT_TARGET unset, pr:<N> or blank being inert;

- shadow: clean, multiset name case, stale_copy, build swap or failed lookup, member and cell mismatches (redacted, no PII in lines), one-side keys, mixed-lineage and Gateway failures degraded and swallowed, legacy failure unchanged;

- gateway: publishes the same rollups and members, fails closed with one operator-safe line, runs the coverage check on the Gateway cells, and publishes the same program name as legacy for a code with two names whatever the row order;

- wiring: the default querySnapshot is the legacy transaction when unset, and routes through the gate via process.env;

- legacy, shadow and gateway send Convex identical rollups and student rows through refreshSisEnrollment.

Other checks:

- sync typecheck clean.

- pnpm lint exit 0 (2 pre-existing warnings elsewhere).

- Full sync vitest run with --maxWorkers=2: 87 files, 1,559 tests passed.

Read-only proof against Redshift (no Gateway, no key, no writes). Legacy pg read vs the gateway path, where the fake Gateway was fed by U19's candidate SQL. That SQL was extracted verbatim from the 2098 diff and run read-only in a BEGIN READ ONLY transaction, with the procedure's mart_row_id and lineage, serialized the way the Data API returns it (BIGINT as a number, TIMESTAMPTZ as text), ordered by mart_row_id. The output is counts and hashes only:

U19 SQL == Aerie SQL (whitespace-normalized): rollups=true members=true

U19 copy rows: rollup_input=1741 member=5303 duplicate_mart_row_id=0/0

shadow aerie-sis-enrollment-rollup-input: clean pg=1741 gw=1741 matched=1741 mismatched=0 pgOnly=0 gwOnly=0

shadow aerie-sis-enrollment-member: clean pg=5303 gw=5303 matched=5303 mismatched=0 pgOnly=0 gwOnly=0

cells: legacy n=1741 sha=fb39a0a284f25b91 | gateway n=1741 sha=fb39a0a284f25b91 | equal=true

rollups: legacy n=214 sha=cbb2d12adee0d63f | gateway n=214 sha=cbb2d12adee0d63f | equal=true

members: legacy n=5303 sha=2e6b2e4f88fe1f23 | gateway n=5303 sha=2e6b2e4f88fe1f23 | equal=true

refresh result: legacy={"count":214} gateway={"count":214}

Convex rollups payload: equal=true (n=214)

Convex student rows payload: equal=true (n=5303, sha=f36d172fedf439e2)

## Stack note

- Base: feat/a8-u02-read-gate-kit. This PR depends on #1566 (U02 kit) and on AI-Builder-Team/Surtr#2098 (U19 marts; their DDL is the contract these specs mirror).

- Live dry-run waits on:

- U19 deployed;

- the slugs registered (U04);

- a granted SURTR_GATEWAY_API_KEY;

- the PII sign-off (plan §9).

- Collision MEDIUM with Vladimir's SIS work (#1299, #1303). The existing-file diff is kept to exports, a split and one wiring line. Keval to give Vladimir a heads-up: while the shadow window is open, changes to rollupsSql / membersSql must be mirrored in Surtr's sp_refresh_aerie_sis_enrollment.

- Rebase hotspot: the .env.example line (other A8 units add neighbours).

## Not covered

- No live Gateway run. The marts are not deployed and the key is not granted. The proof above uses U19's SQL, not the deployed procedure.

- refreshSisEnrollment still requires Redshift to be configured in gateway mode; that guard is unchanged. Dropping Redshift from the worker is later A8/A9 work.

- The pg_class-OID build-marker lookup duplicates U20's (#1580). Hoisting it into the kit is a possible follow-up.

- The upstream writer of staging_education_ai_horizons.raw_* is still unresolved (plan §8), and so is the sandbox_education ownership by a personal user.

- The flip sequence (shadow, then a clean window, then gateway, then 7 days, then delete legacy) is operational and is not part of this PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2125 — feat(aerie-a8): Forecast V2 copy, forecast + grade operands in one transaction (SURTR-1594) @kevalshahtrilogy  approvedmercy-allow-critical

A8 unit U24. Linear: [SURTR-1594](https://linear.app/builder-team/issue/SURTR-1594/a8-u24-forecast-v2-dbt-publication-copy-forecast-grade-operands)

Hold the merge. The forecast's owner says its shape is about to change (see "The forecast shape is still changing"). This PR does not schedule the new procedure: it ships on demand only, so the live schedule cannot call it before its DDL exists (see "Pre-merge DDL").

## What this does

Publishes Surtr copies of Aerie's two Forecast V2 reads, so the Forecast V2 refresh can move from direct Redshift onto the Surtr Gateway. Both Gateway slugs are already registered; this PR gives them their tables.

| Mart | Gateway slug | Aerie read | Rows today |

|---|---|---|---|

| mart_education.aerie_admissions_forecast_v2 | aerie-admissions-forecast-v2 | forecastSql | 116 |

| mart_education.aerie_admissions_forecast_v2_grade_operand | aerie-admissions-forecast-v2-grade-operand | the grade operand read | 1,062 |

One procedure, mart_education.sp_refresh_aerie_admissions_forecast_v2, publishes both marts in one transaction under one source_run_id and one source_build_marker. It follows the SIS pair on the same runner (092_sp_refresh_aerie_sis_enrollment.sql).

## Files

- pipelines/cdk/sql/mart_education/100_aerie_admissions_forecast_v2.sql: the forecast table. 141 columns (the ones forecastSql selects, of the 187 in sandbox_education.mart_admissions_forecast) plus the 6 lineage columns.

- .../101_aerie_admissions_forecast_v2_grade_operand.sql: the grade operand table. 6 columns plus lineage.

- .../102_sp_refresh_aerie_admissions_forecast_v2.sql: the procedure, sole writer of both.

- Runner mart-aerie-dbt-publication-refresh: a new ON_DEMAND_PROCEDURES list that holds the procedure until it is scheduled, one CONTRACT_MARTS entry, three DDL_FILES, the reconciliation SQL, tests and README.

## The procedure is deployed, not scheduled

The runner's 10-minute schedule is live and its DDL is applied out of band. Mercy flagged that appending the procedure to REFRESH_PROCEDURES lets the schedule call it before the DDL exists, with only a README telling people the order. So the handler now reads a second list:

- REFRESH_PROCEDURES: what every scheduled run calls. Unchanged by this PR.

- ON_DEMAND_PROCEDURES: deployed, but called only by a run that names the procedure in params.procedures. The Forecast V2 procedure is here.

Scheduling it is a follow-up PR that moves the name across, citing the applied DDL, the on-demand run and the reconciliation. That is the same gate the runner itself used (it shipped with its schedule disabled, and Surtr PR 2105 enabled it with that evidence).

## How it differs from the SIS pair

- Two source relations, not one. The build marker names both builds: pg_class_oid:<forecast oid>,<grade operand oid>. A new build of either relation republishes both marts. Each mart's source_published_at is its own relation's creation time.

- The no-op is decided on the rows. The procedure builds its candidates on every run (about 1,200 rows) and publishes nothing only when both marts carry the current marker from one run and their rows are exactly the candidates, compared as whole rows in both directions. A mart with a row deleted, added or changed outside the procedure is published again on the next run. The SIS and pipeline-detail procedures decide on lineage alone; they are live and not in this PR.

- A copy taken between the two dbt builds is published as-is. dbt builds the forecast, then the grade operands 13 and 28 seconds later in the two builds seen on 2026-10-02. A run inside that gap copies the new forecast with the previous grade operands, which is what a legacy read at that moment returns. Aerie's own rule (use the grade operands only when they name the forecast's latest calculated_at) then applies to the Gateway rows unchanged, and the next run copies both. I chose this over failing the run, because Aerie deliberately publishes the forecast without grade operands in that state; failing closed would hold back a fresh forecast whenever the grade model fails to build.

- A NaN or infinite rate is refused. The three rates are DOUBLE PRECISION. JSON cannot carry NaN or Infinity, so the Gateway would deliver NULL and Aerie's parse, which rejects NaN on the legacy read, would accept it.

- No candidate comparison in the reconciliation. Both candidates are Aerie's SQL unchanged, so comparing them with Aerie's SQL would compare a query with itself. All five queries check the published marts.

## The forecast shape is still changing

Vladimir (2026-10-02): the mart shape is not stable and the forecast will change a lot shortly. Two things here make that cheaper and keep it loud:

- NUMERIC columns pin their scale, not their precision. dbt infers 18, 21 or 25 digits from each formula. The marts hold NUMERIC(38,6) and NUMERIC(38,4), and the coupling guard compares the scale only, as it already ignores VARCHAR width. A precision change then needs no migration. A scale or type change still fails the procedure.

- The procedure checks that each mart has exactly the columns it fills (147 and 12). A table migrated without its procedure fails the run instead of publishing the new column as NULL.

The contract tests check every place that lists the columns (both tables, the candidates, both inserts and their selects, the row-id ordering, the guards, the reconciliation) against the Aerie SQL pinned in tests/test_sql_contracts_forecast_v2.py. The README has a "When the Forecast V2 shape changes" checklist.

A dbt change that reaches production before a migration here is not published wrong: a dropped, renamed or retyped column fails the procedure and the previous publication stays.

Open Aerie PR 1609 ("Polish forecast calculation rollups") touches two chat/ UI files only, so it does not change this contract. Open Aerie PR 1372 (a draft, last updated 2026-09-18) changes only a comment in the reader, not its SELECT list.

## Read-only evidence (production, SELECT only)

No DDL was applied and nothing was called. Where a check names a query, guard or candidate, it ran the committed SQL text.

| Check | Result |

|---|---|

| Aerie's forecastSql on sandbox_education.mart_admissions_forecast | 141 columns, 116 rows |

| Aerie's grade operand read | 6 columns, 1,062 rows |

| Reconciliation Queries 2 and 3, with each mart replaced by the procedure's own candidate SELECT cast to the table's types | aerie_minus_mart 0 and mart_minus_aerie 0 for both |

| Same, with two BIGINT columns swapped on purpose | 75 rows differ each way |

| mart_row_id | 116 of 116 and 1,062 of 1,062 distinct |

| Table types against the live dbt catalog, by the guard's rule | 0 drift over 127 and 5 raw columns |

| The guard and column-count SQL, run against the existing aerie_admissions_pipeline_detail mart | 0 drift; 6 when its cast exemptions are removed; 59 columns |

| NUMERIC(38,s) against the dbt value, as Data API text | 576 of 576 identical |

| NaN or infinite rates today | 0 |

| The no-op's whole-row comparison (SELECT * ... EXCEPT SELECT * ..., all 147 and 12 columns) on candidate stand-ins | 0 both ways when identical; 1 when a row is removed |

| Table comment lengths on the sibling marts in production (pg_description) | 1,229, 1,572 and 1,841 characters, so Redshift has no 256-character comment limit; the new table comments are 1,729 and 1,802 |

| Build marker query (Query 1) | pg_class_oid:21037866,21037884 |

Not verified, because it needs DDL: the CREATE TABLE, CREATE PROCEDURE and the CALL itself.

## Business Value

Forecast V2 is the admissions enrollment forecast leadership reads in Aerie. Today its refresh reads Aerie's dbt relations straight from Redshift on an EC2 worker. This is the Surtr half of moving that read onto the Surtr Gateway, part of the A8 migration of Aerie's analytics worker reads onto Surtr: the two Forecast V2 Gateway sources get their tables. Aerie can then run its shadow comparison on Forecast V2 and cut over, so the forecast read goes through one governed, keyed path instead of a direct warehouse connection. The copy keeps the single-snapshot guarantee Aerie has today, so the cutover cannot pair a forecast with grade operands read at another moment.

## Manual Effort Estimate

About 3 days of focused work by hand, no AI. Proposed number; Keval to confirm or adjust.

- 0.5 day: read the Aerie reader and dbt models, and work out which 141 of the 187 columns it selects and their catalog types.

- 1 day: the two tables and the 1,300-line procedure, including the two-relation marker.

- 0.5 day: reconciliation SQL and read-only parity checks against production.

- 1 day: contract tests, runner wiring (including the on-demand gate) and README.

## Pre-merge DDL

The runner's 10-minute schedule is live. In the first version of this PR the procedure was appended to REFRESH_PROCEDURES, so the DDL had to be applied before merge. It is now on demand only, so the order no longer matters for the schedule: the DDL can be applied before the merge, as planned, or after it. It must be applied before the first on-demand run and before the follow-up PR that schedules the procedure.

Apply these three files, in this order (tables before the procedure):

1. pipelines/cdk/sql/mart_education/100_aerie_admissions_forecast_v2.sql

2. pipelines/cdk/sql/mart_education/101_aerie_admissions_forecast_v2_grade_operand.sql

3. pipelines/cdk/sql/mart_education/102_sp_refresh_aerie_admissions_forecast_v2.sql

cd pipelines/runners/mart-aerie-dbt-publication-refresh

REDSHIFT_CLUSTER_IDENTIFIER=redshift-cluster-1 REDSHIFT_DATABASE=finance_dw REDSHIFT_DB_USER=CQL_download_OM \

uv run python scripts/apply_ddl.py \

../../cdk/sql/mart_education/100_aerie_admissions_forecast_v2.sql \

../../cdk/sql/mart_education/101_aerie_admissions_forecast_v2_grade_operand.sql \

../../cdk/sql/mart_education/102_sp_refresh_aerie_admissions_forecast_v2.sql

Add --dry-run to print the 44 statements (26, 14 and 4) without sending them. Passing the three paths applies only these files; with no paths the script re-applies the runner's earlier files too, which is idempotent but replaces live procedures.

Then, once this PR is deployed and the DDL is applied:

1. Run on demand with {"procedures": ["mart_education.sp_refresh_aerie_admissions_forecast_v2"]}. Expect two published results, then unchanged on a second run.

2. Run reconciliation/aerie_admissions_forecast_v2_vs_aerie_sql.sql with psql. Query 5 must report forecast_parity, grade_operand_parity and one_publication as PASS.

3. Open the follow-up PR that moves the procedure from ON_DEMAND_PROCEDURES to REFRESH_PROCEDURES. Until it merges, the copy is refreshed only when someone runs it.

## Tests

- uv run pytest in mart-aerie-dbt-publication-refresh: 229 passed. 55 are the new contract file; 17 are new handler, manifest and DDL tests.

- mart-aerie-admissions-refresh (its tests glob the same DDL directory): 556 passed. mart-aerie-expenses-refresh: 91 passed.

- ruff check pipelines and ruff format --check pipelines (0.15.22): clean.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1580 — feat(a8): G2 gate ADMISSIONS_PIPELINE_READ — admissions pipeline detail + tenant crosswalk via Surtr Gateway, stale_copy rule (AERIE-2620) @kevalshahtrilogy  approved

## Summary

A8 unit U20 (AERIE-2620): the read gate for the G2 Admissions Pipeline report. refreshAdmissionsPipeline reads two sources once per cycle:

- every row of the Aerie dbt mart sandbox_education.mart_admissions_pipeline_dtl (30,929 rows on 2026-09-29, key pipeline_key, PII);

- the EduCRM Finalsite tenant → program_code crosswalk (57 pairs).

This PR lets both reads run in shadow against Surtr's copies, and later cut over, without changing what legacy publishes.

- Split, no behaviour change. queryAdmissionsPipelineRows becomes admissionsPipelineRowsSql(relation) + mapAdmissionsPipelineRows, and queryAdmissionsPipelineCrosswalk becomes ADMISSIONS_PIPELINE_CROSSWALK_SQL + mapAdmissionsPipelineCrosswalkRows. The SQL text is byte-identical, the strict parse is unchanged, and the dbt relation is still resolved per call. This is the only edit to the legacy query file.

- Gateway readers (sync/src/analytics/queries/admissions-pipeline-gateway.ts, on the U02 kit), for the two sources in AI-Builder-Team/Surtr#2096 (U16):

- aerie-admissions-pipeline-detail: every legacy output column is typed as the copy stores it: 5 BOOLEAN, 3 NUMERIC(18,2), and everything else VARCHAR, including the six ::text date casts. source_build_marker is read as the kit's build marker. Population floor: 10,000 rows. Freshness uses the kit's 6h default on source_published_at (when dbt created the relation; dbt builds hourly).

- Gateway rows are re-sorted to the legacy ORDER BY source_system, program_code, stage_id (binary UTF-8, NULLS LAST) before the mapper, as the SQL does.

- aerie-admissions-pipeline-tenant-crosswalk: both columns VARCHAR. Floor: 20 pairs.

- Gate: ADMISSIONS_PIPELINE_READ=legacy|shadow|gateway (default legacy), with the overrides …_PIPELINE_DETAIL and …_TENANT_CROSSWALK. It is wired into refreshAdmissionsPipeline (admissions/pipeline-refresh.ts:388) as an optional argument that resolves the env once per cycle.

- Inert unless DBT_TARGET=production. Both sources are declared dbt-backed. The crosswalk is EduCRM, but it only exists to key the dbt detail rows, so the report keeps one transport rule.

- In gateway mode a failure fails the pipeline domain the same way a legacy read failure does: the run is not published and the last published run stays.

- Shadow compare.

- Detail, by pipeline_key over the mapped records (every column the TS reads). PII (names, emails, phones, child DOB and gender), so there is no value allowlist: every value and key in an example reads [redacted], and the a8_shadow_check line carries field names, value types, counts and lineage only.

- dbt stale_copy rule. The pg-side build marker is the relation's pg_class_oid:<oid>, the same lookup and text form Surtr's copy procedure records. When it differs from the copy's source_build_marker, the copy lags a newer dbt build: the outcome is stale_copy (skipped, not a mismatch). The marker lookup brackets the unchanged legacy read: dbt gives every build a new OID, so the same marker before and after proves which build the rows came from. If a build is swapped in mid-read, or the lookup fails, only the shadow check degrades; the published read never changes.

- Crosswalk, as a set of pairs, with the EduCRM source_advanced re-read. Tenant ids and program codes are organisations, not people, so examples may show them.

- Dry-run: sync/src/scripts/dry-run-admissions-pipeline-shadow.ts (needs DBT_TARGET=production).

## Business Value

- The Admissions Pipeline report can move onto a lineage-stamped copy. Today the EC2 worker reads the Aerie dbt relation directly. After this change it can read a Surtr-owned copy that records which dbt build it holds, through one env var: shadow, then gateway, with legacy as the rollback. That is one more A8 group ready for its shadow window on the way to AERIE-445.

- The shadow window will not cry wolf on the copy's lag. The copy trails each hourly dbt build by up to 10 minutes. The stale_copy rule skips those cycles instead of reporting them as mismatches, and it is proven on live data below.

- No personal data in logs. The detail compare is built so that no row value can reach a log line.

## Manual Effort Estimate

About 12 hours of focused time to build by hand without AI. That covers:

- reading the kit, the U05/U09 pattern and the U16 mart contracts;

- the split;

- the readers, gate and marker bracket;

- about 35 tests with transport-shaped fixtures, including the cross-mode Convex payload check;

- the dry-run;

- the read-only parity proofs.

Keval: please confirm or adjust this number.

## Testing / evidence

- Unit tests: sync/src/analytics/queries/admissions-pipeline-gateway.test.ts, 31 tests, on a 10,050-row transport-shaped fixture (above the 10,000 floor). They cover:

- Reader type parity: the same records, in the same order, as the pg read. NUMERIC decimal text, booleans and the ::text dates reach the mapper as pg's values, and shadow_appointments still parses. The spec's columns equal the SQL's select list; the boolean, numeric and ::text column sets are pinned.

- Reader fail-closed cases: an absent column; a wrong-typed value (the error names no value); the strict mapper's one-bad-row throw; a snapshot below the floor, older than 6h, from two source_run_ids, with two source_build_markers, or with a duplicate mart_row_id. The crosswalk floor.

- Order comparator: NULLS LAST, code point (binary UTF-8) order including an astral-vs-BMP case, and prefixes.

- Build marker lookup: the SQL and its parameters (the production relation), the pg_class_oid: text form, and null for no relation or a blank marker.

- Gate modes: unset (exactly the legacy calls; no Gateway call, marker lookup or log line). DBT_TARGET inertness: unset, pr:1433 and blank all keep both sources legacy in shadow and gateway, with a warning per source. With production: global mode, per-source overrides, pg as legacy, and typo capping.

- Shadow, detail: clean when the builds match; stale_copy when the pg read saw a newer build (not reported as a mismatch even though the data differs); a build swapped mid-read, or a failed marker lookup, degrades the check only; a PII mismatch is reported by field name and count, with every key and value [redacted] and a PII-marker absence check on the full log line; a one-sided key and a duplicate pipeline_key are never clean; a Gateway failure is degraded and swallowed; a legacy failure fails exactly as in legacy mode.

- Shadow, crosswalk: the same pairs in another order are clean; a pair on one side only is a mismatch that shows the pair; source_advanced when the re-read matches.

- Gateway: publishes the Gateway records in the legacy order and never runs the legacy reads; a failure rethrows after one operator-safe line.

- refreshAdmissionsPipeline wiring: it reads through the gate it is given; unset, it runs exactly the two legacy SQL texts; legacy, shadow and gateway send Convex identical payloads, in identical order (10,050 rows on real stage ids, shadow using the default marker lookup, both shadow checks clean); a gateway failure sends nothing to Convex.

- admissions-pipeline.test.ts has two new tests: each query is its SQL mapped by its mapper, and the strict parse lives in the mapper.

- Typecheck: tsc --noEmit (sync) passes, as does the pre-commit typecheck-sync. Lint: pnpm lint passes; its only 2 warnings are pre-existing, in chat/skill/forge-api/scripts/sindri.mjs.

- Sync tests: vitest run --maxWorkers=2, 87 files, 1,554 tests pass. The chat suite was not run.

- Read-only pg proof that the split is behaviour-identical (default_transaction_read_only; counts and hash prefixes only). The pre-split module (kit head d6f5958) and the split module were run against Redshift within one dbt build (the same pg_class_oid marker before and after):

- The SQL text is byte-identical for both queries.

- Detail: 30,929 vs 30,929 rows, multiset hash cf29a42f2bfc5658 on both, and 0 of 30,929 positions differ in the ORDER BY key.

- Crosswalk: 57 vs 57 pairs, hash c1ad71c594711793 on both.

- Old code vs the gateway path, fed by U16's mart SQL run read-only. The two aerie-sql candidate blocks from AI-Builder-Team/Surtr#2096 (072, 074) were run as plain SELECTs, shaped as the Gateway returns the copy (lineage columns added, the real pg_class_oid marker and relcreationtime, paged in a hash order unrelated to the legacy order) and fed through the real readers:

- Detail: 30,929 U16 rows → 30,929 Gateway-path records. Keyed compare against the legacy read: 30,929 matched, 0 mismatched, 0 one-sided, 0 duplicates, no field mismatches. After the re-sort, 0 positions differ from the legacy ORDER BY key sequence.

- Crosswalk: 57 matched, 0 one-sided.

- What each path would publish, built by the unchanged refresh functions: resolved rows 30,929 on both (0 dropped), the same resolved-row multiset hash (51bca6ecb2d69eca) and the same Physical-mode cells (2562799d919e8258). The Program-matrix cells are the same 778 cells (multiset hash 7b9bc52bf55a8df2 on two legacy reads and on the gateway path). Only their insertion order differs, inside 24 ORDER BY tie groups (rows sharing source_system, program_code, stage_id that resolve to different cells): Redshift leaves ties unordered, and the copy pages them by mart_row_id. Convex writes the cells as upserts keyed by (refreshRunId, programCode, columnId, schoolYear), so the published run is the same.

- The worker's shadow path end to end (real pg reads with the marker bracket; Gateway served from the U16 rows): detail clean (30,929 matched), crosswalk clean (57 matched). With the Gateway copy stamped with an older build marker, the detail is stale_copy (not compared) and the crosswalk stays clean.

- No live Gateway run. mart_education.aerie_admissions_pipeline_detail and …_tenant_crosswalk are not deployed yet (absent in Redshift on 2026-09-29).

- Source types checked read-only against pg_attribute: the dbt relation has 5 boolean, 3 numeric(18,2), 5 date and 1 timestamp columns (the six cast to text), and everything else varchar, matching U16's DDL and the reader spec. pipeline_key is unique and non-null (30,929 distinct of 30,929).

## Stack note

- Stacked on #1566 (the U02 kit; base branch feat/a8-u02-read-gate-kit). If #1566 merges first, this will be rebased onto main and retargeted.

- The mart contracts come from AI-Builder-Team/Surtr#2096 (U16: aerie_admissions_pipeline_detail and aerie_admissions_pipeline_tenant_crosswalk).

- Collision: HIGH. Vladimir's pipeline report work touches these files often. The diff to existing files is kept to the split (admissions-pipeline.ts), one wiring hunk (pipeline-refresh.ts) and one .env.example line. Keval: please give Vladimir a heads-up. While the shadow window runs, a change to the report's SQL or its columns needs a lockstep change in the Surtr copy (U16's coupling rule).

- Rebase hotspot: .env.example. U05, U09 and U10 add lines at the same spot.

## Not covered

- A live Gateway dry-run. It needs U16 deployed (AI-Builder-Team/Surtr#2096 DDL apply), the U04 sources and a key granted aerie-a8, and Keval's PII sign-off for exposing the detail copy (plan §9 D2).

- Flipping ADMISSIONS_PIPELINE_READ in any environment. This PR ships inert (legacy).

- The Convex write path: insertAdmissionsPipelineDetail, …Columns and …PhysicalColumns are unchanged.

- A dbt_invocation_id build marker (plan §9 D4). The marker is the relation OID. If Vladimir adds the column, only the Surtr procedure and ADMISSIONS_PIPELINE_BUILD_MARKER_SQL change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1587 — feat(a8): G3 marketing seam + Gateway readers D1–D4 + gate ADMISSIONS_MARKETING_READ, PII-redacted shadow compare (A8 U22) @kevalshahtrilogy  approved

Linear: AERIE-2626 (A8 unit U22; related AERIE-445)

## Summary

This is A8 group G3, marketing. It gates the four per-program EduCRM marketing reads behind ADMISSIONS_MARKETING_READ, using U14's seam pattern:

- D1 event aggregates;

- D2 event contacts;

- D3 shadow-day events;

- D4 weekly deposits.

It adds Gateway readers for all four, and a shadow compare that follows the EduCRM skew rule. The default is today's behaviour.

- Split, no behaviour change. Each of the four reads is now a SQL constant plus a mapper. The legacy query calls the mapper unchanged.

- SQL constants, byte-identical to the old text: MARKETING_EVENT_AGG_SQL, MARKETING_EVENT_CONTACTS_SQL, SHADOW_EVENT_AGG_SQL, DEPOSIT_AGG_SQL.

- Mappers: mapMarketingEventAggRows, mapMarketingEventContactRows, mapShadowEventAggRows, mapDepositAggRows.

- D1 and D2 keep their strict throws. D3 and D4 keep safeParseRows.

- Seam AdmissionsMarketingSources (queries/admissions-marketing-sources.ts). Each member has the legacy query's exact signature. PG_ADMISSIONS_MARKETING_SOURCES is the default everywhere.

- The core-file edits are threading only.

- admissions/per-program-refresh.ts: an optional marketingSources argument, passed to the four helpers. That is 4 call-site lines plus the argument.

- admissions/marketing-events-refresh.ts: one defaulted sources parameter per helper.

- admissions/refresh-orchestrator.ts: one injectable dep, resolved once per cycle. Its runShadowChecks() runs right after the per-program loop.

- The D1→D2 identity-manifest dependency is untouched: D2 still runs only after that program's D1 publishes.

- Gate ADMISSIONS_MARKETING_READ=legacy|shadow|gateway (default legacy), on the U02 kit. The overrides are …_MARKETING_EVENT, …_MARKETING_EVENT_CONTACT, …_SHADOW_DAY_EVENT and …_WEEKLY_DEPOSIT. If D1 and D2 would publish from different transports, the gate warns.

- Gateway mode (queries/admissions-marketing-gateway.ts). Sources are typed exactly as Surtr #2095 and #2099 store them.

- Each mart is read once per cycle and sliced by filter_program_name, with the exact legacy predicate.

- Each slice is re-sorted to the legacy ORDER BY before the unchanged mapper. Order decides what gets published:

- D2's (eventId, contactId) dedupe keeps the last row it sees;

- D4's Convex upsert is keyed by weekLabel, and 9,450 (program, week label) pairs collapse more than one SQL row today.

- SUPER strings sort by string value, as Redshift sorts them. This was checked on 146,830 rows: 0 order violations by value, 6 by JSON text.

- A failed read fails that source for every program this cycle.

- Floors are about a fifth of today's counts. The freshness bound is 24h, as for G2.

- Shadow mode publishes the legacy reads untouched and records each program's mapped output. After the loop it logs one a8_shadow_check line per source for the whole cycle, reading each mart once:

- D1 and D3 by (program, event name). Event names are free text, so they stay redacted.

- D2 as a multiset before the dedupe, with no allowlist: every value and key is [redacted]. The line carries field names, types and counts only.

- D4 at the full SQL grain (program, week label, week start, week end, school year), before the Convex collapse.

- EduCRM skew rule: on a mismatch, only the programs whose records differ are re-read from pg, once. If they now match, the result is source_advanced.

- A program whose legacy read failed is left out of the compare. If no program was read, the result is degraded, never clean.

- Dry-run sync/src/scripts/dry-run-admissions-marketing-shadow.ts. It loads dotenv first, then uses await import(). It runs the worker's shadow path for every listed program and exits 0 only if all four sources are clean.

## Business Value

- About 360 Redshift queries per cycle become 4 reads. Once cut over, 4 marketing queries × 90 programs collapse to one paged Gateway read per mart. That moves G3 off the worker's direct EduCRM access (AERIE-445).

- The cutover is provable and PII-safe. Shadow mode compares what would be published, before the order-dependent steps, and tolerates EduCRM's 30-minute republishes. The contact data carries children's dates of birth, and none of it reaches a log line.

- No behaviour change or prod risk now. The default path runs the same SQL text, the same mappers and the same Convex writes, as proved below.

## Manual Effort Estimate

About 14 hours of focused work by hand, without AI. Keval, please confirm or adjust. It covers:

- tracing the marketing refresh path and the D1/D2 manifest coupling;

- reading the U15 and U18 mart contracts;

- empirically checking Redshift's SUPER and ORDER BY semantics;

- the split, seam, gate, readers and cycle-level shadow compare;

- about 30 tests;

- the read-only proofs.

## Testing / evidence

- Read-only proof on real rows, for all 90 programs queryPrograms lists. The script was a throwaway and is not committed. It printed counts and hash prefixes only: no row values or program names. It compares three paths:

- OLD: this branch's base educrm.ts, a temporary copy.

- PG: the new split functions.

- GW: the new Gateway path. It was fed by the U15/U18 procedures' candidate SQL, run read-only against EduCRM through a raw-text pg client and reshaped into Data API values (BIGINT as numbers, DATE/TIMESTAMP as text, booleans), then passed through the kit's reader, the slicing, the re-sort and the mappers.

- The EduCRM run did not change during any source's reads.

| check | D1 events | D2 contacts | D3 shadow days | D4 deposits |

|---|---|---|---|---|

| mart SQL rows (NULL program dropped) | 1,041 (0) | 146,830 (0) | 75 (0) | 56,700 (0) |

| SQL text, OLD vs new constant | identical | identical | identical | identical |

| OLD = PG, per program | 90/90 | 90/90 as a multiset; 86/90 in order (see note) | 90/90 | 90/90 |

| OLD = GW, identical including order | 84/90 | 40/90 | 87/90 | 90/90 |

| the rest: same multiset, and the ORDER BY key sequence is identical, so they differ only among ties | 6 | 50 | 3 | 0 |

| what publishes (D2 after its dedupe; D4 after the Convex upsert by weekLabel) | n/a | 90/90 identical | n/a | 90/90 identical |

| whole-cycle shadow check (GW fed as above) | clean, 866 keys | clean, 63,931 records / 30,118 keys (22,648 repeated pairs, compared as a multiset) | clean, 73 | clean, 56,700 |

- Note on D2's 86/90. Redshift returns rows that tie on D2's ORDER BY event_id, funnel_state_code, last_name in a different order from run to run. So two back-to-back runs of the identical SQL can differ in order only, which is why the multiset result is 90/90. The ties never decide what publishes: 0 (event, contact) groups have a dedupe winner that ties on the full ORDER BY with a differing row.

- SUPER ordering, checked separately. All 146,830 D2 source rows were ordered by Redshift. Comparing adjacent rows with the reader's comparator found 0 violations. Comparing them as JSON text found 6.

- New tests (about 30):

- queries/admissions-marketing-gateway.test.ts (21):

- gate modes and overrides;

- the mixed D1/D2 transport warning;

- read-once, and exact slicing, including a trailing blank and a NULL key;

- D2 type parity against the pg shape;

- D2's strict throw kept;

- D1 and D4 legacy order;

- the D2 SUPER and D4 NULL-ordering comparators;

- failure fan-out without a retry.

- Shadow tests in the same file:

- publish-invariance: it returns the legacy object untouched and reads no Gateway in the loop;

- D3 clean by (program, event);

- the skew rule re-reads only the differing program, giving source_advanced;

- a persistent mismatch;

- failed legacy reads give degraded;

- D2 multiset with full redaction: no email, program name or id in the result;

- D4 at full grain with allowlisted examples.

- queries/educrm.test.ts (+4): each split read runs its constant with [programName] and returns its mapper's output.

- admissions/marketing-events-refresh.test.ts (+2), per-program-refresh.test.ts (+1), refresh-orchestrator.test.ts (+1): the sources are threaded, and the shadow checks run right after the loop and change no result.

- Checks:

- sync tsc --noEmit is clean.

- vitest run --maxWorkers=2: 89 files, 1,591 tests passed.

- pnpm lint (biome) is clean.

- The full chat suite was not run, because nothing in chat changed.

- Live Gateway dry-run: not run. U15/U18's marts aren't deployed, and the aerie-a8 sources aren't registered or granted.

## Stack

- Base: #1576 (U14 seam, AERIE-2619), which sits on #1566 (U02 kit). If those merge first, this is rebased and retargeted.

- AI-Builder-Team/Surtr#2095 (U15, D1/D2 marts) and AI-Builder-Team/Surtr#2099 (U18, D3/D4 marts), both open. Column names, types, lifted keys and lineage match their DDL, and their candidate SQL fed the evidence above.

- Touches admissions/per-program-refresh.ts, where collision is MEDIUM (U21 and Vladimir), with threading-only edits.

## Not covered

- A live Gateway run. The marts aren't deployed, and the aerie-a8 sources aren't registered or granted (U04, key).

- Mart refresh lag. The skew rule re-reads pg, so it covers pg being *older* than the mart. If EduCRM republishes mid-loop and the mart hasn't refreshed yet when the check runs, the programs read after the republish can show a mismatch. The line carries the mart's source_run_id and source_published_at to diagnose this. Judge the shadow window on a run of cycles, not a single line.

- D2 is the largest read, about 147K rows. In gateway and shadow modes the worker holds one D2 snapshot, plus the recorded per-program outputs in shadow mode, for the cycle.

- Out of scope here:

- D5 DOP weekly (U01 retires it);

- the AERIE-2353 behaviour, which is kept for parity;

- the G2 per-program shadow (U21);

- no Convex changes.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1586 — feat(a8): G2 per-program shadow compare + dry-run for ADMISSIONS_PROGRAM_DETAIL_READ (A8 U21) @kevalshahtrilogy  approved

Linear: AERIE-2625 (A8 unit U21; related AERIE-445, blocked by AERIE-2621)

## Summary

A8 group G2, per-program reads, part 3: the shadow compare for ADMISSIONS_PROGRAM_DETAIL_READ. shadow is now released for this gate. gateway stays behind the existing ADMISSIONS_PROGRAM_DETAIL_READ_NONPROD_OPT_IN convention.

- Shadow publishes exactly legacy (queries/admissions-program-shadow.ts).

- Each shadowed seam member runs the legacy query's SQL and binds once, through the legacy mapper, and returns (or throws) that result. It also keeps what it read.

- Nothing is read from the Gateway during the cycle.

- Checks run once the cycle has published. The orchestrator calls finishShadowChecks after the retention sweeps, before the non-admissions interlude. It never throws.

- Each shadowed mart is read once, through gateway mode's own row feed. createAdmissionsProgramGatewayRows is the Gateway path up to the mapper, extracted from U17's readers, so shadow checks exactly what gateway would publish.

- Results land as one a8_shadow_check line per source per cycle, via the kit's runA8ShadowCheck / logA8ShadowCheck.

- Compared per program, before any order-dependent step. programKey is the value the legacy predicate bound.

| Source | Records compared | Grouped by |

|---|---|---|

| Q5 pipeline students | mapped, without the #n recordKey suffix | recordKey base, as a multiset |

| Q6 enrollment cohort, Q7 deposits, Q8 transfers | the mapper's own parse, before pickBetterCohortRow | natural key, as a multiset |

| Q3 coming-year projections | the mapper's own parse, before first-row-per-year | year, code, version, as a multiset |

| Q4 community metrics, Q2 projections, app conversion | mapped | natural key (duplicates never clean) |

A failed call is recorded as a readFailed record: alone if its read or parse failed, beside the parsed rows if only the mapper refused them. Failing alike on both sides, on the same rows, matches; anything else does not.

- Derived outputs that publish are also compared, computed with the refresh's own functions from each side's mapped reads:

- derived:enrollment-snapshot;

- derived:pipeline-funnel (year-aware);

- derived:projection (enrollment and coming-year).

This proves the TS derivations get identical inputs. Programs whose inputs failed are skipped and counted.

- EduCRM skew rule. On a mismatch, only the calls whose records differ are re-read from pg, once, after the Gateway read. A match then is source_advanced, which counts as clean. Derived outputs recompute from those re-reads without further queries.

- PII. Values appear only for non-PII fields: program keys, cohorts, stages, school years and aggregate counts. Names, emails, ids and dates are redacted. Lines carry field names, value types, counts and lineage.

- Minimal core-file edits, no behaviour change:

- educrm.ts exports the four strict parsers its mappers already ran inline.

- enrollment.ts moves the cohort/deposit/transfer assembly into assembleEnrollmentCohortStudents (same order and errors) and exports deriveYearAwarePipelineFunnel.

- The cycle resolver moves to the shadow module, because it needs both transports.

- Dry-run: sync/src/scripts/dry-run-admissions-program-detail-shadow.ts. It loads dotenv first, then await import(). It runs every per-program read the refresh makes, for every program, plus app conversion, in shadow, then the checks. The refresh's Convex writes are not run.

## Business Value

- Starts the clean-shadow window for A8's longest chain (U03 → U08 → U14 → U17 → U21). That window is the gate before the worker's roughly 450 per-cycle EduCRM queries for these sources collapse into 8 Gateway reads (AERIE-445). The sources include enrollmentProjections, which feeds public API v2.

- Proves parity where it can actually break. Comparing before the order-dependent dedupes rules out false alarms from ties that legacy already leaves unspecified. The derived compares show the published snapshots, funnel and projections would be byte-identical.

- Zero publish risk. shadow publishes exactly the legacy reads (tested, and mutation-checked). The default stays legacy. gateway still refuses to start in production.

## Manual Effort Estimate

About 14 hours of focused work by hand, without AI. Keval, please confirm or adjust. It covers:

- reading the kit, U14's seam and U17's readers;

- deciding, per source, which step is order-dependent and what to compare before it;

- the per-call skew re-read and the derived-output compares;

- extracting the Gateway row feed without changing gateway mode;

- 13 new tests on a two-transport fixture;

- the dry-run;

- the read-only proof on real rows.

## Testing / evidence

- Read-only shadow proof on real rows. A throwaway script, not committed. default_transaction_read_only=on was set and checked. Output is counts and field names only: no row values, no program names.

- It runs this PR's shadow path end to end: resolveAdmissionsProgramSources in shadow, every per-program read for all 90 programs the worker refreshes, then app conversion, then finishShadowChecks.

- The Gateway side is a fake client whose marts are Surtr #2090/#2092/#2094's candidate SQL, run read-only against EduCRM on first read (after the pg reads, as in a real cycle). Values are reshaped as the Data API returns them (TIMESTAMP/DATE as text, BIGINT as numbers) and rows are reversed.

| line | outcome | pg / Gateway records | keys matched | re-read calls |

|---|---|---|---|---|

| aerie-admissions-pipeline-student (Q5) | clean | 33,585 / 33,585 | 33,585 | 0 of 90 |

| aerie-admissions-community-metric (Q4) | clean | 900 / 900 | 900 | 0 of 90 |

| aerie-admissions-enrollment-cohort (Q6, pre-dedupe) | clean | 9,238 / 9,238 | 9,238 | 0 of 90 |

| aerie-admissions-pipeline-deposit (Q7, pre-dedupe) | clean | 113 / 113 | 113 | 0 of 90 |

| aerie-admissions-enrollment-transfer (Q8, pre-dedupe) | clean | 36 / 36 | 36 | 0 of 90 |

| aerie-admissions-program-projection (Q2) | clean | 180 / 180 | 180 | 0 of 90 |

| aerie-admissions-coming-year-projection (Q3, pre-first-row) | clean | 180 / 180 | 180 | 0 of 90 |

| aerie-admissions-app-conversion | clean | 4,295 / 4,295 | 4,295 | 0 of 1 |

| derived:enrollment-snapshot | clean | 339 / 339 | 339 | 90 programs, 0 skipped |

| derived:pipeline-funnel | clean | 1,101 / 1,101 | 1,101 | 90 programs, 0 skipped |

| derived:projection | clean | 360 / 360 | 360 | 90 programs, 0 skipped |

There were 0 failed legacy reads, 8 mart reads (one per mart) and no mismatched, one-sided or duplicate keys. The pg reads took 290s, and the whole run 416s.

- New tests: queries/admissions-program-shadow.test.ts (15).

- The fixture is written once as pg rows and converted to Data API rows by each mart's column types. The real reader, parity layer and slicer then run.

- Publish invariance:

- each shadow member issues the legacy query once and returns or throws exactly its result;

- a whole cycle through refreshEnrollmentPipelineDetailed + refreshAppConversionFacts makes identical results, queries and Convex writes to legacy, while the Gateway serves different rows and one mart fails;

- the checks make no Convex write.

- One line per source: a matching cycle gives 8 source lines plus 3 derived lines, all clean. Each mart is read once and nothing is re-read. A per-source override shadows one source, and a derived output needs all its inputs.

- Compare granularity:

- Q5 order is ignored but a lost duplicate is not;

- dropping Q6's losing duplicate or Q3's second row per year is a mismatch even though the published output is identical.

- Skew rule:

- a republish between reads gives source_advanced for the source and its derived snapshot, re-reading only the differing programs;

- a pg read that failed during the cycle gives source_advanced;

- a Gateway-only difference stays mismatch.

- Failures and PII:

- a mapper failure on both sides matches only on the same rows, and its program is skipped for derived outputs;

- a check that throws outside the kit logs its own degraded line, and later checks still run;

- lines hold field names and non-PII keys, and no fixture PII value appears anywhere.

- Mutation-checked: comparing Q6 after dedupe, re-reading every call, or publishing from other rows each fails the suite.

- Updated tests:

- admissions-program-gateway.test.ts: shadow is released with no opt-in; gateway (alone or mixed with shadow) is still refused.

- refresh-orchestrator.test.ts (+1): the checks run once, with the cycle's programs, after the retention sweeps.

- Checks:

- sync tsc --noEmit clean;

- vitest run --maxWorkers=2: 90 files, 1,591 tests passed;

- pnpm lint clean apart from 2 existing warnings in chat/skill/forge-api.

The full chat suite was not run (no chat changes).

- Live Gateway dry-run: not run. The U12/U13 marts aren't deployed, and the slugs aren't registered or granted (U04). The dry-run script is ready for Keval's X5 step.

## Stack

- Base: #1578 (U17, AERIE-2621), which sits on #1576 (U14) and #1566 (kit). If #1578 merges first, this is rebased onto its base and retargeted.

- Mart contracts: Surtr #2090 (Q5/Q4), #2092 (Q6-Q8), #2094 (Q2/Q3/app conversion).

- Heads-up needed: this touches Vladimir's core files (queries/educrm.ts, queries/enrollment.ts, admissions/refresh-orchestrator.ts). The edits are extractions with no behaviour change. Keval to give Vladimir a heads-up before merge.

## Not covered

- gateway release: it stays behind the non-production opt-in until a clean shadow window. The plan asks for 48 cycles here, because enrollmentProjections feeds API v2.

- Published cohort-student rows and pipeline-student #n keys are not compared after dedupe. Legacy is itself nondeterministic on full ties there, so they're compared before it, per the plan.

- The Gateway-side mapper runs as the legacy one does, so its existing warnings (for example the Q6 duplicate-collapse line) print once per side in shadow.

- Q1 (pipeline_agg) isn't in the seam; U01 retires it.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2123 — Enable on-demand QuickBooks Core validation (SURTR-1593) @ashwanth1109  approved

## Business Value

Enable financial mapping repairs to rebuild from an accepted immutable QuickBooks snapshot without repeating the roughly 45-minute extraction of 68 companies. Live validation measured Core at 4m 9s and the successful financial report stage at 2m 26s. Today's full elapsed validation was 23m 11s because an unrelated scheduled HubSpot admissions transaction blocked competing financial writers and the first report attempt timed out.

## Change

Enable quickbooks-core-tables.on_demand_enabled and remove only Core from the application query layer's matching legacy hold. Explicit production registry metadata already enables the live route. Empty runner parameters, the disabled schedule, raw-success subscriptions, run records, runner/image, IAM, timeouts and financial contract checks are preserved.

The candidate was deployed before merge only to the isolated Core production stack and its required registry stack, each exclusively with [1/1]; both reached UPDATE_COMPLETE. Registry parity confirms the other 170 pipelines' controls are unchanged. No other pipeline stack was deployed.

## Validation

- All required CI checks, all 15 existing on-demand control tests, scoped Biome and whitespace checks pass at c0ca0a073ecbf50a11e26b955c7b29373d6bcd2c.

- Supported Core ON_DEMAND execution core-fast-financial-validation-20261002-105212 succeeded, using {} parameters and accepted raw run 6cad5602-4634-4681-9b31-d926b29b33d4; new Core generation is fe6f4160-39f5-4116-b2cd-4b9b052a9e65:posting.

- Core success automatically triggered the financial mart. That first mart confirmed cancellation after a 520-second lock wait. After HubSpot and both queued School Performance writers succeeded and fresh source/lineage/lock preflight passed, report-only ON_DEMAND retry financial-validation-lock-retry-20261002-111257 succeeded; no extra Core or raw run was needed.

- All 24 Core audits pass. Coordinated reconciliation preserves all 539 posting IDs, dates, signed amounts and categories; verifies nine exact School assignments and 30 existing School scopes; confirms 59 Schools, complete QTD shapes, zero GL/lineage/revenue/Timeback mismatches, $270,595.89 dedicated-book QTD Facilities expenses and the separate −$9,900 workshop credit. This validation includes separately applied Finance policies and Facilities/Guide procedure fixes; they are not changes in this configuration PR.

Future validation must wait for relevant financial and shared-School writers to be idle. The successful-stage total is not a promise of contention-free end-to-end timing.

## Implementation Effort

Approximately 2–3 hours for an engineer to inspect controls, deploy the isolated configuration/registry changes and perform live validation; source coordination and pipeline waiting time are additional.

## Linear

[SURTR-1593](https://linear.app/builder-team/issue/SURTR-1593/enable-on-demand-quickbooks-core-validation-from-accepted-snapshots), a sub-issue of Finance escalation SURTR-1496.

#1578 — feat(a8): G2 per-program Gateway readers Q6, Q7, Q8, Q2, Q3 + app conversion (A8 U17) @kevalshahtrilogy  approved

Linear: AERIE-2621 (A8 unit U17; related AERIE-445, blocked by AERIE-2619)

## Summary

A8 group G2, per-program reads, part 2: Gateway readers for the remaining sources behind U14's ADMISSIONS_PROGRAM_DETAIL_READ gate (Q6 enrollment cohort, Q7 pipeline deposits, Q8 transfers, Q2 program projections, Q3 coming-year projections) and the once-per-cycle app conversion read. Release state is unchanged: only legacy is accepted in production until the U21 shadow compare.

- Split, no behaviour change (queries/educrm.ts, commit 1). Each read becomes its SQL (byte-identical text, checked) plus a mapper holding the old body:

- Q6: ENROLLMENT_DETAIL_WITH_COVERAGE_SQL + mapEnrollmentDetailWithCoverageRows

- Q7: PIPELINE_DEPOSIT_DETAIL_SQL + mapPipelineDepositDetailRows

- Q8: ENROLLMENT_TRANSFER_DETAIL_SQL + mapEnrollmentTransferDetailRows

- Q2: allProgramProjectionsSql(sourceColumn) + mapAllProgramProjectionRows

- Q3: comingYearProjectionAliases + comingYearProjectionsSql(aliases) + mapComingYearProjectionRows

- app conversion: APP_CONVERSION_FACTS_SQL + mapAppConversionFactRows

The strict parses, pickBetterCohortRow dedupe, first-day Re/New split, transfer pairing and Q3's first-row-per-year rule all stay in the mappers, unchanged.

- Gateway readers (queries/admissions-program-gateway.ts), one per mart, typed as Surtr's DDL stores them (EduCRM source types, SUPER kept):

- aerie-admissions-enrollment-cohort, -pipeline-deposit, -enrollment-transfer (Surtr #2092);

- aerie-admissions-program-projection, -coming-year-projection, -app-conversion (Surtr #2094).

- Each mart is read once per cycle and sliced by its lifted key (filter_program_name, filter_program_id, filter_program_code_lower), which is dropped before mapping. Floors 1,000 / 10 / 1 / 50 / 50 / 400 rows (9,238 / 113 / 36 / 180 / 180 / 4,295 today); freshness 24h, as for Q4/Q5.

- Legacy order is re-applied by legacyOrder, in Redshift's semantics: BIGINT numerically (pg returns it as text), VARCHAR by UTF-8 bytes, SUPER by value (probed read-only today: booleans < numbers < strings < NULL, so "a" < "a b"), and DESC with NULLs first. Q6's order decides pickBetterCohortRow ties and output order; Q7, Q8 and Q2 are ordered the same way.

- Q3 carry-over from U13. Surtr #2094's mart drops Q3's ORDER BY, because it compares with the program's first alias. comingYearProjectionSlice gathers the IN (aliases) rows and re-applies projection_year, first-alias rank, then projection_version DESC before the first-row-per-year rule. A program with no alias fails, as legacy IN () does.

- App conversion wiring (minimal core-file edits): the seam gains appConversion. It is threaded from the orchestrator (one added property), through refreshAdmissionsAppConversionFacts (an optional programSources arg, passed on), into refreshAppConversionFacts (queries/enrollment.ts, one defaulted parameter replacing the direct queryAppConversionFacts() call). Override ADMISSIONS_PROGRAM_DETAIL_READ_APP_CONVERSION.

- Gate text (.env.example, module docs): every source now has a reader, so the "no reader until U17" warning path is removed. Refusal without the opt-in is unchanged.

## Business Value

- Finishes the per-program read surface for A8's longest chain. With U14 plus this unit, all eight sources behind the gate can read Surtr's parity marts. U21 (shadow compare) can then start the cutover that retires the worker's direct EduCRM reads (AERIE-445), including enrollmentProjections, which feeds public API v2.

- About 450 queries per cycle collapse into 6 reads once cut over: 5 per-program queries across about 90 programs, plus app conversion.

- Zero behaviour change and zero prod risk now. The default path issues the same queries in the same order with the same Convex writes (proved below), and a premature shadow/gateway in prod still fails loudly at config resolution.

## Manual Effort Estimate

About 12 hours of focused work by hand, without AI. Keval, please confirm or adjust. It covers:

- reading Surtr #2092/#2094's contracts;

- splitting six reads without touching their SQL;

- probing Redshift's ORDER BY semantics for SUPER, BIGINT and NULLs;

- six readers plus the Q3 re-sort and the app conversion wiring;

- about 30 new tests with two-transport fixtures;

- the read-only proof.

## Testing / evidence

- SQL identity. A script extracted each of the six SQL templates from U14's educrm.ts and from this branch: all six are byte-identical, and no old template is missing.

- Read-only pg proof on real rows (throwaway script, not committed; default_transaction_read_only=on; counts and hash prefixes only, no row values or program names).

- OLD = U14's per-program reads (a temporary copy of its educrm.ts); PG = this branch's seam.

- GW = the new Gateway readers, fed by Surtr #2092/#2094's candidate SQL run read-only against EduCRM and reshaped into Data API values (TIMESTAMP/DATE as text, BIGINT as numbers). Rows are reversed, so each reader must restore the legacy order itself.

| check | programs | result |

|---|---|---|

| OLD = PG, Q6, Q7, Q8, Q2, Q3 (ordered) | 6 (with Q6, Q7 and Q8 rows) | 30/30 equal |

| OLD = GW, the same five (ordered) | the same 6 | 30/30 equal |

| OLD = GW, whole cycle, ordered | all 90 programs | Q6 90/90 (8,547 records), Q7 90/90, Q8 90/90, Q2 90/90, Q3 90/90 |

| OLD = PG = GW, app conversion (as a set; the SQL has no ORDER BY) | whole table | 4,295 = 4,295 = 4,295 |

No program today has a source code differing from its code, so Q3's two-alias order is proven by the unit tests, not by real rows.

- New tests:

- queries/admissions-program-readers.test.ts (5). One two-transport fixture, hand-written in each legacy ORDER BY (BIGINT 9 < 10, SUPER "in" < "out", DESC NULLs first, Q3 first-alias rank beating a higher version). It checks:

- each Gateway source equals its legacy read per program, order included;

- a whole refreshEnrollmentPipelineDetailed cycle over 3 programs makes identical results and Convex writes from one read per mart;

- app conversion makes the same facts and writes.

Mutation-checked: dropping any one order (Q6, Q7, Q8, Q2, Q3 alias rank, Q3 version DESC) fails it.

- queries/admissions-program-gateway.test.ts (+7, 1 updated): the comparators against the probed Redshift order, legacyOrder with DESC, comingYearProjectionSlice, Q3 with no alias, app conversion read-once and value-free failure; gateway now reads every source with no warnings.

- queries/admissions-program-sources.test.ts (+1): the seam equivalence test covers appConversion.

- admissions/forecast-refresh.test.ts (+1) and admissions/refresh-orchestrator.test.ts (+1 assertion): the cycle's sources reach refreshAppConversionFacts.

- Checks: sync tsc --noEmit clean; vitest run --maxWorkers=2 89 files, 1,576 tests passed (existing educrm.test.ts, enrollment.test.ts, forecast-refresh.test.ts unchanged and green). pnpm lint clean apart from 2 existing warnings in chat/skill/forge-api. Full chat suite not run (no chat changes).

- Live Gateway dry-run: not run. The marts aren't deployed and the slugs aren't registered or granted (U04); U21 ships the shadow dry-run.

## Stack

- Base: #1576 (U14, AERIE-2619). If #1576 merges first, this is rebased onto its base and retargeted.

- AI-Builder-Team/Surtr#2092 (U12 marts for Q6-Q8, open) and AI-Builder-Team/Surtr#2094 (U13 marts for Q2, Q3 and app conversion, merged): column names, types, lifted keys and lineage match their DDL.

- Heads-up needed: this touches Vladimir's core files (queries/educrm.ts, queries/enrollment.ts, admissions/forecast-refresh.ts). Keval to give him a heads-up before merge.

## Not covered

- The shadow compare and its dry-run (U21). The opt-in stays a documented convention, as in U14.

- Q6 and Q8 ties that the legacy ORDER BY leaves unspecified keep the mart's mart_row_id order. Legacy is itself nondeterministic there; U21 compares before the order-dependent dedupe.

- Arrays and objects in the SUPER order-by columns (cohort_id, direction, projection_version) are ordered by JSON text after strings. That is unprobed, but today those columns hold only strings (9,238 / 36 / 180 rows checked).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1576 — feat(a8): G2 per-program source seam + gate ADMISSIONS_PROGRAM_DETAIL_READ + Gateway readers Q5, Q4 (A8 U14) @kevalshahtrilogy  approved

Linear: AERIE-2619 (A8 unit U14; related AERIE-445)

## Summary

A8 group G2, per-program reads: a seam at the one call site where the analytics worker runs its per-program Redshift queries, a gate, and the first two Gateway readers (Q5 pipeline students, Q4 community metrics). The gate defaults to today's behaviour and, until the shadow compare lands (U21), only legacy is accepted in production.

- Split, no behaviour change. queryPipelineDetail / queryCommunityMetrics are now PIPELINE_DETAIL_SQL / COMMUNITY_METRICS_SQL (byte-identical text) plus mapPipelineDetailRows / mapCommunityMetricRows (the old bodies).

- Seam AdmissionsProgramSources (queries/admissions-program-sources.ts) for Q2-Q8. Each member has the legacy query's exact signature. PG_ADMISSIONS_PROGRAM_SOURCES calls today's queries unchanged and is the default everywhere.

- Q1 (queryPipelineAgg) is deliberately left as a direct call, so U01's removal of it (#1565) doesn't conflict.

- Core-file edits are the call-site list plus an injected parameter only:

- queries/enrollment.ts: 7 call-site lines swap queryX(...) for sources.x(...); one defaulted sources parameter on refreshEnrollmentPipelineDetailed and its inner runner; imports.

- admissions/per-program-refresh.ts: optional programSources arg, passed through.

- admissions/refresh-orchestrator.ts: one injectable dep, resolved once per cycle before the run starts, passed to runPerProgram.

- Gate ADMISSIONS_PROGRAM_DETAIL_READ on the U02 kit. Overrides …_PIPELINE_STUDENT, …_COMMUNITY_METRIC; reserved (validated, read legacy until U17) …_ENROLLMENT_COHORT, …_PIPELINE_DEPOSIT, …_ENROLLMENT_TRANSFER, …_PROGRAM_PROJECTION, …_COMING_YEAR_PROJECTION, …_APP_CONVERSION.

- Only legacy is released. shadow or gateway on any source throws AdmissionsProgramDetailConfigError at resolution, before the admissions run starts, unless ADMISSIONS_PROGRAM_DETAIL_READ_NONPROD_OPT_IN=true (exact). Documented in .env.example: never set in production. With the opt-in, gateway serves Q5/Q4 from the Gateway (the rest warn and read legacy), and shadow warns and reads legacy (nothing is compared until U21).

- Gateway readers (queries/admissions-program-gateway.ts) for aerie-admissions-pipeline-student and aerie-admissions-community-metric, typed as Surtr #2090's DDL stores them (EduCRM source types, SUPER kept):

- each mart is read once per cycle, on first use, and sliced by its lifted key (filter_program_code / filter_program_name), which is dropped before mapping;

- the slice is exactly the legacy WHERE <expr> = $1: an exact string match (checked read-only on Redshift today: case and trailing blanks are significant), and a NULL key matches no program;

- Q4 is re-sorted to ORDER BY metric_name in Redshift's byte order (checked: B < _ < a), NULLs last, stable. Q5 has no ORDER BY, so its slices keep the mart's mart_row_id order;

- a failed read fails that source for every program in the cycle, once (as a Redshift outage does on legacy), with a value-free error;

- floors 1,000 / 100 rows (33,585 / 1,140 today); freshness 24h, matching the publisher's own bound and G1.

## Business Value

- Unlocks the longest A8 chain. The per-program group (U08 → U14 → U17 → U21) is the critical path to retiring the worker's direct Redshift reads (AERIE-445). Every later per-program reader (U17) and the shadow compare (U21) plug into this seam without touching Vladimir's files again.

- Roughly 1,400 queries per cycle become 7 reads. Once cut over, about 8 queries × 90 programs per hourly cycle collapse to one paged read per mart.

- Zero behaviour change and zero prod risk now. The default path is the same queries in the same order with the same Convex writes (proved below), and a premature shadow/gateway in prod fails loudly instead of half-working.

## Manual Effort Estimate

About 11 hours of focused work by hand, without AI. Keval, please confirm or adjust. It covers tracing the per-program refresh path and the kit, reading Surtr #2090's contracts, checking Redshift's equality and ORDER BY semantics, the seam, gate and readers, about 45 tests with two-transport fixtures, and the read-only proof.

## Testing / evidence

- Read-only pg proof on real rows (throwaway script, not committed; counts and hash prefixes only, no row values or program names). OLD = origin/main's per-program queries (a temporary copy of main's educrm.ts); PG = the new seam; GW = the new Gateway path, fed by Surtr #2090's mart candidate SQL run read-only against EduCRM and reshaped into Data API values (TIMESTAMP as text, BIGINT as numbers, rows reversed so the reader must restore Q4's order).

| check | programs | result |

|---|---|---|

| OLD = PG for Q2, Q3, Q4, Q5 (as a set), Q6, Q7, Q8 | 6 (of 88 with Q4 and Q5 rows) | 42/42 equal |

| OLD = GW, Q5 (as a set) and Q4 (ordered) | the same 6 | 12/12 equal |

| OLD = GW, whole cycle | all 90 programs | Q5 90/90 (33,585 rows), Q4 90/90 (900 rows) |

- New tests:

- queries/admissions-program-sources.test.ts (13): the seam equivalence test. Each pg member issues exactly the legacy query (SQL and params) and returns its result; refreshEnrollmentPipelineDetailed with the default and with the explicit pg seam issues the same queries in the same order and the same Convex writes; each member is called once with the argument the legacy call took; and, end to end over 3 programs, the Gateway path for Q5/Q4 produces the same results and Convex writes as legacy from one read per mart, including case, trailing-blank and NULL keys.

- queries/admissions-program-gateway.test.ts (25): every gate mode and override, the refusal without the opt-in, non-exact opt-in values, reserved override names, read-once, exact slicing, Q4 ordering, failure fan-out without row values, the kit's stale and floor rules, and the helpers.

- admissions/refresh-orchestrator.test.ts (+2): resolved once before startRun and threaded; a refused configuration throws before any run write. admissions/per-program-refresh.test.ts (+1): sources reach every program's refresh.

- Checks: sync tsc --noEmit clean; vitest run --maxWorkers=2 88 files, 1,562 tests passed (existing enrollment.test.ts, per-program-refresh.test.ts and refresh-orchestrator.test.ts unchanged and green). pnpm lint clean. Full chat suite not run (no chat changes).

- Live Gateway dry-run: not run. U08's marts (Surtr #2090) aren't deployed and the aerie-a8 sources aren't registered or granted (U04). U21 ships the shadow dry-run.

## Stack

- Base: #1566 (U02 read-gate kit, AERIE-2612). If #1566 merges first, this is rebased onto main and retargeted.

- AI-Builder-Team/Surtr#2090 (U08 marts for Q5/Q4, open): column names, types, lifted keys and lineage match its DDL.

- Independent of #1565 (U01): Q1 stays out of the seam.

- Heads-up needed: this touches Vladimir's core files (queries/enrollment.ts, admissions/per-program-refresh.ts). Keval to give him a heads-up before merge.

## Not covered

- Shadow compare and its dry-run (U21), and Gateway readers for Q2, Q3, Q6-Q8 and app conversion (U17).

- The opt-in is a convention: there is no reliable production marker in the worker's env (NODE_ENV=production is also set in local images), so it is documented as never-in-production rather than enforced.

- A refused configuration aborts the whole analytics cycle (the non-admissions interlude runs inside the admissions window), by design: it is loud and happens before any admissions write.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1637 — fix(funnel): relay community funnel sync refusals as userErrors the worker can show (AERIE-2680) @kevalshahtrilogy  approved

Linear: AERIE-2680 (Mercy's non-blocking finding on the release PR, against code from AERIE-2677; related AERIE-445)

## Summary

The finding. previewCommunityFunnelPopulation threw raw Errors for its two refusals (a published run over the read bound; a published run with no rollups). They cross the enrollment sync route, where a raw Error can arrive redacted or wrapped in Convex framing, so the worker's log would not name the reason.

What I found when I traced the message to the operator. Changing the two throws alone would not have helped:

- The route answered with err.message. For a ConvexError that crossed ctx.runMutation, the runtime rewrites message and keeps the original text in data.

- The worker never receives a ConvexError at all. It reaches Convex over HTTP, so every preview failure was a plain Error, and describeA8Error printed Error (details withheld) for it whatever Convex had said.

So the fix has three parts.

1. Convex: the refusals are userErrors.

- Both preview refusals in fullFunnel.ts.

- Family check on the same route (below).

2. Convex: the enrollment route relays a userError by its own message.

- New userErrorMessage(error) in chat/convex/lib/errors.ts: the string data of a ConvexError, or null.

- handleEnrollmentSync answers a userError with { "error": "<message>", "userError": true }. Any other failure keeps today's body, { "error": err.message }.

- The flag is additive: an older worker ignores it.

3. Worker: the reason reaches the operator's line.

- New SyncOperationError (sync/src/analytics/sync-operation-error.ts). sendAnalyticsSyncOperation throws it for a non-2xx answer with the same message as before, plus status, code (HTTP_<status>) and userMessage.

- userMessage is set only when the route flagged the body. The text is flattened to one line and capped at 300 characters.

- previewCommunityFunnelPopulationInConvex rethrows a flagged refusal as Convex refused: <reason>. Anything else is rethrown untouched, and describeA8Error shows SyncOperationError (HTTP_400) (details withheld): more than before, and still never the body.

- The hint is now right for each case: "deploy Convex first" only when Convex did not answer usefully; "set the gate to legacy until the published run can be counted" when it did.

What the operator sees for the no-rollups case, before and after:

before: ... (previewCommunityFunnelPopulation: Error (details withheld)); ... deploy Convex first, or set ADMISSIONS_COMMUNITY_FUNNEL_READ=legacy

after: ... (previewCommunityFunnelPopulation: Convex refused: Published community funnel run has no full-funnel rollups; its deal population cannot be counted); ... Set ADMISSIONS_COMMUNITY_FUNNEL_READ=legacy until the published run can be counted

Family check: the other throws in the two funnel modules that cross this route.

| File | Throws | Changed | Left |

|---|---|---|---|

| fullFunnel.ts | 11 | All 11: the 2 preview refusals and 9 write refusals (deal and stage count rules, batch size, program identity, school year) | None |

| fullFunnelValidators.ts | 1 | 1: the stage-count check fullFunnel.ts calls on every insert | None |

| communityFunnel.ts | 17 | 14: stage count rules, batch size, missing program identity or deal id | 3 |

The three left as plain Errors each quote a deal id (duplicate deal id in a batch; a named enrollment with no program identity; a named enrollment with no stage code). A userError is now relayed to the worker verbatim, so it must be safe to show as is, and this gate treats a deal id as a value that is never logged. The worker also checks those same three conditions itself before it sends anything, so they should not fire from this worker.

No other module on the enrollment route uses userError, so only these messages gain the flag. I did not touch other routes or modules.

## Business Value

- When the community funnel gate refuses a Gateway snapshot because Convex cannot count the published run, the on-call person sees why in the worker's log, with the right next step, instead of a generic error and advice to redeploy.

- It clears the last open review note on the community funnel gate before it can be switched to the Gateway (AERIE-445).

- Other sync callers get a little more too: any failed sync operation described through the A8 kit now shows its HTTP status.

## Manual Effort Estimate

About 3 hours of focused work by hand, without AI. Keval, please confirm or adjust this number. It covers:

- tracing the message from the mutation, through the route and the HTTP transport, to the operator's line;

- reading the 29 throws in the two funnel modules and deciding which are safe to relay;

- the route and transport changes, and 34 new tests.

## Testing

- chat/convex/lib/errors.test.ts (new, 8): userErrorMessage returns a userError's message; reads data, not message, for an error shaped as the runtime rebuilds it across runMutation; is null for a plain Error, object data, blank data and non-errors.

- chat/convex/admissions/fullFunnelDashboards.test.ts (30 to 37):

- each preview refusal is a userError with its exact message (the over-bound case seeds 10,001 rollups);

- both reach the caller through the HTTP route as { error, userError: true } with status 500;

- four converted write refusals are relayed and flagged the same way;

- the duplicate-deal-id refusal is not flagged.

- sync/src/analytics/sync-operation-error.test.ts (new, 13): the flag must be exactly true and the error a string; unflagged, plain-text, array and null bodies give null; line breaks and control characters are flattened; long text is capped; describeA8Error shows the name and HTTP status and never the body.

- sync/src/analytics/sync-http.test.ts (+1): a non-2xx answer is a SyncOperationError with an unchanged message; userMessage is set only when flagged.

- sync/src/analytics/queries/admissions-community-funnel-gateway.test.ts (61 to 66): the preview rethrows a flagged refusal with its reason and rethrows anything else untouched; on the worker path the operator's message is pinned in full for a deliberate refusal, an older Convex (HTTP 400) and an unusable acknowledgement.

- Checks:

- sync: tsc --noEmit clean; vitest run --maxWorkers=2: 114 files, 2534 tests passed.

- chat: tsc --noEmit and the Convex typecheck clean. Five targeted files were run (lib/errors, fullFunnelDashboards, communityFunnelDashboards, admissions, analyticsSyncSecurity: 198 tests passed). The full chat suite wasn't run locally; CI runs it.

- pnpm lint (boundaries, convex paths, read bounds, test architecture, knowledge, biome) is clean. Its two warnings are pre-existing ones in chat/skill/forge-api/scripts/sindri.mjs.

## Deploy order

- None required. An older worker ignores the new flag. A newer worker against an older Convex gets no flag and shows SyncOperationError (HTTP_<status>) (details withheld).

- The gate still defaults to legacy; the preview is only called in gateway mode.

## Not covered

- Whether production Convex redacts a raw Error across ctx.runMutation was not tested against a production deployment. The fix does not depend on it: a ConvexError's data is carried either way, and the route now reads it.

- convex-test does not rewrite an error's message across runMutation, so the route tests cannot show the difference between message and data. The unit test for userErrorMessage covers it with an error built the way the runtime builds one.

- Other sync routes are unchanged. They still answer with err.message (the forecast route answers with a fixed generic body). SyncOperationError.userMessage is null for all of them.

- The three deal-id refusals stay plain Errors, as explained above.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1636 — fix(gateway): fail closed when a mapper drops a Gateway row (AERIE-2679) @kevalshahtrilogy  approved

Linear: AERIE-2679 (fixes a blocking Mercy finding on the release PR, against code from AERIE-2677; related AERIE-445)

## Summary

The finding. In ADMISSIONS_COMMUNITY_FUNNEL_READ=gateway, the legacy mapper drops a Gateway row with a null deal_id or one its schema rejects. Gateway mode only logged the count and then ran the population check on the reduced count. One dropped row in 5,403 is far inside the 5% allowance, so a partial snapshot would publish as a success.

The rule this PR applies to every Gateway read: if the mapper returns fewer records than the Gateway rows it was given, the read is refused.

- gateway mode: the read throws before the population check and before any write. The last published data stays. The message carries counts only, no row values.

- shadow mode: the same refusal makes the check degraded, never clean, so the shadow window shows it before gateway would refuse it. This matters because the mart is a verbatim copy: pg drops the same row, so the keyed compare alone would have called the cycle clean.

- Legacy (pg) paths are unchanged. Production publishes exactly what it publishes today.

What changed.

- Kit, sync/src/analytics/a8/gateway-table-reader.ts: assertNoA8RowsDropped(source, { rowsRead, recordsMapped }, reason). It throws the reader's own A8GatewayReadError, which is on the kit's value-free list, so the message reaches the logs as written.

- Community funnel, admissions-community-funnel-gateway.ts:

- readCommunityConversionFromGateway refuses any dropped row.

- Removed: the droppedRows and sourceRows fields (always 0 and N now), the warn-and-publish branch, and the second floor on the mapped count (dead once any drop throws).

- The module doc and the floor comment now describe the new behaviour.

- Expense vendor classifications, expenses-gateway.ts: the same refusal. It replaces the second floor on the mapped count, which only refused a read once it fell below 500 of about 3,479 rows.

- XO contractor identity, xo-contractor-identity-gateway.ts: a Gateway row whose aliases all trim away now fails the read. Before, it was filtered out and the rest published.

Every Gateway reader under sync/src, and what I found.

| Lane | Reader (Gateway source) | Can it drop a Gateway row? | Change |

|---|---|---|---|

| A1 | site-operational-metadata-gateway.ts | No. z.array(...).safeParse, then throw. Shadow only, never published. | None |

| A2 | schools-data-sheet-gateway.ts | No. Any invalid row, duplicate cell or partial grid refuses the snapshot; every record is published. | None |

| A3 | school-calendar-gateway.ts | No. Any invalid row or duplicate slug refuses the read; every row is sent. Convex skips a row whose site slug Aerie doesn't have, and the worker records that as an error on the run. | None |

| A4 | school-source-directories-gateway.ts (QuickBooks, SIS, Finalsite) | No. Strict parse; the snapshot builders map one to one and throw. | None |

| A5 | camps-full-gateway.ts (7 entities), camps-gateway.ts (parity checks) | No. Strict parse; loadCampSourceSnapshotViaGateway refuses the whole snapshot on a blank or duplicated id, mixed lineage or a dangling reference. | None |

| A6 | rebl3-sites-gateway.ts | No. Any invalid row, repeated site or mixed lineage refuses the snapshot. | None |

| A7 | matterport-discovery-shadow.ts (xref) | No. Strict; shadow only, never published. | None |

| F1 | xo-contractor-identity-gateway.ts | Yes: filtered out rows with no usable alias. | Fixed |

| F2 | xo-contractor-package-gateway.ts | No. Strict parse, one record per row. | None |

| A8 G1 | admissions-reference-gateway.ts (programs, Program directory) | No. mapProgramRows and mapHubspotProgramRows throw on any bad row. | None |

| A8 G2 | admissions-community-funnel-gateway.ts | Yes: null deal_id or schema-rejected. | Fixed |

| A8 G3 | community-deposits-gateway.ts | No. mapCommunityDepositRows maps every row or refuses the read. | None |

| A8 G4 | expenses-gateway.ts, transactions | No. Strict array parse; a conflicting repeated key is refused. An exact repeated row is published twice, as legacy does. | None |

| A8 G4 | expenses-gateway.ts, vendor classifications | Yes: safeParseRows skips a row it cannot parse. | Fixed |

| A8 G4 | expenses-gateway.ts, metadata | Not a row mapper. It derives two DISTINCT lists and leaves out NULLs exactly as the legacy SQL's WHERE ... IS NOT NULL does. | None |

## Business Value

- Unblocks the Aerie production release. Mercy will not approve the release while this finding is open on main.

- A snapshot that lost rows can no longer be published as a clean success on the admissions funnels, the expense vendor categories or the contractor identity index. Each of those would otherwise go quietly stale or short for the affected families, vendors or contractors.

- The shadow window now means what it says. A source that would be refused in gateway mode can no longer pass its shadow window clean.

## Manual Effort Estimate

About 4 hours of focused work by hand, without AI. Keval, please confirm or adjust this number. It covers:

- reading every Gateway reader and its publish path (13 reader files, plus the refresh code each one feeds) for drops;

- the kit helper and the three fixes;

- reworking the tests that pinned the old drop-and-publish behaviour, plus 18 net new tests.

## Testing

- a8/gateway-table-reader.test.ts (+4): equal counts pass; one dropped row of 5,403 is refused with the exact count-only message; the error carries its source; more records than rows is refused too.

- queries/admissions-community-funnel-gateway.test.ts (55 to 61):

- reader: one null-deal_id row and one schema-rejected row each refuse a 1,201-row read, with the exact message and no PII; 250 dropped rows are refused the same way; a clean read returns one record per row;

- gateway mode: one dropped row fails the read with A8GatewayReadError, the population preview is never called, nothing is logged as read, and pg is never called;

- shadow mode: pg and the Gateway both carry the same null-deal row. Legacy publishes as today, and the check is degraded, countsAsClean: false, with the count-only reason;

- end to end through refreshCommunityFunnel: a dropped Gateway row fails the refresh with no Convex call at all; legacy mode still publishes the pg read minus the row its mapper drops.

- Two tests that pinned the old behaviour (drop, warn, publish; floor after drops) were rewritten.

- queries/expenses-gateway.test.ts (31 to 34): the legacy mapper still drops a bad pg row; the Gateway reader refuses the same row; gateway mode fails only the vendors read with a value-free message while transactions still publish; shadow mode is degraded for vendors and clean for the other two sources.

- tests/redshift/xo-contractor-identity-gateway.test.ts (16 to 19): three shapes of an alias-less row each refuse the read with a count-only message and no name; blank entries between delimiters are still trimmed without refusing the row.

- tests/analytics/xo-contractor-identity-refresh.test.ts (+2): through the real Gateway reader. gateway mode returns an error and sends nothing; shadow mode publishes from pg and reports check failed.

- Checks:

- sync: tsc --noEmit clean; vitest run --maxWorkers=2: 113 files, 2515 tests passed.

- pnpm lint (boundaries, convex paths, read bounds, test architecture, knowledge, biome) is clean. Its two warnings are pre-existing ones in chat/skill/forge-api/scripts/sindri.mjs.

- The chat suite wasn't run: this PR changes no chat code.

## Does this refuse anything on today's data?

- Community funnel: no. The production mart was checked today (by Keval's session, not re-queried here): 5,403 rows, 0 null deal_id, 5,403 distinct.

- Vendor classifications: no, on the last evidence I have. The read-only check for the expenses gate on 2026-09-29 found 3,479 rows and none with a NULL name.

- XO contractor identity: not verified against production. The mart's SQL only emits aliases of 6 or more characters, so a row with none needs an all-whitespace contractor name. XO_CONTRACTOR_READ defaults to pg, and shadow mode would report such a row as a failed check before any flip.

- All three gates default to the legacy read, so nothing changes in production until a gate is set.

## Not covered

- The legacy mappers still drop rows on the pg path. That is deliberate (production behaviour is unchanged), but it means pg and the Gateway now differ on a bad row: pg publishes without it, the Gateway refuses.

- A persistent bad row in a mart keeps its shadow line degraded every cycle until Surtr fixes the row, since the Gateway read is refused before the compare runs.

- The held A8 PRs are not on main and were not audited: the per-program readers, the admissions pipeline gate, marketing and SIS enrollment. The pipeline refresh already counts droppedNoColumn and droppedNoProgram, so its gate needs this same rule when it lands.

- No live Gateway run. Gateway mode cannot be exercised end to end until the Surtr marts and the key grant are live.

- safeParseRows still logs a Zod message for the first three rows it skips. Unchanged here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2121 — fix(collections): skip 42DS in CollectIQ v2 @sanketghia  approved

## Summary

- Skip 42DS as an out-of-scope forecast view in CollectIQ v2, consistent with the weekly forecast pipeline.

- Add normalization and parser regression coverage for case and surrounding whitespace.

## Verification

- uv run --extra dev --frozen pytest — 64 passed.

- Local --dry-run against the configured CollectIQ sheet parsed 8 rows for 2026-Q4 and performed no S3 or Redshift writes.

#1633 — fix(rhodes-shadow): exact-cent tuition compare; count roster failures as mismatched (AERIE-2676) @kevalshahtrilogy  approved

Linear: AERIE-2676 (follow-up to PR 1563 / AERIE-2610; related AERIE-445)

## Summary

Follow-up to two non-blocking Mercy notes on the already-merged PR 1563 (A1 rhodes-merger shadow mode). Both are in sync/src/upstream/rhodes/site-metadata-shadow-compare.ts, the read-only compare behind RHODES_MERGER_READ=shadow.

1. Tuition is compared to the cent without float error.

- Before: Math.round(a * 100) === Math.round(b * 100). The product is rounded in binary: 1.005 * 100 is 100.49999999999999. So near a half-cent boundary:

- a real one-cent difference read as clean (1.005 against 1.00);

- an equal pair read as different (1.005 against 1.01, which is what Surtr's NUMERIC(12,2) stores for 1.005).

- After: a new helper, sync/src/analytics/exact-cents.ts:

- exactCents(value) takes the value's shortest decimal text (String(1.005) is "1.005") and rounds it to integer cents with integer arithmetic, half away from zero. No float multiplication. Exponent forms (1e+21, 1.5e-7) are expanded; NaN and the infinities have no cents.

- equalToTheCent(a, b) compares the two.

- Every other field is still compared with ===.

2. mismatchedSiteCount counts every site that is not clean.

- Before: compared - clean. A slug on one side only, or duplicated on a side, made the run not clean but was not counted, so a roster-only failure logged status: "mismatch" with mismatchedSiteCount: 0.

- After: field-level divergences + incumbent-only slugs + Surtr-only slugs + duplicated slugs, each slug once (a slug duplicated on both sides is one site).

- clean is exactly mismatchedSiteCount === 0.

- cleanSiteCount + mismatchedSiteCount is the number of distinct slugs across both sides.

- New field fieldMismatchedSiteCount keeps the previous number (compared sites with a divergent field), so nothing is lost from the log line.

- The dry-run's summary line prints both.

Same pattern elsewhere. I grepped the other shadow compares on main for Math.round(x * 100): XO contractor, camps, school source directories, schools data sheet, REBL3 sites, school calendar and the A8 keyed compare. None use it; they compare with === / !==.

## Business Value

- The Rhodes shadow window decides whether Aerie can stop computing the site-metadata merge itself and read Surtr's instead (part of AERIE-445). A compare that can call a real tuition difference clean, or report zero mismatched sites on a run that isn't clean, weakens that decision.

- Fewer false alarms. An equal tuition pair near a half-cent boundary no longer shows up as a divergence someone has to chase.

- A reusable exact-cents helper for the money comparisons still to come in the Gateway migration.

## Manual Effort Estimate

About 2.5 hours of focused work by hand, without AI. Keval, please confirm or adjust this number. It covers:

- reproducing the half-cent float cases and choosing a rounding that matches Redshift's NUMERIC(12,2);

- the helper, including exponent notation and negatives;

- reworking the count so each slug is counted once, and checking the sibling compares;

- about 40 new test cases.

## Testing

- New sync/src/analytics/exact-cents.test.ts (32 tests):

- amounts already in cents are exact;

- five half-cent cases where Math.round(value * 100) is one cent low, each asserted against both the old expression and the new helper;

- agreement with plain decimal rounding elsewhere, negatives, -0, exponent notation, NaN and the infinities.

- sync/src/upstream/rhodes/site-metadata-shadow-compare.test.ts (20 tests, 8 new):

- three one-cent differences at a half-cent boundary are now mismatches, and three equal pairs are now the same;

- the existing "40000.004 equals 40000" case still passes;

- one-sided slugs and a duplicated slug each count as mismatched sites while fieldMismatchedSiteCount stays 0;

- a mixed case (clean, field-level, one-sided on each side, duplicated on each side, duplicated on both, duplicated and absent from the other side) counts each slug once: 8 mismatched of 9;

- clean === (mismatchedSiteCount === 0) across seven shapes;

- the WARN log line for a roster-only failure now carries mismatchedSiteCount: 1.

- Checks:

- sync: tsc --noEmit clean; vitest run --maxWorkers=2: 113 files, 2429 tests passed.

- pnpm lint (boundaries, convex paths, read bounds, test architecture, knowledge, biome) is clean. Its two warnings are pre-existing ones in chat/skill/forge-api/scripts/sindri.mjs.

- The chat suite wasn't run: this PR changes no chat code.

## Not covered

- No live shadow run. The compare is pure and read-only; a live run would only show today's data, which has no half-cent tuition that I know of.

- mismatchedSiteCount changes meaning on the rhodes_merger_shadow_check line. Anything that read it as "field-level only" should read fieldMismatchedSiteCount. The shadow mode merged on 2026-10-01, and I found no consumer in the repo other than the dry-run script, which is updated.

- The rounding assumes the incumbent's decimal value is the intended amount (1.005 means 1.005, not the double just below it). That is how Redshift rounds the same text into NUMERIC(12,2). If Surtr ever built tuition from a float column instead, a half-cent value could round the other way there.

- Sibling compares' mismatchedRowCount fields keep their field-level meaning. They already report one-sided and duplicate counts as separate fields on the same line, and their summaries name each count; renaming them is a wider change than this follow-up.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1634 — fix(a8): refuse a partial community funnel Gateway snapshot by a proportional check (AERIE-2677) @kevalshahtrilogy  approved

Linear: AERIE-2677 (follow-up to PR 1572 / AERIE-2616; related AERIE-445)

## Summary

Follow-up to a non-blocking Mercy note on the already-merged PR 1572 (A8 U09, ADMISSIONS_COMMUNITY_FUNNEL_READ).

The gap. Gateway mode accepted any correctly lineaged snapshot with at least 1,000 mapped rows. A partial snapshot such as 5,000 of the usual ~5,403 deals would be published with no comparison and no degraded signal. A community funnel publish is run-scoped, so the short run would replace the last good dataset behind both the Community and the Established funnel.

The fix: a proportional check against what is currently published.

- Kit, sync/src/analytics/a8/purge-guard.ts: checkA8PopulationDrop and assertA8PopulationRetained.

- The purge guard bounds a purge by id. This is the same bound by count, for a publish that replaces a whole dataset and so has no purge to preview.

- A snapshot may be at most floor(published x fraction) rows short of the published dataset. A snapshot that is not smaller always passes; nothing published allows anything.

- It reuses checkA8PurgeBound for the rule and A8PurgeGuardError for the refusal, so messages stay count-only.

- Convex, previewCommunityFunnelPopulation (chat/convex/admissions/analytics/fullFunnel.ts), a read-only operation on the bearer-gated /sync/analytics/enrollment route.

- It returns the published run's deal count: the sum of deals over that run's full-funnel rollups. Every record lands in exactly one bucket, and the worker already refuses to publish a run whose buckets don't add up to its records. So the sum is the whole population, read from a few hundred rollup rows, not the detail rows.

- null only when no run is published yet. A published run with no rollups throws: "unknown" is never read as "nothing to protect".

- Worker, admissions-community-funnel-gateway.ts, gateway mode only:

- After the Gateway read and before anything is written, the mapped record count is checked against the published count.

- More than ADMISSIONS_COMMUNITY_FUNNEL_READ_MAX_DROP_PCT short (default 5%; unset or invalid means 5%) is refused. The domain fails as any gateway read failure does, and the last published run stays.

- The message is operator-safe (counts only) and says how to roll back or raise the bound.

- Fails closed if the published population cannot be read.

- legacy and shadow are unchanged and never call the new operation.

The other A8 Gateway readers on main. I checked each for the same floor-only weakness:

| Reader | A short snapshot can replace or delete published rows? | Result |

|---|---|---|

| Reference: programs | Yes, through purgeStalePrograms | Already bounded by the purge guard (ADMISSIONS_REFERENCE_READ_MAX_PURGE_PCT), re-checked in the purge's transaction. No change. |

| Reference: Program directory | Yes, through purgeStaleProgramDirectory | Already fraction-guarded inside that Convex mutation. No change. |

| Community funnel | Yes, the whole published run | Fixed here. |

| Community deposits | No: insert or patch per key, never a delete | No change. See "Not covered". |

| Expenses: transactions, vendors | No: upserts by key | No change. See "Not covered". |

## Business Value

- A half-loaded Surtr snapshot can no longer quietly replace the admissions funnels. Admissions leaders read conversion and commitment numbers off these two dashboards; a silent 7% shortfall would look like real families dropping out.

- It removes the last open review concern on the community funnel gate, so the gate can go through its shadow window and flip to the Gateway. That is a step toward taking the EC2 worker off direct warehouse reads (AERIE-445).

- Reusable. The count-level check sits in the A8 kit next to the purge guard, for the remaining A8 gates whose publish replaces a dataset.

## Manual Effort Estimate

About 5 hours of focused work by hand, without AI. Keval, please confirm or adjust this number. It covers:

- working out where a "currently published" population can be read cheaply (the full-funnel rollups) and that it equals the record count;

- the kit check, the Convex operation and its route, and the worker wiring;

- auditing the three sibling readers' write paths;

- about 50 tests across sync and Convex.

## Testing

- sync/src/analytics/a8/purge-guard.test.ts (+16):

- 5,000 of 5,403 is refused (403 short, limit 270), with the exact message pinned;

- the limit is exact (270 short passes, 271 does not);

- a snapshot that is not smaller passes even at 0%;

- nothing published allows anything; small datasets round the allowance down;

- impossible counts or fractions throw.

- sync/src/analytics/queries/admissions-community-funnel-gateway.test.ts (+26):

- a full-size snapshot publishes and logs the check;

- a partial snapshot above the floor is refused with the exact operator message, never falls back to pg, and logs no row value;

- the limit boundary, a larger snapshot, and "nothing published yet";

- the env override (10, 0, and an invalid value that falls back to 5% with a warning);

- fail closed when the preview throws, with the failure's own text withheld;

- a failed Gateway read never reaches Convex;

- legacy and shadow never ask for the published population;

- the acknowledgement parser refuses eleven unusable shapes (missing field, string, negative, fraction, NaN, array, bare number) and does not coerce them;

- end to end through refreshCommunityFunnel on the worker path: the preview is the first Convex call, and a partial snapshot fails the refresh with the preview as the only Convex call.

- chat/convex/admissions/fullFunnelDashboards.test.ts (+5):

- null with nothing published;

- the seeded run sums to its record count and nothing is written;

- an unpublished newer run's rollups are not counted;

- a published run with no rollups throws;

- through the HTTP route: 401 without the bearer token, 200 with it.

- Checks:

- sync: tsc --noEmit clean; vitest run --maxWorkers=2: 112 files, 2431 tests passed.

- chat: tsc --noEmit and the Convex typecheck clean. Only three targeted files were run (fullFunnelDashboards, communityFunnelDashboards, analyticsSyncSecurity: 100 tests passed). The full chat suite wasn't run locally; CI runs it.

- pnpm lint (boundaries, convex paths, read bounds, test architecture, knowledge, biome) is clean. Its two warnings are pre-existing ones in chat/skill/forge-api/scripts/sindri.mjs.

## Deploy order

- Convex before the worker runs in gateway mode. The worker needs previewCommunityFunnelPopulation; against an older Convex, gateway mode fails closed and names the fix. The gate's default is still legacy, so nothing changes until someone sets it.

- No schema change and no Surtr change.

## Not covered

- It is a count, not a set of ids. It catches a partial snapshot. It does not catch a snapshot of the right size holding different deals; the shadow window's keyed compare is what covers that.

- A legitimate drop of more than 5% (for example a bulk archive in EduCRM) is refused until someone raises ADMISSIONS_COMMUNITY_FUNNEL_READ_MAX_DROP_PCT for a cycle. The default matches the other A8 bounds; I have no production history of how much this population moves day to day.

- The check runs before the writes, not inside the publish transaction. The analytics worker is the only publisher of this domain, and Convex already refuses to let an older run replace a newer one.

- Community deposits and expenses get no proportional check. Their writes only insert or upsert, so a short snapshot cannot remove a published row. Two narrow replace-style spots remain, both already noted in those readers' own comments:

- the single expense metadata row, where a short transactions snapshot could shorten the school and account dropdown lists;

- the deposits report's newest-day view, where the first read of a day could show a short list until the next cycle.

Bounding either safely needs production numbers: deposits roll off by week, so a 5% bound could refuse a normal day. I left them for Keval to decide.

- No live Gateway run. The Surtr mart and the aerie-a8 key grant are not live yet, so gateway mode cannot be exercised end to end.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1632 — fix(xo-contractor): read a NULL trailing-52-week sum as 0 on both transports (AERIE-2675) @kevalshahtrilogy  approved

Linear: AERIE-2675 (follow-up to PR 1499; related AERIE-445)

## Summary

Follow-up to a non-blocking Mercy note on the already-merged PR 1499: a contractor with a regular invoice but no invoice in the trailing 52-week window has no 52-week sum, and that NULL must not fail the strict, all-or-nothing package refresh for every contractor.

What main already does. The pg SQL already selects COALESCE(t.paid_52w, 0) AS trailing_52w_paid_usd, and Surtr's sp_refresh_aerie_xo_contractor_package uses the same expression into a NOT NULL column. So neither transport returns NULL today, and the note is not reachable with the current SQL. Nothing pinned that, though, and the two row schemas each rejected a NULL.

What this PR changes.

- One shared parser, zeroWhenNullFiniteNumber (sync/src/redshift/xo-contractor-fields.ts). An explicit null reads as 0 ("paid 0 in the window"). Everything else is exactly requiredFiniteNumber.

- Both package row schemas use it for trailing_52w_paid_usd: the pg reader (xo-contractor-package.ts) and the Gateway reader (xo-contractor-package-gateway.ts). The two transports give the absent aggregate the same meaning by construction, so shadow compares like with like.

- Still rejected: an absent column (undefined), "", and any non-numeric value. Reading those as 0 would publish "paid 0" for every contractor if a column were renamed or dropped.

- The SQL is unchanged. A new test pins the COALESCE, the expression Surtr's mart copies.

No published row changes for current data: both sources already return a number, never NULL.

Surtr: no change needed. The mart procedure already has the same COALESCE, the column is NUMERIC(14,2) NOT NULL, and Surtr's test_sql_contracts.py pins the expression.

## Business Value

- One odd contractor can no longer take the whole headcount package refresh down. The refresh is all-or-nothing on purpose, so a single rejected row would leave about 3,000 contractors' package data stale until someone noticed.

- Closes the last open review note on the XO contractor Gateway work, which is part of moving the sync workers' warehouse reads onto the Surtr Gateway (AERIE-445).

- Keeps shadow evidence trustworthy. Both transports read the absent aggregate identically, so a future NULL cannot show up as a false parity mismatch.

## Manual Effort Estimate

About 1.5 hours of focused work by hand, without AI. Keval, please confirm or adjust this number. It covers:

- tracing the note through the pg SQL, the Gateway reader and Surtr's mart procedure and DDL;

- the shared parser and wiring both schemas;

- the unit and end-to-end tests.

## Testing

- New and changed tests:

- sync/tests/redshift/xo-contractor-fields.test.ts: zeroWhenNullFiniteNumber reads null as 0, is otherwise requiredFiniteNumber, and still rejects undefined, "", non-numeric text, booleans, NaN and Infinity.

- sync/tests/redshift/xo-contractor-package.test.ts:

- a SQL regression guard for COALESCE(t.paid_52w, 0) AS trailing_52w_paid_usd and the LEFT JOIN;

- a NULL row maps to 0 and the other row is untouched;

- "", an absent column and "n/a" still reject the whole read.

- sync/tests/redshift/xo-contractor-package-gateway.test.ts: the same NULL-to-0 mapping on the Gateway reader. null moved out of the "rejected" list; "" and an absent column were added to it.

- sync/tests/analytics/xo-contractor-package-refresh.test.ts: end to end through the real readers (only the transport is faked), one contractor with a NULL sum among 500, in pg, gateway and shadow mode. The refresh succeeds with no error and no shadowIssue, that contractor publishes 0, every other contractor is unchanged, and each record satisfies the Convex validator contract.

- Checks:

- sync: tsc --noEmit clean; vitest run --maxWorkers=2: 112 files, 2415 tests passed.

- pnpm lint (boundaries, convex paths, read bounds, test architecture, knowledge, biome) is clean. Its two warnings are pre-existing ones in chat/skill/forge-api/scripts/sindri.mjs.

- The chat suite wasn't run: this PR changes no chat code.

## Not covered

- No live Redshift or Gateway run. The change is not reachable with today's data (neither source returns NULL), so there is nothing a live run would exercise. The count of contractors currently in this state (0 of 3,046, from the follow-up brief) was not re-queried here.

- A NULL from the Gateway is now accepted as 0, where it used to fail the read. Surtr's column is NOT NULL, so this can only happen if the mart contract changes; an absent column still fails closed.

- Other XO numeric columns are unchanged. weekly_base_usd stays required and markup_from_data stays nullable: neither is a LEFT JOINed aggregate.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1627 — Add dated backup site need periods and DSS reporting @YibinLongTrilogy  approved

## Summary

Add dated backup site need periods for each Site. Each period records need dates, progress (Requested, Searching, Signing Contract, or Acquired), and the selected backup location and contract term when acquired. Dated DSS questions use these periods; the existing operating confirmation continues to answer where a Site is actually operating now.

## Screenshots

<img width="1288" height="883" alt="Screenshot 2026-10-01 at 4 21 03 PM" src="https://github.com/user-attachments/assets/81331a05-224e-492b-b2e4-7188828ebc3f" />

## Changes

- Add the need period schema, validation, Convex mutations, read model, and shared contract. Need periods can have gaps or an open end, but cannot overlap for one Site.

- Add period editing and status controls to the Site backup card, with guidance when today's plan and the confirmed operating location differ. Clarify the operating control in Backup Sites administration.

- Include need periods in the existing per-Site public API, agent, and Rhodes MCP read paths, and update the DSS instructions for dated questions and operating handoffs.

- Add focused tests for the period rules, UI, public API, agent, and Rhodes reads.

## Design decisions

- Need dates describe planned coverage; contract dates describe a location's agreement; the operating checkbox confirms actual current operation. These remain distinct so a planned handoff does not silently change the operating answer.

- Existing backup records continue to work without a backfill. This PR does not add an API write endpoint or a portfolio-wide aggregate.

## Business value

Teams can see and report which upcoming stretches have secured backup space and which still need work, including a return to an earlier location.

## Estimated manual effort

Approximately 3–4 focused engineering days without AI.

## Test plan

- [x] Focused Convex tests: 199 passed across four files.

- [x] Shared contract tests: 4 passed.

- [x] Rhodes worker backup site tests: 9 passed.

- [x] UI and public API tests passed during implementation.

- [x] Chat typecheck and pre-commit checks passed.

- [ ] Manual: add multiple periods with a gap and an open end; verify overlap rejection, acquired location and term selection, operating handoff guidance, and dated DSS answers.

#1629 — Add Preview datasets for Praxis product verification @caina-barbosa  approved

## Summary

Add the smallest domain-owned Preview dataset needed for Praxis to prepare and verify the main Aerie product surfaces against a real Convex Preview. Operations reconciles a stable three-site portfolio first; Admissions then reconciles stable programs, ontology links, published dashboard generations, enrollment cohorts, and forecast facts. Both paths fail closed unless the deployment is an isolated WorkOS Preview.

This preserves the remaining work in draft PR #1358. It does not bring over the deterministic Assistant smoke driver, System Health, identity pilot, auth bootstrap changes, revision helpers, or unrelated fallback machinery.

## Why

The current Preview preparation script already invokes portfolio/previewData:reconcileMainAppSmoke followed by admissions/previewData:reconcileMainAppSmoke, but those functions are absent from main. That leaves Praxis unable to reach product verification even when the real Preview infrastructure succeeds.

## Business value

Praxis can prepare representative, repeatable product state through the same domain writers and read models used by Aerie. This gives the capable verifier a real Portfolio, Admissions Pipeline, and Enrollment/Forecasting surface to investigate without relying on production data or provider mocks.

## How

- Add domain-owned Operations and Admissions Preview reconcilers.

- Reuse one Preview-only isolation guard from Praxis readiness and both reconcilers.

- Require Operations to run before Admissions so the canonical Site → School → Program relationship exists.

- Derive the operating school's 144-seat capacity from a completed Buildout phase using the current schema.

- Publish one coherent Admissions generation with two school years and stable synthetic identities.

- Split the 153-person enrollment cohort into 72- and 81-person mutations, preserving the proven transaction bound.

- Keep generated Convex declarations current.

- Add focused integration coverage for isolation, ordering, idempotency, published reads, stable person references, and readiness facts.

- Add Preview verification recipes to the School Portfolio, Admissions Pipeline, and Enrollment/Forecasting Journey Suites. Their status remains fresh-verification-pending until a live run proves them.

## Scope

This PR is a focused prerequisite for Praxis Preview preparation. It does not claim to supersede all of #1358.

## Validation

- Focused Preview dataset, Praxis readiness, and capacity tests

- Full Chat TypeScript check

- Architecture boundaries, Convex path, read-bound, test-architecture, and knowledge-hygiene checks

- Biome on changed Convex files

- git diff --check

## Testing contract

### What this PR delivers

A real Convex Preview can be prepared with a stable three-site Operations portfolio and a coherent published Admissions dataset. Reconciliation is repeatable and uses current Aerie domain writers and read models.

### Who uses it and where

Praxis and Aerie contributors use it while verifying the School Portfolio, Admissions Pipeline, Enrollments, and Forecast dashboards in an authenticated Aerie Preview.

### Conditions needed

- A Convex Preview configured for the exact pull request commit

- WorkOS staging authentication with the pinned Preview principal

- VERCEL_ENV=preview

- notification auto-dispatch and automation auto-drain disabled

- Operations reconciliation completed before Admissions reconciliation

### Expected behavior and examples

- Operations creates or updates exactly three stable sites: one diligence site, one buildout site, and one operating school.

- The operating school's 144-seat capacity resolves from its completed Buildout phase.

- Admissions creates or updates the stable Preview programs and canonical Site → School → Program links.

- Admissions publishes Pipeline, Enrollment, Funnel, and Forecast facts from one refresh generation across the 2026 and 2027 school years.

- Repeating both reconcilers updates the same stable records and preserves admissions person references.

- Reconciliation fails closed outside an isolated WorkOS Preview or when either isolation control is disabled.

- Admissions reconciliation fails with a clear setup error when Operations has not prepared the operating site first.

- The dashboard investigator remains free to choose its navigation and evidence based on the testing contract; this PR adds no fixed clickthrough or artifact scorer.

### Limits and unanswered questions

The three Journey Suite recipes remain fresh-verification-pending until live Praxis verification records current screenshots and values. Draft PR #1358 still owns its remaining Assistant, System Health, identity, and broader Preview work.

#2118 — docs(ai-spend): spec 14 - gpt-audio-2025-08-28 pricing, applied @kevalshahtrilogy  approved

## Summary

- gpt-audio-2025-08-28 (a deprecated dated snapshot of gpt-audio) had no row in core_finance.ai_spend_token_pricing, so its 3 scattered historical usage rows ($0.04 billed total) loaded at $0 pricing-derived cost. Run reported PARTIAL.

- Already applied to prod (2026-10-01): this PR documents the migration, per the same convention as specs 7-9, 12 and 13.

- Pricing: $2.50 in / $10.00 out per 1M text tokens, no cached-input support — identical to the already-priced gpt-audio-1.5 (spec 8), verified 2026-10-01 against developers.openai.com/api/docs/models/gpt-audio.

- Deviation from the usual reprice-via-pipeline step: with only 3 rows scattered across 4+ months and the oldest one old enough that OpenAI's usage API might not return it anymore, the standard delete-then-refetch reprice risked losing that row outright. Corrected the 3 existing rows' cost columns directly instead, using the exact formula copied from pricing.py::calculate_costs — nothing deleted, billed_cost_dollars untouched. Full reasoning in spec.md.

- Scope: features/surtr/ai-spend-pipeline/specs/ only — a new spec directory. No runner code, no infra, no CDK.

- Post-merge: no further action. The DDL and correction are already applied and verified.

## Business Value

Restores calculated spend visibility for a legacy model snapshot still seeing occasional use, clearing openai-usage-pipeline's PARTIAL status. Trivial dollar impact ($0.04) but the same silent-$0 class of bug as specs 9/12/13.

## Manual Effort Estimate

About 30 minutes: recognize the same failure shape, look up the rate (confirming it's a deprecated snapshot of an existing priced model), read the pipeline's cost formula to replicate it exactly rather than guess, and reason through why a full reprice invocation was the wrong call here given the old data's retention risk. Proposed by Claude, Keval to confirm or adjust.

## Test plan

- [x] Pricing row inserted, verified present.

- [x] All 3 existing rows corrected, verified: 0 zero-cost rows remain with real token volume.

- [ ] Confirm openai-usage-pipeline's next scheduled run no longer lists gpt-audio-2025-08-28 in unexpected_unpriced_models.

Linear: SURTR-1585

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1626 — docs(documents): clarify controlled managed relink (AERIE-2136) @marcusdAIy  approved

## Summary

- Qualify the document lifecycle restriction: a managed Due Diligence Drive relink is supported with If-Match, while unmanaged destination changes remain unavailable.

- Name relinkSiteDocument alongside updateSiteDocument as a consumer of the document revision.

- Assert the served dictionary and enablement agree on the controlled relink.

## Validation

- pnpm --dir chat exec vitest run lib/public-api/v2/domains/documents.node.test.ts --testTimeout 30000 --maxWorkers 1 (12 passed)

- pnpm exec biome check chat/lib/public-api/v2/domains/documents.ts chat/lib/public-api/v2/domains/documents.node.test.ts

- pnpm --dir chat typecheck

- git diff --check

No runtime relink or storage behavior changed.

Linear: AERIE-2136

#162 — Release: Shipyard 0.6.11 @ashwanth1109  no labels

## Summary

- Bump Shipyard to 0.6.11.

- Add the approved public release notes.

## Business Value

- Delivers the approved branch workflow, Pi conversation controls, repository history pagination, task timing visibility, and Pi image diagnostics.

## Implementation Effort

- Low: metadata-only release change; CI performs the signed build and publication.

## Test Plan

- [x] pnpm test:release

- [x] git diff --check

- [x] Verified the diff contains only package.json and releases/0.6.11.md.

#1625 — docs(admissions): align portfolio status with Milestone 10 postOpen @marcusdAIy  approved

## Summary

- AERIE-2320: correct forecast capacity-outlook portfolioStatus and the related Program Portfolio Status definition. A Program is operating when it has an active Site with a valid Milestone 10 (postOpen) completedDate, not merely when Milestone 9 is completed.

- Align the cited source inputs and provenance. Preserve the separate Milestone 9 rule that unlocks capacity exposure.

- Add a served-dictionary regression test covering both leaves and guarding the distinct capacity rule. No calculation or wire response changes.

## Validation

- pnpm --dir chat exec vitest run lib/public-api/v2/domains/admissions.node.test.ts --testTimeout 30000 --maxWorkers 1 — 30 passed.

- pnpm --dir chat typecheck — passed.

- Biome on changed files, git diff --check, and pre-commit checks — passed.

Closes AERIE-2320.

#161 — AI-960: Remove Pi permissions info box and acknowledgement checkbox from task creation @ashwanth1109  no labels

## Demo

![Smoke test demo](https://github.com/AI-Builder-Team/Shipyard/blob/ash/ai-960-remove-pi-ack/.github/smoke-test-evidence/AI-960/demo.png?raw=true)

## Summary

- Removes the amber Pi info box and the "I understand and want to use Pi for this task." checkbox from the queue's create-task form. Pi tasks are now created the same way as Codex tasks.

- Removes piPermissionsAcknowledged / pi_permissions_acknowledged end to end (createTask, the create_task command and create_with_engine), including the backend rejection branch.

- Removes PI_ACKNOWLEDGEMENT_REQUIRED and the unused acknowledgement field from agent/engine/info.

- Pi readiness still blocks creation. Add stays disabled while the check runs, with no text shown. If credentials are unavailable, the existing task-workspace__error line shows "{error} Open Agent setup & diagnostics to configure them." as soon as the check fails.

- Moves the OS-permissions disclosure to a static note on the Pi readiness card in Agent setup & diagnostics.

- Removes the unused .task-workspace__agent-warning* CSS.

## Linear

https://linear.app/builder-team/issue/AI-960/remove-pi-permissions-info-box-and-acknowledgement-checkbox-from-task

## Test plan

- [x] pnpm build (tsc + vite)

- [x] pnpm theme:check

- [x] cargo test --lib workflow:: (81 passed)

- [x] cargo test --lib agent (38 passed)

- [ ] Manual: choose Pi in the create-task form → no info box; Add is enabled once readiness passes; a task is created.

- [ ] Manual: without TrueFoundry credentials → Add is disabled and the readiness error appears under the form; switching back to Codex clears it.

#160 — AI-959: Allow switching a project's Git branch from the Projects list @ashwanth1109  no labels

## Demo

![AI-959 branch picker smoke test](https://github.com/AI-Builder-Team/Shipyard/blob/991b8551299f0bece0d54a532e6e23119050c14c/docs/smoke-evidence/ai-959-branch-picker.png?raw=true)

## Linear

https://linear.app/builder-team/issue/AI-959/allow-switching-a-projects-git-branch-from-the-projects-list

## Summary

- Backend (src-tauri/src/git.rs, lib.rs)

- list_repository_branches(path, query?, limit?, fetch?): runs git fetch --prune origin (unless fetch=false), detects the default branch (origin/HEAD, then main/master), and merges refs/heads + refs/remotes/origin by name. The default branch comes first, then the rest sorted newest committer date first. Supports a case-insensitive name search with a bounded limit (max 50). If the fetch fails, it falls back to local refs and returns fetchError. It also returns working-tree counts, so detached repos get a dirtiness check too.

- switch_repository_branch(path, branch, strategy: none|stash|discard): validates the branch name with git check-ref-format --branch and rejects option-like names. Works from ready or detached. Checks dirtiness directly. stash runs git stash push --include-untracked with a Shipyard message and never re-applies the stash. discard runs reset --hard HEAD plus clean -fd, which keeps ignored files. Local branches use the existing checkout_branch() (now ---terminated). Origin-only branches use checkout --track -b <b> origin/<b>. Returns the refreshed RepositoryStatus. Selecting the current branch does nothing.

- Frontend

- New shared RepositoryBranchPicker: the branch name / "Detached HEAD" in the BRANCH block is now a trigger with a chevron. It opens a searchable popover (SearchInput) showing a "Fetching origin…" state, the current branch marked, default/origin tags, a search debounced at 180 ms (limit 20, no re-fetch), an empty state, a fetch-failure warning, and keyboard navigation.

- ProjectDetail: new row state (isSwitchingBranch, pendingBranchSwitch). The picker is disabled while the row is loading, syncing, or switching. A dirty tree opens an inline alertdialog with Cancel / Stash & switch / Discard & switch. Success shows a toast (it mentions tracking, stash, or discard). On failure, an error toast appears and the row is re-inspected. History resets on branch change and reloads if it was open.

## Test plan

- cargo test --lib git::: 25 passed (11 new: default-first and recency order, search, remote-only listing, fetch-failure fallback, clean/stash/discard/remote-tracking/detached switches, dirty-without-strategy refusal, missing/invalid branch, checkout failure under index.lock, no-op).

- node --test scripts/test-project-releases.mjs: 17 passed (5 new picker/switch UI tests).

- pnpm build, pnpm theme:check, pnpm test:architecture, pnpm test:search-input: all pass.

Note: switching, stashing, or discarding changes the shared project folder used by any task or agent that points at the same path.

#159 — AI-954: Show task creation time and time taken on task detail @ashwanth1109  no labels

## Demo

![AI-954 smoke test demo](https://github.com/AI-Builder-Team/Shipyard/blob/040a229e5e7e7568b55239d62958928ff28ac427/docs/smoke-tests/ai-954/demo.png?raw=true)

## Summary

- Adds a nullable Task.created_at (Unix seconds) via an idempotent migration. Existing rows stay null, with no backfill. Both task-creation paths set it to now().

- Task payloads now include createdAt, activeAgentSeconds, runningAgentTurns, and agentRuntimeMeasuredAt. Agent runtime is the sum of TaskTraceTurn durations. A turn that is still open counts up to now, or up to completed_at if the task is complete.

- Task detail header: the agent label row shows, right-aligned, a clock icon with the wall-clock duration and a bot icon with the agent runtime. Tooltips and accessible names give the full labels (Took 2h 14m or Running 45m, and Agent runtime 38m). For in-progress tasks these refresh every 30s, and the count includes turns that are still running.

- Task context sidebar: shows Created 3 days ago in a <time> element under the linked item/repository summary. Hovering shows the exact local date and time.

- Legacy tasks (createdAt null) show none of these values.

## Linear

https://linear.app/builder-team/issue/AI-954/show-task-creation-time-and-time-taken-on-task-detail

## Tests

- cargo test --lib: 340 passed. New tests cover the idempotent migration, legacy rows staying null, created_at on new tasks, and agent-runtime aggregation with open turns.

- node --test scripts/test-task-timing.mjs (new, pnpm test:task-timing): duration and relative-time formatting, completed/running/legacy summaries, and the UI hiding values for legacy tasks.

- node --test for task-workspace, workflow, task-trace, pr-review, and github-reviews: all pass.

- pnpm theme:check passes. tsc --noEmit reports only an existing, unrelated missing @tauri-apps/plugin-notification type in a worktree that has no local install.

#2117 — fix(education): reconcile Alpha OKC QuickBooks Class replacement @caina-barbosa  approved

## Summary

This is a standalone production incident-repair slice for the Core Education Ontology refresh; no parent or child Linear issue was supplied. It adds a read-only reconciliation for the governed replacement of the deleted Alpha Oklahoma City QuickBooks Class with the active Alpha Edmond Class, plus a contract test that pins the exact source identities and prevents warehouse mutation.

Production effect: dormant/additive. Merging this PR does not mutate Aerie, invoke a pipeline, or write warehouse data. The authorized Aerie repair was performed separately and has already flowed through the normal scheduled Rhodes → Core publication path.

---

## Why

An active Aerie SchoolLink referenced a QuickBooks Class that no longer existed in the accepted source directory, so Core correctly failed closed for 29 consecutive runs. The repair must preserve the old relationship as history, represent the successor with a new link, and prove that both the source-faithful and governed warehouse layers agree without introducing a direct Redshift patch.

---

## Business Value

- Restores the hourly Core Education Ontology refresh.

- Preserves the historical Alpha Oklahoma City Class relationship for auditability.

- Makes the financially active Alpha Edmond Class mapping explicit instead of relying only on the broader company fallback.

- Leaves a repeatable, source-controlled proof of the exact production repair.

---

## How does it work

1. The reconciliation declares the exact School, old SchoolLink/Class identity, and replacement SchoolLink/Class identity.

2. It obtains rhodes_run_id from the one current dim_school row at the table's school_id primary-key grain, then binds the source-faithful Rhodes rows and governed Core rows to that same publication. The aggregate is scalar normalization over at most one row; it does not choose among historical UUID runs.

3. It proves that the old relationship is archived and historical across raw_school_links, bridge_school_link, and xref_school_source.

4. It proves that the replacement relationship is active and current across those same surfaces, including the exact QuickBooks realm/Class source key.

5. It emits only violations; zero rows is success. The contract test also rejects warehouse write or schema-change statements in this artifact.

---

## Scope

### Included in this phase

- Exact, read-only post-publication reconciliation for the Alpha Oklahoma City Class replacement.

- Contract coverage for identities, warehouse surfaces, expected outcomes, and read-only behavior.

- Exact final diff paths:

pipelines/runners/core-education-ontology-refresh/reconciliation/2026-10-01_alpha_oklahoma_city_qbclass_replacement.sql

pipelines/runners/core-education-ontology-refresh/tests/test_sql_contracts.py

### Deliberately excluded for later phases

- Direct Redshift DML or a warehouse migration — Aerie is authoritative and the Core stored procedure is the sole warehouse writer.

- Core procedure/schema changes — the existing fail-closed behavior worked as designed.

- Pipeline scheduling, retry, notification, UI, and API changes — no runtime code defect was found.

---

## Test plan

### Automated validation

- SQL contract suite — 23/23 passed (uv run pytest -q tests/test_sql_contracts.py)

- Full runner suite — 89/89 passed (uv run pytest -q)

- Ruff lint — passed (uv run ruff check src tests scripts)

- Ruff format — passed (uv run ruff format --check src tests scripts)

- DDL application dry run — 271 statements parsed (uv run python scripts/apply_ddl.py)

- git diff --check origin/main...HEAD — passed

- Exact-head diff scope — only the two authorized paths listed above

### Time for Implementation

Approximately 1–2 engineer days without AI assistance, including source/warehouse investigation, governed source repair, scheduled publication monitoring, reconciliation authoring, and review evidence.

### Manual QC

The authorized source repair was performed in production Aerie: SchoolLink x57g2vvwr1af2esn30sw2jzmvh8exzj4 was archived and replacement SchoolLink x57jvf105j336m018pqp94awdx8fewm5 was added for QuickBooks Class identity qbcl_01M35AHRZ85Q8W0AF2DSQCMWZA.

The next scheduled Rhodes run 52238b12-c585-4af4-919a-48028a263487 succeeded with 18/18 entities, then automatically triggered a successful Core refresh. The repository verifier reported 113 schools, 495 links, 572 xrefs, zero unresolved active links, zero duplicate current xrefs, and matching active-link/current-xref counts for every external target type. Running the new reconciliation against that publication returned zero violation rows.

## Review repairs and contract clarifications

- [Mercy review round 1](https://github.com/AI-Builder-Team/Surtr/pull/2117#pullrequestreview-5381128962) at head 2ea29146852fcacaedbe0563636a7d615b64f035 was classified as a false-premise blocker; no repair ticket was opened because no reachable production defect exists.

- The finding assumes MAX(rhodes_run_id) chooses among multiple historical UUIDs. It does not: core_education.dim_school is a current-state table with PRIMARY KEY (school_id), and this query filters one exact school_id before applying the scalar aggregate. Its input cardinality is therefore zero or one.

- core_education.sp_refresh_aerie_ontology atomically removes and republishes dim_school, bridge_school_link, and xref_school_source from the same selected p_rhodes_run_id. These tables do not retain older publication snapshots.

- Live Redshift verification found one row for sch_01KXS9P7YG09H3HMNAPFKSX838, one distinct rhodes_run_id across all 113 dim_school rows, and one distinct run ID in each source/Core surface used by this reconciliation. The alleged arbitrary-old-snapshot path is therefore unreachable.

- Scope and production effect are unchanged: this remains a read-only, dormant/additive reconciliation with no deployment, traffic, migration, or warehouse write.

- Validation remains 23/23 focused tests and 89/89 full runner tests, with Ruff lint/format, the 271-statement DDL dry run, all hosted CI checks, and live zero-row reconciliation passing.

#158 — AI-953: Support queueing and steering messages in Pi conversations @ashwanth1109  no labels

## Demo

![AI-953 smoke test: Pi queue and steer](https://github.com/AI-Builder-Team/Shipyard/blob/0b3798b9f95586e7777b214686ec2d888045185d/.github/smoke-evidence/AI-953/demo.png?raw=true)

## Summary

Pi conversations now support the same queue and Steer behavior as Codex.

- Queue: while Pi is working, a composer submit clears the draft and adds the message to Queued messages. After a turn succeeds, queued messages auto-send in order.

- Steer: sends a queued message into the running Pi turn using Pi's native RPC steer command. Pi delivers it after the current tool calls and before the next LLM call. The message appears as a user message inside that turn, both live and after reload.

- Remove, pause after stop/failure, and Retry behave as they do in Codex.

## Implementation

Backend

- AgentEngine::steer_turn(request, expected_turn_id) is a new trait method. Its default returns UnsupportedCapability(Steering).

- agent/turn/steer (agent_steer_turn) is a new provider-neutral command. It:

- checks the engine's steering capability;

- rejects companion threads;

- checks that the accepted turn ID matches the expected one;

- records the steer as a task-trace correction using the new engine-aware capture_direct_user_message_for_engine.

- Pi now reports steering: true.

Pi adapter (pi_agent.rs)

- Turn gating: a steer is only sent when the expected turn ID matches the turn the Pi process is running. A per-process gate rejects steers once the turn has settled or been stopped. A steer already in flight finishes before the turn's undelivered list is collected.

- Live delivery: a delivered steer is published as its own item, {turn}:steer:{clientCommandId}, carrying clientId. It reconciles with the optimistic message and never overwrites the prompt's {turn}:user item.

- History: Pi doesn't record whether a user entry came from a prompt or a steer. Shipyard stores steered entry IDs in shipyard-steers.json inside the session directory, and turns_from_entries_with_steers uses it to keep those entries in the running turn.

- Undelivered steers: if the turn ends before Pi delivers a steer (it settled, was stopped, or failed), Shipyard handles it as follows:

- on Stop, it calls clear_queue before abort;

- the terminal turnCompleted event returns the steer as unsentSteers;

- the frontend puts it back at the front of the queue and drops its optimistic transcript message.

Frontend

- PI_CAPABILITIES.steering is now true.

- Steer uses commandFor(engineId, "codex_steer_turn", "agent_steer_turn").

- Notices name the engine (Pi or Codex).

- New helpers: unsentSteersFromPayload, restoreUnsentSteers (idempotent across replayed events), and removeOptimisticCodexCommands.

Not changed: the Codex steer and queue paths.

## Tests

- cargo test --lib: 346 passed. New tests cover:

- default unsupported steer;

- steer turn-ID parsing;

- Pi capability;

- steer gating and closing;

- rejected steers not being tracked;

- steered history grouping;

- entry resolution;

- live steer item events;

- private steer records.

- node --test scripts/test-codex-messages.mjs: 34 passed (Pi notices; undelivered steers restored once and in order).

- node --test scripts/test-codex-conversation-store.mjs: 31 passed (optimistic steer removal; live Pi steer item reconciliation).

- node --test scripts/test-thread-recovery.mjs: 50 passed, 1 failed. The failure is Pi omission history restores the local image…, which fails the same way on origin/main. New hook tests cover:

- Pi queue, then Steer through agent_steer_turn, then auto-send after completion;

- an undelivered steer returning to the queue and pausing after Stop.

- python3 -m unittest discover -s scripts/smoke -p 'test_*.py': OK. Adds a Pi fixture test for steering into a held turn, plus rejection when the session is idle.

- test-companion, test-chat-content, test-codex-detail, test-task-trace, and test-architecture: pass.

- No theme tokens changed.

## Linear

https://linear.app/builder-team/issue/AI-953/support-queueing-and-steering-messages-in-pi-conversations

#1624 — docs(directory): explain typed relationship endpoint IDs @marcusdAIy  approved

## Summary

- AERIE-2114: document from and to relationship endpoints in the served Directory dictionary. Site endpoints nest their canonical ID at from.site.id/to.site.id and have no flat endpoint id; School, Program, Market, and Metro endpoints use flat id.

- Give directory.resolve-site-program the concrete to.site.id match and from.id Program lookup for operatesAt edges.

- Add a served dictionary/enablement regression test. No wire shape, query, or authorization changes.

## Validation

- pnpm --dir chat exec vitest run lib/public-api/v2/domains/directory.node.test.ts --testTimeout 30000 --maxWorkers 1 — 11 passed.

- pnpm --dir chat typecheck — passed.

- Biome on the changed files, git diff --check, and pre-commit checks — passed.

Closes AERIE-2114.

#1623 — fix(admissions): return allowed portfolio statuses for invalid capacity filters @marcusdAIy  approved

## Summary

- AERIE-2341: remove the capacity-outlook portfolioStatus OpenAPI regex that rejected malformed filters before the existing handler could explain the allowed values.

- Retain a bounded string schema and the handler's status validation. No valid filter or capacity calculation changes.

- Add an authenticated HTTP regression test for unknown tokens, a trailing comma, and embedded whitespace; each receives a 400 with portfolioStatus and the allowed statuses.

## Validation

- Focused convex/publicApi/v2/admissions.test.ts invalid-filter test — passed.

- pnpm --dir chat typecheck — passed.

- Biome on changed files and git diff --check — passed.

- Pre-commit Convex paths, Biome, and chat typecheck — passed.

Closes AERIE-2341.

#1622 — fix(admissions): explain legacy appointment-driven Shadowing @marcusdAIy  approved

## Summary

- Restore a detail-panel explanation only for retained pipeline rows with stageResolutionReason: "shadow_appointment".

- State that this is a legacy snapshot reason; current pipeline stages use Finalsite status.

- Test both the legacy row and a status-driven row, without changing stage calculation, storage, validators, or dbt.

## Validation

- pnpm --dir chat exec vitest run components/dashboards/admissions/pipeline/pipeline-record-panel.test.tsx --testTimeout 30000 --maxWorkers 1 — 37 passed.

- pnpm --dir chat typecheck — passed.

- Biome on changed files, git diff --check, and pre-commit hooks — passed.

Closes AERIE-2350.

#1621 — fix(insights): clarify overdue UTC default and portfolio-health status prose @marcusdAIy  approved

## Summary

- AERIE-2118: state the UTC default and date-only consequence for overdue work in the served dictionary and enablement workflow.

- AERIE-2079: replace retired open/closed terms in portfolio-health response methodology with the active/paused membership vocabulary; bump the Insights methodology version and align its response example.

- Add focused tests for the served documents and response metadata. No membership or calculation behavior changes.

## Validation

- pnpm --dir chat exec vitest run lib/public-api/v2/domains/insights.node.test.ts convex/publicApi/v2/insights.test.ts --testTimeout 30000 --maxWorkers 1 — 25 passed.

- pnpm --dir chat typecheck — passed.

- pnpm exec biome check on the four changed files — passed.

- pre-commit Convex paths, Biome, and chat typecheck — passed.

Closes AERIE-2118, AERIE-2079.

#157 — AI-952: Fail fast when the bundled Pi runtime cannot see images @ashwanth1109  no labels

## Demo

![AI-952 smoke test: Pi receives attached Retina screenshot](https://github.com/AI-Builder-Team/Shipyard/blob/c44bd8ec777f846e2888a721291e588563509303/docs/smoke-evidence/AI-952/demo.png?raw=true)

## Summary

Pi turns every image into [Image omitted: could not be resized below the inline image size limit.] if it can't import Photon (@silvia-odwyer/photon-node). That covers both inline attachments and read on saved files. 0.6.9 shipped without Photon. AI-946 added the packaging, and 0.6.10 includes it: the installed resources/pi is layoutVersion: 2 with photonVersion: 0.3.4. This PR adds the regression checks and the copy fix the ticket asks for, so this failure can't ship without anyone noticing again.

- Real image probe (scripts/pi-image-probe.mjs): uses the bundled Node to import Pi's own resizeImage from the staged or packaged bundle, which resolves Photon the same way the RPC process does. It resizes a committed 2624×1644 Retina fixture (scripts/fixtures/pi-retina-screenshot.png) and exits non-zero if Pi gets null.

- Build/release gates: the probe runs in pnpm stage:pi (staging itself fails), in pnpm test:pi-package (now also run in the release workflow), in scripts/release/audit.py against the packaged Shipyard.app (the result is recorded in the private audit JSON), and in pnpm smoke build against the smoke bundle. Each gate has a negative test that removes Photon and expects a failure.

- Runtime check (pi_agent.rs): before strict validation, a layoutVersion: 1 manifest, a missing photonVersion, a mismatched Photon version or integrity, or a missing or unlisted photon_rs.js / photon_rs_bg.wasm / package.json now fails with Pi image support unavailable: … Reinstall or rebuild Shipyard so the bundled Pi runtime includes Photon 0.3.4.

- Fallback notice: now reads "Pi could not see this image. It is shown here only; Pi received no visual input." It no longer suggests re-reading the saved file, which goes through the same pipeline.

- README: added a Pi runtime note (pnpm install + pnpm stage:pi for local debug builds).

## Verification

- pnpm test:pi-package: 2/2. This includes staging with the real pinned Node: 2624×1644 → 2000×1253. Without Photon it fails with "Pi image support unavailable".

- python3 -m unittest in scripts/release (test_audit): 7/7. pnpm test:release: all passing.

- pnpm test:smoke: 33/33.

- cargo test --lib pi_agent::: 21/21, with new layout-1 and missing-Photon tests.

- pnpm test:messages: 32/32. pnpm exec tsc --noEmit: clean.

- Ran verify_pi_image_support against the installed /Applications/Shipyard.app (0.6.10): 2624×1644 → 2000×1253 PNG, so the released package gives Pi real image input.

Not done here: no new release version bump, since 0.6.10 already ships Photon. There's also no live model end-to-end chat check, because that needs a TrueFoundry credential and an interactive session.

## Linear

https://linear.app/builder-team/issue/AI-952/pi-omits-image-attachments-because-the-bundled-runtime-is-missing

#156 — AI-951: Add Prev/Next pagination to project repository commit history @ashwanth1109  no labels

## Demo

![AI-951 smoke test: commit history paging (Commits 21–30) with Prev/Next](https://github.com/AI-Builder-Team/Shipyard/blob/codex/ai-951-commit-history-pagination/docs/smoke-evidence/ai-951/demo.png?raw=true)

Smoke test: PASS (user-verified in the codex/ai-951-commit-history-pagination dev build).

## Summary

- get_repository_history now accepts optional offset/limit (default 10, clamped 1–100), runs git log --skip=<offset> --max-count=<limit+1>, and returns offset, limit, and hasMore.

- Project Detail commit panel gains a footer with Prev / range label / Next; the header shows the visible range ("Commits 11–20") instead of the per-page count.

- Prev disabled on the first page; Next disabled when hasMore is false or while loading/erroring. Footer hidden when history fits on one page.

- Retry reloads the failed page; collapsing/reopening (or branch change) returns to page 1. Stale page responses are dropped.

- Accessibility: "Show newer commits" / "Show older commits" labels, aria-live="polite" range, toggle label now "Show/Hide commit history for …".

## Linear

https://linear.app/builder-team/issue/AI-951/add-prevnext-pagination-to-project-repository-commit-history

## Tests

- cargo test --lib git:: (new repository_history_pages_older_commits_with_skip_and_has_more)

- node --test scripts/test-project-releases.mjs (new paging, loading-disabled, and retry-failed-page tests)

- pnpm theme:check, pnpm build

#2116 — fix(education): accept GuidePlatform audio_recordings/audio_transcripts additive schema drift (SURTR-1582) @kevalshahtrilogy  approved

## Summary

guide-platform-raw-sync's GuidePlatform Postgres source drifted again — the 6th schema-drift incident on this runner in under two weeks:

- audio_recordings gained room_levels (nullable text[]) — SURTR-1582

- audio_transcripts gained request_shape (nullable jsonb), speech_model_used (nullable text), and response_meta (nullable jsonb) — SURTR-1582

- audio_utterances gained speaker_confidence and min_word_confidence (both nullable numeric(4,3)) — found in the same live introspection, filed separately as SURTR-1583 since it's a distinct addition on an unrelated relation

Every scheduled run has failed exact-shape validation before extraction, landing, or publication since ~2026-09-29/09-30, observed still failing through 2026-10-01 05:35 UTC. The whole-source shape check fails every table until every drifted table is caught up, so both SURTR-1582 and SURTR-1583 are fixed in this one PR/release — fixing only the first two tables (the original scope) would have left the pipeline failing closed on audio_utterances alone.

The migrations. ddl/007_audio_recordings_audio_transcripts_schema_additions.sql (two relations, one atomic batch) and ddl/008_audio_utterances_schema_additions.sql (one relation, its own atomic batch, applied independently), both following the 006-hardened pattern exactly (SURTR-1519, PR #2060 then #2080): a plain CREATE VIEW after DROP for each view, a live owner/ACL snapshot-and-replay per relation instead of a hardcoded baseline, and same-transaction pre/postcondition guards. scripts/run_ddl.py's two-relation machinery is generalized into an N-relation relation_schema_additions_migration_execution_statements helper, reused by both migrations (one with a 2-tuple of relations, one with a 1-tuple).

Contract propagation. SOURCE_CONTRACT_VERSION moved twice in building this PR: 1f882787… → aa2da2e2… (audio_recordings/audio_transcripts only) → bf4df98c… (final, also covering audio_utterances). ddl/001_create_staging.sql is regenerated directly from live introspection via the runner's own generate_contract.py (not hand-patched) for the final state. Every consumer hash pin was re-grepped from scratch (not assumed from the first commit's list) and updated: mart-education-guide-roster-refresh's src/handler.py, scripts/verify_source_inventory.py, ddl/001_guide_roster_evidence.sql, tests/integration/redshift_harness.py, tests/test_sql_contracts.py, plus two dated consumer migrations — ddl/006_20261001_source_contract.sql (intermediate, now historical) and ddl/007_20261001_source_contract.sql (final, the one apply_ddl.py now points at) — mirroring exactly how ddl/005/ddl/006 record the September 24 incident's own intermediate/final split. The three workforce-contracts spec docs that pin this hash are updated to the final value only.

A real bug caught by the runner's own tooling. While building the audio_utterances migration, scripts/generate_contract.py --check (run against live data, not my own file) reported the checked-in contract as stale even after I believed I'd fixed it — it turned out speaker_confidence/min_word_confidence were in the wrong order relative to the live source. I'd inferred their order from an earlier unordered set-difference instead of querying information_schema.columns directly; the true order is speaker_confidence (ordinal 9), then min_word_confidence (ordinal 10) — reversed from what I'd first written. Fixed in contract.py/ddl/001 (regenerated straight from a fresh live introspection, not hand-edited) and in ddl/008. The already-applied live view (from an earlier apply in this same session, before the bug was caught) was corrected via a second application of the same grant-snapshot-and-replay primitives run_ddl.py --apply itself uses — its own precondition guard correctly refused to blindly re-run (it expects a from-scratch "nothing migrated yet" precondition, and the live state was "already migrated once, just in the wrong order"), so the correction invoked those same functions directly, with the live state (owner, grants, column presence) independently queried and verified both before and after rather than embedded as an atomic SQL guard. Full before/after verification is in the Test plan below.

ddl/007 header fix (Mercy's non-blocking finding). The header said the migration was "applied from a state where neither view yet exists," which doesn't match its first statements being unconditional DROP VIEW (which requires the view to already exist). Corrected to describe the real precondition: both views exist in their pre-migration shape, and plain CREATE VIEW's job is to refuse if something unreviewed recreated a view between its DROP and its CREATE, not to tolerate a no-view-at-all bootstrap. ddl/008 was written with the corrected wording from the start.

Whether/how applied live. Both ddl/007 and ddl/008 applied live via scripts/run_ddl.py --apply --apply-migration <name> (REDSHIFT_DB_USER=admin), each after confirming sys_query_history was clear of conflicting DDL immediately beforehand. Re-verified with scripts/generate_contract.py --check (exit 0) and a direct clean_catalog_errors() check across all 57 tables (zero errors) that the producer's checked-in contract now exactly matches live reality, end to end. Did NOT apply either consumer-side migration (ddl/006_20261001_source_contract.sql or ddl/007_20261001_source_contract.sql): schema-drift-2026-09-24.md documents an unresolved, unrelated access-visibility gap (MIGRATION_DATABASE_FUNCTION_ACL_SQL preflight can't see a database-level grant as the non-superuser migration owner) that blocked the equivalent consumer migration for the September 24 incident; this PR doesn't attempt to work around it. Applying the consumer migration is a separate, gated operational step per contracts/schema-drift-2026-09-29.md's recovery gates.

Scope. pipelines/runners/guide-platform-raw-sync and pipelines/runners/mart-education-guide-roster-refresh (contract-hash pins only, no business logic changes there) plus the workforce-contracts spec docs that pin the same hash.

Post-merge. Deploy the producer image so the next scheduled run (05:35 UTC) stops failing. The consumer's hash pin only matters once it reads a ledger row stamped with the new contract version — safe to deploy on its own timeline, but blocked on the access-visibility gap above until that's resolved separately.

## Business Value

Restores the daily GuidePlatform raw sync, which has been failing closed on every scheduled run for 2+ days — no new data for any of the 57 tables in this shadow sync has landed since the drift began, since the whole-source shape check fails before any table is extracted. This is infrastructure reliability work: it doesn't ship a new feature, but it un-blocks an existing pipeline every downstream consumer of this shadow sync depends on, and — after a reviewer's catch mid-PR — it fully resolves the drift in one release instead of leaving a third table's drift to cause the exact same failure mode again on the very next scheduled run.

## Manual Effort Estimate

Proposed: 3.5–4.5 days of focused senior-engineer time (revised up from the original PR's 2.5–3.5-day estimate for the two-table fix alone). The added half-to-one day reflects: generalizing the two-relation migration machinery into an N-relation one without regressing either call site's tests, re-tracing and re-propagating a *second* contract-hash bump across two runners and three spec docs, and — the most expensive part — catching and correctly recovering from a live column-order bug mid-build: diagnosing it from the runner's own --check tooling, fixing the contract/migration/tests, and safely correcting an already-applied live view without forcing past the safety guard that (correctly) refused to let the buggy migration simply re-run.

Proposed by Claude, Keval to confirm or adjust.

## Test plan

- [x] cd pipelines/runners/guide-platform-raw-sync && uv sync --all-extras -q && uv run pytest — 199 passed, 1 skipped (up from 190 after the first commit; 9 new tests for the audio_utterances migration mirroring the same rigor as the two-relation suite: canonical-view match with no grants, grant-replay ordering, preflight/guard pinning, client-token binding and length, a fake-Data-API harness proving a full drop+recreate from the documented precondition, CREATE VIEW refusing against an unexpectedly-present view, non-submission when live state differs from the snapshot, and a dry run touching no warehouse)

- [x] cd pipelines/runners/mart-education-guide-roster-refresh && uv run pytest — 78 passed, 5 skipped, including the self-mutating test that verifies every consumer pin surface moves together

- [x] uvx ruff check pipelines && uvx ruff format --check pipelines (CI-pinned ruff==0.15.22) — clean

- [x] scripts/generate_contract.py --check — exit 0: the checked-in contract now exactly matches a fresh live introspection (this is what caught the column-order bug in the first place, and confirms it's genuinely fixed, not just internally self-consistent)

- [x] Direct clean_catalog_errors() check across all 57 tables against the live Redshift catalog — zero errors (confirms exact column name/type/ordinal-position match, not just presence)

- [x] ddl/007 applied live via scripts/run_ddl.py --apply; owner and the complete relacl/svv_relation_privileges set verified identical before/after on both relations

- [x] ddl/008 applied live via scripts/run_ddl.py --apply; then, after the column-order bug was found, corrected via a second application of the same grant-snapshot-and-replay functions (not run_ddl.py --apply, whose own from-scratch precondition correctly refused to re-run against an already-migrated relation). Owner and the complete grant set verified identical across all three applications; final column order independently confirmed against information_schema.columns on the live GuidePlatform source

- [ ] Either consumer-side source-contract migration apply — not run, see Summary (pre-existing, documented, unrelated blocker)

Linear: SURTR-1582, SURTR-1583

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1571 — feat(a8): G3 gate COMMUNITY_DEPOSITS_READ + keep every community deposit (AERIE-2615, AERIE-2353) @kevalshahtrilogy  approved

Linear: AERIE-2615 (A8 unit U10). Fixes AERIE-2353. Related: AERIE-445.

## Summary

A8 group G3, community deposits. The analytics worker's queryCommunityDeposits read can now come from the Surtr Gateway parity mart instead of Redshift, behind a gate that defaults to the pg read. This PR also fixes AERIE-2353 in the mapper both transports share, so the Admissions → Community deposits report now shows every deposit.

### AERIE-2353: what changes in the report

On live data (360 has_fact rows, read-only check on 2026-10-01):

| | before | after |

|---|---|---|

| deposits published | 353 (7 with no community silently dropped) | 360 |

| deposits with a made-up parentId: 0 | 8 | 0 (stored with no parentId) |

| other rows the mapper rejects | dropped, rest published | 0 today; any would fail the refresh |

The report gains an "Unassigned" row holding the 7 deposits EduCRM has no community for. Its total counts in the matrix totals.

### The three decisions (made on Keval's behalf, 2026-10-01)

mapCommunityDepositRows (shared by the legacy pg read and the gateway read) is now all or nothing.

1. A NULL community is kept under an explicit bucket, UNASSIGNED_COMMUNITY = "Unassigned".

- No existing convention covered a missing community or program bucket. The repo only has ad hoc inline "Unknown" fallbacks for one contact's program name, and "Unassigned" is already the app's word for a missing owner. So this is a new named constant.

- The report groups by community dynamically (community-view.tsx, getCommunityDepositsData, the v2 summary endpoint), with no fixed list, so the bucket shows with no UI change. A UI test covers it.

- The mart's x_total_deposits is the community's has_fact row count. This held on all 42 named communities, and the value is NULL for the NULL-community rows. So the bucket's total is its row count in the read, not a coerced 0.

2. A NULL parent_id is absent, never 0.

- Convex communityDeposits.parentId is now v.optional(v.number()) in the schema and in the insertCommunityDeposits validator. Existing documents keep validating.

- CD deploys Convex before the EC2 worker, so the worker never sends a missing parentId to the old validator.

- The mutation's same-day patch now clears a parentId that an earlier cycle wrote for the same key (for example the old 0). patch leaves an absent field alone, so this needs an explicit undefined.

- Consumers: no UI or public API (v1 or v2) projects parentId. No join or dedupe keys on it: Convex keys deposits by (community, childId, snapshotDate), the shadow compare by (community, childId), and the detail panel groups families by email, then name.

- Documents written before this fix, on earlier snapshot dates, still carry parentId: 0. Nothing reads them.

3. Any other row the schema rejects fails the refresh. Nothing is published.

- The read throws a CommunityDepositsValidationError that names columns and counts only, e.g. 1/360 row(s) failed validation (child_id on 1 row(s)).

- Columns that z.coerce would have turned from NULL or a malformed value into a plausible one now reject it:

- BIGINTs (weeks_ago, child_id, parent_id, and x_total_deposits on a named community) accept only a non-negative safe integer, as pg's decimal string or a number.

- The three date columns accept only a valid Date or non-empty text.

- An empty read is refused too, as the Gateway already did. It now fails the domain instead of reporting a successful zero count.

- The refresh's existing catch fails only the communityDeposits domain. There are 0 such rows today.

- The set is written in one insertCommunityDeposits call, so one Convex transaction. It used to go in 100-row batches, so a failure after the first batch left a partial snapshot for the day. 360 rows is far inside Convex's per-transaction limits.

Both transports run the same fixed mapper. The Gateway reader's AERIE-2353 parity special-casing is gone: the known-drop check and the non-NULL column list now live in the shared schema. A mapper refusal on the Gateway side is rethrown as an A8GatewayReadError with the same message, so the shadow degraded line and gateway mode's failure line say why. The shadow compare still compares mapped records, and the NULL-community rows now take part, under Unassigned, on both sides.

### The gate (unchanged from the earlier revision)

- Gate COMMUNITY_DEPOSITS_READ=legacy|shadow|gateway (default legacy). The rules come from the U02 kit (a8/read-mode.ts).

- queries/community-deposits-gateway.ts reads aerie-admissions-community-deposit through the kit, typed as the Surtr DDL stores it (SUPER, BIGINT, DATE, TIMESTAMP). The kit's type-parity layer runs first, then the shared mapper.

- One publish order for both transports (orderForConvexPublish). Convex's insertCommunityDeposits is a last-write-wins patch by (community, childId, snapshotDate), and legacy SQL ties and Gateway mart_row_id order are both unspecified. So shadow and gateway publish in (community, weeksAgo, canonical encoding) order, and a clean shadow means the same Convex outcome. legacy keeps the SQL's own order.

- Shadow publishes the pg records and compares after mapRows, before the Convex collapse, as a multiset keyed by (community, childId), with the source_advanced skew rule.

- PII: counts only. The compare spec allowlists no value. Log lines and refusal messages carry counts, column names and value types.

- Failure semantics. A failed read in any mode fails only the communityDeposits domain; the cycle carries on. Population floor 1, freshness bound 6h.

- Dry-run: sync/src/scripts/dry-run-community-deposits-shadow.ts runs the worker's shadow path read-only. It now prints the mapper's (value-free) refusal reason.

## Business Value

- Families stop disappearing from the Community deposits report. 7 of 360 deposits (about 2%) had been missing for as long as the mart has left them without a community. Admissions now sees all of them, in an Unassigned row they can chase.

- No more fake data. 8 deposits carried a made-up parent id 0. Any future malformed row now fails loudly instead of quietly shrinking the report.

- It moves the third A8 group onto the governed path. The read can move off the EC2 worker's direct Redshift access onto the Surtr Gateway with per-row lineage, a step toward retiring the worker (AERIE-445). Both transports now share one correct mapper, so cutover changes the transport only.

- PII stays out of logs by construction.

- The gate carries no risk until it is switched on. The default is legacy, and rollback is one env var. The AERIE-2353 fix does change the legacy path on deploy, by design.

## Manual Effort Estimate

About 10 hours of focused work by hand, without AI. Keval, please confirm or adjust this number. It covers:

- the gate and Gateway reader (about 7h, as estimated for the earlier revision);

- about 3h more for the AERIE-2353 fix:

- tracing every consumer of the table and of parentId;

- the Convex schema and patch change;

- the all-or-nothing mapper;

- about 25 more tests;

- the read-only before/after proof against Redshift.

## Testing / evidence

- Read-only pg proof on live data (2026-10-01). A throwaway script ran COMMUNITY_DEPOSITS_SQL and the mapper before and after the change. It printed counts only.

| check | before | after |

|---|---|---|

| has_fact rows | 360 | 360 |

| NULL community / parent_id / x_total_deposits | 7 / 8 / 7 | same |

| published records | 353 | 360 |

| dropped by the mapper | 7 | 0 |

| published parentId: 0 | 8 | 0 (8 with no parentId) |

| published under Unassigned | n/a | 7 |

| other rejects (NULL or malformed) | n/a | 0 (no throw) |

| distinct Convex keys (community, childId) | 353 | 360 (no collapse) |

Also checked:

- The NULL x_total_deposits rows are exactly the NULL-community rows.

- On all 42 named communities, x_total_deposits equals the community's has_fact row count.

- No real parent_id is 0, and no community is literally named "Unassigned".

- Unit tests:

- educrm.test.ts:

- NULL community → Unassigned, with the bucket's row count as its total;

- NULL parent_id → null (and a real 0 stays 0);

- each other rejected column throws with the column-and-count message and no values. The cases are NULL, plus malformed values like true, "-1", "1.5", "", "1e3", an unsafe integer and an invalid Date;

- per-column counting;

- an empty read throws;

- nothing is logged.

- community-deposits-gateway.test.ts (19 tests):

- Gateway fixtures map to exactly the pg records, Unassigned rows included, and NULL parentId stays null;

- each rejected column refuses the read as an A8GatewayReadError;

- shadow reports a rejected Gateway row as degraded and publishes pg;

- the shadow compare is clean with the NULL-community row included;

- the gateway-mode log line reports the Unassigned and no-parent counts;

- the earlier coverage (order, ties, modes, redaction, freshness) is kept.

- community-deposits-gate-refresh.test.ts (14 tests), each through a whole runRefreshCycle, in both legacy and gateway modes:

- the Unassigned row is published;

- a NULL parent is published with no parentId;

- a rejected row fails the domain and publishes nothing;

- 250 rows go in exactly one insertCommunityDeposits call;

- a failed write fails the domain;

- an empty read fails the domain.

- communityDashboards.test.ts (+3):

- insertCommunityDeposits accepts a deposit without parentId, and the report returns it under Unassigned;

- a same-day patch clears an old parentId: 0 (this test fails without the patch change);

- a known parentId still stores and patches.

- community-view.test.tsx (+1): an Unassigned row renders with its total.

- Checks:

- sync: tsc --noEmit is clean, and vitest run --maxWorkers=2 passed 88 files and 1584 tests.

- chat: tsc --noEmit and tsc -p convex/tsconfig.json are clean. The touched and adjacent chat tests passed (7 files, 98 tests: community dashboards, community components, public API v2 admissions and domains). The full chat suite was not run.

- pnpm lint is clean. The 2 warnings are pre-existing, in chat/skill/forge-api/scripts/sindri.mjs.

- A live Gateway dry-run is still pending. It needs AI-Builder-Team/Surtr#2089 deployed and an A8 key, which is Keval's step X5.

## Stack

- Base: #1566 (the U02 read-gate kit, AERIE-2612). When it merges, this PR retargets to main and main is merged in. This branch is never rebased.

- AI-Builder-Team/Surtr#2089 (U06, SURTR-1540) supplies mart_education.aerie_admissions_community_deposit. It copies the SQL verbatim, NULL-community rows included, so no Surtr change is needed for the fix.

- Siblings: #1569 (U05) and the other A8 Aerie units each add one .env.example line after SCHOOL_SOURCE_READ, so expect a trivial merge conflict there.

## Not covered

- The live Gateway dry-run and the G3 shadow window. Both wait for U06 to be deployed, the key to be minted and PII sign-off (plan §9 D2).

- Old parentId: 0 documents on past snapshot dates. No consumer reads past snapshots or parentId, so they are left as they are.

- insertCommunityDeposits is a public Convex mutation (plan §8 item 10), like three marketing worker inserts. This is pre-existing and not widened here. It is tracked as AERIE-2663: move all four behind the bearer-gated /sync/analytics routes. Mercy withdrew the finding on that basis.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1573 — feat(a8): G4 gate EXPENSES_READ — expense transactions + vendor classifications via Surtr Gateway (AERIE-2617) @kevalshahtrilogy  approved

Linear: AERIE-2617 (A8 unit U11; related AERIE-445)

## Summary

A8 group G4: the analytics worker's three expense reads can now come from Surtr Gateway parity marts instead of Redshift, behind a gate that defaults to today's behaviour.

- Split, no behaviour change. queryExpenseTransactions, queryExpenseMetadata and queryVendorClassifications now issue exported SQL and map through exported mappers:

- SQL: EXPENSE_TRANSACTIONS_SQL (built by expenseTransactionsUpdatedAtSql), EXPENSE_METADATA_SCHOOLS_SQL, EXPENSE_METADATA_ACCOUNTS_SQL, VENDOR_CLASSIFICATIONS_SQL. The text is byte-identical.

- Mappers: mapExpenseTransactionRows, mapExpenseMetadataRows, mapVendorClassificationRows.

- The updated-time read still falls back to the transaction-date read only when the query itself fails.

- Gate EXPENSES_READ=legacy|shadow|gateway (default legacy), with per-source overrides EXPENSES_READ_TRANSACTIONS and EXPENSES_READ_VENDORS. The rules come from the U02 kit: for example, an unrecognized override never selects gateway.

- New queries/expenses-gateway.ts reads aerie-expense-transaction and aerie-expense-vendor-classification through the kit.

- The specs type each column as the Surtr DDL stores it.

- Rows go through the kit's type-parity layer, then the unchanged legacy mappers.

- Metadata has no mart. It follows the transactions mode.

- In gateway mode, it is queryExpenseMetadata's two DISTINCTs evaluated over the same transactions snapshot, then passed to the legacy mapper:

- NULLs are excluded;

- DISTINCT is exact-string;

- sorting follows Redshift's UTF-8 byte order, with NULL names last.

- Those three semantics were probed read-only against Redshift.

- One transactions read per cycle serves both the transactions and the metadata. If that read fails, the metadata fails too; it never falls back to pg.

- Shadow publishes the pg reads and logs one a8_shadow_check line per source:

- transactions, compared by sourceRowId;

- vendors, compared by vendorName as a multiset, because Convex upserts vendors by name;

- metadata, compared as sets.

It uses the kit's source_advanced skew rule. vendor_name, memo and line_amount values are on no allowlist, so they read [redacted] in every example.

- Gateway failures are value-free. They reach refresh.ts's per-domain catch as an A8GatewayReadError built from describeA8Error, so neither the FAILED log line nor the DomainResult can carry a row value.

- Freshness bounds replace the kit's 6h default (see the evidence below):

- 36h for transactions. mart-education-quickbooks-refresh is daily. Over the last 60 days its median gap was 24.0h and every gap but one was at most 34.8h, so 36h is the daily cadence plus 12h. Only the one missed-days gap (71.8h, 08-07 to 08-10) would have been refused.

- 8 days for vendors. quickbooks-expense-ai-generation is weekly (Sundays 08:00 UTC, 3-15 minutes), so 8 days is the cadence plus one day. It refuses the two windows where a failed Sunday run was recovered by hand 1-3 days late (10.3 days at worst); that surfaces the upstream outage instead of treating last week's classifications as current.

- Refusing is safe: each expense domain fails on its own, and nothing is deleted.

- Floors: 10,000 transactions and 500 vendors, about 12% and 14% of today's counts. There is no purge, so no purge guard. A short transactions read would still overwrite the one metadata row with shorter dropdown lists.

- refresh.ts changes stay inside interlude block 8: one import, four lines to resolve the gate once per cycle, and three changed calls.

- Dry-run: sync/src/scripts/dry-run-expenses-shadow.ts runs the worker's own shadow path, read-only.

## Business Value

- Moves the expense dashboard off the worker's direct Redshift reads. Its transactions, vendor classifications and filter metadata can now come through the governed Surtr Gateway. This is one of the A8 steps toward retiring the EC2 worker's warehouse access (AERIE-445).

- Low-risk cutover. Every expense write is an upsert, so nothing is purged. Each domain fails on its own. The metadata needs no new mart: one read serves both it and the transactions.

- Parity is provable before switching. The shadow compares every row and redacts personal free text. A test pins the SQL the Surtr marts copy.

- Zero risk until it is switched on. The default stays legacy, and rollback is one env var.

## Manual Effort Estimate

About 10 hours of focused work by hand, without AI. Keval, please confirm or adjust this number. It covers:

- reading the kit and the two Surtr mart contracts;

- the split and the gate module;

- the metadata derivation, including probing Redshift's DISTINCT and ORDER BY semantics;

- measuring both upstream cadences for the freshness bounds;

- about 45 tests with dual-transport fixtures;

- the read-only pg proof and the dry-run.

## Testing / evidence

- Split and gate parity on real rows (read-only). A throwaway script ran the pre-split expenses.ts (from d6f595810) and the post-split one against Redshift, reporting counts and booleans only. Everything matched on the first attempt.

| | rows | old = new | old = mapRows(sql) |

|---|---|---|---|

| transactions | 80,709 | yes, including watermarkMode/watermarks | yes |

| vendor classifications | 3,479 | yes | yes |

| metadata (68 schools, 2 accounts) | — | yes | — |

- Metadata derived from the 80,709 transaction rows equals the legacy metadata: school order is exact, account-name order is exact, and the account set is the same.

- The gate in all three modes was run over the same real rows. The "Gateway" was an in-process fake serving them in Data API shape plus lineage columns, so no Gateway traffic was generated.

- legacy, shadow and gateway each published exactly the old transactions, vendors and metadata.

- Shadow logged clean for all three sources: 80,709, 3,479 and 70 matched keys.

- Source shape (counts only): source_row_id is unique (80,709 of 80,709) and vendor_name is unique today (3,479 of 3,479). There are no trailing-blank or NULL-name edge cases in today's data.

- Redshift semantics probed read-only with literal-only queries, no table reads:

- VARCHAR DISTINCT keeps 'a' and 'a ' apart;

- ORDER BY is UTF-8 byte order ('' < '1' < 'B' < 'Z' < '_x' < 'a' < 'á' < '€' < '😀');

- NULLs sort last.

- Upstream cadence, from staging_other.pipeline_runs_prod over the last 60 days, read-only:

- mart-education-quickbooks-refresh: 68 successes, median gap 24.0h, gaps at most 34.8h except one of 71.8h.

- quickbooks-expense-ai-generation: weekly. Today's vendor snapshot is from 09-27 08:03.

- New unit tests:

- queries/expenses-gateway.test.ts (29 tests):

- fixtures in both transport shapes: NUMERIC as text, the ::text dates, '' versus NULL;

- fail-closed cases: NUMERIC delivered as a JSON number, an absent column, the legacy all-or-nothing parse, the freshness bounds at 35h/37h and 7.9/8.1 days, and the floors;

- metadata derivation: DISTINCT, NULL exclusion, byte order including a character where JS UTF-16 order differs, trailing blanks, and account ties;

- every mode and override: one transactions read shared by transactions and metadata; gateway failures that are value-free and never fall back; shadow degraded on Gateway failure;

- compare: field mismatches with vendor, memo and amount redacted, source_advanced, duplicate and one-sided keys, the vendor multiset, and the metadata sets.

- expenses-gate-refresh.test.ts (4 tests) runs a whole runRefreshCycle, using a stub orchestrator that runs the interlude:

- gateway mode publishes the Gateway snapshot and never reads the expense tables;

- a transactions-mart failure fails only transactions and metadata, with a value-free DomainResult error;

- shadow publishes exactly the pg reads and logs three clean lines;

- unset stays legacy.

- queries/expenses.test.ts (+4):

- pins the SQL text the Surtr marts copy;

- checks that each function equals mapRows applied to its exported SQL.

- Checks:

- sync: tsc --noEmit clean; vitest run --maxWorkers=2: 88 files, 1558 tests passed.

- pnpm lint (boundaries, convex paths, read bounds, test architecture, knowledge, biome) is clean. Its two warnings are pre-existing ones in chat/skill/forge-api/scripts/sindri.mjs.

- The chat suite wasn't run: this PR changes no chat code.

- Live Gateway dry-run: not run. The marts aren't deployed yet: AI-Builder-Team/Surtr#2088 is open and its DDL isn't applied. The aerie-a8 Gateway sources aren't registered or granted yet either (U04). Once they are, run cd sync && pnpm exec tsx src/scripts/dry-run-expenses-shadow.ts with the A8 key (Keval's step X5). With no key set, the script exits 1 and names the missing variable (checked).

## Stack

- Base: #1566 (the U02 read-gate kit, AERIE-2612). If #1566 merges first, this PR is rebased onto main and retargeted.

- AI-Builder-Team/Surtr#2088 (U07, SURTR-1539): the mart-aerie-expenses-refresh runner plus aerie_expense_transaction and aerie_expense_vendor_classification. The column names, types, lineage columns and mart_row_id here match those DDLs. The pinned SQL test here matches that PR's test_sql_contracts.py.

- No Convex change and no deploy-order constraint.

## Not covered

- The live Gateway dry-run and the G4 shadow window. Both wait for U07 to be deployed, U04 to be seeded and the key to be minted.

- A copy that lags its source reads as mismatch, not skipped. The kit's source_advanced rule covers a republish *between* the pg read and the Gateway read (the same rule as G1's directory). It doesn't cover the reverse: the upstream republishes and the A8 runner hasn't copied it yet. That window is normally a minute or two after each daily or weekly upstream run, so an hourly cycle rarely lands in it. If the runner is broken, the mismatch persists until the 36h / 8-day bound turns it into degraded.

- Transactions category freshness. It comes from the vendor classification joined at the mart's refresh, but the mart's lineage is the transactions run. If the runner fails on the weekly vendor trigger, categories lag until the next daily transactions republish re-joins them (at most about a day). The freshness bound can't see this; shadow shows it as category mismatches.

- No gateway-mode env precheck in the worker scheduler (the A5 pattern). A missing key fails each expense read with the client's value-free error.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1572 — feat(a8): G2 gate ADMISSIONS_COMMUNITY_FUNNEL_READ — community funnel via Surtr Gateway, PII-redacted shadow compare (AERIE-2616) @kevalshahtrilogy  approved

## Summary

A8 unit U09 (AERIE-2616): the read gate for the G2 community funnel. refreshCommunityFunnel reads the whole EduCRM community_conversion_dtl once per cycle (5,403 deals on 2026-09-29) and builds both the Community and the Established funnel from it. This PR lets that read run in shadow against Surtr's parity mart, and later cut over, without changing what legacy publishes.

- Split, no behaviour change. queryCommunityCommitmentEnrollments becomes COMMUNITY_COMMITMENT_ENROLLMENTS_SQL + mapCommunityCommitmentEnrollmentRows. The SQL text is unchanged. The null-deal_id drop and the committed-without-stage throw stay in the mapper, so they fire on both transports.

- Gateway reader (sync/src/analytics/queries/admissions-community-funnel-gateway.ts, on the U02 kit). It reads aerie-admissions-community-conversion (AI-Builder-Team/Surtr#2089):

- Every legacy output column is typed as the mart stores it: BIGINT ids, SUPER JSON text, and the nine ::text date casts as text.

- The kit reshapes the rows into pg values, then the unchanged mapper runs.

- The records are re-sorted to the legacy ORDER BY program_code, deal_id. The rollups take each program's first row, and Convex receives the enrollments in that order, so gateway mode publishes byte-for-byte what legacy does.

- The population floor is 1,000 rows. It is applied to the raw Gateway rows and again after the mapper's drops. Freshness uses the kit's 6h default: EduCRM republishes every 30 minutes, so 6h is 12 missed runs.

- Rows the legacy mapper drops (a null deal_id or a schema reject) are the same rows it drops from the pg read. Gateway mode reports them with a counts-only WARN before publishing.

- Gate: ADMISSIONS_COMMUNITY_FUNNEL_READ=legacy|shadow|gateway (default legacy). The override is …_COMMUNITY_CONVERSION. It is wired into refreshCommunityFunnel as an optional argument that defaults to resolving the env once per cycle.

- In gateway mode, a failure fails the community funnel domain the same way a legacy read failure does: the run is not published and the last published run stays.

- The shadow result logger defaults explicitly to the kit's logA8ShadowCheck, so the worker path records every outcome.

- The rollout order is documented in the module and in .env.example.

- Shadow compare, PII-safe. Mapped enrollments are compared by dealId. These rows carry children's and parents' names, emails, phone numbers, home addresses, dates of birth and free-text notes, so:

- the compare has no value allowlist;

- every value and key in an example reads [redacted];

- the a8_shadow_check line carries only field names, value types, counts and the snapshot's lineage.

- EduCRM skew rule. On a mismatch, pg is re-read once after the Gateway read. If the re-read matches, the outcome is source_advanced, which counts as clean. Shadow always publishes the first pg read.

- Dry-run: sync/src/scripts/dry-run-admissions-community-funnel-shadow.ts.

## Business Value

- The funnels move onto a lineage-stamped copy. The Community and Established funnels feed the admissions dashboards and the v2 community/established-funnel endpoints. After this change they can be served from a Surtr-owned mart that carries lineage, instead of the worker reading EduCRM directly.

- One env var per step. Shadow, then gateway, then legacy as the rollback. That is one more A8 group ready for the shadow window on the way to AERIE-445 (removing the analytics worker's direct Redshift reads).

- The PII compare logs no personal data. This is the heaviest PII source in A8, and its shadow compare is built so that no row value can reach a log line.

## Manual Effort Estimate

About 8 hours of focused time to build by hand without AI. That covers:

- reading the kit, the U05 pattern and the mart DDL;

- the split;

- the reader and gate;

- about 30 tests with transport-shaped fixtures;

- the dry-run;

- the read-only parity proof.

Keval: please confirm or adjust this number.

## Testing / evidence

- Unit tests: sync/src/analytics/queries/admissions-community-funnel-gateway.test.ts, 29 tests. They cover:

- Reader type parity: the same records in the same order as pg. BIGINT arrives as a Data API number, and one past 2^53 arrives as text. SUPER stays text and the ::text dates stay text. The spec's columns equal the SQL's select list.

- Reader fail-closed cases: an absent column, SUPER delivered as an object, a snapshot below the population floor (raw, and after the mapper's drops), a snapshot older than 6h, and a torn source_run_id.

- Legacy mapper behaviour on Gateway rows: the null-deal drop and the committed-without-stage throw.

- Order: the order comparator. A mutation check confirms the order test fails without the re-sort.

- Gate modes: unset (legacy exactly, with no Gateway call and no log line); shadow (clean, Gateway failure degraded, mapper rejection degraded, pg failure before any Gateway read); gateway (publishes Gateway records, reports mapper-dropped rows by count, failure rethrows with an operator-safe line); the worker path with no deps logs one value-free line per outcome; the override rollback and typo capping.

- Compare: a PII field mismatch reports field name and type only, with a PII-marker absence check on the full log line. The skew rule gives source_advanced when the re-read matches, mismatch when it still differs, and mismatch when the re-read fails. A one-sided key, or the same deal twice on the Gateway, is never clean.

- refreshCommunityFunnel wiring: Convex receives identical payloads, in identical order, in legacy, shadow and gateway modes. A Gateway failure sends nothing to Convex.

- educrm.test.ts has one new test: the query is the SQL constant mapped by the mapper.

- Typecheck: pnpm typecheck passes for every package. Lint: pnpm lint is clean; its only 2 warnings are pre-existing, in chat/skill/forge-api/scripts/sindri.mjs.

- Sync tests: vitest run --maxWorkers=2, 87 files, 1,551 tests pass. The chat suite was not run.

- Read-only pg proof on real rows (default_transaction_read_only; counts and booleans only). The pre-split queryCommunityCommitmentEnrollments (at kit head d6f5958) and the new one both returned 5,403 of 5,403 records (717 committed), identical in value and order. Also identical in order:

- mapCommunityCommitmentEnrollmentRows(query(SQL));

- the gate in legacy mode;

- the gate in shadow mode (its Gateway side degraded offline, with no key).

Further checks in the same run:

- The order comparator, applied to a shuffled copy, reproduces Redshift's own ORDER BY across all 79 program codes.

- Offline transport check: the real pg rows were reshaped the way the Data API delivers them (BIGINT as numbers, plus lineage, shuffled by row id) and fed through the real reader and mapper. It returned the same 5,403 records in the same order. No BIGINT exceeded 2^53.

- Source types checked. Read-only svv_columns confirms the EduCRM source column types match the mart DDL: 5 bigint, 2 integer, 13 boolean, 46 super, 1 date and 9 timestamp, the last two cast to text.

- No live Gateway run. mart_education.aerie_admissions_community_conversion is not deployed yet (it is absent in Redshift). The live dry-run waits for AI-Builder-Team/Surtr#2089's DDL apply, the aerie-a8 key, and PII sign-off.

## Stack note

- Stacked on #1566 (the U02 kit; base branch feat/a8-u02-read-gate-kit). If #1566 merges first, this will be rebased onto main and retargeted.

- The mart contract comes from AI-Builder-Team/Surtr#2089 (U06, aerie_admissions_community_conversion).

- Rebase hotspot: .env.example. U05 (#1569) and U10 add lines at the same spot.

## Not covered

- A live Gateway dry-run. This needs the mart deployed (AI-Builder-Team/Surtr#2089 DDL apply), the U04 sources and a key granted aerie-a8, and Keval's PII sign-off for exposing this mart (plan §9 D2).

- Flipping ADMISSIONS_COMMUNITY_FUNNEL_READ in any environment. This PR ships inert (legacy).

- The Convex-side write path: insertCommunityFunnelEnrollments and the rollups are unchanged.

- The ordering of non-ASCII program codes. The comparator is verified on today's 79 ASCII codes. A mis-order could only change which row's schoolStatus a rollup takes, and the Convex insert order.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1569 — feat(a8): G1 gate ADMISSIONS_REFERENCE_READ — programs + Program directory via Surtr Gateway (AERIE-2614) @kevalshahtrilogy  approved

Linear: AERIE-2614 (A8 unit U05; related AERIE-445)

## Summary

A8 group G1: the analytics worker's two admissions reference reads can now come from Surtr Gateway parity marts instead of Redshift, behind a gate that defaults to today's behaviour.

- Split, no behaviour change. queryPrograms / queryHubspotPrograms are now PROGRAMS_SQL / HUBSPOT_PROGRAMS_SQL (byte-identical text) plus mapProgramRows / mapHubspotProgramRows. Every identity, duplicate and alias throw stays in the mapper.

- Gate ADMISSIONS_REFERENCE_READ=legacy|shadow|gateway (default legacy), with per-source overrides ADMISSIONS_REFERENCE_READ_PROGRAMS and ADMISSIONS_REFERENCE_READ_PROGRAM_DIRECTORY. Rules come from the U02 kit (a8/read-mode.ts): an unrecognized override never selects gateway.

- New queries/admissions-reference-gateway.ts reads the marts through the kit:

- aerie-admissions-program and aerie-admissions-program-directory, typed as the Surtr DDL stores them;

- the kit's type-parity layer, then the unchanged legacy mappers.

- Shadow publishes the pg rows. It compares mapped ProgramRecords by sourceProgramId and directory records by programId, using the EduCRM source_advanced skew rule. It logs one a8_shadow_check line per source. The allowlist keeps staff names, contact points, addresses and free text redacted.

- Failure semantics are unchanged. In gateway mode:

- a programs failure still aborts the cycle;

- a directory failure stays isolated (programDirectoryError).

- Purge guard for purgeStalePrograms in gateway mode. It had no fraction guard, so a short but plausible Gateway read could delete real Programs and trigger an ontology rebuild.

- A new read-only Convex previewPurgeStalePrograms counts what the purge would delete. It uses the purge's own matching rule, now shared through staleProgramRows.

- The worker refuses more than ADMISSIONS_REFERENCE_READ_MAX_PURGE_PCT (default 5%) through the kit's assertA8PurgeBound. That happens before any write, so the cycle aborts cleanly.

- Freshness bound: 24h for both marts instead of the kit's 6h default.

- The directory's lineage is a HubSpot publication that mart-aerie-hubspot-refresh picks up 6-hourly. It was 6h41m old at 09:57 UTC today, so 6h would fail routinely.

- The programs publisher itself refuses observations older than 24h.

- Dry-run: sync/src/scripts/dry-run-admissions-reference-shadow.ts runs the worker's own shadow path, read-only.

- refresh.ts is edited only inside loadAnalyticsReferenceData (+import) and in a comment above the purge. That matches the plan's collision note.

## Business Value

- Unblocks the first A8 cutover. Programs is the hard prerequisite of every hourly admissions refresh. Moving it onto the governed Surtr Gateway is the first step in retiring the worker's direct Redshift reads (AERIE-445).

- Proves type parity first. all_program has 10 SUPER columns. G1's shadow window is the first live test of the kit's type-parity layer, so every later A8 group starts on a proven path.

- Closes a latent data-loss gap. Before this PR, a partial read could silently purge Programs and rebuild the ontology. Gateway mode now refuses that before writing anything.

- Zero risk until it is switched on. The default stays legacy, and rollback is one env var.

## Manual Effort Estimate

About 12 hours of focused work by hand, without AI. Keval, please confirm or adjust this number. It covers:

- reading the kit and the two Surtr mart contracts;

- the split, the gate module and the Convex preview;

- the purge-guard design;

- about 35 tests with dual-transport fixtures;

- the read-only pg proof and the dry-run.

## Testing / evidence

- Split parity on real rows (read-only). A throwaway script ran the pre-split functions (from d6f595810) and the post-split ones against Redshift, reporting counts only. Everything matched.

| | rows | old = new | old = mapRows(sql) | legacy gate = old | shadow gate = old |

|---|---|---|---|---|---|

| programs | 90 | yes | yes | yes | yes |

| program directory | 113 | yes | yes | yes | yes |

With no Gateway key set, shadow logged degraded for both sources and still published exactly the legacy rows. No Gateway traffic was generated.

- New unit tests:

- queries/admissions-reference-gateway.test.ts (29 tests):

- fixtures in both transport shapes: SUPER JSON text, BIGINT school_year as a Data API number and as text, NUMERIC text, '' versus null;

- DATE: the directory's date columns are ::varchar in the SQL, so they stay text; typing them date would hand the mapper a Date it rejects;

- fail-closed cases: absent column, parsed-JSON SUPER, two school years, stale snapshot, below floor;

- legacy identity and duplicate throws on Gateway rows;

- every mode and override, including unrecognized values;

- compare: field mismatch with redaction, source_advanced, one-sided and duplicate keys;

- the purge guard at 5% and 6%, custom and invalid pct, empty table, and the Convex ack validation;

- loadAnalyticsReferenceData failure isolation per source.

- admissions-reference-gate-refresh.test.ts (2 tests) runs a whole runRefreshCycle in gateway mode:

- over the bound, the only Convex call is the preview: nothing is upserted, purged or published;

- within the bound, it publishes and purges with the Gateway's codes and never touches Redshift.

- chat/convex/analyticsReference.test.ts (+1): the preview count equals what purgeStalePrograms then deletes, including quote-stripping.

- reference.test.ts / hubspot.test.ts (+1 each): the split, i.e. the function equals mapRows applied to the exported SQL.

- Checks:

- sync: tsc --noEmit clean; vitest run --maxWorkers=2 88 files, 1554 tests passed.

- chat: pnpm typecheck clean; only convex/analyticsReference.test.ts was run (19 passed), not the full chat suite.

- pnpm lint (boundaries, convex paths, read bounds, test architecture, knowledge, biome) clean.

- Live Gateway dry-run: not possible yet. The marts aren't deployed: AI-Builder-Team/Surtr#2084 is open and its DDL isn't applied. The aerie-a8 Gateway sources aren't registered or granted (U04). Once they are, run cd sync && pnpm exec tsx src/scripts/dry-run-admissions-reference-shadow.ts with the A8 key. This is Keval's step X5.

## Stack

- Base: #1566 (the U02 read-gate kit, AERIE-2612). If #1566 merges first, this PR is rebased onto main and retargeted.

- AI-Builder-Team/Surtr#2084 (aerie_admissions_program + runner, open) and AI-Builder-Team/Surtr#2085 (aerie_admissions_program_directory, merged into 2084's branch). Column names, types, lineage columns and mart_row_id match those DDLs.

- Deploy order: Convex previewPurgeStalePrograms ships in this PR. Setting programs to gateway before that Convex deploy fails closed: the cycle aborts and nothing is purged.

## Not covered

- The live Gateway dry-run and the G1 shadow window. Both wait for U03 to be deployed, U04 to be seeded and the key to be minted.

- A live proof of the purge preview against a deployed Convex. It is covered by the Convex unit test only.

- The shadow check adds one Gateway read per source per cycle. Each is one page at today's sizes, bounded by the client's 30s timeout.

- No gateway-mode env precheck in the worker scheduler (the A5 pattern). A missing key fails the read with the client's value-free error, which aborts the cycle for programs.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1566 — AERIE-2612: A8 read-gate kit (read mode, Gateway reader with type parity, keyed shadow compare, purge guard) @kevalshahtrilogy  approved

Linear: [AERIE-2612](https://linear.app/builder-team/issue/AERIE-2612/a8-u02-a8-read-gate-kit-read-mode-gateway-table-reader-with-type) (related: AERIE-445)

## Summary

A8 unit U02: the shared plumbing every later A8 Aerie unit uses to move one runRefreshCycle Redshift read onto a Surtr Gateway parity mart. New files only, under sync/src/analytics/a8/, plus a dry-run template. It doesn't touch A4 or A5 files or any refresh path, and nothing reads the kit until a group unit (U05, U09–U11, U14, U17, U20–U23, U25) wires it in.

Generalised from A4 (school-source-directory-shadow-compare.ts, school-source-directories-gateway.ts, the per-source override in PR 1531) and A5's purge bound (PR 1530).

### Kit API (what later units copy)

read-mode.ts: the gate

- resolveA8ReadModes({ gate, sources, dbtBackedSources? }, env?, warn?) → Record<source, "legacy" | "shadow" | "gateway">. Call once per cycle.

- <GATE> sets every source; <GATE>_<SOURCE> (a8SourceOverrideVar) overrides one.

- Unset/blank: legacy. pg is accepted as a spelling of legacy (A4/A5 muscle memory).

- Unrecognized global → legacy + WARN. Unrecognized override → ignored + WARN, and capped at shadow, so a typo never selects gateway. Case-sensitive, like A4.

- dbt-backed sources stay legacy (with a WARN) unless DBT_TARGET is exactly production.

- readsGateway(mode), publishesFromGateway(mode).

gateway-table-reader.ts: the reader and the type-parity layer

- readA8GatewayTable({ source, columns, minRows, maxSourceAgeMs?, rowIdColumn?, buildMarkerColumn?, pageSize?, maxPages? }, { client?, now? }) → { rows, lineage: { source, sourceRunId, sourcePublishedAt, buildMarker }, pages }.

- columns maps each legacy SQL output column to its Redshift type (varchar | super | smallint | integer | bigint | numeric | real | double | boolean | date | timestamp | timestamptz). Only these columns are returned, in pg shape, so the legacy query's own mapRows runs unchanged.

- Type parity: each Gateway (Data API) value is rebuilt as the text Redshift sends over the pg wire and run through the pg driver's own text parser (pg.types.getTypeParser). BIGINT → string, NUMERIC → string, DATE/TIMESTAMP → local-time Date, TIMESTAMPTZ → Date, SUPER → its JSON text as returned, REAL/DOUBLE → the 6/15-significant-digit number pg parses. Parity holds by construction, including pg's local-time DATE handling.

- Rules, each a thrown A8GatewayReadError: Zod on every row; an absent column fails; integers must be within their Redshift range; date/time text must name a real calendar day; population floor; one non-empty source_run_id and one source_published_at across all pages; max age (default 6h); unique non-empty mart_row_id (default; rowIdColumn: null opts out); optional single build marker for dbt copies.

- The default client is built per call, not at import, so a script that loads dotenv first sees the key.

- assertSharedLineage(lineages): one source_run_id across several reads (SIS rollups + members, Forecast V2 + operands).

- gatewayValueToPg(type, value): the parity function on its own.

keyed-shadow-compare.ts: the shadow check

- compareA8Keyed({ source, keyFields, keyOf?, duplicateKeys?, valueAllowlist? }, pgRecords, gatewayRecords): keyed multiset compare of mapped records, order-free and type-strict ("5" ≠ 5, Date ≠ ISO string, absent ≠ undefined).

- duplicateKeys: "never_clean" (default) or "multiset" for pre-dedupe row sets (Q5 without #n, D2 before last-row-wins, vendors).

- Output: counts, per-field mismatch counts, value *types*, and up to 10 examples. Values appear only for valueAllowlist fields; a key appears only when every key field is allowlisted, and always as those fields' values (a custom keyOf's output is never shown). Everything else is "[redacted]".

- runA8ShadowCheck({ spec, pgRecords, readGateway, skew?, log? }): never throws, never publishes, logs exactly one a8_shadow_check line (info when it counts as clean, WARN otherwise). Outcomes:

- clean;

- source_advanced: EDUCRM rule, skew: { rule: "source_advanced", rereadPg }. After a mismatch, pg is re-read once; if the re-read equals the Gateway, the result counts as clean;

- stale_copy: dbt rule, skew: { rule: "stale_copy", pgBuildMarker }. Different build markers mean the check is skipped, not a mismatch;

- mismatch;

- degraded: the Gateway read failed (swallowed).

- countsAsClean and compared fields drive the shadow-window count. A check whose log line couldn't be written is degraded.

purge-guard.ts

- resolveA8MaxPurgeFraction("<GATE>_MAX_PURGE_PCT") (default 5%, invalid → 5% + WARN), countA8WouldPurge(existingIds, incomingIds), checkA8PurgeBound(label, counts, fraction), and assertA8PurgeBound (throws A8PurgeGuardError). Counts only in messages. U05 uses it for purgeStalePrograms.

operator-safe-error.ts

- describeA8Error(error): the kit's and the Gateway client's own messages as they are (trusted by instanceof, never by error.name); Zod reduced to codes and paths; anything else to its name and code, each shown only if it looks like an identifier. A kit-local copy of PR 1531's describeErrorForOperator, which isn't on main yet.

sync/src/scripts/dry-run-a8-shadow.ts: the template

- Loads dotenv, then await import()s every app module.

- Ships with one live canary: A4's SIS organization directory read through the kit (varchar + TIMESTAMPTZ parity).

- Exit 0 only when every check counts as clean.

- Run: cd sync && pnpm exec tsx src/scripts/dry-run-a8-shadow.ts. No package.json script is added, to keep this PR new-files-only; each group's copy can add one.

### One deliberate reading of the recipe

The rule "only "" and null map to null" is applied to typed columns (numbers, dates, booleans), where a real value can never be empty. Text columns (varchar, super) keep "": pg returns "" for an empty string, and the shared mapper must see the same input on both transports. Mapping it to null would make the Gateway path diverge from legacy (e.g. a non-nullable z.string() in a legacy mapper). Legacy mappers that already turn "" into null keep doing so on both sides.

## Business Value

A8 is the largest slice of the Aerie EC2 → Surtr migration: about 20 Redshift reads in runRefreshCycle. This PR is the foundation for the 11 Aerie units that follow. Building it once:

- makes each group unit smaller and consistent, so Mercy reviews less and every cutover behaves the same way;

- puts the plan's biggest technical risk (transport type parity, §8 risk 1) in one tested place instead of 11 hand-rolled readers;

- builds in the safety rules (fail-closed reads, PII redaction, purge bound, never-throw shadow) once, so a later unit can't forget one.

Together these move the worker's direct Redshift credential and analytics-worker toward deletion.

## Manual Effort Estimate

Proposed: ~12 focused hours for Keval by hand, without AI. Keval, please confirm or adjust.

- Type-parity research (Data API field union, pg-types/postgres-date behaviour, Redshift float text): ~2.5h

- Reader and parity layer: ~2h

- Keyed multiset compare, redaction, classifiers, orchestrator: ~2.5h

- Read-mode gate and purge guard: ~1.5h

- Tests (158): ~3h

- Dry-run template and PR write-up: ~0.5h

## Testing / evidence

All commands ran in the worktree on this branch, rebased on origin/main 3d4fe1a97 (after Mercy round 3).

| Check | Result |

|---|---|

| cd sync && pnpm typecheck | pass |

| pnpm lint (repo root: boundaries, convex-paths, read-bounds, test-architecture, knowledge, biome) | exit 0. The 2 warnings are pre-existing, in unrelated chat/skill/forge-api/scripts/sindri.mjs |

| cd sync && pnpm test --maxWorkers=2 | 86 files, 1,521 tests passed |

| New kit tests alone, under TZ=Asia/Kolkata, TZ=UTC and TZ=America/Los_Angeles | 5 files, 158 tests, pass in all three |

| lefthook pre-commit (biome + typecheck-sync) | pass |

Type parity (the §8 risk 1 fixtures). One mart row is defined in both transport shapes:

- Data API: BIGINT and INTEGER as numbers; NUMERIC, DATE, TIMESTAMP, TIMESTAMPTZ and SUPER as strings; DOUBLE as the full double.

- node pg: BIGINT and NUMERIC as strings; DATE and TIMESTAMP as local Dates; TIMESTAMPTZ as an absolute Date; SUPER as JSON text.

The tests assert that:

- the reader turns the Gateway shape into exactly the pg shape;

- the pg fixture is what pg's own parsers return, so the fixture itself is pinned;

- a legacy-style mapper built from Aerie's superString / dateToDateString helpers gives identical records from either transport, and the shadow compare is clean;

- every type has accept and reject cases, and no reject reason echoes the value.

Other coverage:

- Reader: every rule above, with messages checked for the absence of a PII fixture value.

- Compare: redaction of values and keys, multiset mode, both classifiers and their edge cases, a throwing key function or log sink, and frozen pg records left unchanged.

- Gate: every global and override value, typos, and the DBT_TARGET precondition.

- Purge guard: bounds, and the "60 of 90 programs from one run" case.

Dry-run template (local only, no prod):

- With no env file: prints the missing vars, exit 1.

- With a scratch env pointing Redshift and the Gateway at 127.0.0.1:1: config loads before the app modules, the pg read fails as Error (ECONNREFUSED) (details withheld), exit 1.

Mercy round 1 (both findings fixed):

- The canonical encoding is now collision-free: every value at every depth is a type-tagged [tag, payload] array, object keys sit inside the payload, and the absent-field marker uses a tag no value produces. A test pins that marker-shaped data ({"$u":1}, {"$absent":1}, ["u"], nested too) never equals undefined, an absent field, a Date or a bigint.

- The dry-run's pg canary no longer has a LIMIT, so both sides see the same snapshot boundary. The template notes now tell each group copy to keep it that way.

Mercy round 2 (all three findings fixed):

- A key from a custom keyOf is never displayed. Examples show the group's keyFields values instead, and only when every key field is allowlisted, so shown data is always allowlisted data.

- describeA8Error trusts kit and Gateway errors by instanceof, not by the spoofable error.name. A name, error code or Zod path segment is shown only when it looks like an identifier. Spoofed-name regression tests are added.

- If the log sink throws, the check becomes degraded (never counted clean) and a best-effort line goes to the default sink.

Mercy round 3:

- Fixed: date, timestamp and timestamptz text must name a real calendar day (a UTC round-trip, so the TZ can't affect it). The pg parser would otherwise roll 2026-02-31 over into a different real date.

- Fixed (nit): SMALLINT/INTEGER/BIGINT values must be within Redshift's ranges, and BIGINT text must be canonical.

- Answered in-thread: a read capped by maxPages can't reach the reader. SurtrGatewayClient.listAll throws when has_more is still true after maxPages, and a new test pins that with the real client.

Throughput (synthetic): 100k rows × 40 columns parse in about 4.6s and compare in about 2s. The largest A8 source is expected to be a few hundred thousand rows, read hourly.

## Not covered

- No live Gateway run. No prod, no keys in this unit. The canary dry-run needs Keval's key and runs as X5. G1 (U05) remains the first live proof of SUPER/DATE/BIGINT parity against real marts (plan §8).

- The float rule is modelled, not yet observed live. REAL/DOUBLE are rounded to 6/15 significant digits, matching Redshift's text output at extra_float_digits=0. The first A8 mart with a float column should confirm it in its shadow run. If it's wrong, the per-field mismatch counts will name the column.

- No wiring, .env.example gate lines or package.json script. Those arrive with each group unit.

- describeA8Error duplicates PR 1531's helper. Fold the two together once PR 1531 merges.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2115 — docs(ai-spend): spec 13 - gpt-6.1-sol pricing insert and reprice, applied @kevalshahtrilogy  approved

## Summary

- gpt-6.1-sol appeared in openai-usage-pipeline's usage from 2026-09-29 with no row in core_finance.ai_spend_token_pricing, loading at $0 pricing-derived cost (billed cost preserved, $1,615 already billed by the time this was caught). Same shape as specs 9 and 12.

- Already applied to prod (2026-10-01): this PR documents the migration, per the same convention as specs 7-9 and 12 (pricing table is governed core_finance; prod INSERT is a manual, documented step).

- Pricing: $2.00 in / $0.10 cached / $10.00 out per 1M tokens — official OpenAI standard rate, verified 2026-10-01 against developers.openai.com/api/docs/models/gpt-6.1-sol. Note: cached input is 5% of the uncached rate for this model, not the usual 10% most other GPT-6-family models use.

- Reprice: one window (09-29 to 10-01, exclusive) via the pipeline's step function, reusing spec 9/12's preflight + single-writer-guard + sequential-polling script.

- Scope: features/surtr/ai-spend-pipeline/specs/ only — a new spec directory. No runner code, no infra, no CDK.

- Post-merge: no further action. The DDL and reprice are already applied and verified (see spec.md's Result section).

## Business Value

Restores calculated spend visibility for a newly-launched GPT-6.1 model so Klair's AI spend reporting doesn't silently show $0 for real usage ($1,615 billed on day one). Clears the daily PARTIAL status on openai-usage-pipeline caused by this model.

## Manual Effort Estimate

About 25 minutes: recognize the same failure shape as the already-solved gpt-6-astra/gpt-6-luna/gpt-6-sol cases, look up the model's official rate (catching the non-standard 5% cache discount), reuse the existing reprice script, apply and verify. Proposed by Claude, Keval to confirm or adjust.

## Test plan

- [x] Pricing row inserted, verified present (01-pricing-insert.sql's sanity check).

- [x] Reprice window SUCCEEDED (arn:aws:states:us-east-1:479395885256:execution:pipeline-openai-usage-pipeline-prod:sol61-reprice-1790857589).

- [x] 03-verification.sql: 0 zero-cost rows with real token volume. Calc vs billed by day recorded in spec.md.

- [ ] Confirm openai-usage-pipeline's next scheduled run no longer lists gpt-6.1-sol in unexpected_unpriced_models.

Linear: SURTR-1575

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1527 — chore(sync): retire gated-off F3–F7 financial-worker tasks (AERIE-2545) @kevalshahtrilogy  approved

Linear: AERIE-2545 (project: Data Layer and Surtr Integration; parent effort AERIE-445 / SURTR-735)

## Summary

This retires the worker side of five financial-worker tasks that have been permanently gated off, with no behaviour change:

| Task | Reads → writes (Convex) |

|---|---|

| F3 enrollment | core_education.fct_enrollment → enrollmentRecords |

| F4 enrollment-by-quarter | fct_enrollment → enrollmentQuarterly |

| F5 mfr-line-items | agg_mfr_line_items_summary → mfrLineItems |

| F6 mfr-vendor-line-items | agg_mfr_line_items_by_vendor → mfrVendorLineItems |

| F7 hc-teamroom-line-items | agg_hc_by_teamroom → hcTeamroomLineItems |

Changes:

- sync/src/financial-worker/index.ts: removed the 5 tasks, the enrollment boot backfills, the LEGACY_* contracts import, the flag-gated startScheduler block and the five _refresh*Fn config fields. Removed the "legacy tasks are disconnected" startup warn; nothing references that string. The F1/F2 XO contractor identity/package tasks are untouched.

- Deleted 8 modules and their colocated tests:

- sync/src/analytics/{enrollment,mfr-line-items,mfr-vendor-line-items,hc-teamroom-line-items}-refresh.ts

- sync/src/redshift/{enrollment,mfr-line-items,mfr-vendor-line-items,hc-teamroom-line-items}.ts

- git grep found no other importer of these modules or their exported symbols, including scripts, configs and docs.

- sync/tests/financial-worker/index.test.ts: removed the 4 legacy-flag / rollback-mode cases. Kept the 2 XO cases: XO tasks run when Redshift is configured, and stay gated when it isn't.

- Fixed two stale comments that pointed at the deleted redshift/enrollment.ts:

- sync/src/analytics/queries/reference.ts: dropped a note about the deleted queryCampusToProgramName.

- chat/components/dashboards/financials/school-year-period.ts: now points at enrollmentQuarterly.quarterKey in the Convex schema. This is a comment-only change.

Why this is safe (no behaviour change):

- Flag is off.

- The tasks only start when LEGACY_EDUCATION_WAREHOUSE_READS_ENABLED is true.

- legacyEducationWarehouseReadsEnabled() in packages/contracts/src/education-warehouse-cutover.ts returns true only for the literal string "true", so the default is false.

- SURTR-740 read the EC2 .env on 2026-08-11: false, with FINANCIAL_WAREHOUSE_READS_ENABLED unset.

- No ungated reader. Checked on current main:

- Every Convex *query* that reads these five tables calls requireLegacyEducationWarehouseReads() first: getEnrollmentSummary, getEnrollmentByQuarter, getVendorBreakdown, getTeamroomRoster, getIncomeStatement, getFilterOptions, getLastSyncedAt.

- The only ungated touches are the internal write mutations (replaceEnrollment*, upsertMfr*, upsertHcTeamroomLineItems, pruneHcTeamroomLineItemsForSyncRun). Only these deleted worker tasks called them.

- Scope is the worker only. The contracts flag, the Convex routes and tables, and the Surtr marts are unchanged.

## Business Value

- Removes ~4.1k lines of dead sync code and 102 tests that have not run in production since the flag was turned off.

- Shrinks the EC2 financial-worker to its two live XO contractor tasks. That puts it one step (F1/F2 moving off-worker) from full removal in the Aerie EC2 → Surtr teardown.

- Removes the one-env-var rollback path that could silently resume writing retiring core_education / mart data into Convex.

## Manual Effort Estimate

Proposed: ~3–4 focused hours, for Keval to confirm or adjust. That covers tracing the five tasks and their flag, grepping for other importers (including scripts and the Convex side), deleting and rewriting the worker and its test, fixing the stale comments, running typecheck, lint and tests for sync and chat, and writing this PR. The migration spec budgets ~0.5 day for this item.

## Testing / evidence

Run locally in a fresh worktree off origin/main (07c8b02ec). No .env was needed.

| Check | Result |

|---|---|

| pnpm --filter @bran/sync typecheck | pass (exit 0) |

| pnpm --filter @bran/sync lint (biome check .) | pass (178 files, no fixes) |

| root pnpm lint (boundaries, convex-paths, read-bounds, test-architecture, knowledge, biome) | pass (exit 0); 2 pre-existing warnings in unrelated chat/skill/forge-api/scripts/sindri.mjs |

| pnpm --filter @bran/sync test before | 81 files / 1363 tests passed |

| pnpm --filter @bran/sync test after | 73 files / 1257 tests passed |

| pnpm --filter @bran/chat typecheck (app + convex) | pass (exit 0) |

| chat vitest: school-year-period.node.test.ts, pl-breakdown-table.test.tsx, containerize.node.test.ts | 3 files / 75 tests passed |

| lefthook pre-commit (biome, typecheck-sync, typecheck-chat) | pass |

The test delta reconciles exactly:

- −8 files: 102 tests (5 + 10 + 15 + 14 + 22 + 11 + 13 + 12).

- −4 flag/rollback cases in financial-worker/index.test.ts (6 → 2).

- Totals: 1363 − 102 − 4 = 1257; 81 − 8 = 73.

Grep: the spec's literal check git grep -E "mfr-line-items-refresh|hc-teamroom-line-items-refresh|enrollment-refresh" still returns 5 lines, and none refer to the deleted modules:

- sync/src/analytics/sis-enrollment-refresh* and its importer sync/src/analytics/refresh.ts. This is the live SIS enrollment refresh in the analytics worker, and it is kept.

- Two Convex comments in chat/convex/admissions/analytics/{enrollment,sisEnrollment}.ts naming an old 03-pipeline-enrollment-refresh pipeline, which is unrelated.

A tighter grep returns nothing in sync/ outside the analytics worker's own separate flag read (analytics/refresh.ts):

git grep -nE "(^|[^a-z-])(enrollment|mfr-line-items|mfr-vendor-line-items|hc-teamroom-line-items)-refresh|redshift/(enrollment|mfr-line-items|mfr-vendor-line-items|hc-teamroom-line-items)\.(js|ts)|_refresh(Enrollment|EnrollmentByQuarter|MfrLineItems|MfrVendorLineItems|HcTeamroomLineItems)Fn" -- sync

## Not covered

- Convex/UI deletion waits on the Education P&L decision (retire the page, or re-point F5–F7 readers at the existing aerie-mfr-* Gateway sources). Not touched here:

- the /sync/analytics/{mfr-line-items,mfr-vendor-line-items,hc-teamroom-line-items} routes;

- dashboards/educationPL.ts and lib/upsertDiff.ts;

- the mfrLineItems, mfrVendorLineItems, hcTeamroomLineItems, enrollmentRecords and enrollmentQuarterly tables, which need a purge migration first;

- the Education P&L UI;

- the platform-error-coverage-inventory.ts rows.

- /sync/analytics/enrollment is shared with the live analytics worker (sync/src/analytics/queries/enrollment.ts). The later Convex PR must remove only the replaceEnrollmentRecords / replaceEnrollmentQuarterly operations, not the route.

- The contracts flag itself stays: the analytics worker and dozens of Convex readers still use it.

- Surtr marts stay (core-education-budget-vintage-refresh consumes the MFR summary).

- F1/F2 XO contractor tasks are unchanged. Draft #1499 touches the XO modules and sync/package.json but not the files changed here.

- The flag reading is ~6 weeks old. Keval should re-confirm LEGACY_EDUCATION_WAREHOUSE_READS_ENABLED on the EC2 .env (read-only) before this ships in a release.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1614 — chore(praxis): follow master during hosted testing (AI-874) @caina-barbosa  approvedmercy-allow-critical

## Summary

This PR is a temporary feedback-loop step in [AI-867 — Praxis Hosted v0](https://linear.app/builder-team/issue/AI-867/vision-agent-hosted-v0-run-author-defined-pr-verification-on-demand).

It makes Aerie follow the private Praxis repository's master branch while the hosted pilot is being repaired and exercised. This step is tracked by [AI-874 — Prove the complete hosted flow and cleanup](https://linear.app/builder-team/issue/AI-874/vision-agent-hosted-7-prove-the-complete-hosted-flow-and-cleanup).

Production effect: controlled rollout or migration.

---

## Why

Immutable action pins are the stable release policy, but each Praxis repair currently requires a second Aerie pin-update PR before the live feedback loop can continue. Following master temporarily removes that delay while the pilot is still under controlled human review; Praxis must be pinned again before this becomes a stable release path. Praxis master currently has no branch protection or ruleset, so this temporary choice relies on private-repository access control, green Praxis CI, and operator discipline rather than an independently enforced approval gate.

---

## Business Value

- Lets Praxis fixes on master reach the hosted pilot without a separate Aerie pin PR each time.

- Shortens the diagnose, repair, and live-verification loop during controlled testing.

- Keeps the existing trigger authorization and limited workflow permissions in place.

---

## How does it work

1. .github/workflows/praxis.yml resolves AI-Builder-Team/Praxis/.github/actions/run@master when an eligible Praxis workflow runs.

2. The action remains in the private AI-Builder-Team/Praxis repository with access granted only to Aerie; Praxis master is not currently branch-protected.

3. Aerie's existing preflight still requires an exact Praxis mention, authorized actor, open pull request, and live head before AWS credentials are used.

4. The workflow keeps its limited contents: read, pull-requests: read, and id-token: write permissions; this PR adds no permission or credential.

5. A human must merge this Aerie change, but later Praxis master updates would take effect without another Aerie review.

6. Restoring an immutable Praxis SHA or adding an independently enforced Praxis approval gate remains required before stable release.

---

## Scope

### Included in this phase

- Temporarily replace the immutable Praxis action SHA with the private repository's master branch.

- Exact final diff paths:

.github/workflows/praxis.yml

### Deliberately excluded for later phases

- Praxis runtime or validator changes — this PR changes only Aerie's action reference.

- Other Aerie workflows, application code, schema, API, agent, migration, deployment, and AWS resources — untouched.

- Permanent mutable action policy — Praxis must return to an immutable reviewed SHA before stable release.

- Merge or live Praxis execution — both remain separately controlled.

---

## Test plan

### Automated validation

- Praxis preflight tests — 2/2 passed (node --experimental-strip-types --test .github/scripts/praxis-preflight.test.ts)

- Praxis branch resolution — master resolved to 604cc6438be951c862b6f611ae0c95a191dd2902 (git ls-remote https://github.com/AI-Builder-Team/Praxis.git refs/heads/master)

- Praxis branch-policy audit — no branch protection, repository ruleset, or effective branch rule is configured for master (gh api repos/AI-Builder-Team/Praxis/branches/master/protection, gh api repos/AI-Builder-Team/Praxis/rulesets, and gh api repos/AI-Builder-Team/Praxis/rules/branches/master)

- git diff --check — passed

- exact-head diff scope — one changed line in .github/workflows/praxis.yml

- Full Aerie application tests and typecheck were not run locally because this change only replaces the action reference; hosted CI is the final gate.

### Time for Implementation

An engineer working without AI assistance would take about half a day to confirm the trust boundary, update the workflow reference, validate the preflight, and prepare this controlled testing PR.

#1499 — fix(analytics): repoint XO contractor sync at the real table + add Gateway parity (F1/F2) @kevalshahtrilogy  approved

> [!IMPORTANT]

> Merge/deploy companion: [AI-Builder-Team/Surtr#2091](https://github.com/AI-Builder-Team/Surtr/pull/2091) (SURTR-1542). It applies the identical package latest-week tie-break (ORDER BY week_start DESC, id DESC) to Surtr's mart_education.sp_refresh_aerie_xo_contractor_package. The two must merge and deploy together; otherwise this repo's query and Surtr's mart pick differently on any future tie, and the shadow compare diverges. Deploy order for this PR: Convex first or together with the worker (see *Partial publishes* below).

> [!WARNING]

> Source freshness: the upstream ledger is about 3 weeks behind. As of 2026-09-29, SELECT MAX(week_start), COUNT(*) FROM staging_finance_xo.raw_contractor_invoices returns 2026-09-07 / 152,718 rows, so the latest invoiced week is 22 days old. Ingest is alive: on 2026-09-25 the max was 2026-08-31 with about 151k rows. Whichever read mode ships, contractor package rates and trailing-52-week totals will be only as current as this table. That also means neither side of the shadow compare can be fresher than 2026-09-07. It's worth confirming with the xo-contractor-invoices-refresh owner whether a ~3-week lag is normal XO invoicing lag or a stalled window.

## The live bug (fixed first, independent of the migration below)

sync/src/redshift/xo-contractor-identity.ts and xo-contractor-package.ts queried core_finance.xo_contractor_invoices_raw, a table that no longer exists in the warehouse. Surtr's seed-gateway-aerie.ts (dated 2026-08-31) flagged that Aerie's contractor sync should be repointed.

These two tasks are not behind LEGACY_EDUCATION_WAREHOUSE_READS_ENABLED (financial-worker's own comment: *"XO contractor tasks remain active"*). They run in production, and since the underlying table moved, every run has failed its query and published nothing. The failure wasn't silent: each tick returned status: "degraded" with error: query: …. The data just never refreshed.

The real, current table is staging_finance_xo.raw_contractor_invoices, and it has an exact column match for these queries. Both queries now point at it.

Verified against real production Redshift when this was first opened: the fixed queries return 3,059 identity and 3,026 package contractors. Surtr's mart_education.aerie_xo_contractor_identity / aerie_xo_contractor_package return 3,058 / 3,026. Surtr's procedures are a verbatim lift of these same queries, so near-equal counts are expected. The counts confirm the repoint reads the right ledger. They are not an independent cross-check.

pg output order changes in this PR. The identity query's alias LISTAGG now has a full sort key (see below), so the aliases array order differs from what the old query would have produced. No consumer has seen the old order recently: the pg path has published nothing since the table moved. The Convex alias index also re-sorts longest-first on read.

## The migration (F1/F2, matching A4's pattern)

Surtr exposes both marts as Gateway sources, aerie-xo-contractor-identity and aerie-xo-contractor-package, already granted to the live Aerie key.

- Gateway readers. redshift/xo-contractor-identity-gateway.ts / xo-contractor-package-gateway.ts are field-complete and Zod-validated. The identity reader drops rows whose aliases all trim away, exactly as the pg reader does.

- Payload shape. Every published row, from pg or the Gateway, is mapped field by field to a new shared wire contract, @bran/contracts/xo-contractor-sync. Package optionals are omitted when null, never sent as null. This matters because Gateway rows carry sourceRunId, and Convex's upsertXoContractor* validators reject undeclared fields. The first version of this PR would have failed every gateway-mode batch. The contract is locked from both sides:

- a chat test reads the registered mutations' own validators via exportArgs()

- sync tests check every record the worker sends

- a convex-test case proves a record carrying sourceRunId is rejected

- Lineage. sourceRunId is used only to validate and log, never published. gateway mode refuses a read whose rows span two mart publishes (mixed sourceRunId); that tick degrades and the next one retries. Otherwise it logs the one source run it published.

- Deterministic package pick (Keval's decision, 2026-09-29). When a contractor has several regular-payment rows in their latest week, the pick is now ROW_NUMBER() OVER (PARTITION BY contractor_id ORDER BY week_start DESC, id DESC): the highest invoice row id wins. week_start DESC alone tied, and Redshift breaks ROW_NUMBER ties nondeterministically, so the published rate could change run to run.

- A read-only check on 2026-09-29 found id unique and non-null: 152,718 rows, 152,718 non-null ids, 152,718 distinct ids.

- A before/after comparison changed no output: 3,046 rows under both orders, 0 rows in only one of them, and 0 contractors with a tied latest week.

- The Surtr side is the companion PR above.

- Partial publishes are reported, not hidden. Publication is not atomic: each 100-row batch is its own Convex mutation, and each mutation isolates every record's write. Both refreshes now publish through analytics/xo-contractor-publish.ts.

- records is the count Convex confirmed it wrote, never the attempted count.

- Any shortfall is an error naming it, for example PARTIAL publish: 200 of 3046 rows confirmed before this batch failed or Convex rejected 3 of 3046 rows. The tick degrades, and the next tick re-sends every row (upserts are idempotent per contractorId).

- To make per-record rejections visible, upsertXoContractorIdentity now returns { errors, total }, as upsertXoContractorPackage already did (XoContractorUpsertResult). This is a small Convex change that ships with this PR.

- Fail-closed: a batch counts as written only if its response acknowledges exactly that batch. A missing, malformed or wrong-total response counts as unacknowledged and is reported. Deploy Convex before, or together with, the worker; otherwise identity ticks report "unacknowledged" (degraded) until Convex catches up, although the rows are still sent.

- External error text (Gateway, Redshift, Convex responses) is flattened to one line and capped at 300 characters before it reaches a log line or tick summary (boundedErrMsg).

- Population floor. A read that "succeeds" with too few rows is refused before anything is published, in every mode (analytics/xo-contractor-population.ts). This covers the direct Redshift path, not just the Gateway: an empty read used to publish nothing and still report success with records: 0. The floor is 500 rows, the same as the Gateway readers' own guard, which stays in place; production has about 3,050 rows per entity. A read below it is an error, nothing is sent, and the tick degrades.

- Strict, all-or-nothing parsing, on both sides.

- z.coerce.number() turned null and "" into 0 and true into 1, and z.coerce.string() turned null into "null". An incomplete row could become a plausible contractor id, dollar amount or date.

- All XO readers (pg and Gateway, identity and package) now use the same parsers from redshift/xo-contractor-fields.ts:

- contractorIdNumber: a positive integer

- requiredFiniteNumber / nullableFiniteNumber: a finite number, or a string that is wholly a plain decimal (so 0x10 and 1e3 are rejected)

- isoDateText: the whole value must be an ISO date (optionally with a time part) naming a real calendar day, normalised to YYYY-MM-DD (so 2026-02-30 and 2026-01-01garbage are rejected)

- Optional team_name / company / currency use one nullableText parser, so "" is treated as absent on both sides; pg used to publish "" where the Gateway omitted it.

- source_run_id must be one printable, whitespace-free token (lineageRunId), because it goes into worker log lines.

- Monetary columns keep their sign. The ledger and the mart DDL put no sign constraint on them, and trailing_52w_paid_usd sums every invoice row.

- The pg readers now use parseRowsStrict instead of safeParseRows. One invalid row fails the whole read (the tick degrades and Convex keeps the last good snapshot) rather than silently dropping the row and publishing a partial snapshot as a success. The Gateway readers already worked this way.

- Shadow compare. analytics/xo-contractor-shadow-compare.ts compares field by field, keyed by contractorId. The pg side has no sourceRunId (it recomputes live on every call), so lineage is checked on the Gateway side only, as in A5.

- A failed or unclean shadow check is visible. In shadow mode, if the Gateway read or compare fails, or the check runs but comes back not clean, the rows are still published from Redshift. The refresh returns shadowIssue, which is PII-free:

- check failed: …, or

- not clean: N mismatched, N pg-only, N Gateway-only row(s), <lineage>

- incomplete: trailing52wPaidUsd not compared (pg day X vs mart publish day Y) when the run was clean but skipped the date-anchored field

The financial worker reports that tick as degraded, so a run without complete, clean parity evidence can't pass for a clean check. The compare result also counts duplicate contractor ids on each side. All shadow-side work sits inside one try, so nothing on the shadow side can block publishing from pg.

- XO_CONTRACTOR_READ typos are visible. An unrecognised value (e.g. gatewayy) still falls back to pg, so data keeps publishing, but the tick is degraded with a config: note.

- Deterministic aliases. The pg identity query now uses Surtr's mart ordering exactly: ORDER BY LENGTH(a.alias) DESC, a.alias ASC. LENGTH DESC alone left same-length aliases tied, and Redshift breaks LISTAGG ties nondeterministically. Aliases are compared order-sensitively, which is only safe because both sides now share that key.

- snapshotDate is not field-compared. It's a run-date stamp: pg gets the worker's CURRENT_DATE on every tick. The mart gets its own publish date, and Surtr's replay-idempotency guard deliberately leaves an already-published source run untouched, so that date can trail the worker's by a day or more. Comparing it would flag every row on any tick that lands on a different UTC day than the mart's last publish, which says nothing about data agreement. Surtr's own replay diff excludes snapshot_date for the same reason. Both dates are still reported as pgSnapshotDate / gatewaySnapshotDate.

- trailing52wPaidUsd is compared only when the days match. It sums a window anchored on that same CURRENT_DATE. When the two days differ it goes into skippedFields instead of producing a false mismatch.

- PII safety. Every compared field is redacted in a mismatch record. Canonical names, aliases and weekly pay are real compensation data tied to real people, so a mismatch shows which field and which id, and never a name or dollar amount in a CloudWatch log line. Dates are not redacted.

- Env gate. XO_CONTRACTOR_READ=pg|shadow|gateway (xo-contractor-read-mode.ts, documented in .env.example) defaults to pg. Both refresh functions share it, since they are the same migration object and always move together.

- Dry-run scripts.

- dry-run-xo-contractor-pg-fix proves the table fix against real Redshift. It exits 1 below the 500-row floor, so an empty read can't pass as confirmation.

- dry-run-xo-contractor-shadow runs a full shadow compare with PII-redacted output. It also prints both snapshot dates, any skipped fields and duplicate counts, and exits 0 only when the run is clean AND complete.

- Both scripts print only bounded error text. It uses await import() rather than a static import for credential-dependent modules, per the ESM/dotenv-ordering fix Mercy flagged on the A5 PR.

The existing regression-guard test that asserted the *old* table name has been fixed. It now asserts the new table and guards against regressing back to the dead one.

## Verified locally

- Real production Redshift: the table fix (counts above) and today's freshness query (warning at the top). Both were read-only.

- Live Gateway shadow dry-run, 2026-09-29. This ran dry-run-xo-contractor-shadow against real Redshift and the real Surtr Gateway. It was read-only: nothing was sent to Convex, and only counts and field names were printed.

- Identity: 3,078 / 3,078 matched, clean.

- Package: 3,046 / 3,046 matched, clean.

- Later runs used the strict, all-or-nothing, whole-value parsers, and no row was rejected on either side. They landed on a day when the pg date and the mart publish date were aligned (2026-09-29), so trailing52wPaidUsd was compared as well. The latest run reported "Clean and complete".

- Verbatim output is in the PR comments.

- Tests:

- sync: 90 files, 1,585 tests, all green

- @bran/contracts: 93 files, 1,259 tests, including the new contract test

- chat, touched files only: financialContractorPackage.test.ts (9 tests, including the validator-contract and upsert-result tests) and analyticsSyncSecurity.test.ts (18 tests)

- the full chat suite was not run

- Typecheck: clean across the workspace.

- Lint: clean. The only output is 2 warnings in chat/skill/forge-api/scripts/sindri.mjs, which this PR doesn't touch.

## Known limitations (not fixed here)

- Gateway mode still needs Redshift credentials. In financial-worker, both XO tasks are still gated on isRedshiftConfigured(), so gateway mode can't run without Redshift env vars. That's fine while pg stays the rollback; it needs revisiting before Redshift creds are torn down.

Linear: [AERIE-2489](https://linear.app/builder-team/issue/AERIE-2489/f1f2-fix-broken-xo-contractor-sync-gateway-parity-identity-package)

## Business Value

This fixes a production task that has failed on every run for weeks. Contractor identity and compensation data feeding Aerie's finance dashboards stopped refreshing when the source table moved; the worker reported degraded ticks, but nothing repointed it. That makes this a real correctness fix, not just migration progress.

It also lands the F1/F2 Gateway migration object, following the A4 pattern. Gateway mode can now actually publish, and the shadow compare is deterministic, so a clean shadow run means real parity rather than noise. Once shadow is confirmed clean, this sync can cut over to Surtr's Gateway the way School Directories already has, moving one more EC2 worker task closer to full teardown.

## Manual Effort Estimate

About 2 days by hand. Proposed; Keval to confirm or adjust.

- Diagnosing the break: trace which table Aerie queries, confirm it no longer exists, find the real table and its pipeline, and check column and shape compatibility.

- Fixing it: repoint both queries and validate against real Redshift.

- Migration half: Gateway readers, the PII-aware shadow compare, the env gate, two dry-run scripts and about 40 tests.

- Fix round:

- the shared Convex payload contract with validator-derived tests

- lineage validation

- alias tie-break alignment against the Surtr procedure

- date-aware compare semantics

- the population floor, strict numeric parsing and shadow-failure observability

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#155 — Release: Shipyard 0.6.10 @ashwanth1109  no labels

## Summary

- Prepare Shipyard 0.6.10 with the approved public release notes.

## Business Value

- Gives users read-only Companion lookups across Shipyard data, more dependable replay evaluations, clearer Pomodoro transitions, and resilient Pi image attachments.

## Implementation Effort

- Metadata-only change: one version bump and one public release-notes file.

## Test Plan

- [x] pnpm test:release

- [x] git diff --check

#1529 — feat(sync): A7 MATTERPORT_DISCOVERY_MODE gate + shadow compare (AERIE-2244) @kevalshahtrilogy  approved

Linear: AERIE-2244 (A7: validate, gate off, retire matterport-discovery). AERIE-2247 is marked as its duplicate.

## Summary

This PR adds MATTERPORT_DISCOVERY_MODE=write|shadow|off to the analytics-worker matterport-discovery scheduler. It is object A7 of the Aerie EC2 → Surtr migration, and the spec recommends retiring this task rather than porting it. Nothing changes until the env var is set.

- write (default, unchanged; also what unset or empty means): posts setMatterportModelId to Rhodes, as today.

- shadow: builds the mapping the task converges on, using the same rules as refreshMatterportFromIsp: lowest model id per Wrike folder, first slug per folder.

- It never calls setMatterportModelId. The shadow code only holds Pick<RhodesClient, "listSites">.

- It reads Gateway source aerie-site-matterport (core_education.xref_site_matterport) via SurtrGatewayClient.listAll.

- It diffs the two sides after URL→id normalisation. There are no lineage checks because the xref has no lineage columns.

- It logs one matterport_discovery_shadow_check line: INFO when clean, WARN when not.

- A Gateway, listSites or ISP-scan failure makes the tick degraded. The run never throws and never writes.

- The line also counts what write mode *would* send: fills, URL→bare-id rewrites, and overrides of a different model.

- It still depends on Rhodes. Shadow calls Rhodes listSites (a read) on every tick, so it needs the same RHODES_CONVEX_SITE_URL + RHODES_API_KEY as write mode.

- Without them the tick skips on the existing env gate.

- A listSites failure makes the tick degraded.

- If legacy Rhodes is gone, shadow can't run. Only off removes the dependency.

- off: shouldRun returns not-ok with the reason MATTERPORT_DISCOVERY_MODE resolves to off, logged once.

- Parsing is fail-safe (changed in the 2026-09-29 fix round):

- The value is trimmed and case-insensitive, so OFF, Shadow and WRITE all work.

- Unset or empty means write, so default behaviour is unchanged.

- Any other non-empty value resolves to off, never write, so a mistyped kill switch stops writing. An error naming the value is logged once per distinct value, not on every tick.

- Before this fix, an unknown value, including an uppercase OFF, fell back to write.

- What counts as clean: the xref deliberately quarantines two ambiguity classes, where the task picks a winner instead:

- class 1: a folder with more than one model;

- class 2: a folder shared by more than one site.

Task-only rows explained by one of those classes are reported but don't break "clean". Differences, xref-only rows, unexplained task-only rows and duplicate Gateway slugs do.

- Dry run: pnpm run dry-run-matterport-discovery-shadow is read-only. dotenv loads before the dynamic await import(). Exit 0 means clean.

- .env.example: new MATTERPORT_DISCOVERY_MODE entry.

- Edits to shared files are minimal: analytics-worker/index.ts +23 lines, .env.example +1.

Legacy code is untouched. Deleting matterport-from-isp-sync.ts, the scheduler and RhodesClient.setMatterportModelId happens after prod validation. isp/fetcher.ts, isp-adapter.ts and listSites stay while A1 shares them.

## Business Value

- Safe off-switch. A task that writes into the archived Rhodes deployment can be switched off safely. Aerie's own Convex has no receiver for these writes, and sites.matterportModelId is now hand-maintained.

- Removes a risk. Write mode would currently rewrite 11 hand-entered URL values to bare ids and overwrite 1 hand-entered model with a different one. If the legacy dual-write back into Aerie is still alive, that is a live whole-record overwrite risk. off removes it, and shadow proves Surtr's xref reproduces everything the task derives before we delete it.

- Evidence behind the switch. The flip comes with evidence (a daily shadow line, the parity SQL below) rather than a guess. That is one of the steps toward emptying the EC2 analytics-worker, which is the migration's definition of done.

## Manual Effort Estimate

Proposed: ~10 focused hours, plus ~1h for the fail-safe fix round. Keval, please confirm or adjust.

- ~2h: reading the legacy task, the ISP fetcher and the Surtr xref procedure to pin down the exact selection rules.

- ~3h: gate + shadow module + URL normalisation.

- ~3h: tests, including the legacy-replay equivalence test.

- ~1h: dry-run script + local verification harness.

- ~1h: parity SQL.

## Testing / evidence

- sync typecheck: clean.

- biome check over sync/: clean.

- Root lint:boundaries, lint:convex-paths, lint:read-bounds, lint:test-architecture, lint:knowledge: all pass.

- sync tests (--maxWorkers=2, after the fix round, rebased on current main): 82 files / 1,420 tests pass. New tests:

- Mode parsing: unset/empty/blank → write; trimmed and case-insensitive (OFF, SHADOW, Off\t, WRITE); unrecognised values (0, false, disabled, of, gateway, pg, true) → off + error; the error is logged once per distinct value across repeated calls; env read.

- URL normalisation: show/?m=, extra query params (&mls=1, &play=1), /models/, /space/, missing scheme, case kept, non-Matterport hosts and id-less URLs left as-is.

- Plan equivalence: replays refreshMatterportFromIsp's actual writes against a recording stub and asserts the plan matches the final values and the written set.

- Diff counts: matched, model/folder differences, quarantined vs unexplained task-only, xref-only, duplicate Gateway slug.

- Shadow never writes: setMatterportModelId is never called in any shadow test.

- Gateway errors: HTTP 403 through the real SurtrGatewayClient, a missing key (no request made), an empty read, a missing column. Each is degraded with exactly one check line.

- Worker level:

- off, OFF and Off run neither path and log one skip.

- An unrecognised value runs neither path and logs exactly one error across several ticks.

- shadow runs only the shadow path.

- Unset, empty, write and WRITE keep the write path with no error.

- Local dry run (read-only):

- Live setup: the ISP DynamoDB Scan + S3 GETs ran against the real Klair-ISP-Jobs data.

- Stub for listSites: a 127.0.0.1 stub serving staging_education_rhodes.raw_sites, Aerie's mirror, not live Rhodes.

- Stub for the Gateway: the same local stub serving core_education.xref_site_matterport in the real Gateway page contract.

- Result, CLEAN:

- Totals: 63 ISP models, 171 sites, 19 task-mapped vs 18 xref.

- Comparison: 18 matched, 0 differences, 0 xref-only.

- Task-only: 1 row, quarantined as class 1 (156-william-st-new-york-ny).

- Write-mode projection: it would send 12 writes, 11 URL rewrites and 1 override.

- Without a Gateway key, the same run reports the Gateway leg as failed and exits 1, as intended.

- Parity SQL (read-only, BEGIN READ ONLY + rollback; the local role can read all three schemas). It rebuilds the task mapping from staging_education_isp.raw_jobs + staging_education_rhodes.raw_sites (lowest id per folder, first slug wins) and diffs it against the xref:

- Inputs: 163 jobs (all completed), 27 models with a Wrike folder, 19 task-mapped sites, 18 xref rows.

- 18 matched, 0 differ, 0 xref-only. 1 task-only row, class-1 conflict count = 1 (that folder has 2 models), 0 class-2. So xref ⊆ task mapping holds exactly.

- Against Aerie's copy (raw_sites.matterport_model_id, 96 non-blank, 61 URL-form), write mode would leave 7 rows as they are, fill 0, rewrite 11 URL-form values and override 1 with a different model (8000-sw-56th-st-miami-fl).

- The full SQL is on AERIE-2244.

## Not covered

- Live Gateway leg: not run, because there is no SURTR_GATEWAY_API_KEY locally. The Gateway grant for aerie-site-matterport on the Aerie key is also untested.

- Live Rhodes listSites: not called; Aerie's Redshift mirror stood in for it.

- Prod behaviour: not verified. Still needed: the prod tick history (updated/unchanged/threw/skipped), whether RHODES_* is set on EC2, and whether legacy Rhodes and dual-write are alive. Keval, per the A7 tab.

- Retire vs port: still undecided (Benji, D3). Whether to validate with a runtime shadow window or a straight off flip is also open.

- Legacy deletion: comes in a follow-up PR on the same ticket.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1563 — feat(rhodes-merger): A1 shadow mode against Surtr's published merge (AERIE-2610) @kevalshahtrilogy  approved

## Summary

Plan object A1 (rhodes-merger, AERIE-2610). The analytics worker's rhodes-merger can now run in shadow mode against Surtr's published merge. Surtr already builds the same per-field merge hourly (core_education.site_operational_metadata, Gateway source aerie-site-operational-metadata). Aerie's prod Gateway key can already read it.

- Gate: RHODES_MERGER_READ=legacy|shadow, default legacy. Merging and deploying changes nothing.

- Shadow: computes the merge and writes to Rhodes exactly as legacy. Afterwards, at most once an hour, it reads the Surtr snapshot and logs one rhodes_merger_shadow_check JSON line, a field-level diff keyed on site slug across the 9 written fields plus their provenance. It never publishes anything read from the Gateway. A Gateway or validation failure is logged as status: "degraded" and swallowed, and the tick's own clean/degraded status is never changed by shadow. The tick summary gets a ; shadow <status> suffix.

- Throttle: the task ticks every 5 min, but Surtr publishes hourly. One throttle per worker allows one Gateway read per hour; a failed read also uses its slot.

- Skips: the compare is skipped (the slot stays free) when the ISP or HubSpot bulk fetch failed. In that case the cycle's merge is partial, and comparing it would report false divergences.

### Files

- sync/src/analytics/queries/site-operational-metadata-gateway.ts: the Gateway reader.

- Zod validation. An absent column fails; only null/"" mean no value.

- One page (pageSize: 1000, maxPages: 1), so a read can't straddle a republish. A full page (possible truncation) is refused.

- Numeric fields must be non-negative, and classrooms and occupancy must be integers.

- A floor of 100 rows.

- A single refreshed_at with a 3 h max age and a future-timestamp guard. The table has no source_run_id, so refreshed_at stands in for it.

- Tuition arrives as a NUMERIC string and is converted to a number; schema_ui is mapped to schemaUi; NULL becomes absent.

- The client is built lazily, so importing the module never captures env before dotenv loads.

- sync/src/upstream/rhodes/site-metadata-shadow-compare.ts: compares the two merges per field. Each field is counted as same / differs / only-incumbent / only-Surtr / both-absent, plus a separate provenance-mismatch count.

- Duplicate slugs are never compared and never clean.

- Tuition is compared to the cent.

- Contact (email/phone) values are redacted in examples; presence is kept.

- Examples and slug lists are capped at 10; counts are exact.

- Also holds the throttle.

- sync/src/upstream/rhodes/sync.ts: the gate, collecting the merge, the shadow run, and the skip rules.

- sync/src/analytics-worker/index.ts: the throttle is created once per worker, and the shadow suffix is added to the summary.

- sync/src/scripts/dry-run-rhodes-merger-shadow.ts (pnpm run dry-run-rhodes-merger-shadow): runs the real merger once in shadow mode with a no-op Rhodes writer. It prints counts and slugs only. It follows the ESM rule: dotenv first, then dynamic import().

- .env.example: documents RHODES_MERGER_READ.

### Why there is no gateway mode

Where merged values should be written after cutover is undecided (Benji + Keval). The options are the external Rhodes deployment (archived per AI-183), an Aerie-local, override-aware apply step (depends on AERIE-2365), or retiring the task. A gateway mode would publish from Surtr, so it has to wait for that decision; building it now would mean guessing the write contract. Shadow doesn't depend on the decision, and its evidence feeds into it.

## Business Value

A1 is the last and most complex object keeping analytics-worker on the EC2 box. This PR lets the Surtr replacement be checked against the live merge in production with no behaviour change and no new write path. The write-target decision can then rest on measured per-field agreement instead of the one-off analysis in closed #1427. The local run below already shows 0 value disagreements on any field where both merges have a value. Every divergence is a value Surtr has and Aerie's merge lacks, and each one traces to a known join difference. That lowers the risk of the eventual cutover or retirement, which in turn removes ~4 merge modules and a 5-minute scheduler from the worker.

## Manual Effort Estimate

~12 hours of focused work by hand, proposed for Keval to confirm or adjust. Breakdown:

- ~2 h: reading the merger, field-merge, the A4 pattern, and the Surtr procedure/table.

- ~3 h: reader and validation.

- ~3 h: compare, gate, and throttle.

- ~3 h: tests.

- ~1 h: dry-run script and the live run.

## Testing / evidence

- pnpm --filter @bran/sync typecheck: clean. biome check and the root pnpm lint (boundaries, convex paths, read bounds, test architecture, knowledge): clean. The 2 warnings are pre-existing and in unrelated files.

- Sync tests with vitest run --maxWorkers=2: 83 files, 1424 tests passed. That includes 61 new tests: 33 for the reader, 12 for compare/throttle, and 16 for the gate/orchestrator.

- New orchestrator tests prove:

- legacy never calls the Gateway.

- shadow makes exactly the same Rhodes upserts as legacy.

- A Gateway error leaves both writes and errors untouched.

- The throttle allows 1 read per hour across 5-min ticks.

- ISP and HubSpot bulk failures skip the compare without using the throttle slot.

- An absent column fails validation; it never becomes NULL.

- Error messages never echo field values.

- Live read-only dry-run (details are in a PR comment): the real Gateway read (172 rows, one page) plus Aerie's merge computed locally (HubSpot from Redshift, ISP from DynamoDB/S3). Nothing was written to Rhodes, Convex or Redshift. Rhodes listSites was stubbed, because no Rhodes credentials exist locally. The stub reads Surtr's Redshift mirror of Aerie sites (staging_education_rhodes.raw_sites) with a read-only SELECT.

## Not covered

- Cutover and the write target: there is no gateway mode (see above).

- Prod shadow window: setting RHODES_MERGER_READ=shadow on EC2 is Keval's step (X3). Shadow only runs where the merger itself runs, i.e. where RHODES_CONVEX_SITE_URL + RHODES_API_KEY are set. Whether A1 runs in prod at all is still the spec's HARD blocker to verify.

- Surtr lineage columns (source_run_id / source_published_at / TIMESTAMPTZ): not added. The reader parses the naive refreshed_at as UTC (GETDATE()).

- Deleting the legacy merge files happens in a follow-up, after the decision.

Linear: AERIE-2610

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1564 — feat(sync): A6 REBL3 shadow mode, REBL3_READ=legacy|shadow|gateway (AERIE-2239) @kevalshahtrilogy  approved

## Summary

Spec object A6 (REBL3 sites). Adds a read-mode gate, REBL3_READ=legacy|shadow|gateway, to the analytics worker's daily rebl3 task. With it, rebl3Sites can be shadow-checked against Surtr's aerie-rebl3-sites Gateway source (mart mart_education.aerie_rebl3_sites, 490 sites) before any cutover. It follows the A4 recipe (school-source-directory-*).

> Context: Benji paused A6 on 2026-09-23 when the earlier read-path PRs (1415, 1417, 1423) were closed. Keval directed this build on 2026-09-29. Merging changes nothing in prod, because the gate defaults to legacy.

| Mode | Reads | Publishes | On a Gateway problem |

|---|---|---|---|

| legacy (default) | REBL3 | REBL3 rows (unchanged) | n/a: never touches the Gateway |

| shadow | REBL3, then Gateway | REBL3 rows only | logged as a degraded rebl3_sites_shadow_check line and swallowed; the publish is untouched |

| gateway | Gateway only (no REBL3 key needed) | Gateway rows, all or nothing | throws before the first write, leaving rebl3Sites as it was |

What's in it

- Gateway reader (sync/src/analytics/queries/rebl3-sites-gateway.ts), built on A4's SurtrGatewayClient:

- A Zod row schema where every mart column is required. An absent key fails validation; only an explicit null maps to no value.

- Each row maps to exactly the Rebl3SiteUpsert that the REBL3 path builds. A test pins parity with mapRebl3SiteToUpsert.

- The whole snapshot is refused on any invalid row, mixed or blank source_run_id / source_published_at, or a repeated site.

- Further refusals: fewer than 200 rows (the mirror sat at 100 for months before Surtr PR 1903), or a publication older than 48h or dated in the future.

- Error text never includes a column value.

- Relative-drop guard: the snapshot must still contain at least 90% of the slugs rebl3Sites already holds. REBL3_GATEWAY_MIN_INVENTORY_RATIO overrides the threshold.

- The baseline comes from a read-only call to the existing GET /sync/analytics/rebl3 route, so no Convex deploy is needed.

- rebl3Sites is upsert-only, so held slugs are a high-water mark.

- Shadow logs the would-be verdict (held / retained / wouldPublish) every run, so the real headroom is measured before any flip.

- Shadow compare (sync/src/upstream/rebl3/rebl3-sites-shadow-compare.ts):

- Keyed on rebl3Slug, comparing every Rebl3SiteUpsert field.

- The legacy side is this run's own REBL3 fetch output, not stored Convex values.

- Differences are counted as real only when REBL3's own updated_at / created_at predates the Surtr publication. Otherwise they are changedSinceSnapshot / newSinceSnapshot.

- Known mart transforms (the same instant written differently, and the N/A to NULL on 8 columns) are counted separately.

- Duplicate keys, or an incomplete REBL3 scan, are never clean.

- The log line carries counts, site ids and field names only.

- Gate in refreshRebl3Sites. Every run emits one structured [rebl3] sync {...} line with outcome: applied | degraded | refused | failed. The scheduler's env check is mode-aware: gateway needs SURTR_GATEWAY_API_KEY, not a REBL3 key.

- Dry-run: cd sync && pnpm run dry-run-rebl3-sites-shadow. It is read-only: every Convex upsert is intercepted, and dotenv loads before the dynamic await import() calls.

Known bugs

- AERIE-2240: fixed here (it sits in the transport this PR reuses).

- sendRebl3SyncOperation now throws on a 200 with no valid result, instead of substituting { upserted: 0, failed: 0 }.

- Both push loops count a batch as failed when its acknowledgement doesn't sum to the rows sent. That reconciliation was in PR 1417, which never merged.

- This is a deliberate change to the legacy path, visible only on a malformed Convex ack: such a run now reports degraded instead of clean.

- Item 2 (per-row isolation of a partially failed batch) needs a Convex change, so it's not included.

- AERIE-2241: not fixed. An upstream NULL still can't clear a stored value. Fixing it needs a Convex mutation contract change, and it affects both modes equally. The shadow compare runs on the upsert payloads, so this bug doesn't distort parity.

## Business Value

A6 is one of the objects keeping the analytics-worker EC2 container alive. It is on the critical path to the Phase-5 worker teardown and to removing the worker's direct REBL3 credential.

This PR makes the move reversible and provable:

- Shadow gives a zero-risk parity window against real production reads.

- gateway is an env flip with legacy as the rollback.

- It also closes a silent failure in today's sync: a Convex 200 without an acknowledgement used to report a clean run that wrote nothing.

## Manual Effort Estimate

About 14 hours of focused work (~2 days) to build this by hand with no AI: the reader and guards, the timing-aware compare, the three-mode gate, the ack fix, about 130 tests and the dry-run script. That assumes the closed PR 1415/1417 branches were available to salvage mapping and guards from. Proposed by Claude — Keval to confirm or adjust.

## Testing / evidence

- sync: pnpm run typecheck passes, pnpm run lint (biome) passes, and npx vitest run --maxWorkers=2 passes: 84 files, 1,448 tests (the chat suite was not run).

- New tests:

- reader/guards: 39

- compare/log lines: 16

- gate, ack fix and held-slug read: 33

- worker gating: 2

- Live read-only Gateway dry-run, 2026-09-29, with no writes anywhere:

- The Gateway snapshot had 490 rows in 5 pages and passed every guard.

- source_run_id=cefc525e-e694-4019-9c7e-86766ce293ec, published 2026-09-29 04:07:50 UTC (4.3h old).

- Shadow (REBL3 side = echo stub; no REBL3 key in the local env): legacy 490, gateway 490, matched 490. Zero real mismatches, zero legacy-only, zero gateway-only, zero duplicates, 0 known transforms.

- Gateway mode would publish 490 rows in 5 batches, with 0 failures.

- The relative-drop baseline was not evaluated locally, because there is no SYNC_API_TOKEN in the local env.

- The echo stub serves the live Gateway rows back through the real legacy loop, mapper, compare and log line. It proves the plumbing on real data shapes, not Surtr-vs-REBL3 parity. That parity comes from the prod shadow window, or from running the dry-run where a REBL3 key is set (it then uses the live REBL3 read automatically).

## Not covered

- The prod flip. Setting REBL3_READ=shadow in the EC2 .env and restarting the worker is Keval's step (spec X3). The spec proposes at least 7 daily runs with zero real differences before moving to gateway.

- The option A/B/C end-state decision (migrate vs retire vs leave) is still open on the A6 tab. This PR builds the option-A path, and its default changes nothing.

- Coupling with the A6 cleanup item. The relative-drop baseline reads slugs through the existing enrichment GET (/sync/analytics/rebl3). The cleanup item that deletes that GET must keep a slug-only read while gateway mode is in use.

- Gateway writes are not atomic. The read is all or nothing, but a batch that Convex rejects mid-publish is recorded as degraded and re-sent next run. rebl3Sites has no stage/publish protocol, and the mutation is upsert-only, so nothing is ever deleted.

- Enrichment. gateway mode never runs the REBL3 enrichment pass (production already skips it).

- AERIE-2241, and AERIE-2240 item 2 (see above).

Linear: AERIE-2239

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#154 — AI-949: Play a sound and show a macOS notification when the Pomodoro timer changes phase @ashwanth1109  no labels

## Demo

![AI-949 smoke test demo](https://github.com/AI-Builder-Team/Shipyard/blob/f0d8fa2146eaabf6b74408ef2650106baad32f56/docs/smoke-test-evidence/AI-949/demo.png?raw=true)

## Summary

Pomodoro phase changes (Focus → Break, Break → Focus) now play a macOS system sound and show a notification banner, including when Shipyard is behind other apps or minimized.

- src/pomodoro.ts: new refreshPomodoroTimer returns { state, phaseChange }. A refresh that crosses one or more boundaries reports a single change for the current phase, so sleeping through several phases gives one alert. advancePomodoroTimer now wraps it.

- PomodoroTray: transitions are computed against a ref, outside React state updaters, so StrictMode double-invocation cannot double-alert. Only the running refresh loop announces. Start, Pause, Resume and Reset never do. Start requests notification permission once.

- src/pomodoroAlerts.ts: plays the sound through a new play_pomodoro_sound command (AppKit NSSound: Glass when a break starts, Hero when focus starts) and sends a silent banner ("Break started · 5 min" / "Focus started · 25 min") via tauri-plugin-notification. If the native sound fails, a short Web Audio tone plays instead.

- Skip control: a new Skip button jumps to the start of the next phase. Its accessible name is "Skip to break" or "Skip to focus". If the timer is running, it keeps running and plays the normal phase-change alert, so you can test alerts without waiting 25 minutes. If the timer is idle or paused, Skip moves to the next phase silently and leaves it idle or paused.

- tauri-plugin-notification ~2.4 (Rust + JS) is registered, and notification:default is added to the capability. It is pinned to 2.4 because 2.5 requires a Tauri upgrade.

- tauri.conf.json: the main window gets "backgroundThrottling": "disabled" so WKWebView keeps the 1 s timer running while hidden or minimized (macOS 14+).

### Implementation note

On desktop, the plugin always reports notification permission as granted and cannot tell whether macOS actually shows the banner. So the sound is played natively and separately from the banner rather than attached to the notification. That way the sound still plays if the user turns off Shipyard banners in System Settings, and the sound doesn't play twice.

## Tests

- pnpm test:pomodoro: covers a single change firing once, Focus → Break → Focus titles, a multi-cycle sleep collapsing to one change, manual controls never firing, Skip behaviour when running, idle and paused, and UI-level alerts under StrictMode (once per change, hidden popover, visibilitychange after a long sleep).

- cargo test --lib pomodoro::

- pnpm build, pnpm theme:check

### Manual check (pending)

> In pnpm tauri dev, the notification plugin sends banners as Terminal (com.apple.Terminal) because the dev binary isn't a bundled app. macOS only shows them if Terminal is allowed to send notifications. To check banners, use a bundled build, or allow Terminal in System Settings → Notifications.

- [ ] Start the timer and press Skip, confirm that Glass and the "Break started · 5 min" banner appear right away.

- [ ] Start Focus, minimise or hide Shipyard, confirm that Glass and the "Break started · 5 min" banner arrive on time.

- [ ] Confirm that Break → Focus plays Hero.

- [ ] Deny notifications in System Settings, confirm that the sound still plays.

## Linear

https://linear.app/builder-team/issue/AI-949/play-a-sound-and-show-a-macos-notification-when-the-pomodoro-timer

#1530 — feat(camps): A5 CAMP_SOURCE_READ publish gate for refreshCampData (AERIE-2548) @kevalshahtrilogy  approved

Linear: AERIE-2548 (blocked by AERIE-2329)

> CI has never run on this PR. Its base is feat/camp-gateway-parity-check (#1478), not main, so the workflows don't trigger. CI will only run after #1478 merges and this PR is retargeted to main. All evidence below is local.

## Summary

Stacked on #1478 (base branch feat/camp-gateway-parity-check). #1478 only watches: it shadow-compares the 7 raw camp entities over the Surtr Gateway against Supabase and never publishes from the Gateway. This PR adds the cutover path, an A4-style read gate on refreshCampData:

CAMP_SOURCE_READ=pg|shadow|gateway (the name and values fixed in the migration spec's A5 tab; pg is the direct Supabase/Postgres read)

| Mode | Publishes from | Behaviour |

|---|---|---|

| pg (default, also any unrecognized value) | Supabase | Today's code path: the same 7 sequential reads, per-table upserts, and purge/projection rules. The read step moved into a helper; its behaviour is unchanged and covered by the existing refresh.test.ts camp tests. |

| shadow | Supabase | Publishes exactly as pg. After that, it compares the rows it just published (no second Supabase read) against the Gateway, using #1478's runCampSourceShadowCheck and its PII redaction. It logs one summary line. Gateway failures are logged and swallowed, and result is never touched. |

| gateway | Surtr Gateway (#1478's 7 readers) | Uses the same upserts, projection snapshot and mark-and-sweep purge as today. The gate is all-or-nothing: if any check below fails, nothing is upserted, purged or published. The previous snapshot stays live and the tick reports degraded. Once the gate passes, writes follow the same contract as pg (see Design choices). |

The gateway checks, all applied before the first write. Any failure refuses the whole cycle:

1. Every entity reads and passes its Zod schema, population floor and non-empty source_run_id (the #1478 readers).

2. No duplicate supabaseId in any entity.

3. Exactly one source_run_id within each entity, and the same one across all 7. Surtr's sp_refresh_camps publishes the 7 core tables atomically from one raw-sync run, so a mismatch means the reads straddled a publish.

4. Forward references resolve (9 relationships: location/week/registration → program, registration → location/parent/child, registration-week → registration/week, child → parent). This catches a short read of a *referenced* entity only.

5. Purge bound (added 2026-09-29). The worker lists each Convex camp table's current supabaseIds through a new read-only, paginated Convex op, listCampSupabaseIdsBatch. It then counts how many rows the purge would delete. If any table would lose more than CAMP_GATEWAY_MAX_PURGE_PCT (default 5%) of its current rows, or the listing fails, the cycle is refused. This covers the short reads that checks 1 and 4 miss: a registration-week, child or parent read that is short but above its floor and breaks no forward reference.

sourceRunId is stripped before the Convex upserts, because the Convex v.object validators reject unknown fields. The summer-camps scheduler requires SURTR_GATEWAY_API_KEY instead of Supabase in gateway mode; pg/shadow still require Supabase.

Also added:

- sync/src/scripts/dry-run-camp-gateway-publish.ts. It runs checks 1 to 4 plus the projection size bounds against the live Gateway, and prints counts, the run id and a verdict. It never talks to Convex, so it does not evaluate the purge bound (check 5).

- CAMP_SOURCE_READ and CAMP_GATEWAY_MAX_PURGE_PCT in .env.example.

### Design choices

- Refusal is all-or-nothing in gateway, even though pg is per-table. In pg, a failed Supabase table still lets the other six upsert. In gateway, a partial snapshot is refused, because the purge that follows would delete real rows.

- After the gate passes, the writes themselves are not one transaction, in either mode. This is unchanged from pg:

- Each table is upserted in atomic batches.

- The purge and the projection switch run only after every batch of every table succeeded.

- The request-path projections flip in one pointer change (publishCampProjectionSnapshot).

So a transient Convex failure mid-write skips the purge and the switch: no row is lost and the previous projection stays live. Detail readers of the raw camp* tables can briefly see this cycle's upserted rows next to the previous cycle's, until the next successful cycle, exactly as in pg today. Making the raw tables switch atomically would need versioned tables and reader changes across Convex for both modes. That's a separate design, flagged for Keval rather than bolted onto this gate.

- Only forward references are enforced; reverse references are not guaranteed by the data. Live Supabase on 2026-09-29 (counts only) had:

- 240 registrations with no week row

- 266 children and 259 parents with no registration

- 105 parents with no child

- 29 weeks with no booking

Requiring those directions would refuse every cycle. The purge bound covers them instead. The forward check is strict, because all 9 forward relationships had 0 dangling references in live Supabase on both 2026-09-26 and 2026-09-29.

- Purge-bound baseline: current Convex rows, not the last published count. The current Convex table is exactly the population the purge deletes from. It is correct on the first cycle after the shadow → gateway flip and after a worker restart. A "last published count" would need to be persisted somewhere; kept in process memory, it has no baseline after every deploy. The bound counts the *rows that would be purged*, not just the net count change, so a short read hidden by new rows is still caught. It implies the plain count-drop check.

- Threshold: 5%, configurable, fail-closed. An invalid value falls back to 5, never to unbounded. An empty Convex table allows anything, because there is nothing to purge. A real upstream bulk deletion over 5% keeps gateway refusing until an operator raises CAMP_GATEWAY_MAX_PURGE_PCT for one cycle or flips back to pg.

- Shadow runs after the publish. It cannot delay or alter what gets published.

- Errors never include row values. Every camp catch, in both modes, goes through describeCaughtCampValue. It reports non-Error values by type only, and it cuts a Convex error body where Convex starts quoting the rejected value (Value: / Object:). Gateway read errors in both modes go through #1478's describeCampReadError: issue count plus the first 3 issues, with invalid_enum_value/custom messages masked to their code. Duplicate, lineage, reference and purge-bound failures report counts and run ids only. The refusal tests check that sentinel PII values never reach console.*. That covers the tested paths, not every possible log line.

## Business Value

This is the missing cutover step for A5, the object that holds the whole summer-camps SaaS data set, including children's health data and parents' contact details. Without it, #1478 can prove parity but can never move Aerie off its direct Supabase credential. With it, the flip is an env change (pg → shadow → gateway) with an instant rollback. The gateway path refuses rather than publishes when a read looks incomplete, and it can't purge more than the configured share (default 5%) of any table in one cycle. That guard is only as good as its checks and its threshold; it is not a guarantee. Cutting A5 over removes one of the dependencies that block analytics-worker teardown.

## Manual Effort Estimate

Proposal for Keval to confirm or adjust: ~11 focused hours (was ~9 before the 2026-09-29 purge-bound round). The extra time went to working out that reverse references aren't guaranteed, adding the Convex listing op and its test, and the production-scale refusal tests.

## Testing / evidence

Local only (see the CI note at the top). Latest run: 2026-09-29 on ea5eca6f0, rebased onto #1478 at b2e8855dc (Mercy-approved).

- Typecheck: sync and chat (a Convex op was added) are clean.

- Lint: pnpm lint exits 0. The 2 warnings are pre-existing, in chat/skill/forge-api/scripts/sindri.mjs.

- Sync tests (--maxWorkers=2): 1,507 passed, including 57 in camp-source-read.test.ts.

- Targeted chat tests: only 4 files were run: campsAtomicUpsert.test.ts (touched here), the sync-route security test, and the two Data Health camp tests from #1478. 67 tests passed. The full chat suite was not run.

- What the gate tests cover:

- Mode parsing, and the default and unrecognized values keeping Supabase.

- The gateway publish path.

- Refusal with zero writes on a reader failure, cross-entity lineage mismatch, dangling reference, or Convex listing failure.

- The reviewer's scenario: 2,000 of 2,704 registration weeks, at live scale, refused with nothing written or purged.

- Children and parents missing only unregistered rows (the forward check alone accepts this snapshot) refused.

- A 5-row genuine deletion still publishing, and CAMP_GATEWAY_MAX_PURGE_PCT=0 refusing a single deletion.

- Shadow never publishing Gateway values.

- Mutation checks: disabling the purge bound fails 4 tests. Earlier rounds showed that disabling the lineage/reference guards fails 4, and making shadow publish from the Gateway fails 2.

- Read-only Supabase checks, counts only:

- Row counts: 27 / 27 / 183 / 2,614 / 2,704 / 2,522 / 1,883.

- 0 duplicate ids, and 0 dangling forward references.

- The reverse-orphan counts listed in Design choices.

- The gate's snapshot checks (1 to 4) accept today's rows as a synthetic single-run snapshot. The purge bound was not exercised against real Convex; nothing touched a Convex deployment.

## Not covered

- CI, see the top of this description.

- The live Gateway leg. There's no SURTR_GATEWAY_API_KEY locally, and the 7 raw camp sources aren't granted to any key yet (Keval's step, X4).

- The purge bound against a real deployment. listCampSupabaseIdsBatch ships in this PR's Convex code, so the worker change and the Convex deploy must go out in the same release. If the op is missing, gateway refuses every cycle (fail-closed), so the risk is a stuck cutover, not data loss.

- Grants, prod shadow, flip. These need Keval's approval, the EC2 .env (X3), an Aerie release (Benji, X2), then CAMP_SOURCE_READ=shadow for about 24 clean hourly cycles before gateway.

- Precondition for gateway: Surtr's Redshift client must follow NextToken (https://github.com/AI-Builder-Team/Surtr/pull/2036, merged 2026-09-23). The readers use 5,000-row pages.

- No freshness guard. The readers don't read refreshed_at, and that column's presence on the Gateway sources is unverified.

- Legacy deletion (Supabase reads and env vars, parity schedulers) comes after cutover.

## Open questions for Keval

1. Is 5% the right default for CAMP_GATEWAY_MAX_PURGE_PCT? Camp rows are rarely hard-deleted upstream, so a tighter value (1 to 2%) is also defensible.

2. Add a staleness guard (refuse if the mart's refreshed_at is older than N hours)? That needs the column confirmed on the 7 Gateway sources.

3. With CAMP_SOURCE_READ=shadow on, CAMP_SOURCE_SHADOW_CHECK_ENABLED duplicates it (a second Gateway read each hour). Should both stay on during the shadow window, or only this one?

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1478 — feat(camps): A5 Gateway-backed parity check for summer camp directories @kevalshahtrilogy  approved

Adds a read-only Gateway-backed validation layer for Aerie's Summer Camps sync, plus a "Summer Camps" tab on the Data Health page -- same pattern as A4 (school-source-directories). Started as registration-count-only parity (Surtr's 3 original camp marts had no field-complete equivalent to Aerie's raw Supabase read); extended below once Surtr registered field-complete Gateway sources for the raw camp entities.

## An honest scope limit, found before writing anything

Checked the real Redshift mart schemas first. Surtr's original 3 Gateway sources for A5 (aerie-camp-registrations, aerie-camp-week-summary, aerie-camp-location-summary) are denormalized, aggregate-shaped mart tables -- registration grain with joined names/emails but none of the PII Aerie's raw Supabase sync carries (allergies, medical conditions, DOB, addresses), plus two rollup tables with no row-level entity data at all. So the first commit here ships observability only (registration-count parity), not a cutover-track shadow mode.

## What's here (original scope)

- sync/src/analytics/queries/camps-gateway.ts -- Gateway reads for the 3 original marts. Registrations returns a minimal, purpose-built shape (registration id/status/source_run_id); week/location summary are population-floor checks only.

- sync/src/analytics/camp-registration-gateway-parity-check.ts -- compares Supabase vs Gateway registration counts, broken down by status. Logs camp_registration_gateway_parity_check, WARN when not clean.

- Wired into analytics-worker as an opt-in scheduler (CAMP_GATEWAY_PARITY_CHECK_ENABLED=true, default off).

- dry-run-camp-gateway-parity.ts -- on-demand verification script.

- "Summer Camps" tab (chat/app/(main)/sync/page.tsx + camp-directory-health.ts + camp-directory-pipeline-health-server.ts + /api/sync/camp-directories route) -- same structure as School Directories, gated behind operations.surtrParityExperimental.read, enforced server-side.

## Extended: full field-level parity across all 7 raw camp entities

Surtr's core_education schema already holds all 7 raw camp entity tables (fed by an existing pipeline chain) -- so rather than a cutover with nowhere to land, this became: register Gateway sources exposing those 7 tables (done, live on Surtr), then extend this PR with a genuine field-level shadow compare.

- sync/src/analytics/queries/camps-full-gateway.ts -- full-shape Gateway reads for all 7 entities (programs, locations, weeks, registrations, the registration<->week bridge, children, parents), Zod-validated against the real Redshift column names/types confirmed via direct schema introspection.

- sync/src/analytics/camp-source-shadow-compare.ts -- field-level compare per entity, keyed by the shared supabaseId. Lineage (source_run_id) is checked Gateway-side only -- Supabase, the live source, has no "run" concept, unlike A4 where both sides read a republished mart.

- PII safety: dim_camp_child (allergies, medical conditions, DOB, email, gender) and dim_camp_parent (email, phone, address) mismatch *values* are redacted to "[redacted]" before they can reach a log line -- only the field name and row id surface, so a mismatch is debuggable without a child's health data or a family's contact details ending up in CloudWatch. The other 5 entities log values in full (no PII).

- sync/src/analytics/camp-source-shadow-check.ts -- orchestrates all 7 entities in parallel; a read failure on either side of one entity is isolated and never blocks the other six.

- Second, independent opt-in scheduler (CAMP_SOURCE_SHADOW_CHECK_ENABLED=true, default off), alongside the original registration-count check, not replacing it.

- dry-run-camp-source-shadow-check.ts -- on-demand verification script for the full compare.

Pending, not in this PR's control: the 7 new Gateway sources are registered and live on Surtr, but not yet granted to a Gateway key -- Aerie's worker can't read them for real until that grant exists. Live dry-run execution against real data is blocked on that; this extension's evidence is unit + full regression coverage plus the confirmed real Redshift schema shapes.

## Verified locally

- Original 3-mart scope: dry-run script against real Supabase + Gateway data (local mock, no production Surtr credentials used) -- 2,614 Supabase registrations vs 2,614 Gateway registrations, clean, 0 status mismatches. Week summary 151 rows, location summary 27 rows. A synthetic 3-row status tamper correctly produced a WARN with the exact mismatch. Summer Camps tab renders correctly against the same mock.

- Full-parity extension: 37 new unit tests covering every query function's field mapping, PII redaction, lineage checks, and per-entity error isolation. Live execution pending the Gateway key grant above.

Full suites green: sync (86 files, 1,422 tests), chat (746 files, 11,267 tests), typecheck clean both sides, biome clean.

Linear: [AERIE-2329](https://linear.app/builder-team/issue/AERIE-2329/a5-surtr-gateway-parity-validation-for-summer-camp)

## Business Value

Extends the same Gateway-backed validation pattern proven on A4 to the summer-camps directories. The original scope surfaced a real finding before any cutover work was assumed possible (the marts couldn't carry PII or base entities); the extension turns that finding into the actual fix -- new Gateway sources exposing the raw entities -- and ships the field-level validation code ready to run the moment the one remaining access step (granting the sources to a key) is approved, instead of leaving "migrate camps" as an open-ended follow-up.

## Manual Effort Estimate

About 2.5-3 days by hand total: ~1.5 days for the original registration-count-parity scope (tracing the Supabase sync, checking mart schemas, building the Gateway client + parity check + scheduler + dry-run + Data Health tab + tests), plus ~1-1.5 days for the full 7-entity extension (registering 7 new Gateway sources, building full-shape query functions against confirmed Redshift schemas, a redaction-aware field-level compare module, orchestration with per-entity error isolation, scheduler wiring, dry-run script, and 37 new tests). Flagged for Keval to confirm/adjust.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2113 — Honor quarter-close window in CollectIQ sync @sanketghia  approved

## Summary

- Accept previous-quarter CollectIQ data through 13:00 UTC on day 2 of the new quarter and keep the snapshot labeled with its source quarter.

- After the cutoff, fail immediately if the source still reports the previous quarter; require Forecast and QTD labels to agree.

- Handle the Q4-to-Q1 year rollover and pass the same run timestamp through the local runner.

## Verification

- uv run --extra dev pytest — 60 passed

- ruff check src tests scripts/run_local.py — passed

- Local live-sheet dry run — 8 business units, 2026-Q3, no Redshift or S3 writes.

## Operational note

- The full-replace table behavior is unchanged. The new quarter replaces the previous-quarter rows when its source data arrives. This change has not been deployed.

#1531 — AERIE-2547: Finalsite tenant directory Gateway reader + shadow comparator @kevalshahtrilogy  approved

Linear: [AERIE-2547](https://linear.app/builder-team/issue/AERIE-2547/a4-finalsite-finalsite-tenant-directory-gateway-reader-shadow) (blocked by [SURTR-1523](https://linear.app/builder-team/issue/SURTR-1523/a4-finalsite-register-gateway-source-aerie-finalsite-tenant-directory))

## Summary

Finalsite was added as a third School source (AERIE-2300) after the QuickBooks/SIS Gateway work (AERIE-2323) merged. So SCHOOL_SOURCE_READ never applied to it: Finalsite always read Redshift directly, and that direct read can't be deleted. This PR adds the missing Gateway path, the same way QuickBooks and SIS work:

- Reader: queryFinalsiteTenantDirectoryViaGateway (sync/src/analytics/queries/school-source-directories-gateway.ts).

- Reads aerie-finalsite-tenant-directory in one page.

- Zod-validates every row: tenant_status must be active (same contract as the pg path), slug and source_run_id must be non-empty (z.string().min(1)), and the timestamp must carry an offset. An absent column fails validation.

- Population floor of 10 rows. Every complete full Finalsite run since 2026-08-11 has had 47–59 tenants.

- Comparator: compareFinalsiteDirectories (school-source-directory-shadow-compare.ts).

- Keyed on finalsiteTenantSlug.

- displayName values are redacted ("[redacted]") in mismatch records; key and field name stay visible. This uses a new optional redactedFields argument on the shared compare(). QuickBooks/SIS output is unchanged, and a test pins that.

- Duplicate keys and mixed lineage are never clean, same as the other two sources.

- Wiring: added to SCHOOL_SOURCE_GATEWAY_QUERIES and SHADOW_COMPARATORS in school-source-directory-refresh.ts.

- shadow still publishes from pg and logs a finalsite school_source_directory_shadow_check line. A Gateway error is logged and swallowed.

- gateway publishes Finalsite from the Gateway and fails closed with no pg fallback, like QuickBooks and SIS, but only once Finalsite is explicitly opted in (next bullet).

- Per-source read mode (fix round, 2026-09-29). SCHOOL_SOURCE_READ used to be one value for every source, so flipping it to gateway would also have cut Finalsite over. Finalsite fails closed in gateway until its key is granted, so QuickBooks/SIS could no longer be flipped on their own. Now:

- An optional SCHOOL_SOURCE_READ_<SOURCE> (_QUICKBOOKS, _SIS, _FINALSITE) overrides the global value for that one source.

- Finalsite is in GATEWAY_REQUIRES_SOURCE_OVERRIDE. With no override, a global gateway runs Finalsite in shadow, and only SCHOOL_SOURCE_READ_FINALSITE=gateway moves it.

- An unrecognized override is logged (one line per refresh) and ignored, and that source is capped at shadow, so a typo never selects gateway.

- With no overrides, QuickBooks/SIS behave exactly as before under every global value (unset, empty, pg, shadow, gateway, unrecognized). A table test pins this.

- Why this shape rather than a Finalsite-only flag: the override name is derived from the source registry, so every source gets the same two knobs, cut over one source or roll back one source, from the same code path. The only Finalsite-specific piece is one set entry, which is removed at legacy deletion along with the whole gate.

- Sibling fixes in the same reader file (the recipe's rules, applied to all three readers):

- QuickBooks/SIS source_run_id must now be non-empty.

- An absent nullable column now fails instead of reading as null. Surtr always serializes a SQL NULL as an explicit null (Surtr/src/db/redshift/client.ts), so real rows are unaffected.

- Mercy round 1 (2026-09-29):

- Every Gateway reader (QuickBooks, SIS, Finalsite) now checks lineage for the whole response, not only per row. It refuses a read that holds more than one source_run_id or source_published_at across all fetched pages, so a read that straddled a mart republish is stopped at the reader. singleLineage in the refresh still guards the publish independently.

- The population floor and the lineage guard throw SurtrGatewayError with count-only messages.

- The dry-run no longer prints raw errors. Both its failure lines go through the new describeErrorForOperator (sync/src/analytics/operator-safe-error.ts), because a ZodError's text includes the received row values:

- SurtrGatewayError messages are shown as they are.

- Zod failures are reduced to issue codes and paths.

- Anything else is reduced to its name and error code.

- Dry-run fix: dry-run-school-source-directory-shadow now covers Finalsite and loads dotenv before importing app modules (dynamic await import()).

- Before this, static imports ran first, so the default SurtrGatewayClient snapshotted an empty SURTR_GATEWAY_API_KEY even when .env.local set one.

- .env.example: the SCHOOL_SOURCE_READ comment now says QuickBooks/SIS/Finalsite, and SCHOOL_SOURCE_READ_FINALSITE is documented, with the _QUICKBOOKS/_SIS pattern. Prod compose.prod.yml uses env_file: .env, so no deploy change is needed.

The direct-Redshift Finalsite read is not deleted. That comes after a clean shadow window and the flip to gateway.

Dependency: the Surtr side ([SURTR-1523](https://linear.app/builder-team/issue/SURTR-1523/a4-finalsite-register-gateway-source-aerie-finalsite-tenant-directory)) registers aerie-finalsite-tenant-directory (mart_education.aerie_finalsite_tenant_directory, ordered by finalsite_tenant_slug). Surtr PR: [Surtr PR 2069](https://github.com/AI-Builder-Team/Surtr/pull/2069), merged. The source only exists once Keval runs the Surtr seed:gateway-aerie seed; until then the Gateway answers 404 for it.

## Flip sequence

Each prod step needs Keval's approval, and the analytics-worker restarts after every .env change.

1. Today: SCHOOL_SOURCE_READ=shadow. QuickBooks/SIS run in shadow. After this PR is released, Finalsite also runs in shadow and logs Finalsite shadow Gateway read failed … (404/403) each hour until steps 2–3.

2. [Surtr PR 2069](https://github.com/AI-Builder-Team/Surtr/pull/2069) is merged. Keval runs Surtr's seed:gateway-aerie seed to register aerie-finalsite-tenant-directory; until then the Gateway answers 404 for it.

3. Grant aerie-finalsite-tenant-directory (read) to the Aerie Gateway key, then verify with one read. This is Keval's step.

4. Watch school_source_directory_shadow_check with domain: finalsite until it is clean for the agreed window (proposal: 24 hourly cycles).

5. Flip QuickBooks/SIS: SCHOOL_SOURCE_READ=gateway, with SCHOOL_SOURCE_READ_FINALSITE unset. Finalsite stays in shadow. This step can run before steps 2–4 finish, because it no longer depends on them.

6. Flip Finalsite: SCHOOL_SOURCE_READ_FINALSITE=gateway, then observe the same window.

7. Delete the legacy direct-Redshift reads, the gate and the overrides (spec tab A4, a separate PR).

Rollback:

- One source: set SCHOOL_SOURCE_READ_<SOURCE>=pg, or =shadow.

- Everything: SCHOOL_SOURCE_READ=pg. Note that an explicit per-source override still wins over this, so unset the overrides too.

## Known conflict

#1533 (A3 school calendar) also edits school-source-directory-shadow-compare.ts. Whichever of the two PRs merges second needs a rebase and a retest. It is not resolved here.

## Business Value

A4 is the reference implementation for the whole Aerie EC2 → Surtr migration. It can't reach "legacy read deleted" while one of its three directories still needs a direct Redshift credential on the EC2 worker. This PR closes that gap.

Once the Surtr source is granted and a shadow window is clean, the information_schema probe and every direct Redshift directory query in Aerie can go. That removes one warehouse credential dependency from the worker, and brings analytics-worker one step closer to being deleted from compose.prod.yml. It also fixes the shadow dry-run script, which otherwise can't run a live comparison for any source.

## Manual Effort Estimate

Proposed: ~6.5 focused hours for Keval by hand, without AI. Keval, please confirm or adjust.

- Reader + schema + population-floor research: ~1h

- Comparator + redaction: ~1h

- Wiring and test updates across three test files: ~2h

- Dry-run ESM fix, real-data check, PR write-up: ~1h

- Fix round: per-source read mode, tests, flip-sequence write-up: ~1.5h

## Testing / evidence

All commands were run in the worktree on this branch. Fix-round numbers are on head 418246b6d, rebased on origin/main (92fd47992). Verbatim output is in the PR comments ([local evidence](https://github.com/AI-Builder-Team/Aerie/pull/1531#issuecomment-5848001014) and the fix-round comment).

| Check | Result |

|---|---|

| cd sync && pnpm typecheck | pass |

| pnpm lint (repo root: boundaries, convex-paths, read-bounds, test-architecture, knowledge, biome) | exit 0. The 2 warnings are pre-existing, in unrelated chat/skill/forge-api/scripts/sindri.mjs |

| cd sync && pnpm test --maxWorkers=2 | 82 files, 1,421 tests passed after Mercy round 1 (1,409 after the per-source fix round, 1,395 before it) |

| lefthook pre-commit (biome + typecheck-sync) | pass |

| cd sync && pnpm run dry-run-school-source-directory-shadow, live with the Aerie key (2026-09-29, [comment](https://github.com/AI-Builder-Team/Aerie/pull/1531#issuecomment-5885435993); re-run after Mercy round 1) | QuickBooks CLEAN 604/604, SIS CLEAN 164/164. Finalsite reports HTTP 404 (not seeded yet) on its own line and doesn't affect the other two |

New or changed tests:

- Gateway reader:

- response-level lineage, for every reader: two run ids refused, two source_published_at values refused, and the same instant in Data API vs ISO format counts as one snapshot

- the population floor is a SurtrGatewayError

- maps rows to the pg record shape, including Data API timestamp normalization

- calls listAll("aerie-finalsite-tenant-directory", { pageSize: 5000 })

- floor: 9 rows rejected, 10 accepted, 0 rejected

- rejects empty or null source_run_id, a non-active status, an empty slug, a null name, a bad or offset-less timestamp, and any absent column (all 5 columns)

- strips unknown columns

- propagates transport errors

- sibling checks: absent nullable column (SIS) and empty source_run_id (QuickBooks/SIS)

- describeErrorForOperator (4 tests): a SurtrGatewayError is shown as-is, while Zod values, response bodies and driver details are withheld.

- Comparator:

- clean case

- displayName redaction: the raw values never appear in the serialized result

- pg-only and gateway-only tenants

- duplicate slugs on either side

- mixed lineage

- republish race

- QuickBooks/SIS values still shown

- Refresh:

- pg mode never calls the Finalsite Gateway reader

- shadow: clean comparison logged, and pg rows published even when the Gateway disagrees (redacted mismatch logged)

- shadow: a Gateway 403 is swallowed, logged, and pg is published

- gateway mode, with the Finalsite override, publishes all three directories from the Gateway

- an opted-in Finalsite: a Gateway error fails closed with no pg fallback

- per-source read modes (fix round, 14 tests):

- every global value with no overrides: QuickBooks/SIS unchanged; Finalsite is capped at shadow under global gateway

- an unrecognized global value still reads pg for every source

- SCHOOL_SOURCE_READ_FINALSITE=gateway opts Finalsite in, with or without the global flip

- per-source rollback (Finalsite pg; QuickBooks pg under global gateway)

- a blank override counts as unset

- an unrecognized Finalsite override under gateway/shadow/pg never selects gateway and logs once

- an unrecognized QuickBooks override never selects gateway and leaves SIS alone

Read-only Redshift checks (Redshift Data API, SELECT only, counts only):

- mart_education.aerie_finalsite_tenant_directory:

- 59 rows, 59 distinct slugs, 1 source_run_id, 1 source_published_at

- all rows active

- 0 blank display names, 0 blank run ids, 0 non-canonical slugs

- 5 columns, all NOT NULL

- The exact SQL Surtr's declarative Gateway executor builds (SELECT * … ORDER BY finalsite_tenant_slug LIMIT 5000 OFFSET 0) compared with Aerie's direct-path query (… LIMIT 2001): 59 vs 59 rows, 0 differences in each direction (EXCEPT both ways).

- Real-data local run of this PR's reader and comparator. A temporary read-only script, not committed, did three things:

- read Finalsite through Aerie's existing pg reader

- read the Gateway-shaped SQL above, with timestamps rendered in the Data API's YYYY-MM-DD HH:MM:SS+00 form, through queryFinalsiteTenantDirectoryViaGateway

- compared the two with compareFinalsiteDirectories

Result: pgRowCount 59, gatewayShapedRowCount 59, matched 59, mismatched 0, pgOnly 0, gatewayOnly 0, duplicates 0/0, sourceRunIdsMatch true, clean true.

Live Gateway read (2026-09-29, [comment](https://github.com/AI-Builder-Team/Aerie/pull/1531#issuecomment-5885435993)).

- QuickBooks and SIS are clean against Redshift over the live Gateway.

- Finalsite answers 404 until the seed run and key grant (below). Shadow and the global flip swallow that error; a premature SCHOOL_SOURCE_READ_FINALSITE=gateway fails only Finalsite.

## Not covered

- Surtr source registration: [Surtr PR 2069](https://github.com/AI-Builder-Team/Surtr/pull/2069) (SURTR-1523) is merged. The seed run that creates the source row is Keval's step, and it has to happen before Finalsite's shadow comparison can succeed.

- Gateway key grant: Keval's step. Grant aerie-finalsite-tenant-directory (read) to the Aerie key, then verify with one read. Adding the source to the seed's aerie entity does not grant it to existing keys.

- Release order.

- Releasing this PR before the source is registered and granted is safe in shadow (prod today). Each hourly tick logs Finalsite shadow Gateway read failed (publishing from Redshift as usual) with a 404/403, and keeps publishing from Redshift.

- Flipping SCHOOL_SOURCE_READ=gateway is safe for QuickBooks/SIS on their own: Finalsite stays in shadow. Setting SCHOOL_SOURCE_READ_FINALSITE=gateway is not safe until the grant exists and Finalsite has its own clean shadow window (see Flip sequence).

- Prod shadow observation, the gateway flip, and deleting the legacy direct-Redshift read are follow-ups (spec tab A4).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1617 — Forecast V2: export the primary table as CSV @vvp-trilogy  approved

## Summary

- add a Forecast V2 CSV export using the executive table's dynamic headers

- export the searched, sorted school rows plus the matching Total row, preserving unavailable values as blank cells

- expose the export on desktop and mobile with the shared hardened CSV downloader

- hide the Program/Physical switch on Forecast V2 while continuing to honor the shared admissions usage preference

## Validation

- pnpm --dir chat exec vitest run components/dashboards/admissions/forecast/v2/__tests__/forecast-v2-view.test.tsx components/dashboards/admissions/forecast/v2/__tests__/forecast-v2-report.test.tsx --maxWorkers=1 (59 passed)

- pnpm --dir chat typecheck

- pnpm lint:test-architecture

- pnpm exec biome check chat/components/dashboards/admissions/forecast/v2/forecast-v2-report.tsx chat/components/dashboards/admissions/forecast/v2/forecast-v2-view.tsx chat/components/dashboards/admissions/forecast/v2/__tests__/forecast-v2-report.test.tsx chat/components/dashboards/admissions/forecast/v2/__tests__/forecast-v2-view.test.tsx

Closes #1616

#1533 — feat(school-calendar): read Surtr Gateway behind SCHOOL_CALENDAR_READ with shadow compare (AERIE-2544) @kevalshahtrilogy  approved

Linear: AERIE-2544 (supersedes AERIE-2242; closed PR #1422). Depends on SURTR-872's prod DDL for a clean shadow.

## Summary

A3 of the Aerie EC2 → Surtr migration: the analytics-worker school-calendar task can now take its data from Surtr's Gateway (aerie-school-calendar = mart_education.aerie_school_calendar) instead of reading the Google Sheet itself. The behaviour is gated by SCHOOL_CALENDAR_READ=legacy|shadow|gateway. The default legacy is today's behaviour, byte for byte. It follows the A4 (#1472) pattern.

- Gateway reader: sync/src/analytics/queries/school-calendar-gateway.ts. querySchoolCalendarViaGateway() validates every row with Zod:

- An absent column fails; only ""/null map to null.

- source_run_id must be non-empty after trim.

- The start date must be a real YYYY-MM-DD date.

- source_published_at is normalised with A4's toIsoInstant, now exported.

- A row carrying no calendar data is rejected.

- Snapshot-level checks: a row floor of 20 (the mart holds 72), a ceiling of 2000, a single lineage and unique slugs.

- Staleness: a publication whose source_published_at is older than 6h (the same limit as A2, #1528; Surtr's chain runs hourly) is refused. So is one stamped more than 5 minutes in the future.

- A validation failure quotes at most 3 Zod issues (path + message, never a value) and counts the rest.

- Shadow compare: sync/src/upstream/school-calendar/shadow-compare.ts.

- It compares the legacy run's own output (result.matched[].sites[], flattened per slug) with the Gateway rows. It does not compare against stored Convex values, which keep never-cleared stale values.

- A field the legacy run omitted matches a Surtr NULL, the same rule as the patch.

- The key is the site slug, and values are redacted.

- It emits one [school-calendar] shadow {"event":"school_calendar_shadow_check",...} line per run: INFO when clean, WARN otherwise.

- It reuses A4's compare / ShadowCompareResult / logShadowCompareResult, exported rather than duplicated. That module gained two opt-in options: Gateway-only lineage, and value redaction. The directory output is unchanged, and its existing tests still pass.

- Read gate: in sync.ts, plus a one-line change to the scheduler in analytics-worker/index.ts.

- shadow publishes from the Sheet as usual and then logs the compare. A Gateway error is logged as degraded and swallowed. No compare runs if the legacy run itself failed.

- gateway does no Sheet read, and shouldRun needs only CONVEX_SITE_URL + SYNC_API_TOKEN + SURTR_GATEWAY_API_KEY. No GSheet env is needed.

- A refused Gateway read (including a stale one) in gateway mode records a failed audit row and sends nothing to Convex. In shadow it logs as degraded.

- An unrecognized SCHOOL_CALENDAR_READ value is reported once per tick: the scheduler resolves the mode in shouldRun and passes it to the refresh.

- Convex: new internal mutation applySchoolCalendarsFromSurtr in chat/convex/analytics/gsheet.ts, registered in the handleGSheetSync dispatch.

- It takes {siteSlug, name, startDate, driveUrl} rows plus sourceRunId / sourcePublishedAt. There is no syncedAt argument.

- It is patch-only: a NULL or blank value is skipped and never cleared. It shares calendarPatchForRecord with the Sheet path.

- Cancelled and unknown slugs are skipped. An unknown slug marks the run degraded.

- Every run writes a schoolCalendarSyncRuns row whose syncedAt is Surtr's sourcePublishedAt, derived by the mutation itself and never the time of the call. So the public API v2 getSchoolCalendar freshness reports when the sheet was last read, and stops advancing if Surtr stalls (tested).

- An unparseable sourcePublishedAt is rejected before any write.

- The row has a new optional surtrSource field {gatewaySource, sourceRunId, sourcePublishedAt, skippedSites}, which is additive; existing rows and readers are unaffected.

- The mutation is code and tests only. Nothing is deployed.

- Dry run: pnpm run dry-run-school-calendar-shadow, plus a SCHOOL_CALENDAR_READ entry in .env.example. It pins readMode: "legacy" with dryRun: true, so it can never take the gateway write path. It exits non-zero when the compare is not clean. Following the ESM gotcha, dotenv loads first and the modules load afterwards with await import().

## Business Value

- Removes a hand-maintained duplicate. Today the campus→site matching for school calendars lives twice: once in Aerie's Convex mutation and once in Surtr's mart, whose README says "keep in manual sync". This PR is the step that lets Aerie stop reading the Sheet and matching itself, so one governed copy (Surtr) feeds every calendar consumer:

- the portfolio fact sheet

- the FTO matrix

- the public API getSchoolCalendar

- the rhodes MCP sites tool

- Retires a worker task. It is one of the tasks keeping the EC2 analytics-worker alive.

- Low-risk rollout. The shadow mode proves parity on live data before anything flips, and legacy stays the one-line rollback.

## Manual Effort Estimate

Proposal for Keval to confirm or adjust: ~16–20 focused hours (about 2.5 days) to build this by hand with no AI:

- reading the A4 pattern and the legacy Sheet/Convex matching code

- the Zod reader

- the run-output shadow compare and refactoring the shared A4 compare

- the three-mode gate and scheduler env logic

- the Convex mutation, the audit-row schema addition and the freshness wiring

- about 70 new tests across two packages

- the dry-run script

- the read-only Redshift parity and the Lexington/"Anywhere" investigation

The spec's own estimate for this PR is about 3 days by hand.

## Testing / evidence

Re-run on 2026-09-29 after the review fix round, on the branch rebased onto origin/main:

| Check | Result |

|---|---|

| cd sync && pnpm typecheck | clean |

| cd sync && pnpm lint (biome, 199 files) | clean |

| cd sync && pnpm exec vitest run --maxWorkers=2 | 83 files, 1427/1427 passed |

| cd chat && pnpm typecheck (app + convex/tsconfig.json) | clean |

| cd chat && pnpm exec vitest run --maxWorkers=2 convex/analyticsGsheet.test.ts | 39/39 passed (9 new) |

| cd chat && pnpm exec vitest run --maxWorkers=2 convex/publicApi/v2/portfolioDomain.test.ts (reads schoolCalendarSyncRuns; not edited) | 22/22 passed |

| biome check on the changed chat files | clean |

New and changed tests cover:

- Gateway reader: row mapping and timestamp normalisation; trimming; blank→null; absent column for each of the 5 fields fails; empty or blank source_run_id; bad or impossible dates; unparseable timestamp; empty or padded slug; an all-empty row; the floor; mixed run id; mixed published-at; duplicate slug; error propagation; accepted at exactly 6h, refused past 6h and when future-stamped; at most 3 issues quoted with no value.

- Shadow compare: clean; omitted field ≡ NULL; redacted mismatch; one-sided slugs both ways; duplicate slug on the legacy side; mixed Gateway lineage; empty run is not clean; the log-line format.

- Shared compare: Gateway-only lineage; redaction; the default is still strict; custom event and prefix.

- Read modes: env parsing and fallback; the default never touches the Gateway; shadow via option and via env; shadow swallows a 403; shadow skips when legacy fails; gateway sends rows plus lineage with no Sheet; dryRun passthrough; unknown-site degrade; failed audit on a read failure and on a push failure (stamped from the publication); the real reader refusing a 7h-old publication yields a failed audit and no Convex call; the apply payload carries no syncedAt; silent resolution; the apply HTTP op and its malformed result.

- Worker: gateway runs without GSheet env; gateway skips without a key; shadow still needs the GSheet env; an unrecognized value is reported exactly once per tick (this test fails against the previous code).

- Convex: patch plus a clean audit with lineage; never clears; cancelled and unknown slugs skipped (unknown → degraded); unchanged counting; dry run; duplicate or missing lineage is rejected before any write; syncedAt and public-API freshness come from the publication (a 3-day-old publication reads 3 days old); an unparseable sourcePublishedAt is rejected; the HTTP dispatch returns 200.

Read-only Redshift checks (2026-09-26). SELECT/WITH only, through the local read-only role. No values are printed; only counts, slugs and campus keys.

- Mart: 72 rows and 72 slugs from 47 campuses, one source_run_id, one source_published_at, refreshed 2026-09-26 16:20. Match methods: 65 site_name, 7 program_code, 0 override. One null field: a Drive URL. There is no Nashville campus in core_education.ref_academic_term (49 campuses), which confirms that the SURTR-872 DDL is still unapplied.

- Reader on real data: all 72 mart rows, in the Data API wire shape (DATE and TIMESTAMPTZ as Redshift strings), pass querySchoolCalendarViaGateway validation.

- Proxy shadow: Convex's *stored* values, read through Surtr's hourly staging_education_rhodes.raw_sites mirror for non-cancelled sites, compared with the mart. Result: 75 stored vs 72 mart. 71 match on every field, 1 mismatch, 3 legacy-only, 0 Gateway-only.

- The mismatch: 3815-washington-st-roslindale-ma.driveUrl. Stored [redacted], mart null.

- The legacy-only slugs: 1704-dorothy-pl-nashville-tn (SURTR-872), 92-hayden-ave-lexington-ma, 180-maiden-ln-new-york-ny.

- Separately, 9 cancelled sites keep old values; they are out of scope.

- Lexington and "Anywhere": both are stale, never-cleared values, not a matching difference; details are on AERIE-2544. Neither site is reachable by today's matching in Surtr or in Convex, per Convex's own school links mirrored in raw_school_links.

- Lexington's school carries only the "Alpha Lexington" program, and no sheet campus targets it. Its stored Drive URL equals the Boston campus row's.

- 180-maiden-ln's old Alpha New York link is archived. Its stored Drive URL matches no campus in today's sheet.

- Because the real shadow compares the run's output, neither should appear in it. Expected run-output diff today: Nashville as legacy-only until SURTR-872 is applied. There is one open question on Roslindale's Drive URL, below.

## Not covered

- The live Gateway half was not run. There is no SURTR_GATEWAY_API_KEY locally (blocker X5), and the Aerie key's grant on aerie-school-calendar is untested (X4). Keval runs pnpm run dry-run-school-calendar-shadow with a key.

- The legacy half of the dry run was not run either. There are no GSHEET_* credentials locally, and the legacy dry run POSTs syncSchoolCalendarsFromGSheet (dryRun) to Convex, which this lane is not allowed to do.

- No prod changes:

- no .env flip

- no convex deploy

- no SURTR-872 DDL apply, which is Keval's approval; the dependency is noted on SURTR-872

- no Mercy review; the PR stays draft

- Known conflict with #1531 (A4 Finalsite reader): both PRs edit sync/src/analytics/school-source-directory-shadow-compare.ts (+ its test). Whichever merges second needs a rebase and a retest of that file's tests plus this PR's shadow-compare tests.

- Follow-up: ambiguous and unmatched visibility in gateway runs. Legacy runs record ambiguous and unmatched campus rows in the audit and mark ambiguity as degraded. Gateway runs can't: the mart excludes ambiguous sites and reports them only in the procedure's RAISE INFO, and unmatched campuses aren't published. This isn't cheap from Aerie's side. It needs a Surtr contract addition (for example a companion Gateway source of excluded/unmatched campus rows), and then the mutation can record and degrade on it. Until then, gateway audit rows carry empty ambiguous/unmatched lists.

- Legacy deletion is a follow-up PR:

- the parser and the GSheet parts of sync.ts

- GSheetClient.getSheetGridData

- the Convex matching block

- GSHEET_SCHOOL_CALENDAR_*

- the gate itself

- Open questions:

- Roslindale: does the Sheet row have a Drive link that the legacy parser extracts but Surtr's raw sync doesn't (smart chip vs hyperlink)? If so, the shadow will show it as a driveUrl mismatch.

- Surtr voids struck-through rows unconditionally; Aerie voids them only when another row shares the campus (spec decision 3).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1528 — feat(sync): A2 Schools Data Sheet via Surtr Gateway, with shadow compare and read gate (AERIE-2543) @kevalshahtrilogy  approved

Linear: AERIE-2543 (supersedes AERIE-2243 and the closed PR #1421)

## Summary

Aerie EC2 → Surtr migration, object A2: the analytics-worker gsheet task that mirrors the exec Schools Data Sheet into Convex schoolsDataSheet. This follows the A4 pattern (#1472): a Gateway reader, a shadow compare and a read-mode gate. The default behaviour doesn't change.

- sync/src/analytics/queries/schools-data-sheet-gateway.ts: reads Surtr's aerie-schools-data-sheet Gateway source (mart_education.aerie_schools_data_sheet, one row per cell) in one listAll page. It folds the rows back into the exact SchoolsDataSheetRecord[] that extractRawSchoolRecords builds today, with the same fields, key order and record order.

- Every mart column is validated with Zod. Every mart column is NOT NULL, so an absent key or a null fails the read and is never published as "". source_run_id and source_extraction_id must be non-blank.

- Guards (each one refuses the whole snapshot, because the Convex write is replace-all):

- zero rows

- single lineage (run id, extraction id and published_at)

- unique cell id and unique (column, row)

- header fields constant within each column

- row_label constant within each row

- full grid

- exactly one well-formed defaults column, at index 1

- plausibility floors: at least 20 school columns and at least 20 data rows (live: 40 and 45)

- staleness limit: 6h

- Error messages name columns, rows and fields, never cell text.

- sync/src/upstream/gsheet/schools-data-sheet-shadow-compare.ts: keyed on sheetColumnIndex. It compares defaultName, location, hiddenName, isDefaultsRow, values.length, and label + value at each position.

- value and hiddenName are always logged as [redacted], with a whitespaceOnly flag so a whitespace-only difference is recognisable.

- It emits one schools_data_sheet_shadow_check JSON line per cycle with status clean, mismatch (WARN) or degraded (WARN).

- Gate SCHOOLS_DATA_SHEET_READ=sheet|shadow|gateway (default sheet = today) in sync.ts:

- shadow publishes from the sheet as usual, then compares. A Gateway failure is logged as degraded and never touches errors or the publish.

- gateway publishes the Gateway records with no Sheets call. The spreadsheetId partition is unchanged (GSHEET_SPREADSHEET_ID), so replace-all still works. syncedAt = source_published_at. Any read or guard failure aborts before the Convex write. This also closes the existing "a sheet with fewer than 2 rows wipes the table and reports clean" hole in that mode.

- Scheduler: buildGSheetScheduler.shouldRun is now mode-aware. gateway needs SURTR_GATEWAY_API_KEY + GSHEET_SPREADSHEET_ID, not the Sheets service account. The edit in analytics-worker/index.ts is a few lines.

- .env.example: one new line, SCHOOLS_DATA_SHEET_READ.

- dry-run-schools-data-sheet-shadow: counts, column indices and field names only; exits 0 or 1. It uses dotenv config() first, then dynamic await import().

- dry-run-gsheet.ts (the legacy Sheets dry-run) now pins readMode: "sheet", so it keeps verifying the Sheets read whatever SCHOOLS_DATA_SHEET_READ says.

- toIsoInstant in the A4 Gateway reader is now exported so it can be reused. Its behaviour is unchanged.

## Business Value

This moves the A2 sync off the EC2 worker's direct Google Sheets read and service-account credential and onto Surtr's governed, hourly-refreshed mart. It is one of the objects that has to move before analytics-worker can be retired. It can be deployed safely right now, because the default is unchanged. The shadow mode then produces a field-level, PII-redacted diff, so the go/no-go on cutover rests on evidence rather than a guess. Gateway mode also refuses thin, torn or stale snapshots instead of wiping Site Detail / API v2 school data, as today's sheet path can when the sheet read comes back short.

## Manual Effort Estimate

Proposed: ~10 hours of focused work for Keval by hand, no AI. That covers reading A4 and the mart DDL/stored procedure, writing the reader and its guards, the compare, the gate and the scheduler change, about 70 tests including an independent TS port of the stored procedure for cross-implementation checks, the dry-run script, and the local Redshift verification. *Proposed number: Keval to confirm or adjust.*

## Testing / evidence

All commands run from sync/ in the worktree at head 1fedf8572, after rebasing on origin/main (2026-09-29):

- npx tsc --noEmit: clean

- npx biome check .: 199 files, no fixes

- npx vitest run --maxWorkers=2: 83 files / 1435 tests passed (+72 new over main)

- schools-data-sheet-gateway.test.ts (43):

- Cross-implementation: mart-shaped rows built by an independent TS port of sp_refresh_aerie_schools_data_sheet (JavaScript .trim() semantics, mirroring the procedure as changed by AI-Builder-Team/Surtr#2068; MD5 ids, shuffled ORDER BY cell_id, Data API types) from values-response.json, sheet-with-tag-column.json and sheet-with-empty-default.json. The reader's output equals extractRawSchoolRecords(fixture) byte-for-byte (JSON.stringify, key order included).

- The untrimmed defaults hiddenName and raw row labels are preserved.

- Three trim-parity tests (non-breaking-space-edged cell; the wider .trim() set with interior whitespace kept; NBSP-padded markers, labels and blank names) assert byte-identical output against the parser. All three fail against the old spaces-only BTRIM semantics.

- Each of 16 absent, null or blank field cases is refused.

- Every guard has a test, and refusal messages are asserted to hold no secret text.

- schools-data-sheet-shadow-compare.test.ts (15): scenarios salvaged from #1421, namely exact mirror, a differing cell (redacted), a duplicate label told apart by position, a cell on one side only, a skipped data row shown as shifted cells, and a differing column header. Plus: a whitespace-only flag, a redacted hidden name, one-side-only columns, duplicate column indexes, the example cap, and the clean/mismatch/degraded log lines.

- sync.test.ts (+13): the gate's default and unknown-value handling. sheet never calls the Gateway. shadow publishes the sheet on clean, mismatch and a Gateway throw. gateway publishes the Gateway records under the same spreadsheetId with no Sheets call and no service account. gateway aborts before Convex on a read or guard failure and on a blank spreadsheetId. Mode-aware config-gap cases.

- tests/analytics-worker/index.test.ts (+1): the scheduler runs in gateway mode with no service account set.

- Live Gateway dry-run (2026-09-29, Gateway leg only) through querySchoolsDataSheetViaGateway on a real SurtrGatewayClient with the Aerie key: 1,845 rows in 1 page, all default guards passed, 41 records × 45 values. It matches the same mart read directly over read-only Redshift exactly (0 mismatched columns and cells, same lineage, same cell ids and order). On the 07:17 UTC publication (created_by = mart-aerie-schools-data-sheet-refresh/v2, i.e. with Surtr PR 2068 applied) the previously flagged trim-gap cell is gone: 0 fields where JS .trim() would change a value. Verbatim output is in the PR comments.

- Read-only Redshift check (2026-09-26, live mart, .env.local read-only role, SELECT only). Read mart_education.aerie_schools_data_sheet, shaped it like the Data API and fed it through querySchoolsDataSheetViaGateway with an injected client:

- 1,845 rows → 41 records (1 defaults + 40 school columns) × 45 values each

- all default guards passed

- one lineage (run 76683f82-…, published 35 min earlier)

- records ascending by sheetColumnIndex

- 1 cell where JS .trim() would differ from the mart (column 20, value position 31; value redacted). This was the BTRIM gap, since fixed on the Surtr side (see the live Gateway dry-run above).

- pnpm run dry-run-schools-data-sheet-shadow itself was not run: it needs the Sheets service account as well, and there are no GSHEET_* credentials locally. The Gateway leg above used a read-only scratch variant.

## Not covered

- The live Sheets side of the shadow compare. No service-account credentials exist locally. The full dry-run needs Keval to run it with the SA key + a gwk_ key.

- Prod shadow window and cutover. These need Keval's approval to set SCHOOLS_DATA_SHEET_READ=shadow on EC2, a release slot from Benji (one object at a time), and about 48 clean hourly checks. Shadow lines only reach container stdout.

- Surtr trim fix: resolved on the Surtr side. Decision 2 went to the stored procedure: AI-Builder-Team/Surtr#2068 (SURTR-1524) was applied in prod (the v2 publication was observed at 07:17 UTC on 2026-09-29) and merged at 08:04 UTC. No reader-side trim is needed, and shadow should no longer show the whitespace-only mismatch.

- Open decisions for Keval, all encoded as current defaults:

- syncedAt = source_published_at (decision 3)

- keep GSHEET_SPREADSHEET_ID as the partition label (decision 4)

- guard values ≥20 / ≥20 / 6h (decision 5)

- Legacy deletion (sheet path, parser.ts, dry-run-gsheet.ts): a separate PR and ticket after cutover.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#142 — Gate auto-approve only on the checks the base branch requires (AI-948) @kevalshahtrilogy  approvedmercy-allow-critical

## Summary

Mercy's CI gate treated any failing check on the head SHA as "required CI checks are failing" and withheld auto-approve. On Aerie, the new optional Praxis workflow (job review) fails AWS OIDC on every pull_request_target sync, so Mercy has been withholding on every Aerie PR even with all seven checks the main ruleset requires green (Aerie PR 1528 @ 6c12c758).

The gate now asks the base branch what it requires and only lets those checks block:

- New harness/ci_gate.py: takes <name>\t<conclusion> lines for the head SHA (same jq, same self-exclusion, same latest-run-per-name), then computes the required set:

- repository rulesets: GET /repos/{o}/{r}/rules/branches/{b}, rules of type required_status_checks → parameters.required_status_checks[].context (org-level rulesets included);

- classic protection: GET .../branches/{b}/protection/required_status_checks. The built-in token has no Administration:read, so on anything but a 200 it reads the same list from the branch summary (GET .../branches/{b}).

- A failing check in the required set → failing (withheld, same message as today). A failing check outside it → listed in the review body as ℹ️ Failing checks not required by the base branch, and does not gate.

- Fail safe: any of these makes every failing check block, same as today:

- a lookup error or an empty required set;

- any ruleset rule type outside an allowlist of types known not to gate on CI. The allowlist covers ref/commit/file/review restrictions and merge_queue. Everything else fails safe: required_deployments, workflows, code_scanning, code_quality, code_coverage, license_compliance_scanning, and any rule type GitHub adds later;

- a partial answer: a 200 that fails to name a required check. That means a malformed rule, a required_status_checks rule with no list or an unnamed entry, a classic object missing contexts/checks, or a branch summary with no readable protection.enabled. Treating these as an empty set would un-gate the check they failed to name. Each fallback logs why.

- Fail closed on unreadable CI: a CI state Mercy can't read is now error, which withholds auto-approve with the reason "CI status could not be read". This covers the check-runs read failing (it used to leave unknown, which didn't block), the status's starting value, an unreadable conclusions file, a counted check whose conclusion is neither a documented failure nor a documented pass, ci_gate.py crashing or raising, and any unrecognised value from ci_gate.py. Keval approved this as part of AI-948: it only makes the gate stricter. unknown stays as the "not supplied" default for local runs, and the workflow never passes it.

- Pending: unchanged. Only a confirmed failure blocks, and ok/pending come from all completed checks, as before.

- decide_review.py gets --ci-nonrequired-failures (informational only, never gates) and records ci_nonrequired_failures in decision.json. Nothing else in the verdict logic changes.

Permissions: none widened. Per GitHub's "Permissions required for GitHub Apps" (the built-in token is an installation token), the endpoints need these scopes:

- GET /repos/{o}/{r}/rules/branches/{b} needs Metadata: read. Every installation token has it.

- GET /repos/{o}/{r}/branches/{b} needs Contents: read, which this workflow grants.

- GET .../branches/{b}/protection/required_status_checks needs Administration: read, which the built-in token can never hold. That is why the gate falls back to the branch summary.

The rules endpoint also returns active rules from every level, repository and organization. includes_parents belongs to the /rulesets listing endpoints and is not a parameter here.

Linear: AI-948

## Business Value

This unblocks Mercy approvals on every Aerie PR. On Aerie, Mercy's approval is the only merge path for self-authored PRs, so one optional, unrelated workflow failing had stalled the whole repo's merge queue. The same fix stops any future optional check (a new bot or experimental workflow) from blocking merges in Surtr, Klair, or Sindri. Required checks still gate exactly as before, so "never approve a red build" keeps its meaning.

## Manual Effort Estimate

~5 hours of focused work. That covers reading the gate across mercy.yml and decide_review.py, learning how the rulesets and branch-protection APIs behave with the built-in token, writing the module with its fail-safe paths, the workflow wiring, and about 40 tests. Keval: please confirm or adjust.

## Testing

- pytest harness/tests → 678 passed. pytest heimdall/tests → 1264 passed, 1 skipped. ruff check harness heimdall and ruff format --check harness (ruff 0.15.22) are clean, and so is actionlint.

- New harness/tests/test_ci_gate.py:

- only a non-required check failing → ok, approve possible (Aerie 1528 replica);

- a required check failing → failing;

- ruleset lookup failing (0/403/404/500/non-list) → failing (fail safe);

- classic lookup failing → failing;

- lookup returning nothing → failing;

- every CI-gating rule type, plus an unknown future one → failing. Every allowlisted type still scopes;

- classic protection via its own endpoint and via the branch-summary fallback, and rulesets ∪ classic both covered;

- partial ruleset, classic and branch-summary answers → failing. Live responses from Aerie, Surtr, Klair, Sindri, trilogy-drones and mercy all parse under the strict rules.

- case-insensitive matching, pagination, and gh HTTP-status parsing;

- the branch encoded as one path segment (including a /) on all three endpoints;

- the Aerie replica against both the built-in token's 403 and an admin token's 404 on the classic endpoint;

- main() end to end.

- Fail closed: compute_guards/e2e withhold on error, and ci_gate.main returns error on unreadable conclusions; and a contract test pins every read-failure branch (start value, else branch, ci_gate.py crash, case validation) to error, and a crash inside ci_gate.main gives error as well. I also ran the extracted shell block against stubs: gh 502 gives error, and garbage from ci_gate gives error.

- test_pr_review.py: e2e APPROVE with the non-required note, a required failure still withheld, and a malformed or missing names file ignored.

- test_workflow_contract.py: BASE_REF is declared, names reach ci_gate.py, the names file reaches decide_review, and the workflow token stays contents/checks read.

- I ran the step's gate snippet locally against Aerie PR 1528 @ 6c12c758. It found 7 required checks from rulesets, review as failing but not required, and set CI_STATUS=ok.

## Rollout

- Aerie, Surtr, Klair, and Sindri call mercy.yml@main with harness_ref: main, so the change is live for them on merge. trilogy-drones pins @v1, so the v1 tag moves to the merge commit afterwards (old SHA recorded for rollback).

- Canary: re-run Mercy's review run on Aerie PR 1528 and confirm it no longer withholds over Praxis.

- Rollback: revert this PR on main (the ref the Aerie/Surtr/Klair/Sindri callers track), and move v1 back to the recorded SHA.

## Not covered

- Praxis's OIDC trust policy (Aerie/IAM). That is owned by the Praxis side and not touched here.

- The heimdall steward's own "red CI" heuristics (ci_fix / nudge_mercy) still look at all checks. They aren't Mercy's verdict and are out of scope.

- Required commit statuses (as opposed to check-runs) were never part of Mercy's gate, and they still aren't.

- Check-runs sharing a name across --paginate pages are still grouped per page. This is a pre-existing limitation.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#151 — AI-942: Evals execute real node replays from immutable entry snapshots @ashwanth1109  no labels

## Demo

![AI-942 smoke-test failure evidence](https://github.com/AI-Builder-Team/Shipyard/blob/75e463a7b05fb6af1fc06b550c004a8506046e2d/.smoke-evidence/ai-942-versioned-case-blocked.png?raw=true)

User-approved POC acceptance: PASS. The attached smoke evidence records the replay as Blocked because the production-only node input is unavailable; this is accepted for this proof of concept and is not a production-readiness claim.

## Linear

https://linear.app/builder-team/issue/AI-942/evals-execute-real-node-replays-from-immutable-entry-snapshots

## Summary

- Capture immutable node-entry snapshots with repository state, prompt/template provenance, runtime settings, and content-addressed evidence.

- Restore those snapshots into owned replay worktrees and execute production node and feature replays through the registered Codex engine.

- Persist real thread, turn, action, artifact, lifecycle, interruption, and recovery evidence while keeping fixture execution explicitly marked.

- Support candidate template byte substitution with hash validation and preserve dynamic/path provenance.

## Business Value

Eval comparisons now exercise the real node runtime from the exact captured entry boundary, so regressions can be attributed to the product or template with reproducible repository and prompt evidence instead of fixture-only behavior or mutable current checkouts.

## Implementation Effort

Large: this spans SQLite capture/export, eval manifest validation, immutable bundle restoration, real agent execution, lifecycle controls, feature replay orchestration, reporting, TypeScript contracts, documentation, and focused regression coverage.

## Test Plan

- [x] cargo test --locked --manifest-path src-tauri/Cargo.toml --lib (318 passed, 2 ignored)

- [x] pnpm test:replay-reports

- [x] pnpm exec tsc --noEmit

- [x] cargo check --locked --manifest-path src-tauri/Cargo.toml

- [x] git diff --check

- [x] Focused replay, eval-contract, feature-replay, eval-scoring, and workflow checks passed.

## Acceptance Limitations

- A live Codex provider turn was not run in this environment; real-agent coverage uses the provider-neutral engine boundary and recorded fixture regression seams. The production path blocks when Codex cannot prove the run-owned workspace and records that limitation as evidence.

- This POC is accepted with its production-node limitation; production readiness is explicitly out of scope for this merge.

## Smoke test result

PASS: User explicitly accepted the blocked POC replay and requested merge; no production-readiness claim.

#1615 — test(financials): pin the clock in QTD reports test (AERIE-2662) @kevalshahtrilogy  approved

## Summary

chat/components/dashboards/financials/qtd-reports-view.test.tsx started failing on main (6df20f88f) on 2026-10-01. Linear: [AERIE-2662](https://linear.app/builder-team/issue/AERIE-2662/fix-qtd-reports-test-broken-by-the-q4-quarter-rollover).

Root cause: QtdReportsView (qtdCurrentPeriodCode()) and QtdContractorTrace (getQuarterBoundaries()) work out the current quarter from the real clock, but the test fixtures describe 2026-Q3 (Jul-Sep 2026). Once the clock entered Q4, the fixture QB transactions and XO invoices fell into the trace's "Previous Quarter" bucket. Four drilldown tests could then no longer find the ... current quarter QuickBooks transactions tables or the XO invoice toggle.

Fix (test-only): the test now fakes only Date (vi.useFakeTimers({ toFake: ["Date"] }), the same convention as app/(main)/sync/__tests__/page.test.tsx). It pins the clock to 2026-08-18T12:00:00Z, which is inside the fixtures' quarter and the day after their 2026-08-17 reporting cutoff, and restores real timers in afterEach. The component is untouched.

## Business Value

This unblocks CI. The Test job has been red on every Aerie PR since the quarter rolled over, so nothing could show a green check. The suite also stops depending on the calendar, so it won't break again at the next quarter boundary.

## Manual Effort Estimate

About 1 hour of focused work by hand: reproduce, trace the quarter logic through the view and contractor trace, pin the clock, sweep the package for other clock-dependent quarter tests, and run checks. *Proposed by Claude. Keval, please confirm or adjust.*

## Testing

- vitest run components/dashboards/financials/qtd-reports-view.test.tsx: before the fix, 4 failed and 30 passed; after it, 34/34 pass. These are the same 4 failures CI shows on recent PRs (for example run 36818755046), and that is the only failing file there.

- Sweep for other clock-dependent quarter/period tests: I ran the 51 other chat test files that touch quarter, period, QTD, YTD or fiscal logic (financials, education P&L, consolidated, school-year-period, the finance Convex dashboards, the public API, and others). All passed (1405 passed, 17 skipped), so nothing else needed fixing.

- biome check on the touched file, pnpm lint:test-architecture, and tsc --noEmit (chat) are clean. The pre-commit hook (biome and chat typecheck) passed.

## Not covered

- No component changes. The quarter-boundary maths in qtd-contractor-trace.tsx and qtdCurrentPeriodCode() is correct across the Q4 to Q1 year wrap.

- Possible follow-up, not changed here: several QTD column labels are hardcoded to the first school-year quarter: "Q1 SY 26/27 Model" and "QTG for Q1" (guide staffing and leadership comparison), "Q1 SY26/27" (facilities), and "Q1 HC" (unit economics). They pair with the q1* budget fields. The page now asks for 2026-Q4 (SY Q2) on its period-scoped sections, so these labels may no longer describe what users see. Someone who owns QTD should check this.

- I did not verify that the warehouse has 2026-Q4 QTD data yet. On the first days of a quarter the page asks for a period that may still be nearly empty.

## Testing contract

### What this PR delivers

A test-only fix that makes the QTD Reports view tests independent of the current date, so CI Test goes green again.

### Who uses it and where

Aerie engineers and CI. There is no change to the product surface (Financials → QTD Reports).

### Conditions needed

None beyond the unit test environment. The test pins the clock itself.

### Expected behavior and examples

qtd-reports-view.test.tsx passes on any real date: inside Q3, after the Q4 rollover, and across the Q4 to Q1 year boundary. Production behavior is unchanged.

### Limits and unanswered questions

See "Not covered" above.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#152 — AI-946: Keep Pi image attachments visible and usable in chat @ashwanth1109  no labels

## Demo

![Pi image attachment smoke test](https://github.com/AI-Builder-Team/Shipyard/blob/codex/ai-946-pi-image-attachments/.smoke-evidence/AI-946/pi-image-attachments.png?raw=true)

## Linear

https://linear.app/builder-team/issue/AI-946/keep-pi-image-attachments-visible-and-usable-in-chat

## Summary

- Package and validate Photon plus its WASM asset in the Pi runtime.

- Restore saved local attachment previews when Pi history omits inline image content.

- Show an honest fallback status and hide recognized transport metadata from chat text.

- Add parser, conversation-store, UI, packaged-runtime, release-audit, Rust, and smoke-fixture coverage.

## Testing

- pnpm build

- pnpm test:messages

- pnpm test:conversation-store

- pnpm test:chat

- pnpm test:recovery

- pnpm test:pi-package

- pnpm test:release

- pnpm test:smoke

- cargo test --manifest-path src-tauri/Cargo.toml --lib pi_agent::tests

The repository-wide cargo fmt check still reports pre-existing formatting drift in unrelated Rust files.

#153 — AI-947: Let the companion look up Shipyard data outside the current view @ashwanth1109  no labels

## Demo

![AI-947 smoke test: companion Shipyard lookup](https://github.com/AI-Builder-Team/Shipyard/blob/8ff7cfb41d6692862561752e71a0eac1a367d894/docs/smoke-test-evidence/ai-947-demo.png?raw=true)

## Linear

https://linear.app/builder-team/issue/AI-947/let-the-companion-look-up-shipyard-data-outside-the-current-view

## Summary

The companion can now answer questions about any Shipyard task, project, release, template, or database table from any view. It does this through a fixed set of allowlisted, schema-validated, read-only backend tools. Foreground context is still the default, and questions about the current view need no tool call.

## What changed

- src-tauri/src/companion_tools.rs (new): the tool gateway.

- Declares one Codex dynamicTools namespace, shipyard, containing search_entities, get_task, list_tasks, get_project, get_release, get_template, list_database_tables, and get_database_table. Every input schema uses additionalProperties: false, enums, and bounds.

- Handles item/tool/call:

- Rejects any thread that isn't the active companion thread (CompanionSession), any other namespace, and any unknown tool.

- Strictly validates arguments: no unknown keys, types checked, ranges and enums enforced.

- All SQL uses prepared statements checked with Statement::readonly() and bound values.

- Table and column identifiers are validated against sqlite_master and PRAGMA table_info before quoting.

- LIKE wildcards are escaped.

- Only fixed operators are allowed: = != < <= > >= contains is_null not_null. Filters are AND-only, up to 5.

- There is no raw SQL.

- Results are bounded:

- Candidates: at most 20. Tasks: at most 50. Rows: at most 50.

- Cells are clipped to 200 characters.

- Each response is capped at about 24 KB and drops trailing page entries with truncatedForSize and correct returned/remaining counts.

- Totals and remaining counts are always reported.

- Answers are explicit instead of guessed:

- search_entities returns resolution (exact / single / ambiguous / multiple / none) with candidate IDs.

- get_* returns found: false or ambiguous plus candidates.

- Errors return success: false with a message.

- Pull-request state comes from the cached TaskGitHubSnapshot and is labeled as possibly stale. There is no live network access.

- codex_app_server.rs:

- Companion thread/start sends dynamicTools.

- item/tool/call is answered in the owning process on a worker thread. It is never stored as a pending request, so no approval card or "needs recovery" notice appears.

- Developer instructions now describe how to use the lookup tools and how to handle ambiguous, no-match, and error cases.

- companion.rs: Codex fixes tools at thread/start, and thread/resume cannot add them. A toolset_version column therefore marks older sessions, which are rolled once to a new thread with the tools on the next open.

- Frontend:

- Shipyard lookups render as Shipyard lookup · <tool> activity with a search icon and a "Read-only" label, and no controls.

- The per-turn prompt and the tray empty copy were updated.

The companion still cannot mutate data, run workflows, approve requests, or access files, the network, raw SQL, or the DOM. The permission profile stays :read-only with approval policy never.

## Protocol verification

I probed codex-cli 0.144.6 app-server directly through the TFY provider:

- thread/start dynamicTools namespaces are accepted.

- The model issues item/tool/call with {threadId, turnId, callId, namespace, tool, arguments}.

- The reply is {contentItems:[{type:"inputText"}], success}.

- Items stream as dynamicToolCall.

- Tools persist across thread/resume for threads created with them.

## Test plan

- [x] cargo test --lib: 322 passed. This includes 11 new companion_tools tests:

- Allowlist and schema declaration.

- Thread, namespace, and tool rejection.

- Argument validation.

- Ambiguous, no-match, and exact search.

- Task, project, release, and template lookups.

- Table listing, paging, ordering, and filtering.

- SQL-injection attempts via values, identifiers, and orderBy.

- LIKE escaping.

- Cell and response size caps.

- Legacy-session migration.

- [x] pnpm test:companion (with a new read-only lookup activity render test), test:messages, test:chat, test:recovery, and test:smoke

- [x] tsc --noEmit, pnpm build, and pnpm theme:check. No color tokens changed.

- [x] Ran the tools against a copy of a real Shipyard database:

- The Pomodoro task resolves with its status.

- Shipyard project tasks are listed.

- Tasks with open PRs are returned.

- CompanionSession returns its columns and row count.

- A nonsense query reports "No match found".

- [ ] Desktop smoke run of a live companion lookup turn. This was not run here.

#2112 — fix(q118): remove unsupported Redshift temp schema @sanketghia  approved

## Summary

- Remove pg_temp from the Q118 Core procedure SECURITY DEFINER search path because Redshift rejects it as a nonexistent schema at call time.

- Add a regression contract preventing the unsupported search-path entry.

## Verification

- Full runner suite: 232 passed.

- Pinned Ruff and format checks passed.

- Corrected procedure DDL applied to Production.

- Core-only recovery succeeded for June source run 748ad378-0c92-4252-9bd7-dafae187d163.

- June recovery now has 419 raw rows, 18 Core rows, and one Core receipt.

- No new QuickBooks pull or raw publication was performed during recovery.

#2110 — feat(q118): route BalanceSheet landing separately @sanketghia  approved

## Summary

- Add Q118_REPORT_LANDING_PREFIX for the Q118 BalanceSheet landing namespace.

- Keep scheduled Q106 AgedReceivableDetail runs on the existing REPORT_LANDING_PREFIX.

- Route Q118 payloads and manifests by report_kind=balance_sheet without changing the Q106 path.

- Add routing, environment, and pipeline-contract regression coverage.

## Verification

- Full runner suite: 232 passed.

- Pinned CI Ruff and format checks passed locally.

- Live Q106 landing-only validation: 68/68 complete under the Q106 prefix.

- Live Q118 landing-only validation for 2026-06-30 and 2026-07-31: 68/68 complete under the Q118 prefix for both dates.

- Read-only Redshift audit found zero rows for all landing-only test run IDs.

Live evidence is stored in the ignored local evidence directory and is not committed. No deployment or Redshift publication was performed by this branch.

#2109 — fix(q118): make Redshift DDL application compatible @sanketghia  approved

## Summary

- Make the Q118 perimeter compatibility migration safe on Redshift by handling the conditional rename in the bounded DDL helper.

- Use Redshift-compatible INSERT INTO ... WITH ... SELECT ordering for the perimeter seed.

- Document owner-specific DDL application and one-time migration verification.

- Add regression coverage for conditional rename translation, connection propagation, and Redshift SQL ordering.

## Verification

- Full runner suite: 230 passed.

- Focused Ruff and format checks passed.

- Applied and verified the live sequence: 003, 014, 004, and 011.

- Verified 012 was already applied: canonical raw table is nullable, legacy recovery table exists, and canonical data is preserved.

- Live Redshift verification confirmed 59- and 68-company perimeters, Core ownership by sanket.ghia, and the updated procedure references.

No pipeline invocation or new raw/Core publication was performed.

#3841 — fix(budget-load): prepare Q4 refresh and sandbox backups @sanketghia  approved

## Summary

- Store the quarterly loader's 20 table backups in sandbox_finance, prefixing each backup name with its source schema.

- Update backup discovery, restore, archive/load guards, and cleanup for the new location; restore remains compatible with legacy backups in the source schemas.

- Set the income-statement refresh call to 2026-Q4 / 2026-10-01.

- Keep the trailing call in sp_update_consolidated_budgets.sql aligned to Q4 and document that executing the full SQL file also runs the load. The preflight-gated step3_load_budgets.py remains the preferred load path.

- Document the installed HC procedure's pinned workforce_run_id contract and the Q4 EDU adjustment scope.

## Validation

- Ruff 0.15.22 format and check passed for changed Python files.

- git diff --check passed.

- The backup dry-run planned all 20 tables in sandbox_finance; the actual backup created and row-count-verified all 20 (15,923,982 rows).

- The Q4 load completed: Step 5 reported 16 passed, 0 failures, and 1 HC forecast warning. The separate workforce HC refresh was not run because it requires a pinned Q4 workforce_run_id.

- Pyright on income_statement.py reported 30 errors and 33 warnings in the router; no diagnostic was reported on the changed lines. No pytest suite was run.

#2107 — fix(q118): accept QuickBooks no-data metadata rows @sanketghia  approved

## Summary

- Accept QuickBooks BalanceSheet responses with NoReportData=true, Account-only columns, and metadata-only rows as valid empty observations.

- Continue failing closed when a no-data response exposes Total together with rows.

- Add a regression fixture/test and document the source contract.

## Verification

- uv run pytest -q: 225 passed

- Focused Ruff and format checks passed for changed Python files.

- Replayed a retained live QuickBooks payload locally through the parser.

- Live landing-only captures for 2026-06-30 and 2026-07-31 completed 68/68 with zero failures.

- Read-only Redshift audit found zero rows for both landing-only test run IDs.

No deployment, DDL application, raw publication, or Core refresh was performed. Live-run evidence is stored in the ignored local evidence directory and is not committed.

#2073 — feat(q118): add deferred-revenue BalanceSheet ingestion @sanketghia  changes requested

## Summary

- Add Q118 QuickBooks BalanceSheet ingestion through immutable landing, raw staging, ledger lineage, and Core publication.

- Support the configured company perimeter through the governed QuickBooks deferred-revenue reference table while preserving historical compatibility.

- Preserve unnumbered accounts and blank source totals without synthesizing zeroes.

- Add source-controlled read-only verification and local-only evidence handling.

## Business Value

- Makes historical deferred-revenue balances reproducible from the QuickBooks BalanceSheet source and immutable landing evidence.

- Prevents incomplete company coverage, mismatched lineage, unsafe replay, and malformed publication results from appearing as successful Core data.

- Separates historical evidence from current-account comparison and preserves source rows outside the approved Core account perimeter for investigation.

## Manual Effort Estimate

- Not separately tracked in the source handoff; this PR includes implementation, review remediation, local source validation, full-path testing, and verification artifacts.

## Tracking

- Linear ticket: not supplied in the source handoff.

## Validation

- Q118 runner tests: 223 passed.

- Exact CI-pinned Ruff checks and formatting pass.

- Live source validation: 18/18 requests for the nine newly discovered companies passed.

- Fresh 68-company local runs for both 2026-06-30 and 2026-07-31 completed with raw/Core/ledger/landing reconciliation.

## Deployment boundary

No production deployment or schedule change is included. The production deployment and any Q118 on-demand execution remain separate Prod deployment steps.

#1613 — fix(praxis): use corrected private source authentication (AI-874) @caina-barbosa  approvedmercy-allow-critical

## Summary

This PR is a repair within [AI-867 — Praxis Hosted v0](https://linear.app/builder-team/issue/AI-867/vision-agent-hosted-v0-run-author-defined-pr-verification-on-demand).

It updates Aerie's immutable Praxis action pin to the private-source authentication repair tracked by [AI-874 — Prove the complete hosted flow and cleanup](https://linear.app/builder-team/issue/AI-874/vision-agent-hosted-7-prove-the-complete-hosted-flow-and-cleanup).

Production effect: compatibility hardening.

---

## Why

The first live Praxis run accepted the tag and testing contract, then stopped before investigation because Git could not download the exact private PR commit with bearer authentication. The new Praxis commit uses Git smart HTTP basic authentication with x-access-token and has green exact-head Praxis CI.

---

## Business Value

- Lets Praxis reach hosted setup after accepting an eligible Aerie tag and contract.

- Preserves exact-commit verification for private Aerie pull requests.

- Removes an operational failure before any investigator work or product verdict.

---

## How does it work

1. .github/workflows/praxis.yml invokes the Praxis action at commit 604cc6438be951c862b6f611ae0c95a191dd2902.

2. That Praxis commit formats the GitHub App installation token for Git smart HTTP as basic authentication with x-access-token.

3. An eligible Praxis run can download the exact private pull request head instead of reporting exact head is unavailable during source staging.

4. The workflow's trigger, permissions, preflight, AWS role, and terminal-result behavior remain unchanged.

5. This PR does not trigger another live product investigation or retag PR #1609.

---

## Scope

### Included in this phase

- Update the immutable Praxis action pin to the reviewed authentication repair.

- Exact final diff paths:

.github/workflows/praxis.yml

### Deliberately excluded for later phases

- Praxis runtime implementation changes — already reviewed and green at the pinned commit.

- A live Praxis rerun — requires a separate explicit tag after this integration repair is merged.

- Aerie application, schema, API, agent, migration, deployment, and AWS changes — untouched.

---

## Test plan

### Automated validation

- Praxis preflight tests — 2/2 passed (node --experimental-strip-types --test .github/scripts/praxis-preflight.test.ts)

- git diff --check — passed

- exact-head diff scope — one changed line in .github/workflows/praxis.yml

- Praxis exact-head CI for 604cc6438be951c862b6f611ae0c95a191dd2902 — passed

- Aerie hosted CI — all eight checks passed on this PR head

- Mercy review — no blocking issues; auto-approval withheld because the PR changes .github/workflows/praxis.yml

### Time for Implementation

An engineer working without AI assistance would take about half a day to confirm the failure, verify the repaired Praxis commit, update the immutable pin, validate the workflow boundary, and prepare this PR.

#1602 — feat(praxis): connect Aerie PRs to Praxis verification (AI-924) @caina-barbosa  approvedmercy-allow-critical

## Title

feat(praxis): connect Aerie PRs to Praxis verification (AI-924)

## Summary

Praxis is an agent that checks whether an Aerie PR delivers the user experience its author promises. It prepares a temporary Aerie test environment, tries the relevant parts of the app, collects evidence, and reports whether the feature works.

This PR connects Aerie to that system. It adds the PR testing-contract template, the @praxis-e2e comment trigger, and the Aerie helpers Praxis needs to sign in and prepare test data. The PR description supplies the default testing contract; an explicit ## Testing contract section in the same triggering comment can override it from that heading onward. It is tracked by [AI-924 — Integrate Praxis into Aerie](https://linear.app/builder-team/issue/AI-924), part of [AI-867 — Praxis Hosted v0](https://linear.app/builder-team/issue/AI-867).

Production effect: controlled rollout or migration. Merging makes the integration available. An authorized user must tag Praxis to request an investigation; merging alone does not start one.

## Why

Code can pass review and automated tests while the feature still fails when someone uses it. Praxis checks the author's promise by using the running app and showing the evidence behind its conclusion. For example, if a PR promises a CSV export of the current table, Praxis can change the table's filters, download the CSV, and compare the two. This PR supplies the Aerie connection needed to run those investigations from GitHub.

## Business Value

- Authors describe what users should be able to do, giving reviewers a clear promise to evaluate.

- Reviewers can use screenshots, downloaded files, and other evidence to see whether that promise was delivered.

- Praxis can prepare data and sign in to a temporary test environment without changing production data.

- Authors request an investigation from their PR and can follow its progress in the GitHub Actions job.

## How does it work

1. The author fills in the testing contract in the PR description: what the change delivers, who uses it, the conditions needed, and expected behavior. The PR description is the default contract. A [completed example from Aerie PR #1233](https://github.com/AI-Builder-Team/Praxis/blob/master/docs/praxis-testing-contract-example.md) shows how to describe an Enrolments CSV export.

2. A user with write, maintain, or admin access comments with @praxis-e2e, the mention for the Praxis E2E GitHub App. A casual trigger comment uses the PR description. If that same comment contains an explicit ## Testing contract section, Praxis uses that section from the heading onward for this request. It does not search older comments, and the result records which source it used. .github/scripts/praxis-preflight.ts checks the user's permission and confirms the PR is open, targets Aerie main, and has a valid current commit.

3. .github/workflows/praxis.yml calls the reviewed Praxis action after those checks pass. The action validates the selected contract, then runs the investigation when accepted. The GitHub job stays open until Praxis publishes a final result.

4. In the temporary test environment, chat/scripts/preview.ts sets up the test user, default roles, and initial Portfolio and Admissions data. The investigator can then prepare any additional data needed for the PR.

5. The Preview login route lets the investigator select a configured test user without receiving that user's password. The Preview-only Convex helpers report which environment is in use, check the setup, and allow a temporary test record to be created and removed.

6. A new commit tells Praxis to stop an investigation of the previous commit. It does not start another investigation automatically; the author tags Praxis again when ready.

The GitHub runner uses workflow code from Aerie main; it never checks out or runs the PR's code. PR code runs later in the isolated investigation environment. The workflow uses temporary AWS credentials, and test-user passwords stay on the server. No credentials are added to GitHub secrets.

## Scope

### Included in this phase

- A PR testing-contract template with a link to a completed example.

- The comment-triggered GitHub workflow and checks on who may start it.

- Cancellation handling when the PR receives a new commit.

- Preview setup, test-user selection, and checks that the test environment is ready.

- Focused tests for those changes.

- Exact final diff paths:

.github/pull_request_template.md

.github/scripts/praxis-preflight.test.ts

.github/scripts/praxis-preflight.ts

.github/workflows/praxis.yml

chat/app/api/auth/preview/bootstrap/__tests__/route.node.test.ts

chat/app/api/auth/preview/bootstrap/route.ts

chat/convex/praxis/readiness.ts

chat/convex/praxis/schema.ts

chat/convex/praxisReadiness.test.ts

chat/convex/schema.ts

chat/scripts/preview.node.test.ts

chat/scripts/preview.ts

### Deliberately excluded for later phases

- The first full hosted investigation of a real Aerie feature. That pilot and its evidence review are tracked separately in AI-925.

- Automatically investigating every new PR. Hosted v0 requires a comment tagging Praxis.

- Changes to the Praxis agent itself or its AWS resources; those are maintained in the Praxis repository.

- Changes to Aerie's normal product screens, production login behavior, or production data.

## Test plan

### Automated validation

- All current Aerie CI jobs passed on head 0e16352b9cbda64980d69941c27202febaf95bac: lint and boundaries, typecheck, tests, both application builds, both Docker builds, secret scanning, and the review workflow. [Fresh CI run](https://github.com/AI-Builder-Team/Aerie/actions/runs/36767425062)

- The latest blocking review repeats a Node-runner premise that its immediately preceding review had already withdrawn. The evidence and classification are recorded in the final section below.

- .github/workflows/praxis.yml pins the Praxis action to immutable runtime commit b65a57137b72ce290321a48b0c5fc0adb53f6683.

- Five focused Aerie test files passed: 27 tests covering Preview login, readiness, setup, and test-user preparation.

- Trigger authorization tests passed: 2 tests (node --test .github/scripts/praxis-preflight.test.ts).

- Chat typecheck (pnpm typecheck), Biome for changed TypeScript files, Convex path checks, and git diff --check passed.

- The Praxis integration-package tests passed: 11 tests. Its 12 proposed files match this PR.

- A full hosted Praxis investigation has not yet been run. These results validate the integration code, not a completed product pilot.

Reproduce the focused Aerie tests from the repository root:

pnpm --dir chat exec vitest run \

app/api/auth/preview/bootstrap/__tests__/route.node.test.ts \

convex/praxisReadiness.test.ts \

scripts/preview.node.test.ts \

scripts/preview-auth-user.node.test.ts \

convex/users/previewAuth.test.ts

### Time for Implementation

Estimated 3–5 working days without AI assistance for an engineer familiar with Aerie, including implementation, tests, integration, and review preparation.

## Testing contract

### What this PR delivers

An Aerie contributor can request a Praxis investigation from a PR. Aerie passes the request to Praxis and provides the test-environment setup and login support it needs. Reviewers can follow the job until Praxis reports its result.

### Who uses it and where

- Aerie contributors fill in the PR description as the default contract and request an investigation through a PR comment. An explicit ## Testing contract section in that same triggering comment overrides the contract from that heading onward; a casual trigger comment continues to use the PR description, and older comments are not searched.

- Reviewers follow the GitHub Actions job and read Praxis's result on the PR.

- The Praxis investigator uses Aerie's temporary Preview environment to sign in, prepare data, and try the feature.

### Conditions needed

- An open Aerie PR targeting main, with a completed testing contract.

- A contributor with repository write access and another user without that access.

- The Praxis action and its hosted configuration available to Aerie.

- Configured Preview test users, including a second identity to check user selection.

- A fresh Preview environment where setup can create the initial data.

### Expected behavior and examples

- An authorized comment containing @praxis-e2e hands the current PR commit to Praxis. Without an explicit ## Testing contract section in that same comment, Praxis uses the PR description. With that section, it uses the comment from the heading onward and records that source. Older comments are not searched. A comment containing @praxis-e2e-test does not.

- A user without repository write access cannot start an investigation. The workflow stops before obtaining AWS credentials.

- The GitHub job remains open until Praxis reports its final result.

- Pushing a new commit stops work on the previous commit without starting another investigation.

- Preview setup creates the configured test user, roles, and initial data, then reports facts about that same Preview. Setup does not report success when required data is missing.

- Selecting a configured test identity signs in as that user. Unknown identities and caller-supplied passwords are rejected.

- Readiness checks can create and remove a temporary record in the assigned Preview. Their results contain no credentials, and the helpers reject use outside the Preview environment.

### Limits and unanswered questions

This PR connects Aerie to the existing Praxis system. It does not change how the investigator decides what to test or whether the product feature passes. The first complete hosted product investigation is still pending.

## Review repairs and contract clarifications

The blocking finding in [review 5371366899](https://github.com/AI-Builder-Team/Aerie/pull/1602#pullrequestreview-5371366899) says the Praxis preflight command will fail because the GitHub runner does not provide a Node version that supports --experimental-strip-types. That premise is false on the reviewed head 0e16352b9cbda64980d69941c27202febaf95bac:

1. The workflow pins runs-on: ubuntu-24.04.

2. GitHub's [Ubuntu 24.04 runner definition](https://github.com/actions/runner-images/blob/14d8569222caf7662f18b6875bf518683db9ff58/images/ubuntu/toolsets/toolset-2404.json) sets Node 22 as the default. GitHub completed that default-version rollout in May 2026, including Ubuntu 24.04, as recorded in [runner-images issue 14029](https://github.com/actions/runner-images/issues/14029).

3. Node documents --experimental-strip-types as available from Node 22.6.0 in its [Node 22 command-line reference](https://github.com/nodejs/node/blob/v22.17.0/doc/api/cli.md#--experimental-strip-types).

4. The workflow therefore reaches the TypeScript preflight script with a runtime that accepts the flag. The claimed failure before authorization is not reachable under the runner version this workflow selects.

5. The [immediately preceding review of the same commit](https://github.com/AI-Builder-Team/Aerie/pull/1602#pullrequestreview-5371270763) had already withdrawn this concern for the same reason.

No code or test change is warranted for this blocker. Adding actions/setup-node would repeat the runtime guarantee already provided by the selected runner without strengthening the authorization boundary. The other findings remain visible as non-blocking readiness, coverage, and deployment-configuration follow-ups; this clarification does not dismiss them.

Validation on the reviewed head is unchanged: all nine reported GitHub checks pass, including the full test job; the five focused Aerie test files pass 27 tests; and the trigger-authorization suite passes 2 tests. This description-only clarification does not activate Praxis, write product data, deploy code, or run a migration.

#1611 — Refresh DSS contract fingerprint @caina-barbosa  approved

## Summary

- refresh Aerie’s pinned SHA-256 for the canonical Data Source Skill contract

- keep the DSS response shape and behavior unchanged

## Verification

Fetched https://data-source-skills.vercel.app/contract three times. Each response was HTTP 200 with version c5ee63b; its X-Doc-SHA256 matched an independently computed SHA-256 of the 2,836-byte response body:

a3c77832675f31f92bed0c3b728e76ac276deb716c86f1ca70c9e054fb21c76d

## Validation

- DSS HTTP and route-manifest tests — 19/19 passed

- combined focused suite — 62/62 passed

- pre-commit full chat TypeScript check passed

- Biome and git diff --check passed

## Release context

This must merge to main before the Sep 30 production release PR is refreshed, so Aerie does not advertise a stale canonical DSS contract fingerprint.

#1610 — Fix Forecast V2 public response projection @caina-barbosa  approved

## Summary

- explicitly project the established public Session 1 fields instead of spreading the evolving internal Forecast object

- keep the public v2 contract unchanged while preventing new internal calculation fields from triggering response-schema 500s

- exercise the endpoint with a production-shaped row containing all 15 newer internal Session 1 fields

## Validation

- pnpm exec vitest run convex/publicApi/v2/admissions.test.ts convex/publicApi/dss/http.test.ts convex/publicApi/routeManifest.test.ts — 62/62 passed

- Admissions API suite — 43/43 passed

- pre-commit full chat TypeScript check passed

- Biome and git diff --check passed

## Release context

This must merge to main before the Sep 30 production release PR is refreshed. It fixes the exact-head smoke finding where production-shaped Forecast rows could fail the closed public response schema with operation_response_schema_mismatch.

#1608 — Present End-of-Year additions and deductions @vvp-trilogy  approved

## Summary

- carry published End-of-Year withdrawal, transfer, departure, net-movement, status, and reason fields through contracts, sync, Convex storage, and readers

- keep legacy stored rows schema-readable while gating incomplete generations from dashboard and API readers

- group the Previous Year cross-check into Roster, Additions, Deductions, and Result with authoritative subtotals and clean unavailable messaging

- identify carried-forward arrivals as preceding End-of-Year net additions

## Validation

- pnpm typecheck

- pnpm lint:test-architecture

- pnpm exec biome check <changed files>

- pnpm --dir packages/contracts exec vitest run src/admissions-forecast-v2.test.ts src/admissions-forecast-v2-physical.test.ts --maxWorkers=1 (60 passed)

- pnpm --dir sync exec vitest run src/redshift/admissions-forecast.test.ts src/analytics/admissions-forecast-refresh.test.ts --maxWorkers=1 (86 passed)

- pnpm --dir chat exec vitest run --project edge convex/admissions/forecastV2.test.ts --maxWorkers=1 (85 passed)

- pnpm --dir chat exec vitest run --project edge convex/publicApi/v2/admissions.test.ts --maxWorkers=1 (43 passed)

- pnpm --dir chat exec vitest run --project browser components/dashboards/admissions/forecast/v2/__tests__/forecast-v2-report.test.tsx --maxWorkers=1 (42 passed)

Built on the End-of-Year mart fields merged in #1605.

Closes #1606

#2106 — fix(q75): verify established service grants and current catalog @marcusdAIy  approved

## Q75 verifier: accept established service-user ACL without widening the view

The deployed mart_education.aerie_enrolled_student_multi_campus_exceptions view has the same ten non-admin direct Surtr_Service_User grants as the established aerie_parent_child_link_exceptions view. The strict Q75 catalog verifier rejected these existing warehouse-standard grants, even though the detector itself returns 38 contacts with zero direct-source parity differences.

- Permit only the complete ten named Surtr_Service_User grants, with identity type user and admin_option=false; reject missing/duplicate/extra/other grants and NULL-shaped catalog entries.

- Add explicit --verify-current catalog-only read mode for aged-out Data API receipts. Require the approved DDL digest, approval reference, fixed account/target; never submit DDL or claim historical receipt verification.

- Leave the immutable DDL statements, receipt-bound --verify-only, and --apply authorization unchanged.

Validation: 25 focused tests pass; ruff and diff checks pass. The rebased exact source ran --verify-current read-only against finance_dw as CQL_download_OM and passed. Canonical DDL digest remains cb390d2565463f2b740ac87e31ae3d6c7ea87735598c357fa6303cc1549e0672. Original September 22 receipt is retained but no longer describable by Data API. No DDL, warehouse mutation, Mart refresh, or release PR was run.

This is the SURTR-1457 technical detector-verification follow-up. It does not adjudicate the separate Q75 enrollment policy/remediation question.

#1607 — Add DSS discovery to Aerie Tools and API Docs @YibinLongTrilogy  approved

## Summary

Make Aerie’s existing Data Source Skill (DSS) discoverable to teammates who need to point an agent at it. Add a dedicated Tools page and a compact API Docs entry for the public DSS endpoint, without changing the endpoint or its access rules.

### Screenshots

<img width="2301" height="883" alt="Screenshot 2026-09-30 at 12 41 44 PM" src="https://github.com/user-attachments/assets/765c5f82-eefb-46aa-92ea-ac4122d28216" />

### Changes

- chat/app/(main)/tools/dss/page.tsx and dss-shell.tsx *(new)* — Add a dedicated DSS page with the standard Tools layout and breadcrumb.

- chat/components/tools/dss-access-section.tsx *(new)* — Show the DSS URL, a copyable agent instruction, and guidance for accessing protected records with a scoped key.

- chat/components/shell/tools-context-panel.tsx — Add DSS below API Keys in the Tools sidebar for users with api.use.

- chat/app/(main)/tools/api-docs/page.tsx — Add a single GET /dss reference above V2 Contract Discovery, using the existing endpoint-row style.

- chat/components/shell/__tests__/tools-context-panel.test.tsx — Verify DSS placement, active state, and api.use visibility.

### Design Decisions

- Keep the full how-to on Tools → DSS. API Docs uses a concise endpoint reference suited to that page; the two surfaces do not link to each other.

- Derive the DSS URL from the configured API base URL, so production and local backing URLs use the same page logic.

- Reuse the existing copyable code block component for the URL and agent instruction.

## Business value

New team members can find the DSS from Tools without already knowing its URL, and API Docs readers can discover the endpoint where they expect API references.

## Estimated manual effort

1–2 hours.

## Test Plan

- [x] Focused Tools navigation test: 10 passed.

- [x] Chat typecheck passed.

- [x] Biome check on all changed files and git diff --check passed.

- [x] Pre-commit Biome and Chat typecheck hooks passed.

- [ ] In the app, confirm the DSS sidebar item appears below API Keys for an api.use user and opens /tools/dss.

- [ ] Confirm the copy buttons and DSS link work, and API Docs shows a single GET /dss row above V2 Contract Discovery.

#1605 — Net historical departures in End-of-Year forecast @vvp-trilogy  approved

## Summary

- count aligned prior-year SIS withdrawals and transfers beside gross expected arrivals

- publish gross departure operands, subtotal, and net movement through live and locked forecast contracts

- subtract expected departures once from End-of-Year additions and document the corrected formula

## Validation

- dbt test --select int_admissions_milestone_metric_observations --resource-type unit_test (10 passed)

- dbt test --select int_admissions_forecast --resource-type unit_test (2 passed)

- dbt test --select int_admissions_end_of_year_forecast_history_coherent_departures (passed)

- targeted End-of-Year reconciliation/status data tests (3 passed)

- PR-schema build of forecast, history, resolved, and mart models (5 passed)

- warehouse comparison: Alpha Austin 16 arrivals, 12 withdrawals, 1 transfer, 13 departures, +3 net; Alpha Atlanta and other zero-history programs preserve zero

Closes #1604

#211 — Settle stalled workflow nodes by progress, not heartbeat (SINDRI-516) @marcusdAIy  approved

## Summary

Stall detection built into the existing lease lifecycle, as agreed in Munawar's review of Sindri #205 (SINDRI-516). It replaces the idea of a blanket per-node cutoff.

- Progress is tracked separately from liveness. Activations get a lastProgressAt field: the last agent turn or tool call. lastHeartbeatAt still only means the process is alive. Claim sets lastProgressAt, and each progress post advances it, including posts past the 1000-event cap.

- Dropped progress posts still count. The runner records the time of every progress event it emits, whether or not the post arrives. It sends that time on each heartbeat. The server only moves the stored value forward, and never past its own clock. The runner now heartbeats at least every five minutes, where before it waited until half the lease had run out (about 30 minutes).

- Stalled nodes are settled. The existing five-minute recovery cron also fails running activations with no progress for 30 minutes (WORKFLOW_RUNNER_STALL_MS). The limit applies to the gap between steps, not to total node time, so healthy long nodes and long single model or tool calls are unaffected. A stalled node fails rather than re-queueing, because recovery has no attempt cap and a node that stalls once tends to stall again.

- The stale task is stopped. Stalled and lease-expired ECS activations now get an ECS StopTask through the existing stopEcsTasks action. Before, lease expiry only queued recovery.

- Late results are fenced. A settled activation is no longer running, so the existing lease check rejects late completion, failure, progress, and heartbeat calls from that attempt. The new tests cover this.

There is no general execution ceiling. No CD006-protected files are touched.

## Test plan

- [x] New convex/__tests__/workflowStalledNodes.test.ts:

- A node 40 minutes into healthy work is left alone.

- One long call just inside the window is allowed.

- A node whose heartbeat keeps renewing with no progress is failed and its task is stopped, with no recovery attempt created.

- A runner-reported progress time rescues a node whose progress posts were dropped.

- A reported progress time can't move backwards or into the future.

- Progress past the event cap still counts.

- Late completion, failure, progress, and heartbeat calls are rejected after both stall settlement and lease-expiry recovery, and lease-expiry recovery stops the task.

- [x] workflowHttp.test.ts: the heartbeat endpoint forwards a numeric lastProgressAt and ignores anything else.

- [x] agent-runner: full suite (199 passed), including the five-minute heartbeat cap and a heartbeat carrying progress from a dropped post.

- [x] Root suite: no new failures. Two tests fail locally both here and on main: the OpenAPI artifact sync check and the Forge schema-shape check.

- [x] pnpm typecheck, pnpm typecheck:runner, Biome on changed files.

#1603 — Capacity: bind the delegated Sindri identity and validate run ownership (AERIE-2582) @marcusdAIy  approved

## Summary

Binds the delegated Sindri identity and validates what Sindri returns before Aerie records a result (AERIE-2582: Yibin #9 on Aerie #1439). This completes the ticket.

- One capacity user. Capacity automation now runs as a single Aerie user: the one in CAPACITY_AUTOMATION_ACTOR_EMAIL (already the user that files record-mode DD requests). That setting is now required, and CAPACITY_SINDRI_ACT_AS_WORKOS_USER_ID is gone.

- Before anything is uploaded, the run looks up that user's Sindri link. If there isn't one yet, it sets it up automatically, the same way Aerie does for anyone on first Sindri use. If setup fails (for example, Sindri is down), the run retries later instead of ending.

- The run then pins that Sindri identity (WorkOS user, Aerie user, org) as sindriIdentity. Every later call for the run acts as the pinned identity, and each dispatch and poll re-checks it. If the link is removed or changes mid-run, the run ends.

- Ownership checks.

- Start: the response's workflowInstanceId and workflow.workflowInstanceId must be the configured instance, or the run is not marked running.

- Poll: once Sindri reports a terminal status, the status response and then the inspect response must report the same run id, the run's workflowInstanceId, and startedBy equal to the pinned WorkOS user. This happens before anything is recorded or any artifact is copied. Sindri's run responses carry no org field; Sindri scopes these reads to the pinned org.

- A run still in progress, or a response with no status, keeps polling as before. The check runs once there is something to record.

- Failures end the run unresolved with a new unresolvedCause: identity. There are no retries, because this is a configuration or routing error.

- Keeping data inside Aerie until the site is cleared (the third part of Yibin #9) was already done in AERIE-2356: capacitySiteAllowed gates at enqueue, before evidence is assembled or sent.

- Not changed:

- Sindri. Its act-as door already pins the org and checks live membership.

Rollout note (AERIE-2499): set CAPACITY_AUTOMATION_ACTOR_EMAIL to the capacity user and drop CAPACITY_SINDRI_ACT_AS_WORKOS_USER_ID. No manual Sindri setup is needed. Sindri's setup script (author-capacity-workflow.mjs) still prints the old variable; ignore that line.

## Test plan

- [x] chat: vitest run convex/capacityAutomation.test.ts convex/capacityAutomation/config.test.ts (190 passed)

- Start: a response naming another instance, a mismatched workflow, or no instance is rejected and never marked running.

- Poll: a finished run reporting another run, instance, or starter, or no owner fields, ends identity without copying artifacts. A failed run with the wrong starter is identity, not sindri. An inspect response for another run is rejected. An in-progress run keeps polling. An owned run is recorded.

- Binding: pins the capacity user's link. Asks for setup when there is no link or only a deactivated one. Refuses an unknown email and a link without an org. Keeps the pinned identity after the configured user changes. Refuses a pinned link that was deactivated, moved org, or relinked. Refuses an inactive run.

- Through the actions: a queued run whose capacity user doesn't exist ends identity with no fetch. A run whose Sindri setup fails is requeued, not ended. A running run whose pinned link was removed ends identity without polling Sindri.

- [x] tsc --noEmit (chat), Biome, and the pre-commit hook.

#2105 — feat(mart-aerie-dbt-publication-refresh): enable the 10-minute schedule (SURTR-1562) @kevalshahtrilogy  approved

## Summary

Turns on the 10-minute schedule for mart-aerie-dbt-publication-refresh: one line in pipeline.json, plus the contract test and README. The runner shipped disabled on purpose until its out-of-band DDL existed; that rollout step is now done.

## Business Value

The admissions pipeline-detail copy and the SIS enrollment copy now stay fresh on their own, every 10 minutes after each dbt build. That's the precondition for turning on Aerie's ADMISSIONS_PIPELINE_READ and SIS_ENROLLMENT_READ shadow gates. Without it, those marts are a one-off snapshot.

## Manual Effort Estimate

About 1 hour: verify the first run, reconcile, flip the flag, update the test. *Proposed by Claude; Keval to confirm or adjust.*

## Testing / evidence

- Prod DDL: 070-072 and 090-092 applied and verified on 2026-09-30.

- On-demand run: manual-a8-first-run-20260930-131812 SUCCEEDED, publishing 3/3 marts (pipeline detail 30,908; SIS rollup 1,741; SIS members 5,299), 0 failed, 0 stale.

- Reconciliation (read-only, reconciliation/*_vs_aerie_sql.sql):

- pipeline detail parity = PASS, 0/0, on the current build marker;

- SIS rollup and members parity = PASS, 0/0;

- rolls_up = PASS, one_publication = PASS.

- Local checks: pytest (157 passed) and ruff 0.15.22 clean.

## Not covered

- Takes effect on the next Surtr release.

- Aerie's gates for these marts are still legacy.

Linear: SURTR-1562

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1600 — Forecast V2: add next-year forecast and calculation details @vvp-trilogy  approved

## Summary

- add the dynamic next-year Start-of-Year forecast column across desktop, mobile, sorting, and coverage-safe totals

- publish and render V3 replacement operands, rate lineage, and End-of-Year provenance in redesigned January/Next Year details

- accept coherent aerie_milestone_v2 and aerie_milestone_v3 snapshots across ingestion and consumers while rejecting mixed generations

- reveal the Next Year tab, including controlled unavailable states and accessible source tracing

## Validation

- pnpm typecheck

- pnpm lint:test-architecture

- pnpm --filter @bran/contracts exec vitest run src/admissions-forecast-v2.test.ts --maxWorkers=1

- pnpm --filter @bran/sync exec vitest run src/analytics/admissions-forecast-refresh.test.ts src/redshift/admissions-forecast.test.ts --maxWorkers=1

- pnpm --dir chat exec vitest run --project edge convex/admissions/forecastV2.test.ts --maxWorkers=1

- pnpm --dir chat exec vitest run --project browser components/dashboards/admissions/forecast/v2/__tests__/forecast-v2-report.test.tsx --maxWorkers=1

Closes #1598

#2098 — feat(aerie-a8): G6 SIS enrollment copy, rollup input + members in one transaction (SURTR-1549) @kevalshahtrilogy  approvedmercy-allow-critical

## Summary

A8 plan unit U19 (SURTR-1549). This PR publishes Surtr copies of Aerie's two SIS enrollment reads, so the SIS report can move onto the Surtr Gateway. Aerie's gate for them is U23.

Aerie reads the rollups and their student members from one dbt relation inside one transaction (querySisEnrollmentSnapshot, sync/src/redshift/sis-enrollment.ts:280-296). That makes both reads one snapshot. Two separately published copies would lose that guarantee, so one procedure publishes both marts in one transaction under one source_run_id.

| Mart | Gateway slug (U04) | Aerie SQL, verbatim at e366e27d0 | Rows today |

|---|---|---|---|

| mart_education.aerie_sis_enrollment_rollup_input | aerie-sis-enrollment-rollup-input | rollupsSql, l.72-81: GROUP BY program, name, year, cohort, x_pipeline | 1,741 |

| mart_education.aerie_sis_enrollment_member (minors' PII) | aerie-sis-enrollment-member | membersSql, l.198-222, plus mart_row_id | 5,303 |

How the copy works

- Source. Both marts read the production relation sandbox_education.mart_enrollment_dtl, never a DBT_TARGET pr<N>_ build.

- SQL changes. The only substitutions are the two Aerie's own template makes: ${relation}, and ${reportCohortList()}, which becomes SIS_ENROLLMENT_REQUIRED_COHORT_IDS. Neither query has an ORDER BY to drop.

- Procedure. sp_refresh_aerie_sis_enrollment follows U16's pattern:

- it locks the shared mutex;

- it no-ops when both marts already carry the current pg_class_oid build marker from one run;

- it builds both candidates in the CALL transaction;

- it publishes both with DELETE + INSERT in that same transaction.

- Fails closed when:

- the raw columns drift (coupling guard);

- a candidate is empty;

- a row of Aerie's query is lost;

- a mart_row_id is duplicated;

- dbt swaps the build mid-copy;

- the members do not roll up to the rollup input. For each group, the members' distinct student_id count must equal student_count, and every rollup group that counts a student must have members. This is the single-snapshot invariant.

- Runner. CONTRACT_MARTS in src/handler.py lets one procedure publish several marts. The runner verifies each mart, then requires one source_run_id and build marker across them. If that check or the CALL fails, both marts are reported failed.

- Checks that stay in Aerie. Aerie's own checks (cohort coverage, the first-day partition, the members' strict parse) run on the Gateway rows unchanged. A build that fails them is published as-is and fails in Aerie exactly as on the legacy read.

- Keys.

- Rollup: MD5 of the 5 group columns. The GROUP BY makes them unique.

- Member: MD5 of (enrollment_id, cohort_id) plus the occurrence number. That pair is unique today (5,303 of 5,303) but dbt doesn't enforce it.

- PII. The member table has a PII: COMMENT on student_id, full_name, first_name, last_name, email, withdrawal_reason and transfer_reason, and owner-only grants (SELECT revoked too). The rollup input is owner-only too, because it counts minors and some of the counts are small.

- DDL. pipelines/cdk/sql/mart_education/090-092 (U19 owns 090-099). They are applied by this runner's scripts/apply_ddl.py.

- U04 registration. Checked against Surtr/src/seed-gateway-aerie-a8.ts. The slugs, table names and orderBy: mart_row_id match; no fix needed. A new contract test pins this.

Coupling to Vladimir's SIS work (collision MEDIUM)

SIS is Vladimir's active area (AI-Builder-Team/Aerie#1299, AI-Builder-Team/Aerie#1303, AI-Builder-Team/Aerie#1307 and AI-Builder-Team/Aerie#1445). A change to any of these alters what Aerie reads:

- a column in rollupsSql or membersSql;

- the cohort list;

- the has_fact filter;

- the type of the raw cohort_id or x_pipeline.

Any such change needs a lockstep change here:

1. Update the candidate SQL in 092_sp_refresh_aerie_sis_enrollment.sql.

2. Update the pinned SQL in tests/test_sql_contracts_sis_enrollment.py.

3. Migrate the marts.

If this is missed, it fails loudly, not silently:

- A dropped or renamed column breaks the copy's SELECT.

- A type change on a raw column trips the coupling guard. The run goes PARTIAL and the last publication stays.

- A pure SQL change on the Aerie side shows up as differences in the U23 shadow.

New dbt columns that Aerie doesn't read, and changed values, need nothing.

## Business Value

- One G6 read off the legacy path. Aerie's SIS enrollment report is one of the reads A8 moves off the analytics worker's direct Redshift connection and onto governed, lineage-stamped Surtr marts. That is a step towards retiring the EC2 worker's warehouse reads.

- The dashboard stays consistent. The copy keeps the report's key guarantee: the rollup counts and the student drill-down always describe the same dbt build. Otherwise the SIS dashboard could show a count that disagrees with its student list.

- A stable row key. It adds the row key and run id that mart_enrollment_dtl lacks. The Gateway can page it, and the shadow compare can tell a stale copy from a real mismatch.

## Manual Effort Estimate

Proposed: about 10 focused hours. Keval, please confirm or adjust.

- Tracing Aerie's SIS reader, contract and dbt model: 1.5h.

- The two tables and the one-transaction procedure, with the roll-up invariant: 3.5h.

- The runner multi-mart contract and its tests: 1.5h.

- The reconciliation SQL, running it and the negative control: 1.5h.

- SQL contract tests and README: 2h.

## Testing / evidence

- pytest (mart-aerie-dbt-publication-refresh): 142 passed. This includes the new test_sql_contracts_sis_enrollment.py and the TestOneTransactionContract handler tests. mart-aerie-admissions-refresh: 199 passed, unchanged.

- ruff 0.15.22: check and format --check are clean.

- scripts/apply_ddl.py --dry-run exits 0 and lists this runner's files in numbered order:

  -- 070_aerie_dbt_publication_refresh_writer_mutex.sql: 5 statement(s)

-- 071_aerie_admissions_pipeline_detail.sql: 26 statement(s)

-- 072_sp_refresh_aerie_admissions_pipeline_detail.sql: 4 statement(s)

-- 090_aerie_sis_enrollment_rollup_input.sql: 14 statement(s)

-- 091_aerie_sis_enrollment_member.sql: 26 statement(s)

-- 092_sp_refresh_aerie_sis_enrollment.sql: 4 statement(s)

- Read-only reconciliation. reconciliation/aerie_sis_enrollment_vs_aerie_sql.sql Queries 1-4 were run against prod Redshift as CQL_download_OM. They were SELECTs only, and printed counts only. The dbt build was pg_class_oid:20602444.

| Check | Aerie | Candidate | Aerie minus candidate | Candidate minus Aerie | Verdict |

|---|---|---|---|---|---|

| Rollup input (Q1) | 1,741 | 1,741 | 0 | 0 | PASS |

| Members (Q2) | 5,303 | 5,303 | 0 | 0 | PASS |

| Roll-up, one statement (Q3) | Rollup rows | Member rows | Members not in rollups | Rollups not in members | Verdict |

|---|---|---|---|---|---|

| Candidate | 1,741 | 5,303 | 0 | 0 | PASS |

- mart_row_id. The procedure's exact expressions, run as a SELECT, are unique: rollup 1,741 of 1,741, members 5,303 of 5,303.

- Negative control. Dropping the withdraw cohort's members makes the roll-up check report 48 unmatched rollup groups, so the check does catch a member set from a different snapshot.

- Not yet run. Queries 5-7 compare against the published marts and need the DDL applied and one run. No DDL was applied and nothing was deployed.

## Keval steps

1. PII sign-off (minors) before the two marts are exposed over the Gateway (A8 plan §9 D2). The member mart holds students' names, emails, SIS ids, enrollment history and free-text reasons.

2. Apply the DDL (090-092). apply_ddl.py applies all of this runner's files, which is idempotent:

REDSHIFT_CLUSTER_IDENTIFIER=redshift-cluster-1 REDSHIFT_DATABASE=finance_dw REDSHIFT_DB_USER=CQL_download_OM uv run python scripts/apply_ddl.py

If the dbt publication schedule (#2096) has been enabled by the time this merges, apply the DDL before merging. Otherwise every run calls a missing procedure and goes PARTIAL.

3. Run on demand with {"procedures": ["mart_education.sp_refresh_aerie_sis_enrollment"]}. Then run reconciliation Queries 5-7: expect parity = PASS twice and one_publication = PASS.

## Stack note

Depends on #2096 (U16, SURTR-1547), which adds the mart-aerie-dbt-publication-refresh runner this unit extends. The PR's base is feat/a8-u16-dbt-publication-runner. If #2096 merges first, this will be rebased onto main and retargeted.

## Not covered

- The Aerie gate SIS_ENROLLMENT_READ, its shadow compare and the dry-run. That is U23.

- The Gateway key mint and seed run. Those belong to U04.

- The unresolved upstream writer of the SIS model's staging_education_ai_horizons.raw_* inputs (A8 plan §8).

- dbt_invocation_id as the build marker (§9 D4). The OID fallback is used; switching is a one-line change in the procedure.

- Enabling the dbt publication schedule. That is #2096's follow-up.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2096 — feat(aerie-a8): mart-aerie-dbt-publication-refresh runner + admissions pipeline detail copy + tenant crosswalk (SURTR-1547) @kevalshahtrilogy  approvedmercy-allow-critical

## Summary

This is A8 plan unit U16 (SURTR-1547). It adds the third A8 runner, for copies of Aerie's dbt marts, and both reads behind Aerie's Admissions Pipeline report. That lets Aerie's queryAdmissionsPipelineRows and queryAdmissionsPipelineCrosswalk move onto the Surtr Gateway; the Aerie gate is U20.

- New runner pipelines/runners/mart-aerie-dbt-publication-refresh (Lambda, bundling: true, src/requirements.txt).

- Schedule: cron(0/10 * * * ? *). Aerie's dbt build is not a Surtr pipeline, so there is no success event to trigger on. Instead, each procedure detects a new build itself.

- The schedule ships disabled (Mercy round 1). Its objects are out-of-band DDL, so it is enabled in a one-line follow-up once the DDL is applied and an on-demand run is verified.

- What it runs: it CALLs each procedure in REFRESH_PROCEDURES with (run_id, force), then checks the mart read-only: non-empty, unique mart_row_id, one source_run_id and one source_build_marker. The run_id must be the platform UUID before it is inlined. U19 (SIS) and U24 (Forecast V2) will append their procedures here.

- Result per mart: published, unchanged (the build was already copied, which is most runs) or failed.

- Failure handling: the same as U03. A failed procedure makes the run partial_failure, and the run fails if every procedure fails or a commit outcome is unknown.

- Freshness: a copy whose dbt build is older than SOURCE_MAX_AGE_MINUTES (180) makes the run partial_failure. The copy is still exact, but dbt has stopped publishing (PIPELINE §5.4).

- 071/072 aerie_admissions_pipeline_detail (slug aerie-admissions-pipeline-detail, PII) is a style-C dbt publication copy.

- The candidate is Aerie's SQL. It is copied verbatim from admissions-pipeline.ts:154-173 (Aerie e366e27d0, unchanged since 92fd47992), bound to the production relation sandbox_education.mart_admissions_pipeline_dtl. The one other edit drops the trailing ORDER BY: a table has no row order, the Gateway pages by mart_row_id, and U20 re-sorts.

- The build marker. source_build_marker = pg_class_oid:<oid>. dbt's table materialization swaps in a new relation on every build, so the OID changes. When the published marker is current, the procedure no-ops unless p_force.

- Switching to dbt_invocation_id (plan §9 D4) changes only the one v_source_build_marker := assignment.

- Lineage: source_run_id is this copy's own run; source_published_at is the relation's creation time (pg_class_info.relcreationtime).

- Coupling guard (the lockstep rule). Before copying, the procedure compares the dbt relation's column types, by OID, with the mart's pinned types. Only the six ::text columns are exempt, and varchar width is ignored. On any drift it raises and keeps the previous publication.

- A mid-copy dbt swap is detected by re-reading the OID after the copy. In that case the procedure publishes nothing, and the next tick copies the new build.

- Other guards: it fails closed on a missing relation, an empty candidate, a candidate count different from the source's, or a duplicate mart_row_id. mart_row_id = MD5 of MD5(pipeline_key) and its occurrence number, ordered by every other column.

- PII: SELECT is revoked as well as writes, and the 11 PII columns carry PII: comments.

- 073/074 aerie_admissions_pipeline_tenant_crosswalk (slug aerie-admissions-pipeline-tenant-crosswalk).

- Aerie's SQL is copied verbatim from admissions-pipeline.ts:201-208.

- It is EduCRM-backed, so it is appended to U03's mart-aerie-admissions-refresh REFRESH_PROCEDURES, and it uses the observed sales-educrm-mart-sync provenance for mart_pipeline_dtl.

- Aerie's 1:1-per-tenant assertion stays in Aerie.

- DDL: pipelines/cdk/sql/mart_education/070-074; U16 owns 070-079. 070-072 are applied by the new runner's scripts/apply_ddl.py, and 073-074 by U03's.

- The U03 DDL test now requires every aerie_admissions file to be applied by exactly one of the two runners.

- Both apply_ddl.py scripts now have no default target (Mercy round 1). They refuse to send a statement unless REDSHIFT_CLUSTER_IDENTIFIER, REDSHIFT_DATABASE and REDSHIFT_DB_USER are all set.

- README: it documents the dbt-copy PIPELINE §13 exception (WAREHOUSE §2.2 and §2.5, and PIPELINE §4 for the external dbt writer; §7 is met through the explicit build marker), the coupling rule and the lockstep steps, the marker, and PII.

- Gateway registration (U04, #2082) checked: both slugs map to exactly these table names, ordered by mart_row_id. No change was needed.

## Business Value

- The G2 Admissions Pipeline report can leave the EC2 analytics worker. It is Aerie's widest PII read, at 30,929 rows and 53 columns. Aerie can read it through the Surtr Gateway with its unchanged row mapper. That is the SURTR-735 quarterly commitment, and a step toward tearing the worker down.

- dbt stays with Vladimir, with no fork. The copy follows each hourly dbt build within 10 minutes and carries lineage to the exact build. A dbt column change that would break Aerie's parity now fails loudly in Surtr, instead of silently drifting.

- U19 (SIS enrollment) and U24 (Forecast V2) reuse this runner, adding only a procedure and one REFRESH_PROCEDURES entry.

## Manual Effort Estimate

About 14 focused hours (roughly 2 days) to build by hand without AI. That covers:

- reading the Aerie reader, the dbt materialization and the catalog to design the build marker, the swap guard and the coupling guard;

- two procedures, the runner, and its freshness reporting;

- tests, reconciliation, and the read-only probes.

Keval: please confirm or adjust.

## Testing / evidence

- uv run pytest: 91 passed (new runner) and 105 passed (mart-aerie-admissions-refresh, including 14 new crosswalk contract tests and the explicit-target apply_ddl tests).

- The SQL contracts pin both Aerie queries. They assert:

- each candidate is exactly that SQL plus the allowed edits;

- the reconciliation uses the same candidate;

- the guards come before the DELETE;

- the coupling guard exempts only the ::text and lineage columns;

- the marker is a single assignment.

- Ruff 0.15.22: ruff check pipelines and ruff format --check pipelines are clean.

- CDK: real-pipeline-configs.test.ts passed (590). The app also synthesized with Docker bundling skipped (CDK_CONTEXT_JSON aws:cdk:bundling-stacks=[]). Pipeline-mart-aerie-dbt-publication-refresh-prod contains:

- one Lambda (handler.handler, python3.11, 900 s);

- one Step Functions state machine;

- the rule pipeline-mart-aerie-dbt-publication-refresh-schedule-prod cron(0/10 * * * ? *) ENABLED;

- 4 alarms.

- Read-only reconciliation was run with psql as CQL_download_OM, SELECT only, reading counts only:

| mart | aerie rows | candidate rows | aerie − candidate | candidate − aerie | parity |

|---|---|---|---|---|---|

| aerie_admissions_pipeline_detail | 30,929 | 30,929 | 0 | 0 | PASS |

| aerie_admissions_pipeline_tenant_crosswalk | 57 | 57 | 0 | 0 | PASS |

- The candidate as the mart stores it (every column CAST to the mart's declared type) is also EXCEPT 0/0 against Aerie's SQL: 30,929 and 57 rows. So the INSERT changes no value.

- Negative controls on the detail compare: dropping a row gives 1 / 0, and duplicating a row gives 1 / 1.

- The procedures' exact mart_row_id expressions give 30,929 and 57 distinct values. pipeline_key is unique and non-null (30,929).

- The source today: sandbox_education.mart_admissions_pipeline_dtl is a table (relkind r) with OID 20602348, created 2026-09-29 11:39:17 UTC (the 11:30 dbt build), owned by vladimir.pikalov. It is readable by CQL_download_OM.

- The coupling guard's pinned columns are 39 character varying, 5 boolean and 3 numeric(18,2). The guard's catalog query returns 0 against the source itself.

- EduCRM mart_pipeline_dtl: the observed run ed781ef8… is the latest run (SUCCESS), and rows_loaded 33,585 = snapshot 33,585.

- Not yet run: Query 3 (detail) and Query 2 (crosswalk) compare against the published marts, and need the DDL.

- scripts/apply_ddl.py --dry-run passes. The statements are the committed SQL files verbatim, in this order:

- new runner: 070_aerie_dbt_publication_refresh_writer_mutex.sql (5 statements), 071_aerie_admissions_pipeline_detail.sql (26), 072_sp_refresh_aerie_admissions_pipeline_detail.sql (4);

- mart-aerie-admissions-refresh: 006-011 unchanged, then 073_aerie_admissions_pipeline_tenant_crosswalk.sql (10) and 074_sp_refresh_aerie_admissions_pipeline_tenant_crosswalk.sql (4).

## Keval steps

1. Crosswalk DDL before merging. A merge reaches production within the hour, and the EduCRM trigger then CALLs the crosswalk every 30 minutes. Until 073-074 exist, that CALL would make each run PARTIAL (amber, throttled); nothing wrong is published. Run:

cd pipelines/runners/mart-aerie-admissions-refresh && REDSHIFT_CLUSTER_IDENTIFIER=redshift-cluster-1 REDSHIFT_DATABASE=finance_dw REDSHIFT_DB_USER=CQL_download_OM uv run python scripts/apply_ddl.py

It applies 006-011 (idempotent) and 073-074.

2. PII sign-off (A8 plan §9 D2) for exposing aerie_admissions_pipeline_detail over the Gateway. It holds parent and child names, emails and phones, and child date of birth and gender.

3. Merge. Mercy withholds auto-approve on pipelines/cdk/ paths. The new runner deploys with its schedule disabled.

4. Apply the detail DDL (070-072):

cd pipelines/runners/mart-aerie-dbt-publication-refresh && REDSHIFT_CLUSTER_IDENTIFIER=redshift-cluster-1 REDSHIFT_DATABASE=finance_dw REDSHIFT_DB_USER=CQL_download_OM uv run python scripts/apply_ddl.py

5. Run mart-aerie-dbt-publication-refresh on demand.

- The first run should report published with about 30,929 rows.

- A second run should report unchanged.

6. Run both reconciliation files. Each comparison against the mart must report parity = PASS. For the detail, first check that the mart's marker equals Query 2's current marker.

7. Enable the schedule: a one-line follow-up PR setting "enabled": true in the new pipeline.json.

8. Optional: ask Vladimir to add {{ invocation_id }} AS dbt_invocation_id to mart_admissions_pipeline_dtl (plan §9 D4). The switch is then one assignment per procedure.

## Not covered

- The Aerie gate and shadow compare (U20), SIS (U19) and Forecast V2 (U24).

- The procedure bodies have not been executed in Redshift, because no DDL was applied. Their candidate SELECTs, the type-projected candidate, the mart_row_id expressions, the catalog queries (source OID, creation time, the coupling guard's shape) and the EduCRM observation were run read-only instead.

- A procedure-only change does not republish by itself. With an unchanged dbt marker, the procedure no-ops. After a lockstep migration or procedure fix, run on demand with force: true, as the README says.

- OID reuse. The marker assumes Redshift does not reuse the relation's OID across builds. OIDs are 32-bit and only wrap after about 4 billion allocations.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2101 — fix(aerie-a8): widen applied parity marts' text columns to the query's width (SURTR-1552) @kevalshahtrilogy  approvedmercy-allow-critical

## Summary

Mercy flagged a bug family on U18 (PR 2099): a parity mart declared a text column narrower than the width its refresh query returns. When that happens, one long source value fails the publication INSERT, and the mart stays stale while Aerie's legacy read moves on.

I audited every A8 parity mart for this. For each mart I compared the declared VARCHAR widths with the candidate query's Redshift result metadata, read with LIMIT 0 and read-only. I also measured today's max OCTET_LENGTH of each column. The two marts already applied in prod have three narrow columns:

| Mart (unit) | Column | Declared | Query returns | Max today |

|---|---|---|---|---|

| aerie_admissions_program (U03) | source_program_id | VARCHAR(256) | VARCHAR(65535) | 11 bytes |

| aerie_admissions_program_projection (U13) | filter_program_id | VARCHAR(32) | VARCHAR(65535) | 11 bytes |

| aerie_admissions_coming_year_projection (U13) | filter_program_code_lower | VARCHAR(256) | VARCHAR(65535) | 32 bytes |

All three are trimmed EduCRM SUPER text. None is a live risk today; the risk is theoretical. The published tables have 0 rows so far.

aerie_admissions_program_directory's school_year_start/end also show as VARCHAR(65535) in the result metadata. They are DATE::varchar, though, so a value is at most 13 bytes. They stay VARCHAR(256), and the test records that bound.

Changes:

- Base DDL: 008, 050 and 052 now create these columns at VARCHAR(65535), so a fresh apply matches.

- Migrations:

- 016_widen_admissions_program_text_columns.sql (U03). 005 is taken and U03's 006–011 are full, so 016 is the first free number that sorts after 008.

- 056_widen_admissions_forecast_input_text_columns.sql (U13).

- scripts/widen_text_columns.py applies the migrations; apply_ddl.py keeps only re-runnable files. For each ALTER the script:

- reads the column from pg_attribute;

- refuses a column that is a sort key, has a default, has a BYTEDICT/RUNLENGTH/TEXT255/TEXT32K encoding, or is not VARCHAR;

- skips a column that is already at least that wide;

- otherwise runs the ALTER as its own Data API statement (Redshift rejects ALTER COLUMN TYPE inside a transaction block) and re-reads the catalog to confirm.

--dry-run reads the catalog and prints the plan without changing anything. There is no default target. The prod columns are all LZO, non-sort-key, with no default and no dependent views, so nothing blocks the ALTER.

- Tests:

- Both U03 and U13 contract suites now assert that no published text column is narrower than the query returns. This is the same pattern as PR 2099, but every published VARCHAR column must be listed.

- test_widen_text_columns.py checks that each migration matches its CREATE TABLE and sorts after it. It also covers the dry run, skip-if-wide, refusals, and the no-default target.

## Business Value

Aerie's Gateway reads of program, program-projection and coming-year-projection data can't go stale because a single upstream value got longer. A stale parity mart would silently serve old numbers while Aerie's legacy read moves on, and that undercuts the A8 cutover off the Aerie EC2 workers (SURTR-735). The same audit covered every open A8 mart PR, so the whole family is fixed before any of them reaches prod.

## Manual Effort Estimate

About 6 hours of focused work without AI (Keval to confirm or adjust). That covers auditing about 20 marts' DDL against their query types, measuring widths, writing the migrations and the guarded apply script with tests, and patching three open PRs.

## Testing

- pytest in mart-aerie-admissions-refresh: 202 passed.

- ruff@0.15.22 check and ruff@0.15.22 format --check: clean.

- Mutation check: reverting 008 to VARCHAR(256) fails both the contract test and the migration/CREATE-agreement test.

- widen_text_columns.py's catalog query ran read-only against prod. It reports the three columns at 256/32/256, lzo, sort-key position 0, no default.

- No DDL was applied anywhere.

## Keval steps

Approve, then apply these two prod DDL files, dry-run first:

1. pipelines/cdk/sql/mart_education/016_widen_admissions_program_text_columns.sql

2. pipelines/cdk/sql/mart_education/056_widen_admissions_forecast_input_text_columns.sql

cd pipelines/runners/mart-aerie-admissions-refresh

export REDSHIFT_CLUSTER_IDENTIFIER=redshift-cluster-1 REDSHIFT_DATABASE=finance_dw REDSHIFT_DB_USER=CQL_download_OM

uv run python scripts/widen_text_columns.py --dry-run # expect 3 "would widen to VARCHAR(65535)"

uv run python scripts/widen_text_columns.py # expect 3 "widened from ... to VARCHAR(65535)"

uv run python scripts/widen_text_columns.py --dry-run # expect 3 "skipped"

The base 008/050/052 edits need no prod apply. CREATE TABLE IF NOT EXISTS is a no-op on the existing tables, and the ALTERs bring those tables to the same width.

## Not covered

- The open A8 PRs are fixed on their own branches: U08 (PR 2090), U12 (PR 2092) and U15 (PR 2095).

- U06 (PR 2089), U07 (PR 2088), U16 (PR 2096) and U19 (PR 2098) needed no change. Their narrower columns are bounded DATE/TIMESTAMP/BIGINT text, or already VARCHAR(65535).

- The README is not updated here, to avoid conflicts with the open A8 PRs that all edit it. The migration headers and the script docstring carry the procedure.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2091 — fix(mart-aerie-xo-contractor-refresh): deterministic latest-week tie-break (SURTR-1542) @kevalshahtrilogy  approved

## Summary

mart_education.sp_refresh_aerie_xo_contractor_package() picks each contractor's latest regular weekly invoice with ROW_NUMBER() OVER (PARTITION BY contractor_id ORDER BY week_start DESC). If a contractor has more than one regular payment in their latest week, that pick is nondeterministic. Aerie's sync/src/redshift/xo-contractor-package.ts has the identical logic, so the AERIE-2489 shadow compare can't catch a divergence between the two.

This PR changes the window to ORDER BY week_start DESC, id DESC, where id is staging_finance_xo.raw_contractor_invoices.id (bigint). Aerie PR 1499 makes the same change on its side.

- ddl/sp_refresh_aerie_xo_contractor_package.sql: adds the tie-break, plus a comment on keeping it in lockstep with Aerie.

- tests/test_sql_contracts.py: updates the existing contract assertion and adds test_latest_week_pick_has_a_fully_deterministic_sort_key. That test pins the single ROW_NUMBER() window to exactly this ORDER BY. The identity mart's LISTAGG tie-break is already pinned the same way.

- README.md: documents the tie-break.

Linear: SURTR-1542 (related: AERIE-2489)

## Business Value

The Aerie-to-Surtr XO contractor migration (F2 package) depends on a shadow compare between Aerie's live query and this mart. That compare is only meaningful if both sides pick the same row every time. Without the tie-break, a contractor paid twice in one week could get a different weekly_base_usd / markup_from_data from run to run. That would do two things:

- Make the compare report false mismatches, or hide real ones, on compensation data used for education cost reporting.

- Trip this procedure's own replay-idempotency guard on a same-run re-invocation. The guard raises "already published with different values".

This change makes the pick reproducible, so the cutover to reading the mart can be trusted.

## Manual Effort Estimate

~1.5 hours of focused work without AI: confirm the key is unique against Redshift, make the one-line SQL change, update and add contract tests, prove before/after equivalence, and write up the PR. Keval: please confirm or adjust this number.

## Testing / evidence

All Redshift work was read-only SELECTs inside BEGIN READ ONLY ... ROLLBACK. Output is counts only, with no contractor names or pay.

- Key uniqueness (staging_finance_xo.raw_contractor_invoices): 152,718 rows, 152,718 distinct id, 0 null id. So id DESC alone is a total tie-break, and the created_at fallback isn't needed.

- Output unchanged: comparing the candidate SELECT before (week_start DESC) and after (week_start DESC, id DESC) gives 3,046 = 3,046 rows. EXCEPT both ways is 0 / 0. 0 contractors have more than one regular payment in their latest week.

- New candidate vs the currently published mart (business columns): 3,046 rows, EXCEPT both ways is 0 / 0.

- Deployed procedure drift check: the pg_proc.prosrc body in prod matches origin/main's file (whitespace-insensitive diff is empty). Applying this file changes only the ORDER BY and regresses nothing.

- uv run pytest: 45 passed.

- ruff 0.15.22 format --check and ruff check on src tests scripts: clean.

- scripts/apply_ddl.py --dry-run (full ddl/, and the procedure file alone): exit 0. The procedure file splits into 4 statements (CREATE OR REPLACE PROCEDURE, ALTER OWNER, REVOKE, GRANT).

## Companion

Companion: AI-Builder-Team/Aerie PR 1499; ship together

## Keval steps

This is a procedure change, so the DDL has to be applied after merge. Nothing has been applied.

1. Merge this together with Aerie PR 1499.

2. Apply the procedure: cd pipelines/runners/mart-aerie-xo-contractor-refresh && uv run python scripts/apply_ddl.py ddl/sp_refresh_aerie_xo_contractor_package.sql. This is CREATE OR REPLACE plus the existing owner and grant statements, with no table changes. The runner code is unchanged, so no redeploy is needed.

3. Optional: on the next xo-contractor-invoices-refresh success, the mart republishes. Output is expected to be identical.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2088 — feat(aerie-a8): mart-aerie-expenses-refresh runner + G4 expense parity marts (SURTR-1539) @kevalshahtrilogy  approvedmercy-allow-critical

## Summary

This is A8 plan unit U07 (SURTR-1539, part of SURTR-735). It adds the runner and the two G4 marts that Aerie's expenses reads will move onto through the Surtr Gateway. It has the same shape as the admissions runner in Surtr PR 2084, so the A8 runners look alike.

- New runner pipelines/runners/mart-aerie-expenses-refresh (Lambda, bundling: true, src/requirements.txt).

- Triggers: on_pipeline_success of mart-education-quickbooks-refresh (about daily) and quickbooks-expense-ai-generation (weekly, Sundays 08:00 UTC), with forward_upstream_execution_context. Both pipeline ids exist, and a contract test checks each one's manifest and the input mart it publishes. The handler checks with describe_execution that the upstream run SUCCEEDED on that pipeline's state machine. On-demand runs are also accepted, and may name a subset of procedures.

- What it runs: it CALLs each procedure in REFRESH_PROCEDURES, then checks the published mart read-only: non-empty, unique mart_row_id, one source_run_id.

- Failure handling: a failed procedure makes the run partial_failure and the others still run. If every procedure fails, or a commit outcome is unknown, the run fails.

- mart_education.aerie_expense_transaction (slug aerie-expense-transaction).

- It is Aerie's queryExpenseTransactions SQL, copied verbatim from expenses.ts:101-118 at Aerie 3d4fe1a97. That is the updated_at read the refresh cycle runs.

- Two template slots are expanded in place: the ${REDSHIFT_TABLES.*} table names, and ${updatedWhere}, which is empty because refresh.ts calls queryExpenseTransactions() with no watermark. The mart is therefore the full snapshot (80,709 rows).

- Key: source_row_id. Lineage: source_run_id / source_published_at = the transactions mart's mart_run_id / mart_refreshed_at, plus a per-row vendor_generation_run_id from the joined classification.

- mart_education.aerie_expense_vendor_classification (slug aerie-expense-vendor-classification).

- It is queryVendorClassifications, copied verbatim from expenses.ts:236-240. It is narrow: the AI SUPER columns (vendor_description, reasoning, web_evidence) and the evidence URIs are not carried. It appends vendor_id (the row key) and generation_run_id.

- Lineage decision: source_run_id = the source's publication_idempotency_key, not generation_run_id. The AI generation carries earlier classifications forward and only classifies new vendors, so generation_run_id varies by row (9 values today) and can't be the single run id the Gateway reader requires. The publication key is the one value every row of a publication shares. It also appears as idempotency_key in the publishing run's output_summary.

- One sp_refresh_ per mart (020-024 in pipelines/cdk/sql/mart_education/).

- Each one locks a shared owner-only writer mutex, then builds the candidate in temp tables.

- It fails closed on an empty candidate, on missing or mixed upstream lineage, and on a duplicate mart_row_id.

- It publishes with DELETE + named-column INSERT in the CALL transaction. There is no TRUNCATE.

- Rows are published as the legacy read returns them. Unmatched vendors keep a NULL category, and any duplicate key reaches Aerie's mapper as it does today.

- The mart_row_id occurrence is ordered by every other published column, so ids are deterministic.

- No metadata mart. queryExpenseMetadata is two DISTINCTs over the transactions table. Aerie derives it from the transaction rows in U11.

- Grain and PII are in the table and column COMMENTs: vendor_name and memo are QuickBooks free text and can name people.

- PIPELINE §13 exception (PIPELINE §7: two independent triggers, not a fan-in barrier) is documented in the README, with owner, risk, controls and follow-on.

- DDL numbering. The block starts at 020, which keeps it clear of PR 2084's 006-011 and of the admissions units stacked on it. A test fails if any of these numbers is reused by another file in the directory.

## Business Value

- It moves Aerie's expense-dashboard reads to Surtr. Today the Aerie EC2 analytics worker queries the two QuickBooks marts directly on every refresh cycle. Moving the worker's reads onto Surtr marts is the SURTR-735 quarterly commitment, and it is what unblocks tearing down the worker.

- Parity is provable. Each mart is Aerie's own SQL with per-row lineage. A test pins that SQL (byte-for-byte equal to Aerie origin/main), and the reconciliation EXCEPT is 0 in both directions.

- It shrinks what the Gateway exposes. The vendor mart drops the AI reasoning and evidence columns, so Aerie's key sees only the two columns it reads.

- It gives the expenses group its own runner. Its QuickBooks triggers and cadence stay separate from the 30-minute EduCRM runner (plan §3.1.4).

## Manual Effort Estimate

About 10 focused hours (roughly 1.5 days) to build by hand without AI. That covers reading Aerie's expense queries and the two upstream marts' publication semantics, the lineage decision for the vendor mart, the two procedures, adapting the runner, the tests, and the reconciliation. The runner shape reuses PR 2084. Keval: please confirm or adjust.

## Testing / evidence

- uv run pytest: 91 passed. This covers the handler, the pipeline contract, the SQL contracts, apply_ddl and the Redshift client.

- The SQL-contract tests pin Aerie's SQL text, template slots included. They assert that each procedure's candidate is exactly that SQL plus the appended lineage columns, and that the reconciliation uses the same candidate.

- A local check confirmed that the pinned text equals Aerie origin/main (3d4fe1a97) byte for byte, for both queries. It also confirmed the tables.ts table names and the no-argument call in refresh.ts.

- Ruff 0.15.22: ruff check and ruff format --check are clean on the pipeline.

- CDK jest: real-pipeline-configs, schema/pipeline-config and schema/owners pass (707 tests) with the new manifest. In the full suite, 905 pass. The 6 local failures are all in pipeline-shared-stack and come from Docker not being available for Lambda bundling on this machine; they are unrelated.

- Read-only reconciliation was run with psql as CQL_download_OM (SELECT only), using query 1 of each reconciliation file (Aerie SQL vs the procedure's candidate):

| Mart | aerie_row_count | candidate_row_count | aerie_minus_candidate | candidate_minus_aerie |

|---|---|---|---|---|

| aerie_expense_transaction | 80,709 | 80,709 | 0 | 0 |

| aerie_expense_vendor_classification | 3,479 | 3,479 | 0 | 0 |

Query 2 (against the published mart) errors with "relation does not exist" as expected, because no DDL has been applied.

- Procedure guards, run read-only as SELECTs against today's data:

- Transactions: 1 mart_run_id, 1 mart_refreshed_at, 0 rows missing lineage, and 80,709 distinct mart_row_ids for 80,709 rows. 7,654 rows have no classification, so category is NULL.

- Vendors: 1 publication key, 1 mart_refreshed_at, 0 rows missing lineage, 3,479 distinct mart_row_ids for 3,479 rows, and 9 generation_run_ids.

- Value widths are well inside the declared columns (for example, source_row_id is at most 14 characters against 256).

- Upstream health, checked read-only in staging_other.pipeline_runs_prod (last 30 days):

- mart-education-quickbooks-refresh: 39 SUCCESS and 4 FAILED. All 4 failures were on 09-01 and 09-02. The latest run is 4171f196, at 09-29 07:26, which is the mart_run_id on every row.

- quickbooks-expense-ai-generation: 5 SUCCESS, and it is weekly (cron(0 8 ? * SUN *)). The latest run is c02fe1f6 on 09-27. Its idempotency_key equals the vendor table's publication_idempotency_key.

- scripts/apply_ddl.py --dry-run applies 40 statements across 5 files: mutex 5, transaction table 15, transaction procedure 4, vendor table 12, vendor procedure 4. The full output is below.

<details><summary><code>uv run python scripts/apply_ddl.py --dry-run</code> (full output)</summary>

-- 020_aerie_expenses_refresh_writer_mutex.sql: 5 statement(s)

-- Owner-only mutex shared by every mart-aerie-expenses-refresh procedure.

-- Each procedure locks it first, so overlapping runs (both upstream triggers

-- can fire close together) publish one at a time. It is locked instead of the

-- target mart, which is locked only for the final DELETE + INSERT.

CREATE TABLE IF NOT EXISTS mart_education.aerie_expenses_refresh_writer_mutex (

lock_scope VARCHAR(64) NOT NULL

)

DISTSTYLE ALL;

COMMENT ON TABLE mart_education.aerie_expenses_refresh_writer_mutex IS

'Owner-only writer mutex for the mart-aerie-expenses-refresh stored procedures. It contains no data and is not a consumer contract.';

REVOKE ALL ON mart_education.aerie_expenses_refresh_writer_mutex FROM PUBLIC;

REVOKE ALL ON mart_education.aerie_expenses_refresh_writer_mutex FROM GROUP team_engineers;

ALTER TABLE mart_education.aerie_expenses_refresh_writer_mutex

OWNER TO "CQL_download_OM";

-- 021_aerie_expense_transaction.sql: 15 statement(s)

-- Canonical DDL for mart_education.aerie_expense_transaction (A8 unit U07,

-- Gateway slug aerie-expense-transaction). Sole writer:

-- mart_education.sp_refresh_aerie_expense_transaction().

--

-- Thin parity mart: columns are exactly the output aliases of Aerie's

-- queryExpenseTransactions SQL (sync/src/analytics/queries/expenses.ts) over

-- mart_education.quickbooks_expense_transactions LEFT JOIN

-- mart_education.quickbooks_vendor_classifications, with source types kept.

CREATE TABLE IF NOT EXISTS mart_education.aerie_expense_transaction (

source_row_id VARCHAR(256),

transaction_id VARCHAR(64),

company_id VARCHAR(128),

txn_date VARCHAR(256),

vendor_name VARCHAR(1024),

line_amount NUMERIC(20, 6),

account_name VARCHAR(1024),

account_id VARCHAR(256),

class_name VARCHAR(1024),

category VARCHAR(128),

memo VARCHAR(4000),

updated_at VARCHAR(256),

vendor_generation_run_id VARCHAR(128),

mart_row_id VARCHAR(32) NOT NULL,

source_run_id VARCHAR(128) NOT NULL,

source_published_at TIMESTAMPTZ NOT NULL,

refreshed_at TIMESTAMP NOT NULL,

created_by VARCHAR(128) NOT NULL,

PRIMARY KEY (mart_row_id)

)

DISTSTYLE AUTO

SORTKEY (mart_row_id);

COMMENT ON TABLE mart_education.aerie_expense_transaction IS

'Purpose: Surtr publication of the rows Aerie''s queryExpenseTransactions reads (mart_education.quickbooks_expense_transactions LEFT JOIN mart_education.quickbooks_vendor_classifications on company_id and vendor_id), so Aerie can read them through the Surtr Gateway with its unchanged row mapper. Grain: one output row of that SQL, normally one accepted QuickBooks Purchase expense line. Key: mart_row_id (MD5 of source_row_id plus its occurrence number); source_row_id (company_id|purchase_id|source_line_index) is expected unique but not enforced, so Aerie sees any duplicate exactly as the legacy read would. Lineage: source_run_id/source_published_at are the transactions mart''s mart_run_id/mart_refreshed_at; vendor_generation_run_id is the joined classification''s generation_run_id. Sensitive data: vendor_name and memo are free text from QuickBooks and can name people (individual payees, staff). Full snapshot replaced atomically by mart_education.sp_refresh_aerie_expense_transaction.';

COMMENT ON COLUMN mart_education.aerie_expense_transaction.source_row_id IS

'company_id || ''|'' || purchase_id || ''|'' || source_line_index, as Aerie builds it; Aerie''s upsert key.';

COMMENT ON COLUMN mart_education.aerie_expense_transaction.txn_date IS

'transaction_date DATE cast to text (YYYY-MM-DD), as Aerie selects it.';

COMMENT ON COLUMN mart_education.aerie_expense_transaction.vendor_name IS

'QuickBooks vendor display name. Can name a person.';

COMMENT ON COLUMN mart_education.aerie_expense_transaction.category IS

'quickbooks_vendor_classifications.classification of the line''s vendor; NULL when the vendor is not classified.';

COMMENT ON COLUMN mart_education.aerie_expense_transaction.memo IS

'QuickBooks line description (free text). Can name people.';

COMMENT ON COLUMN mart_education.aerie_expense_transaction.updated_at IS

'quickbooks_updated_at TIMESTAMP cast to text, as Aerie selects it for its watermark.';

COMMENT ON COLUMN mart_education.aerie_expense_transaction.vendor_generation_run_id IS

'generation_run_id of the joined vendor classification (the quickbooks-expense-ai-generation run, or baseline import, that produced category); NULL when unmatched. Lineage addition; not read by Aerie.';

COMMENT ON COLUMN mart_education.aerie_expense_transaction.mart_row_id IS

'Deterministic row key: MD5 of source_row_id and its occurrence number. Gateway orderBy for total-order paging.';

COMMENT ON COLUMN mart_education.aerie_expense_transaction.source_run_id IS

'quickbooks_expense_transactions.mart_run_id (the mart-education-quickbooks-refresh run); the procedure requires exactly one value per snapshot.';

COMMENT ON COLUMN mart_education.aerie_expense_transaction.source_published_at IS

'quickbooks_expense_transactions.mart_refreshed_at (UTC) of that publication.';

ALTER TABLE mart_education.aerie_expense_transaction

OWNER TO "CQL_download_OM";

-- Writer protection only. Reader access is provisioned by the Redshift DBA

-- (PIPELINE §13); the Surtr Gateway reads as the owner.

REVOKE INSERT, UPDATE, DELETE, TRUNCATE

ON mart_education.aerie_expense_transaction FROM PUBLIC;

REVOKE INSERT, UPDATE, DELETE, TRUNCATE

ON mart_education.aerie_expense_transaction FROM GROUP team_engineers;

-- 022_sp_refresh_aerie_expense_transaction.sql: 4 statement(s)

-- Sole writer for mart_education.aerie_expense_transaction (WAREHOUSE §7.1).

--

-- The candidate is Aerie's queryExpenseTransactions SQL (the updated_at read

-- its refresh cycle runs), copied verbatim from

-- sync/src/analytics/queries/expenses.ts:101-118 at Aerie 3d4fe1a97, with

-- three lineage columns appended to the select list. Two template slots are

-- expanded in place: ${REDSHIFT_TABLES.*} with the two table names, and

-- ${updatedWhere} with the empty string, because Aerie's refresh cycle calls

-- queryExpenseTransactions() without a watermark (refresh.ts). Lineage is

-- carried from the upstream Surtr marts: every transactions row carries the

-- one mart-education-quickbooks-refresh run that published it, and every

-- classification row its quickbooks-expense-ai-generation run.

--

-- Fails closed on: an empty candidate, missing or mixed transactions lineage,

-- and a duplicate mart_row_id. The DELETE + INSERT publish stays inside the

-- CALL transaction; never TRUNCATE (it commits implicitly).

CREATE OR REPLACE PROCEDURE mart_education.sp_refresh_aerie_expense_transaction()

AS $$

DECLARE

v_candidate_count BIGINT;

v_lineage_run_count BIGINT;

v_lineage_published_count BIGINT;

v_lineage_invalid_count BIGINT;

v_duplicate_count BIGINT;

v_source_run_id VARCHAR(128);

v_source_published_at TIMESTAMPTZ;

v_refreshed_at TIMESTAMP;

BEGIN

LOCK TABLE mart_education.aerie_expenses_refresh_writer_mutex;

v_refreshed_at := GETDATE();

DROP TABLE IF EXISTS tmp_aerie_expense_transaction_query;

CREATE TEMP TABLE tmp_aerie_expense_transaction_query AS

-- aerie-sql:begin

SELECT

t.company_id || '|' || t.purchase_id || '|' || t.source_line_index::text AS source_row_id,

t.purchase_id AS transaction_id,

t.company_id,

t.transaction_date::text AS txn_date,

t.vendor_name,

t.line_amount,

t.account_name,

t.account_id::text,

t.class_name,

vc.classification AS category,

t.line_description AS memo,

t.quickbooks_updated_at::text AS updated_at,

-- A8 lineage additions (not in Aerie's select list):

vc.generation_run_id AS vendor_generation_run_id,

t.mart_run_id AS transaction_mart_run_id,

t.mart_refreshed_at AS transaction_mart_refreshed_at

FROM mart_education.quickbooks_expense_transactions t

LEFT JOIN mart_education.quickbooks_vendor_classifications vc

ON t.company_id = vc.company_id AND t.vendor_id = vc.vendor_id

ORDER BY t.quickbooks_updated_at DESC NULLS LAST, t.transaction_date DESC

-- aerie-sql:end

;

SELECT COUNT(*) INTO v_candidate_count FROM tmp_aerie_expense_transaction_query;

IF v_candidate_count = 0 THEN

RAISE EXCEPTION 'aerie_expense_transaction: candidate is empty';

END IF;

SELECT COUNT(DISTINCT transaction_mart_run_id),

COUNT(DISTINCT transaction_mart_refreshed_at),

SUM(CASE

WHEN NULLIF(BTRIM(transaction_mart_run_id), '') IS NULL

OR transaction_mart_refreshed_at IS NULL THEN 1

ELSE 0

END),

MIN(transaction_mart_run_id),

MIN(transaction_mart_refreshed_at)

INTO v_lineage_run_count, v_lineage_published_count, v_lineage_invalid_count,

v_source_run_id, v_source_published_at

FROM tmp_aerie_expense_transaction_query;

IF v_lineage_run_count <> 1 OR v_lineage_published_count <> 1 OR v_lineage_invalid_count <> 0 THEN

RAISE EXCEPTION

'aerie_expense_transaction: transactions lineage is mixed or incomplete (% run id(s), % refreshed_at value(s), % row(s) missing lineage)',

v_lineage_run_count, v_lineage_published_count, v_lineage_invalid_count;

END IF;

DROP TABLE IF EXISTS tmp_aerie_expense_transaction;

CREATE TEMP TABLE tmp_aerie_expense_transaction (LIKE mart_education.aerie_expense_transaction);

INSERT INTO tmp_aerie_expense_transaction (

source_row_id,

transaction_id,

company_id,

txn_date,

vendor_name,

line_amount,

account_name,

account_id,

class_name,

category,

memo,

updated_at,

vendor_generation_run_id,

mart_row_id,

source_run_id,

source_published_at,

refreshed_at,

created_by

)

SELECT

q.source_row_id,

q.transaction_id,

q.company_id,

q.txn_date,

q.vendor_name,

q.line_amount,

q.account_name,

q.account_id,

q.class_name,

q.category,

q.memo,

q.updated_at,

q.vendor_generation_run_id,

MD5(

'aerie_expense_transaction|'

|| COALESCE('v' || q.source_row_id, 'n')

|| '|'

|| (ROW_NUMBER() OVER (

PARTITION BY q.source_row_id

ORDER BY q.transaction_id, q.company_id, q.txn_date, q.vendor_name, q.line_amount,

q.account_name, q.account_id, q.class_name, q.category, q.memo, q.updated_at,

q.vendor_generation_run_id

))::VARCHAR

),

v_source_run_id,

v_source_published_at,

v_refreshed_at,

'mart-aerie-expenses-refresh/v1'

FROM tmp_aerie_expense_transaction_query q;

SELECT COUNT(*) INTO v_duplicate_count

FROM (

SELECT mart_row_id

FROM tmp_aerie_expense_transaction

GROUP BY mart_row_id

HAVING COUNT(*) > 1

) duplicates;

IF v_duplicate_count <> 0 THEN

RAISE EXCEPTION

'aerie_expense_transaction: candidate has % duplicate mart_row_id value(s)',

v_duplicate_count;

END IF;

LOCK TABLE mart_education.aerie_expense_transaction;

DELETE FROM mart_education.aerie_expense_transaction;

INSERT INTO mart_education.aerie_expense_transaction (

source_row_id,

transaction_id,

company_id,

txn_date,

vendor_name,

line_amount,

account_name,

account_id,

class_name,

category,

memo,

updated_at,

vendor_generation_run_id,

mart_row_id,

source_run_id,

source_published_at,

refreshed_at,

created_by

)

SELECT

source_row_id,

transaction_id,

company_id,

txn_date,

vendor_name,

line_amount,

account_name,

account_id,

class_name,

category,

memo,

updated_at,

vendor_generation_run_id,

mart_row_id,

source_run_id,

source_published_at,

refreshed_at,

created_by

FROM tmp_aerie_expense_transaction;

IF (SELECT COUNT(*) FROM mart_education.aerie_expense_transaction) <> v_candidate_count THEN

RAISE EXCEPTION 'aerie_expense_transaction: post-publication row count mismatch';

END IF;

RAISE INFO 'aerie_expense_transaction: published % row(s) from QuickBooks mart run %',

v_candidate_count, v_source_run_id;

DROP TABLE tmp_aerie_expense_transaction;

DROP TABLE tmp_aerie_expense_transaction_query;

END;

$$ LANGUAGE plpgsql SECURITY INVOKER;

ALTER PROCEDURE mart_education.sp_refresh_aerie_expense_transaction()

OWNER TO "CQL_download_OM";

REVOKE ALL ON PROCEDURE mart_education.sp_refresh_aerie_expense_transaction()

FROM PUBLIC;

GRANT EXECUTE ON PROCEDURE mart_education.sp_refresh_aerie_expense_transaction()

TO "CQL_download_OM";

-- 023_aerie_expense_vendor_classification.sql: 12 statement(s)

-- Canonical DDL for mart_education.aerie_expense_vendor_classification (A8

-- unit U07, Gateway slug aerie-expense-vendor-classification). Sole writer:

-- mart_education.sp_refresh_aerie_expense_vendor_classification().

--

-- Thin, narrow parity mart: columns are exactly the output aliases of Aerie's

-- queryVendorClassifications SQL (sync/src/analytics/queries/expenses.ts) over

-- mart_education.quickbooks_vendor_classifications, with source types kept.

-- The source's AI SUPER columns (vendor_description, reasoning, web_evidence)

-- and evidence URIs are deliberately not carried.

CREATE TABLE IF NOT EXISTS mart_education.aerie_expense_vendor_classification (

vendor_name VARCHAR(255),

ai_category VARCHAR(128),

vendor_id VARCHAR(64),

generation_run_id VARCHAR(128),

mart_row_id VARCHAR(32) NOT NULL,

source_run_id VARCHAR(128) NOT NULL,

source_published_at TIMESTAMPTZ NOT NULL,

refreshed_at TIMESTAMP NOT NULL,

created_by VARCHAR(128) NOT NULL,

PRIMARY KEY (mart_row_id)

)

DISTSTYLE ALL

SORTKEY (mart_row_id);

COMMENT ON TABLE mart_education.aerie_expense_vendor_classification IS

'Purpose: Surtr publication of the rows Aerie''s queryVendorClassifications reads (mart_education.quickbooks_vendor_classifications for company alpha with a classification), so Aerie can read them through the Surtr Gateway with its unchanged row mapper. Grain: one output row of that SQL, one classified QuickBooks vendor. Key: mart_row_id (MD5 of vendor_id plus its occurrence number); vendor_id is unique in the source, vendor_name is expected unique but not enforced, so Aerie''s name-keyed upsert sees any collision exactly as the legacy read would. Lineage: source_run_id/source_published_at are the classification publication''s publication_idempotency_key/mart_refreshed_at; generation_run_id is the run that classified each vendor. Sensitive data: vendor_name is a QuickBooks payee name and can name a person. Full snapshot replaced atomically by mart_education.sp_refresh_aerie_expense_vendor_classification.';

COMMENT ON COLUMN mart_education.aerie_expense_vendor_classification.vendor_name IS

'QuickBooks vendor display name. Can name a person.';

COMMENT ON COLUMN mart_education.aerie_expense_vendor_classification.ai_category IS

'quickbooks_vendor_classifications.classification.';

COMMENT ON COLUMN mart_education.aerie_expense_vendor_classification.vendor_id IS

'Canonical QuickBooks vendor ID (company alpha). Lineage addition and row key; not read by Aerie.';

COMMENT ON COLUMN mart_education.aerie_expense_vendor_classification.generation_run_id IS

'quickbooks-expense-ai-generation run (or baseline import) that classified this vendor. Varies by row because earlier classifications are carried forward. Lineage addition; not read by Aerie.';

COMMENT ON COLUMN mart_education.aerie_expense_vendor_classification.mart_row_id IS

'Deterministic row key: MD5 of vendor_id and its occurrence number. Gateway orderBy for total-order paging.';

COMMENT ON COLUMN mart_education.aerie_expense_vendor_classification.source_run_id IS

'quickbooks_vendor_classifications.publication_idempotency_key: the one identity every row of a classification publication shares; it is the idempotency_key in the publishing quickbooks-expense-ai-generation run''s output_summary. The procedure requires exactly one value per snapshot.';

COMMENT ON COLUMN mart_education.aerie_expense_vendor_classification.source_published_at IS

'quickbooks_vendor_classifications.mart_refreshed_at (UTC) of that publication.';

ALTER TABLE mart_education.aerie_expense_vendor_classification

OWNER TO "CQL_download_OM";

-- Writer protection only. Reader access is provisioned by the Redshift DBA

-- (PIPELINE §13); the Surtr Gateway reads as the owner.

REVOKE INSERT, UPDATE, DELETE, TRUNCATE

ON mart_education.aerie_expense_vendor_classification FROM PUBLIC;

REVOKE INSERT, UPDATE, DELETE, TRUNCATE

ON mart_education.aerie_expense_vendor_classification FROM GROUP team_engineers;

-- 024_sp_refresh_aerie_expense_vendor_classification.sql: 4 statement(s)

-- Sole writer for mart_education.aerie_expense_vendor_classification

-- (WAREHOUSE §7.1).

--

-- The candidate is Aerie's queryVendorClassifications SQL, copied verbatim

-- from sync/src/analytics/queries/expenses.ts:236-240 at Aerie 3d4fe1a97, with

-- four lineage columns appended to the select list and ${REDSHIFT_TABLES.*}

-- expanded to the table name. Lineage is carried from the upstream Surtr mart

-- (quickbooks-expense-ai-generation), which replaces the whole classification

-- state in one publication and stamps every row with that publication's

-- publication_idempotency_key and mart_refreshed_at.

--

-- Fails closed on: an empty candidate, missing or mixed upstream publication

-- lineage, and a duplicate mart_row_id. The DELETE + INSERT publish stays

-- inside the CALL transaction; never TRUNCATE (it commits implicitly).

CREATE OR REPLACE PROCEDURE mart_education.sp_refresh_aerie_expense_vendor_classification()

AS $$

DECLARE

v_candidate_count BIGINT;

v_lineage_run_count BIGINT;

v_lineage_published_count BIGINT;

v_lineage_invalid_count BIGINT;

v_duplicate_count BIGINT;

v_source_run_id VARCHAR(128);

v_source_published_at TIMESTAMPTZ;

v_refreshed_at TIMESTAMP;

BEGIN

LOCK TABLE mart_education.aerie_expenses_refresh_writer_mutex;

v_refreshed_at := GETDATE();

DROP TABLE IF EXISTS tmp_aerie_expense_vendor_classification_query;

CREATE TEMP TABLE tmp_aerie_expense_vendor_classification_query AS

-- aerie-sql:begin

SELECT vendor_name, classification AS ai_category,

-- A8 lineage additions (not in Aerie's select list):

vendor_id,

generation_run_id,

publication_idempotency_key AS vendor_publication_key,

mart_refreshed_at AS vendor_mart_refreshed_at

FROM mart_education.quickbooks_vendor_classifications

WHERE company_id = 'alpha'

AND classification IS NOT NULL

ORDER BY vendor_name

-- aerie-sql:end

;

SELECT COUNT(*) INTO v_candidate_count FROM tmp_aerie_expense_vendor_classification_query;

IF v_candidate_count = 0 THEN

RAISE EXCEPTION 'aerie_expense_vendor_classification: candidate is empty';

END IF;

SELECT COUNT(DISTINCT vendor_publication_key),

COUNT(DISTINCT vendor_mart_refreshed_at),

SUM(CASE

WHEN NULLIF(BTRIM(vendor_publication_key), '') IS NULL

OR vendor_mart_refreshed_at IS NULL THEN 1

ELSE 0

END),

MIN(vendor_publication_key),

MIN(vendor_mart_refreshed_at)

INTO v_lineage_run_count, v_lineage_published_count, v_lineage_invalid_count,

v_source_run_id, v_source_published_at

FROM tmp_aerie_expense_vendor_classification_query;

IF v_lineage_run_count <> 1 OR v_lineage_published_count <> 1 OR v_lineage_invalid_count <> 0 THEN

RAISE EXCEPTION

'aerie_expense_vendor_classification: classification lineage is mixed or incomplete (% publication key(s), % refreshed_at value(s), % row(s) missing lineage)',

v_lineage_run_count, v_lineage_published_count, v_lineage_invalid_count;

END IF;

DROP TABLE IF EXISTS tmp_aerie_expense_vendor_classification;

CREATE TEMP TABLE tmp_aerie_expense_vendor_classification (LIKE mart_education.aerie_expense_vendor_classification);

INSERT INTO tmp_aerie_expense_vendor_classification (

vendor_name,

ai_category,

vendor_id,

generation_run_id,

mart_row_id,

source_run_id,

source_published_at,

refreshed_at,

created_by

)

SELECT

q.vendor_name,

q.ai_category,

q.vendor_id,

q.generation_run_id,

MD5(

'aerie_expense_vendor_classification|'

|| COALESCE('v' || q.vendor_id, 'n')

|| '|'

|| (ROW_NUMBER() OVER (

PARTITION BY q.vendor_id

ORDER BY q.vendor_name, q.ai_category, q.generation_run_id

))::VARCHAR

),

v_source_run_id,

v_source_published_at,

v_refreshed_at,

'mart-aerie-expenses-refresh/v1'

FROM tmp_aerie_expense_vendor_classification_query q;

SELECT COUNT(*) INTO v_duplicate_count

FROM (

SELECT mart_row_id

FROM tmp_aerie_expense_vendor_classification

GROUP BY mart_row_id

HAVING COUNT(*) > 1

) duplicates;

IF v_duplicate_count <> 0 THEN

RAISE EXCEPTION

'aerie_expense_vendor_classification: candidate has % duplicate mart_row_id value(s)',

v_duplicate_count;

END IF;

LOCK TABLE mart_education.aerie_expense_vendor_classification;

DELETE FROM mart_education.aerie_expense_vendor_classification;

INSERT INTO mart_education.aerie_expense_vendor_classification (

vendor_name,

ai_category,

vendor_id,

generation_run_id,

mart_row_id,

source_run_id,

source_published_at,

refreshed_at,

created_by

)

SELECT

vendor_name,

ai_category,

vendor_id,

generation_run_id,

mart_row_id,

source_run_id,

source_published_at,

refreshed_at,

created_by

FROM tmp_aerie_expense_vendor_classification;

IF (SELECT COUNT(*) FROM mart_education.aerie_expense_vendor_classification) <> v_candidate_count THEN

RAISE EXCEPTION 'aerie_expense_vendor_classification: post-publication row count mismatch';

END IF;

RAISE INFO 'aerie_expense_vendor_classification: published % row(s) from classification publication %',

v_candidate_count, v_source_run_id;

DROP TABLE tmp_aerie_expense_vendor_classification;

DROP TABLE tmp_aerie_expense_vendor_classification_query;

END;

$$ LANGUAGE plpgsql SECURITY INVOKER;

ALTER PROCEDURE mart_education.sp_refresh_aerie_expense_vendor_classification()

OWNER TO "CQL_download_OM";

REVOKE ALL ON PROCEDURE mart_education.sp_refresh_aerie_expense_vendor_classification()

FROM PUBLIC;

GRANT EXECUTE ON PROCEDURE mart_education.sp_refresh_aerie_expense_vendor_classification()

TO "CQL_download_OM";

</details>

## Keval steps

1. Apply the DDL to prod before merging. A merge reaches production within the hour, and the QuickBooks trigger then fires daily. Run cd pipelines/runners/mart-aerie-expenses-refresh && uv run python scripts/apply_ddl.py. It runs as CQL_download_OM and applies pipelines/cdk/sql/mart_education/020-024 in order: mutex, then each table before its procedure. Applying those files through the usual numbered out-of-band SQL deploy is equivalent.

2. Merge. Once released, run mart-aerie-expenses-refresh on demand. Expect status: success, with results[0].rows = the transactions count (80,709 today) and results[1].rows = 3,479.

3. Run the reconciliation files. Query 2 must show aerie_minus_mart = mart_minus_aerie = 0 for both marts.

4. Reader access (optional). If anyone other than the owner needs to read the new objects, that is a DBA grant.

## Not covered

- Other A8 units: the Aerie EXPENSES_READ gate (U11) and the key mint. The Gateway registration for both slugs already landed in U04 (SURTR-1533).

- Freshness limits for U11. The inputs publish about daily (transactions) and weekly (vendors). The U02 reader's default source_published_at max age of 6 hours would reject both, so U11 needs per-source limits of at least about 36 hours and 8 days.

- The procedure bodies have not been executed in Redshift, because no DDL was applied. Their candidate SELECTs, lineage guards and mart_row_id stamping were run read-only instead.

- Aerie's fallback read is not copied. It is the txn_date read used only when the updated_at read throws (expenses.ts:156-172), and it returns the same columns minus updated_at. The same goes for the backfill-expense-transactions script's filtered read. The gateway path serves the refresh cycle's unfiltered read only.

- Behaviour inherited from Aerie: a vendor that first appears in transactions between weekly AI generation runs has a NULL category until the next run, exactly as in the legacy join.

- Procedure failures are PARTIAL (amber, throttled), not paging, per the plan. Only a failure of every procedure pages.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2099 — feat(aerie-a8): G3 marketing parity marts, D3 shadow-day events + D4 weekly deposits (SURTR-1550) @kevalshahtrilogy  approvedmercy-allow-critical

## Summary

A8 plan unit U18 (SURTR-1550). Adds the last two G3 marketing parity marts to the mart-aerie-admissions-refresh runner (U03, PR 2084), so Aerie's per-program shadow-day and weekly-deposit reads can move onto the Surtr Gateway. Sibling of U15 (PR 2095, D1/D2).

| Mart | Gateway slug | Aerie query (Aerie e366e27d0) | Rows 09-29 |

|---|---|---|---|

| mart_education.aerie_admissions_shadow_day_event | aerie-admissions-shadow-day-event | D3 queryShadowEventAgg, educrm.ts:1699-1716 | 75 (37 programs) |

| mart_education.aerie_admissions_weekly_deposit | aerie-admissions-weekly-deposit | D4 queryDepositAgg, educrm.ts:1760-1769 | 56,700 (90 programs) |

Line numbers were re-verified against Aerie origin/main (educrm.ts last changed in cd9494749); they are the SQL text inside each query(...) call.

What the candidates change in Aerie's SQL. Only the lifted key (A8 plan §3.1 rule 1):

- TRIM(BOTH '"' FROM program_name::varchar) = $1 is each query's only filter. It becomes the filter_program_name column, so the WHERE goes away.

- The key leads the GROUP BY, which ran after the filter:

- D3: (TRIM(program_name), TRIM(event_name));

- D4: (TRIM(program_name), week_label, week_start_date, week_end_date, school_year).

- Neither query filters on has_fact; neither does the mart. Aerie's ORDER BY is kept.

- Output types are kept: week_label stays SUPER, the week dates DATE, school_year INTEGER, counts BIGINT.

- tests/test_sql_contracts_marketing_d3_d4.py rebuilds each candidate from Aerie's pinned SQL by those edits alone, so any other difference fails.

Writers (081, 083, appended to REFRESH_PROCEDURES contiguously after U13's):

- Observed EduCRM provenance, as in U03/U13/U15: the latest sales-educrm-mart-sync run must also be the latest started run, its rows_loaded must equal the snapshot, and it must be under 24 hours old.

- Shared writer mutex first, then candidate in temp tables, then the target lock and an atomic DELETE + named-column INSERT in the CALL transaction. No TRUNCATE.

- Fails closed on an empty candidate, a duplicate mart_row_id, and a missing, superseded, stale or count-mismatched observation.

- NULL program keys can't be selected by = $1. They're counted in the run log and not published (0 today).

- mart_row_id: MD5 of the group key, which is unique by construction (text parts and the serialized week_label hashed, so each part has a fixed shape), plus the occurrence number.

No strict-parse assertion, on purpose. Both queries use safeParseRows, which drops a bad row instead of failing the program (unlike D1/D2's strict parse). Refusing to publish would diverge from legacy, so rows are published as-is, like U13's app conversion:

- D3's only droppable row is a NULL event_name group (0 today); Aerie drops it itself.

- No D4 value can fail its schema: every field goes through String() or z.coerce.number() on a COUNT.

The tests derive both facts from Aerie's pinned row schemas.

Gateway registration (U04). Checked, and correct as registered: slugs, table names, orderBy: mart_row_id, no dateColumn, and descriptions naming shadow_event_dtl / deposit_dtl. No change; a test pins the slug-to-table mapping.

## Business Value

- Completes the Surtr side of G3 marketing (with U15). This carries on the retirement of Aerie's EC2 analytics worker, Benji's Q3 priority 1 (A8, SURTR-735).

- Removes about 180 per-program EduCRM queries per hourly cycle. Today D3 and D4 run once per program (90 programs × 2). After U22's gate, Aerie reads each mart once per cycle, with Surtr lineage on every row.

- Feeds public Aerie tables (insertShadowDayEvents, insertWeeklyDeposits). Parity is proven, not asserted: every program reconciles EXCEPT = 0 both ways against Aerie's own SQL.

## Manual Effort Estimate

About 7 hours of focused work by hand. That covers:

- reading D3/D4, their Zod schemas and parse modes;

- two DDL files and two provenance-pinned procedures;

- the contract tests;

- profiling the sources (types, widths, NULL keys, PII);

- the per-program reconciliation against Redshift.

_Proposed by Claude. Keval, please confirm or adjust._

## Testing / evidence

- uv run pytest: 243 passed, 60 of them new in test_sql_contracts_marketing_d3_d4.py. Five mutations are caught:

- swapping the GROUP BY order;

- COUNT(deal_id) for COUNT(DISTINCT ...);

- dropping the NULL-key filter;

- a non-SUPER week_label;

- a wrong reconciliation predicate.

- ruff check and ruff format --check (0.15.22): clean.

- uv run python scripts/apply_ddl.py --dry-run: files in order 006-011, 050-055, then 080 (13 statements), 081 (4), 082 (13), 083 (4).

- Read-only reconciliation (psql, SELECT only, counts only). This is Query 1 of each reconciliation/*_vs_aerie_sql.sql: Aerie SQL vs the candidate, per program, multiset EXCEPT both ways.

| Mart | Programs | Aerie rows | Candidate rows | aerie − candidate | candidate − aerie |

|---|---|---|---|---|---|

| D3 shadow_day_event (3 busiest programs) | 3 | 7 | 7 | 0 | 0 |

| D3 shadow_day_event (every program) | 37 | 75 | 75 | 0 | 0 |

| D4 weekly_deposit (3 busiest programs) | 3 | 1,890 | 1,890 | 0 | 0 |

| D4 weekly_deposit (every program) | 90 | 56,700 | 56,700 | 0 | 0 |

- Simulated publication (read-only). This runs each procedure's candidate and its tmp INSERT ... SELECT exactly as written, cast to the mart's column types:

| Mart | Rows | Row diff vs candidate | Duplicate mart_row_id | NULL-key groups |

|---|---|---|---|---|

| D3 | 75 | 0 / 0 | 0 | 0 |

| D4 | 56,700 | 0 / 0 | 0 | 0 |

Each simulation took about 6 s, well inside the 480 s statement timeout.

- Provenance today: v_aerie_educrm_observed_publication returns one observation per source. It is the latest run (SUCCESS, 13:06 UTC), and rows_loaded equals the snapshot: 2,620 for shadow_event_dtl, 58,254 for deposit_dtl.

- Types and widths. Output types come from Redshift result metadata (node-pg, LIMIT 0). Every text column is at least as wide as the query returns it, so no value the legacy read accepts can fail the publication:

- the trimmed SUPER text (event_name, custom_event_type, filter_program_name) is VARCHAR(65535);

- shadow_date is VARCHAR(64), holding its VARCHAR(52) cast.

A test pins this. Largest values today: event name 70 bytes, category 10, timestamp text 19, program key 32.

- PII check (counts only):

- Neither query selects a name, email or date of birth. Student and deal ids appear only inside COUNT(DISTINCT ...).

- I checked whether shadow-day event names embed student names. 3 of 2,620 rows contain a 3-letter first name. Each is on an event name shared by 8 students, so the match is coincidental.

## Keval steps

1. Apply the DDL before merging. A merge reaches production within the hour, and the trigger then fires every 30 minutes. Until the DDL exists, the two new CALLs fail and make the run PARTIAL. Run uv run python scripts/apply_ddl.py. Its list is 006-011, 050-055 and 080-083; the earlier files re-apply idempotently.

2. After release, run the pipeline on demand. Then run Query 2 of both reconciliation files for a few programs; it must return 0 both ways.

3. No PII sign-off needed. Both marts are aggregate counts, and their DDL is writer-protected only (reader grants stay with the DBA). Say if you'd rather gate the Gateway exposure anyway.

## Not covered

- The Aerie reader, shadow compare and gate for D1–D4 (ADMISSIONS_MARKETING_READ) are U22. Notes for U22:

- Re-apply the legacy ORDER BY per program. D3's shadow_date is the text of MIN(milestone_date) and sorts the same.

- week_label is SUPER. Aerie stores String(value) as pg returns it, so the type-parity layer must match that.

- Lifted-key matching caveat, as in U15. Redshift compares VARCHAR ignoring trailing blanks. There are no such keys today: 37 and 90 keys, all byte-distinct.

- There's no behavioral rollback test for the DELETE + INSERT path. There's no Redshift in CI; this is the same deferral as U03/U13/U15.

- REFRESH_PROCEDURES, apply_ddl.py's list, test_pipeline_contract.py and the README each conflict on one line with U15 (PR 2095). Whichever merges second rebases.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2095 — feat(aerie-a8): G3 marketing parity marts, D1 event aggregates + D2 event contacts (SURTR-1546) @kevalshahtrilogy  approvedmercy-allow-critical

## Summary

A8 plan unit U15 (SURTR-1546). Adds the first two G3 marketing parity marts to the mart-aerie-admissions-refresh runner (U03, PR 2084), so Aerie's per-program marketing event reads can move onto the Surtr Gateway.

| Mart | Gateway slug | Aerie query (Aerie e366e27d0) | Rows 09-29 |

|---|---|---|---|

| mart_education.aerie_admissions_marketing_event | aerie-admissions-marketing-event | D1 queryMarketingEventAgg, educrm.ts:1484-1517 | 1,041 (55 programs) |

| mart_education.aerie_admissions_marketing_event_contact (PII) | aerie-admissions-marketing-event-contact | D2 queryMarketingEventContacts, educrm.ts:1593-1629 | 146,830 (55 programs) |

What the candidates change in Aerie's SQL. Only the lifted key (A8 plan §3.1 rule 1):

- D1: TRIM(BOTH '"' FROM program_name::varchar) = $1 becomes the filter_program_name column. It leads the GROUP BY: (TRIM(program_name), TRIM(event_name)). The has_fact = true filter stays.

- D2: the predicate is lifted in both places. The event_identity_counts CTE groups by (TRIM(program_name), TRIM(event_name)) and is joined on both.

- D2 keeps Aerie's quirks as-is:

- neither the CTE nor the outer query filters on has_fact;

- the event_planning subquery isn't program-filtered, and is joined on the raw event_id / contact_id columns. These are BIGINT in both tables; the plan said SUPER.

- Rows with ambiguous event names (event_identity_count > 1, 50,905 rows) are published. Aerie withholds them in its mapper.

- tests/test_sql_contracts_marketing.py rebuilds each candidate from Aerie's pinned SQL by those edits alone, so any other difference fails.

Writers (061, 063, both appended to REFRESH_PROCEDURES, contiguous):

- Observed EduCRM provenance, as in U03. D2 reads two tables, so both observations must be the latest sales-educrm-mart-sync run.

- Atomic publish. DELETE + named-column INSERT in the CALL transaction; no TRUNCATE.

- Fails closed on:

- an empty candidate;

- a duplicate mart_row_id;

- for D2, a join fan-out: candidate rows + NULL-key rows must equal the source's contact rows;

- any row Aerie's strict parse rejects. Derived from Aerie's pinned Zod schemas:

- D1: a NULL event_name;

- D2: a NULL event_id, event_name, contact_id or event_identity_count, a blank event name, or a non-positive count.

- NULL program keys can't be selected by = $1. They're counted in the run log and not published (0 today).

- mart_row_id:

- D1: MD5 of (filter_program_name, event_name), unique by construction, plus the occurrence number;

- D2: MD5 of every published column plus the occurrence number among identical rows. Each part has a fixed shape ('n', or 'v' + an MD5 / digits / a date literal), so no value can collide two rows.

Gateway registration (U04) fix. Slugs, table names, orderBy: mart_row_id and the no-dateColumn setting already match. Only the two marketing descriptions were wrong: the contact rows come from marketing_event_dtl, not event_planning_dtl. Surtr/src/seed-gateway-aerie-a8.ts now describes both correctly. It takes effect on the next seed run; it isn't needed before then.

## Business Value

- Carries on the retirement of Aerie's EC2 analytics worker. This is Benji's Q3 priority 1 (A8, SURTR-735).

- Covers the largest per-program read in the refresh cycle. Today D1 and D2 run once per program: about 55 programs × 2 queries per hourly cycle, straight against EduCRM tables. After U22's gate, Aerie reads each mart once per cycle, with Surtr lineage on every row.

- Parity is provable, not asserted. Every program reconciles EXCEPT = 0 both ways against Aerie's own SQL.

- The strict-parse guard makes the mart fail before Aerie would. It never publishes a row that would break a program's marketing refresh.

## Manual Effort Estimate

About 12 hours of focused work by hand. That covers:

- reading D1/D2 and their Zod schemas;

- designing the two-predicate lift for D2;

- two provenance-pinned procedures;

- DDL with PII comments;

- the contract tests;

- the per-program reconciliation against Redshift.

_Proposed by Claude. Keval, please confirm or adjust._

## Testing / evidence

- uv run pytest: 144 passed, 57 of them new in test_sql_contracts_marketing.py. A mutation that drops D2's program join is caught.

- ruff check and ruff format --check (0.15.22): clean. biome check (2.4.4) on the seed file: clean.

- uv run python scripts/apply_ddl.py --dry-run: files in order 006-011, then 060 (14 statements), 061 (4), 062 (18), 063 (4).

- Read-only reconciliation (psql, SELECT only, counts only), Query 1 of each reconciliation/*_vs_aerie_sql.sql, Aerie SQL vs the candidate:

| Mart | Programs | Aerie rows | Candidate rows | aerie − candidate | candidate − aerie |

|---|---|---|---|---|---|

| D1 marketing_event (first 3 programs) | 3 | 72 | 72 | 0 | 0 |

| D1 marketing_event (every program) | 57 | 1,041 | 1,041 | 0 | 0 |

| D2 marketing_event_contact (first 3 programs) | 3 | 14,782 | 14,782 | 0 | 0 |

| D2 marketing_event_contact (every program) | 57 | 146,830 | 146,830 | 0 | 0 |

57 program names exist in the source; 2 have no contact or fact rows.

- Simulated publication (read-only). This runs each procedure's candidate and its tmp INSERT ... SELECT exactly as written, cast to the mart's column types:

| Mart | Rows | Row diff vs candidate | Duplicate mart_row_id | Strict-parse violations | NULL keys |

|---|---|---|---|---|---|

| D1 | 1,041 | 0 / 0 | 0 | 0 | 0 |

| D2 | 146,830 | 0 / 0 | 0 | 0 | 0 |

D2 row accounting: 146,830 source contact rows = 146,830 candidate rows + 0 NULL-key rows.

- Refresh runtime estimate. Candidate queries took 1.5 s (D1) and 1.7 s (D2). The full simulated publication, with keys and checks, took 4.0 s (D1) and 9.2 s (D2). With the DELETE + INSERT of 147K rows, I expect each D2 CALL to take about 20–40 s and D1 under 10 s. That's well inside the 480 s statement timeout and the 900 s Lambda.

- Column widths: the largest values today are an event name of 119 bytes (column is 1,024) and a program key of 32 (column is 256).

## Keval steps

1. PII sign-off (A8 plan §9 D2) for exposing aerie_admissions_marketing_event_contact on the Gateway. It holds contact and parent names, parent emails and dates of birth (23,223 rows have one; contacts are often minors). The DDL makes it owner-only: REVOKE ALL from PUBLIC and team_engineers.

2. Apply the DDL before merging. A merge reaches production within the hour, and the trigger then fires every 30 minutes. Run uv run python scripts/apply_ddl.py. Its list is 006-011 and 060-063; the U03 files re-apply idempotently.

3. After release, run the pipeline on demand. Then run Query 2 of both reconciliation files for a few programs; it must return 0 both ways.

4. Optional: re-run pnpm seed:gateway-aerie-a8 to pick up the corrected descriptions.

## Not covered

- The Aerie reader, shadow compare and gate for D1–D4 (ADMISSIONS_MARKETING_READ) are U22.

- D3 shadow-day and D4 weekly deposits are U18.

- Lifted-key matching caveat for U22. Redshift compares VARCHAR ignoring trailing blanks, so Aerie's = $1 would also match a program name with a trailing space. An exact-match TS index wouldn't. There are no such keys today: 57 keys, 57 byte-distinct. U22's shadow compare would surface it as pg-only rows.

- There's no behavioral rollback test for the DELETE + INSERT path. There's no Redshift in CI. This is the same deferral as the U03/U06/U08/U12 reviews.

- REFRESH_PROCEDURES, apply_ddl.py's list, test_pipeline_contract.py and the README will conflict on one line each with the sibling PRs 2089/2090/2092 and U13. Whichever merges second rebases.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2092 — feat(aerie-a8): per-program enrollment parity marts, Q6 cohort + Q7 pipeline deposit + Q8 transfer (SURTR-1543) @kevalshahtrilogy  approvedmercy-allow-critical

## Summary

A8 plan unit U12 (Linear: SURTR-1543, related to SURTR-735). It adds the three G2 per-program enrollment parity marts to the mart-aerie-admissions-refresh runner from U03 (PR 2084, merged). Aerie's per-program enrollment reads can then move onto the Surtr Gateway.

| Mart (Gateway slug) | Aerie query | Rows (09-29) |

|---|---|---|

| aerie_admissions_enrollment_cohort (aerie-admissions-enrollment-cohort) | queryEnrollmentDetailWithCoverage, Q6 (educrm.ts:686-707) | 9,238 |

| aerie_admissions_pipeline_deposit (aerie-admissions-pipeline-deposit) | queryPipelineDepositDetail, Q7 (educrm.ts:885-893) | 113 |

| aerie_admissions_enrollment_transfer (aerie-admissions-enrollment-transfer) | queryEnrollmentTransferDetail, Q8 (educrm.ts:1011-1030) | 36 |

All line numbers are at Aerie e366e27d0 (origin/main, 2026-09-29).

What each mart does

- The candidate is Aerie's SQL verbatim, with the lifted key. The TRIM(BOTH '"' FROM program_name::varchar) = $1 predicate becomes a filter_program_name column, and the key is added to every window, GROUP BY or join that ran after the filter:

- Q6: the key leads the coverage_rank partition, so each program keeps one placeholder per cohort and year.

- Q8: the enrollment-detail subquery groups by (TRIM(program_name), deal_id) and is joined on both deal_id and the key.

- Q7: no window or aggregate, so the key is only selected.

- Strict parse is asserted. Aerie parses these rows with parseRowsStrict, where one bad row fails the program. Each procedure refuses to publish a NULL in any column Aerie's row schema requires:

- Q6: program_code, cohort_id, cohort_label, session_school_year, has_fact.

- Q7: program_code, session_school_year.

- Q8: session_school_year, direction, program_code.

The tests derive these column sets from Aerie's pinned schemas.

- NULL keys are not published. = $1 can never select a NULL program name, so those rows are counted in the run log and left out. Q7 and Q8 also check that published rows plus NULL-key rows equal the source row count, which proves the Q8 join does not fan out.

- Observed EduCRM provenance, as in U03: the observed run must be the latest started run, no older than 24h, and its rows_loaded must equal the snapshot count. Q8 reads two EduCRM tables, and both must come from that same run.

- Sole-writer procedures. Each procedure is the only writer of its mart. It publishes with an atomic DELETE+INSERT (no TRUNCATE), fails closed, and rejects a duplicate mart_row_id.

- mart_row_id is an MD5 of the row key plus an occurrence number. Each text or SUPER key part is hashed separately, so a | in a value cannot make two keys collide. The occurrence number is ordered by every other published column.

- Where the code lives:

- DDL: pipelines/cdk/sql/mart_education/040-045. U12 owns 040-049; U06 uses 012-015 and U08 uses 030-033.

- The files are added to apply_ddl.py.

- The three procedures are appended to REFRESH_PROCEDURES in one contiguous change.

## Business Value

Aerie's admissions dashboard and the public v2 enrollment API are built from these three per-program reads. Today they are about 3 × 90 direct EduCRM Redshift queries every refresh cycle. With these marts, the Aerie shadow gate can swap them for three paged Gateway reads. Each read has provable lineage, and each is reconciled to 0 differences against Aerie's own SQL. This is one step toward taking Aerie off direct warehouse access (A8, SURTR-735), which was Benji's quarter priority 1. The strict-parse guard also means Surtr never serves Aerie a row that would silently fail a program's refresh.

## Manual Effort Estimate

About 14 hours of focused work by hand, with no AI. That covers:

- reading the three Aerie queries and their Zod schemas;

- designing the lifted-key edits, with the Q6 partition and Q8 join subtleties;

- writing 6 DDL files, 3 reconciliation scripts and the contract tests, patterned on U03;

- running the per-program reconciliation across all programs.

Keval: please confirm or adjust this number.

## Testing / evidence

- uv run pytest: 169 passed. This includes the new tests/test_sql_contracts_enrollment.py, which rebuilds each candidate from Aerie's pinned SQL by the lifted-key edits alone, and derives the strict-parse columns from Aerie's pinned schemas.

- ruff check and ruff format --check with ruff 0.15.22: clean.

- uv run python scripts/apply_ddl.py --dry-run applies 006-011, then:

  -- 040_aerie_admissions_enrollment_cohort.sql: 13 statement(s)

-- 041_sp_refresh_aerie_admissions_enrollment_cohort.sql: 4 statement(s)

-- 042_aerie_admissions_pipeline_deposit.sql: 14 statement(s)

-- 043_sp_refresh_aerie_admissions_pipeline_deposit.sql: 4 statement(s)

-- 044_aerie_admissions_enrollment_transfer.sql: 13 statement(s)

-- 045_sp_refresh_aerie_admissions_enrollment_transfer.sql: 4 statement(s)

- Read-only reconciliation. Query 1 of reconciliation/<mart>_vs_aerie_sql.sql compares Aerie's SQL, bound to each program, against the procedure's candidate filtered on filter_program_name, as multisets with SUPER serialized. It was run for every program in each source, with SELECTs only and counts only.

| Mart | Programs | Aerie rows | Candidate rows | Aerie − candidate | Candidate − Aerie |

|---|---|---|---|---|---|

| enrollment_cohort | 85 | 9,238 | 9,238 | 0 | 0 |

| pipeline_deposit | 36 | 113 | 113 | 0 | 0 |

| enrollment_transfer | 14 | 36 | 36 | 0 | 0 |

The per-program sums equal the full candidate counts, so every mart row falls within some program's slice.

- Read-only candidate check. The procedures' candidate and mart_row_id expressions were evaluated in a SELECT:

- rows: 9,238 / 113 / 36;

- distinct mart_row_id: 9,238 / 113 / 36;

- NULL-key rows: 0;

- strict-parse violations: 0.

- Observed provenance today: all three EduCRM tables report rows_loaded equal to the snapshot counts (9,238 / 113 / 36).

## Keval steps

1. PII sign-off (A8 plan §9). All three marts hold student (mostly minor) names and email addresses. The DDL has PII COMMENTs and revokes all privileges from PUBLIC and team_engineers; the Gateway reads as the owner.

2. Apply the DDL before merging. Run uv run python scripts/apply_ddl.py, which applies 006-011 and 040-045 in order. A merge reaches prod within the hour, and the runner then calls these procedures every 30 minutes.

3. After release, run the pipeline on demand. Then run Query 2 of each reconciliation file for a few programs; it must show 0 differences in both directions.

## Not covered

- Behavioural rollback test. Nothing tests a failure after the DELETE. Redshift is not reachable from CI, so the transaction shape is asserted textually, as in U03, U06 and U08.

- Aerie's mapper-level checks stay in Aerie. A fact row with an unusable contact_id, an unusable student_id, or a transfer direction other than in/out is published as-is, and Aerie's mapper throws exactly as it does on the legacy read. This is the U03 convention. Only schema-level (parse) failures are asserted here.

- The Aerie reader and shadow gate for these slugs are a separate Aerie unit, as is the U04 Gateway key.

- Rebase conflicts. Expect trivial conflicts with U06 (PR 2089) and U08 (PR 2090) on REFRESH_PROCEDURES, apply_ddl.py, test_pipeline_contract.py and the README.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2090 — feat(aerie-a8): per-program parity marts part 1, Q5 pipeline students + Q4 community metrics (SURTR-1541) @kevalshahtrilogy  approvedmercy-allow-critical

## Summary

A8 unit U08 (SURTR-1541). It adds the first two G2 per-program parity marts to the mart-aerie-admissions-refresh runner from PR 2084 (U03, now merged):

| Mart (Gateway slug) | Aerie query | Lifted key | Rows (09-29) |

|---|---|---|---|

| mart_education.aerie_admissions_pipeline_student (aerie-admissions-pipeline-student), PII | queryPipelineDetail, Q5 (sync/src/analytics/queries/educrm.ts:498-507) | TRIM(BOTH '"' FROM program_code::varchar) → filter_program_code | 33,585 |

| mart_education.aerie_admissions_community_metric (aerie-admissions-community-metric) | queryCommunityMetrics, Q4 (educrm.ts:1848-1861) | TRIM(BOTH '"' FROM program_name::varchar) → filter_program_name | 1,140 |

Line numbers were re-verified on Aerie origin/main 3d4fe1a97; educrm.ts is unchanged since 92fd47992.

How the per-program lift works (A8 plan §3.1 rule 1)

- Aerie runs each query once per program with WHERE <key> = $1. Each mart holds every program instead, and one program's legacy rows are exactly WHERE filter_<key> = <value>.

- Q5 has no window or aggregate after the filter, so the key is only selected.

- Q4: the key is added to the ROW_NUMBER partition, which becomes (TRIM(program_name), metric_name), so the latest snapshot is still picked per program. The raw SUPER metric_name stays in the partition, as in legacy. A read-only probe confirmed that metric_name inside the window resolves to the source column, not the trimmed alias.

- Nothing else changes. tests/test_sql_contracts_per_program.py rebuilds each candidate from Aerie's pinned SQL by those edits alone, so any other difference fails. A mutation check (dropping the Q4 partition key) fails 3 tests.

Writers: sp_refresh_aerie_admissions_{pipeline_student,community_metric} follow U03's pattern:

- observed EduCRM provenance through v_aerie_educrm_observed_publication: it must be the latest started run, rows_loaded must equal the snapshot count, and it must be under 24 hours old;

- the candidate is built in temp tables;

- the publish is an atomic DELETE + named-column INSERT, with no TRUNCATE.

It fails closed on:

- an empty candidate;

- a duplicate mart_row_id;

- for Q5, a candidate row count that differs from the snapshot count.

mart_row_id hashes each text key part on its own ('v' || MD5(text)), so a | inside a value can't make two keys collide.

Other changes

- PII: aerie_admissions_pipeline_student holds parent and child names and email addresses, parent program interest, and free-text notes and details.

- The table and column COMMENTs say so.

- The DDL revokes SELECT as well as writes from PUBLIC and team_engineers, so the table is owner-only. The Gateway reads as the owner.

- Q4 retirement: Aerie PR 1565 (U01) keeps Q4 communityMetrics, and whether to retire it is still open. If Q4 is retired, drop aerie_admissions_community_metric and its procedure. The table comment and README say so.

- Wiring: both procedures are appended to REFRESH_PROCEDURES and to the DDL_FILES list in apply_ddl.py.

## Business Value

- This moves Aerie's two highest-volume per-program reads onto a Surtr-owned, lineage-stamped contract. Pipeline students (Q5) are about 33.6K rows re-queried once per program, about 88 times per cycle.

- Aerie's G2 per-program cutover (U14 → U17 → U21) can then replace about 176 pg queries per refresh cycle with two paged Gateway reads.

- It is on the A8 critical path: U03 → U08 → U14 → U17 → U21. It is also a step toward tearing down the analytics-worker's direct Redshift access.

- The parity is provable (EXCEPT = 0 per program), and PII stays owner-only until sign-off.

## Manual Effort Estimate

About 7 hours of focused time, proposed; Keval to confirm or adjust. That covers:

- reading and verifying the two Aerie queries;

- designing the lifted key and partition change, including probing how metric_name resolves inside the window;

- 4 DDL files (about 680 lines) in U03's pattern;

- 2 per-program reconciliation scripts;

- the contract test module;

- running per-program reconciliation across every program;

- the README and PR write-up.

## Testing / evidence

- uv run pytest: 131 passed. U03's 87 tests, plus 44 new contract tests in tests/test_sql_contracts_per_program.py.

- ruff check and ruff format --check (ruff 0.15.22, as CI pins) at both the runner and the repo root: clean.

- uv run python scripts/apply_ddl.py --dry-run: exit 0. It prints only this runner's files, in numbered order:

  -- 006_v_aerie_educrm_observed_publication.sql: 3 statement(s)

-- 007_aerie_admissions_refresh_writer_mutex.sql: 5 statement(s)

-- 008_aerie_admissions_program.sql: 14 statement(s)

-- 009_sp_refresh_aerie_admissions_program.sql: 4 statement(s)

-- 010_aerie_admissions_program_directory.sql: 13 statement(s)

-- 011_sp_refresh_aerie_admissions_program_directory.sql: 4 statement(s)

-- 030_aerie_admissions_pipeline_student.sql: 17 statement(s)

-- 031_sp_refresh_aerie_admissions_pipeline_student.sql: 4 statement(s)

-- 032_aerie_admissions_community_metric.sql: 10 statement(s)

-- 033_sp_refresh_aerie_admissions_community_metric.sql: 4 statement(s)

- Read-only reconciliation, Query 1, run through psql as CQL_download_OM, SELECT only, printing counts only. It compares Aerie's per-program SQL with $1 bound against the procedure's candidate filtered on the lifted key, as a multiset with EXCEPT both ways.

| Mart | Program | Aerie rows | Candidate rows | aerie − candidate | candidate − aerie |

|---|---|---|---|---|---|

| pipeline_student | Alpha Austin | 2,319 | 2,319 | 0 | 0 |

| pipeline_student | Alpha Denver | 1,646 | 1,646 | 0 | 0 |

| pipeline_student | Alpha Orange County | 1,029 | 1,029 | 0 | 0 |

| community_metric | Alpha Austin | 10 | 10 | 0 | 0 |

| community_metric | Alpha Boston | 10 | 10 | 0 | 0 |

| community_metric | Alpha Denver | 10 | 10 | 0 | 0 |

- Every program, beyond the required 3: Q5 is clean for 88/88 programs, and the slices sum to 33,585, the full snapshot. Q4 is clean for 114/114 programs, and the slices sum to 1,140, the full candidate.

- Full-candidate checks (read-only SELECT):

- mart_row_id is unique and non-null: 33,585 of 33,585 for Q5 and 1,140 of 1,140 for Q4.

- The observed provenance resolves for both tables: the latest run is the observed run, and rows_loaded equals the snapshot count (33,585 and 71,680).

- Data shape:

- Q5 has 0 NULL keys, 0 fully identical rows, and 1,800 key groups that share (program, deal, parent, child). Those groups are why the occurrence number orders by every column.

- Q4 has no ties at the top snapshot_timestamp, so rn = 1 is deterministic.

- No DDL was applied and nothing was written.

## Keval steps

1. PII sign-off (A8 plan §9 D2) for exposing aerie-admissions-pipeline-student on the Gateway.

2. Apply the DDL before merge: uv run python scripts/apply_ddl.py. It is idempotent. U08's new files are 030–033, which need 006/007 from PR 2084 applied first (the script applies them in order either way).

3. After release, run the pipeline on demand, then run reconciliation Query 2 for a few programs, for example psql -v program_code='<code>' -f reconciliation/aerie_admissions_pipeline_student_vs_aerie_sql.sql. Every *_minus_* must be 0.

4. Decide Q4's fate with U01 (Aerie PR 1565). If Q4 is retired, a follow-up drops the community-metric mart.

## Stack note

- This was built stacked on PR 2084 (U03), which adds the runner, the provenance view and the mutex. 2084 merged while this was in progress, so this PR is rebased onto main and targets main. The U03 files are byte-identical to 2084's tip.

- U06 (SURTR-1540) also stacks on 2084, uses DDL 012–015, and appends to the same REFRESH_PROCEDURES line and DDL_FILES tuple. U08 takes a separate range, 030–039, and keeps its additions contiguous, so the conflict is one line in pipeline.json, apply_ddl.py and test_pipeline_contract.py.

- The new contract tests are in their own module, so they don't conflict with U06's edits to test_sql_contracts.py.

## Not covered

- Aerie-side reads, gate and shadow compare: U14 (seam + Q5/Q4 gateway paths) and U21 (per-program shadow).

- The remaining per-program marts: U12 (Q6–Q8) and U13 (Q2, Q3, app conversion).

- Gateway registration was already done in U04. The key mint and seed run are Keval steps there.

- A behavioural rollback test for the DELETE + INSERT path. It is deferred for the same reason as on PR 2084, since CI can't reach Redshift.

- Retiring Q4 is U01's decision, not this PR's.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2089 — feat(aerie-a8): community conversion + deposit parity marts (SURTR-1540) @kevalshahtrilogy  approvedmercy-allow-critical

## Summary

A8 plan unit U06 (SURTR-1540): the two community parity marts, appended to the mart-aerie-admissions-refresh runner from U03. Aerie can then read its G2 community-funnel and G3 community-deposit rows through the Surtr Gateway with its unchanged row mappers.

- DDL is pipelines/cdk/sql/mart_education/012-015, continuing 2084's numbering (U07 already uses 020-024). scripts/apply_ddl.py now applies 006-015 in order.

- 012/013 aerie_admissions_community_conversion (slug aerie-admissions-community-conversion). The candidate is Aerie's queryCommunityCommitmentEnrollments SQL, copied verbatim from sync/src/analytics/queries/educrm.ts:282-316 at Aerie 3d4fe1a97, ORDER BY included, with nothing changed or appended. Grain: one EduCRM mart_community_conversion_dtl row (deal_id); mart_row_id = MD5(deal_id + occurrence number, ordered by every other column).

- 014/015 aerie_admissions_community_deposit (slug aerie-admissions-community-deposit). The candidate is Aerie's queryCommunityDeposits SQL, copied verbatim from educrm.ts:2149-2166, nothing appended.

- NULL-community rows are kept. AERIE-2353's drop happens in Aerie's safeParseRows mapper, not its SQL, so parity keeps them here.

- The SQL has no unique key, so mart_row_id = MD5(an injective encoding of every published column + the row's occurrence number among identical rows).

- Observed provenance, per table. Each procedure reuses 2084's run-identity rule through v_aerie_educrm_observed_publication for its own EduCRM table. It requires exactly one observation, which must also be the latest started sales-educrm-mart-sync run (so no later writer can have replaced the table), be under 24 hours old, and have rows_loaded = the full-table snapshot count. For the deposit mart that is the whole table, has_fact or not, before Aerie's predicate narrows it. The conversion procedure also requires the candidate to equal the whole snapshot.

- Publication matches U03: shared writer mutex, temp-table candidate, fail closed on empty candidate / duplicate mart_row_id, atomic DELETE + named-column INSERT in the CALL transaction, no TRUNCATE.

- Runner: both procedures appended to REFRESH_PROCEDURES. No handler change.

- PII. Both marts hold heavy PII.

- Conversion: children's and parents' names, emails, phone numbers and home addresses, the child's date of birth and gender, and free-text admissions, shadow-day and loss-reason notes.

- Deposit: children's and parents' names and parents' emails.

- Each table has a Sensitive data: PII table comment and PII: column comments. The DDL makes them owner-only: REVOKE ALL (reads included) from PUBLIC and team_engineers. The Gateway reads as the owner.

- Keval's PII sign-off (A8 plan §9 D2) is required before either mart is exposed on the Gateway.

- Tests are a new tests/test_sql_contracts_community.py (it reuses U03's helpers) so U03's and U08's test files stay conflict-free.

## Business Value

- It moves two more Aerie admissions domains off direct EduCRM reads from the EC2 analytics worker. The community funnel (G2) feeds the community and full admissions funnels. Community deposits (G3) feeds a public Convex mutation. Both then become readable through the Surtr Gateway with lineage on every row. That is the SURTR-735 quarterly commitment, and the step that lets the worker be torn down.

- Parity is provable before cutover: Aerie's SQL is pinned, and the reconciliation EXCEPT is 0 both ways against both the candidate and a simulated publication.

- The PII is handled explicitly: owner-only tables, PII comments, and a named sign-off gate before exposure, rather than widening access by accident.

## Manual Effort Estimate

About 9 focused hours (just over 1 day) to build by hand without AI, on top of U03's runner: reading both Aerie queries and mappers (including AERIE-2353), typing the 60-column conversion mart, two procedures, the deposit mart's no-key row id, the reconciliation files, the contract tests, and the read-only evidence. Keval: please confirm or adjust.

## Testing / evidence

- uv run pytest: 135 passed (48 new in test_sql_contracts_community.py).

- Each candidate is asserted to be exactly Aerie's pinned SQL, and the pinned text matches Aerie origin/main (3d4fe1a97; educrm.ts is unchanged since 92fd47992).

- Table columns are asserted to equal the aliases parsed from Aerie's SQL, then lineage.

- The deposit row-id encoding is asserted to cover every column once, in order.

- Also asserted: observed-provenance guards come before the candidate, the snapshot count is unfiltered, has_fact appears only inside Aerie's SQL, and PII comments and owner-only grants are present.

- Ruff 0.15.22: ruff check and ruff format --check are clean.

- Read-only reconciliation was run with psql as CQL_download_OM in a default_transaction_read_only session (SELECT only; counts only, because the rows are PII).

Query 1 of each reconciliation file (Aerie SQL vs the procedure's candidate):

| mart | aerie rows | candidate rows | aerie − candidate | candidate − aerie |

|---|---|---|---|---|

| aerie_admissions_community_conversion | 5,403 | 5,403 | 0 | 0 |

| aerie_admissions_community_deposit | 360 | 360 | 0 | 0 |

Simulated publication. The procedure's exact tmp INSERT … SELECT was run as a SELECT: temp tables became CTEs, variables became literals, and every column was cast to the mart's declared type. It was then compared with Aerie's SQL as multisets:

| mart | aerie rows | simulated mart rows | aerie − simulated | simulated − aerie | distinct mart_row_id |

|---|---|---|---|---|---|

| conversion | 5,403 | 5,403 | 0 | 0 | 5,403 |

| deposit | 360 | 360 | 0 | 0 | 360 |

Query 2 (against the published marts) errors with "relation does not exist", as expected: no DDL has been applied.

- Negative controls on the deposit multiset shape:

- dropping 1 row gives 1 / 0;

- duplicating 1 row gives 1 / 1;

- dropping the NULL-community rows (Aerie's mapper behaviour) gives 7 / 0. This shows the mart must keep them for SQL parity.

- Source facts (09-29, counts only):

- community_conversion_dtl has 5,403 rows, all with a non-NULL, distinct deal_id.

- community_deposit_dtl has 1,622 rows, of which 360 are has_fact; 7 have a NULL community, and none are exact duplicates.

- Both tables' observed run (3c055bcf…, 10:06 UTC) is the latest started run, and rows_loaded equals the snapshot (5,403 and 1,622).

- The plan's "1,622 rows" for the deposit mart is the table count; the mart publishes the 360 has_fact rows Aerie reads.

- scripts/apply_ddl.py --dry-run: 10 files, 78 statements, in apply order (006-011 unchanged from 2084):

-- 006_v_aerie_educrm_observed_publication.sql: 3 statement(s)

-- 007_aerie_admissions_refresh_writer_mutex.sql: 5 statement(s)

-- 008_aerie_admissions_program.sql: 14 statement(s)

-- 009_sp_refresh_aerie_admissions_program.sql: 4 statement(s)

-- 010_aerie_admissions_program_directory.sql: 13 statement(s)

-- 011_sp_refresh_aerie_admissions_program_directory.sql: 4 statement(s)

-- 012_aerie_admissions_community_conversion.sql: 17 statement(s)

-- 013_sp_refresh_aerie_admissions_community_conversion.sql: 4 statement(s)

-- 014_aerie_admissions_community_deposit.sql: 10 statement(s)

-- 015_sp_refresh_aerie_admissions_community_deposit.sql: 4 statement(s)

013 and 015 are each one CREATE OR REPLACE PROCEDURE … $$ LANGUAGE plpgsql SECURITY INVOKER; plus ALTER PROCEDURE … OWNER, REVOKE ALL … FROM PUBLIC and GRANT EXECUTE … TO "CQL_download_OM". Their bodies are the diff.

<details><summary>apply_ddl.py --dry-run output: 012 and 014 (table DDL)</summary>

-- 012_aerie_admissions_community_conversion.sql: 17 statement(s)

-- Canonical DDL for mart_education.aerie_admissions_community_conversion (A8

-- unit U06, Gateway slug aerie-admissions-community-conversion). Sole writer:

-- mart_education.sp_refresh_aerie_admissions_community_conversion().

--

-- Query-shaped parity mart: columns are exactly the output aliases of Aerie's

-- queryCommunityCommitmentEnrollments SQL

-- (sync/src/analytics/queries/educrm.ts), in its select-list order. Raw

-- columns keep their EduCRM source type (SUPER included) so pg and Gateway

-- readers serialise them the same way; the ::text casts are VARCHAR(256), the

-- type Redshift gives TEXT.

--

-- PII: this mart holds children's and parents' names, emails, phone numbers,

-- home addresses, the child's date of birth, gender, and free-text admissions

-- notes. Exposing it on the Surtr Gateway requires Keval's PII sign-off (A8

-- plan §9 D2).

CREATE TABLE IF NOT EXISTS mart_education.aerie_admissions_community_conversion (

deal_id BIGINT,

child_id BIGINT,

child_first_name SUPER,

child_last_name SUPER,

child_email SUPER,

parent_id BIGINT,

parent_first_name SUPER,

parent_last_name SUPER,

parent_email SUPER,

parent_phone SUPER,

deal_name SUPER,

child_gender SUPER,

child_street SUPER,

child_city SUPER,

child_state SUPER,

child_zip SUPER,

child_country SUPER,

parent_primary_campus SUPER,

parent_program_interest SUPER,

parent_street SUPER,

parent_city SUPER,

parent_state SUPER,

parent_zip SUPER,

parent_country SUPER,

program_code SUPER,

program_name SUPER,

school_status SUPER,

session_school_year INTEGER,

session_start_date VARCHAR(256),

has_committed BOOLEAN,

committed_date VARCHAR(256),

student_date_of_birth VARCHAR(256),

age_at_session_start BIGINT,

is_eligible BOOLEAN,

enrollment_grade SUPER,

has_attended_shadow BOOLEAN,

shadow_date VARCHAR(256),

shadow_status SUPER,

shadow_feedback SUPER,

has_applied BOOLEAN,

application_date VARCHAR(256),

has_guide_approved BOOLEAN,

finalsite_status SUPER,

finalsite_application_url SUPER,

finalsite_student_url SUPER,

has_deposit BOOLEAN,

deposit_paid_date VARCHAR(256),

offer_sent_date VARCHAR(256),

has_contract_signed BOOLEAN,

contract_signed_date VARCHAR(256),

has_enrolled BOOLEAN,

enrollment_date VARCHAR(256),

deal_stage SUPER,

admissions_notes SUPER,

loss_reason_category SUPER,

loss_reason_notes SUPER,

tuition_profile SUPER,

is_founding_family BOOLEAN,

furthest_stage_code SUPER,

furthest_stage_name SUPER,

mart_row_id VARCHAR(32) NOT NULL,

source_run_id VARCHAR(128) NOT NULL,

source_published_at TIMESTAMPTZ NOT NULL,

refreshed_at TIMESTAMP NOT NULL,

created_by VARCHAR(128) NOT NULL,

PRIMARY KEY (mart_row_id)

)

DISTSTYLE ALL

SORTKEY (mart_row_id);

COMMENT ON TABLE mart_education.aerie_admissions_community_conversion IS

'Purpose: Surtr publication of the rows Aerie''s queryCommunityCommitmentEnrollments reads (every row of EduCRM mart_community_conversion_dtl), so Aerie can read them through the Surtr Gateway with its unchanged row mapper for the community and full admissions funnels. Grain: one output row of that SQL; normally one admissions deal (deal_id). Key: mart_row_id (MD5 of deal_id plus its occurrence number). deal_id is expected unique but not enforced: NULL or duplicate deal_id rows are published as-is, because Aerie''s mapper, not its SQL, drops NULL deal_id rows. Lineage: source_run_id/source_published_at are the observed sales-educrm-mart-sync run (v_aerie_educrm_observed_publication), not an atomic publication link. Sensitive data: PII. Children''s and parents'' names, emails, phone numbers and home addresses, the child''s date of birth and gender, and free-text admissions, shadow-day and loss-reason notes. Owner-only; Gateway exposure requires Keval''s PII sign-off. Full snapshot replaced atomically by mart_education.sp_refresh_aerie_admissions_community_conversion.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_conversion.child_email IS

'PII: the child''s email address.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_conversion.parent_email IS

'PII: the parent''s email address.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_conversion.parent_phone IS

'PII: the parent''s phone number.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_conversion.child_street IS

'PII: the child''s home street address; child_city, child_state, child_zip and child_country complete it.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_conversion.parent_street IS

'PII: the parent''s home street address; parent_city, parent_state, parent_zip and parent_country complete it.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_conversion.student_date_of_birth IS

'PII: the child''s date of birth, EduCRM TIMESTAMP cast to text (YYYY-MM-DD HH:MI:SS) as Aerie selects it.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_conversion.shadow_feedback IS

'PII: free-text shadow-day feedback about the child.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_conversion.admissions_notes IS

'PII: free-text admissions notes about the family.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_conversion.loss_reason_notes IS

'PII: free-text notes on why the deal was lost.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_conversion.mart_row_id IS

'Deterministic row key: MD5 of deal_id and its occurrence number. Gateway orderBy for total-order paging.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_conversion.source_run_id IS

'Observed sales-educrm-mart-sync run_id whose mart_community_conversion_dtl rows_loaded equalled the snapshot this publication read.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_conversion.source_published_at IS

'End time (UTC) of the observed sales-educrm-mart-sync run.';

ALTER TABLE mart_education.aerie_admissions_community_conversion

OWNER TO "CQL_download_OM";

-- PII: owner-only. REVOKE ALL (reads included) so the table stays owner-only

-- even if a different user applies this file. Reader access, beyond the

-- Gateway reading as the owner, is a DBA grant (PIPELINE §13) after Keval's

-- PII sign-off.

REVOKE ALL

ON mart_education.aerie_admissions_community_conversion FROM PUBLIC;

REVOKE ALL

ON mart_education.aerie_admissions_community_conversion FROM GROUP team_engineers;

-- 014_aerie_admissions_community_deposit.sql: 10 statement(s)

-- Canonical DDL for mart_education.aerie_admissions_community_deposit (A8

-- unit U06, Gateway slug aerie-admissions-community-deposit). Sole writer:

-- mart_education.sp_refresh_aerie_admissions_community_deposit().

--

-- Query-shaped parity mart: columns are exactly the output aliases of Aerie's

-- queryCommunityDeposits SQL (sync/src/analytics/queries/educrm.ts), in its

-- select-list order, with EduCRM source types kept (SUPER included) so pg and

-- Gateway readers serialise them the same way.

--

-- PII: this mart holds children's and parents' names and parents' emails.

-- Exposing it on the Surtr Gateway requires Keval's PII sign-off (A8 plan §9

-- D2).

CREATE TABLE IF NOT EXISTS mart_education.aerie_admissions_community_deposit (

community SUPER,

weeks_ago BIGINT,

week_label SUPER,

week_start_date DATE,

week_end_date DATE,

child_id BIGINT,

child_first_name SUPER,

child_last_name SUPER,

parent_id BIGINT,

parent_first_name SUPER,

parent_last_name SUPER,

parent_email SUPER,

community_deposit_paid_date TIMESTAMP,

x_total_deposits BIGINT,

mart_row_id VARCHAR(32) NOT NULL,

source_run_id VARCHAR(128) NOT NULL,

source_published_at TIMESTAMPTZ NOT NULL,

refreshed_at TIMESTAMP NOT NULL,

created_by VARCHAR(128) NOT NULL,

PRIMARY KEY (mart_row_id)

)

DISTSTYLE ALL

SORTKEY (mart_row_id);

COMMENT ON TABLE mart_education.aerie_admissions_community_deposit IS

'Purpose: Surtr publication of the rows Aerie''s queryCommunityDeposits reads (EduCRM mart_community_deposit_dtl rows with has_fact = true), so Aerie can read them through the Surtr Gateway with its unchanged row mapper. Grain: one output row of that SQL; one deposit fact, which the SQL does not make unique. Key: mart_row_id (MD5 of every published column plus the row''s occurrence number among identical rows); no business key is guaranteed. Rows with a NULL community are published as-is: Aerie''s mapper, not its SQL, drops them (AERIE-2353), so parity keeps them. Lineage: source_run_id/source_published_at are the observed sales-educrm-mart-sync run (v_aerie_educrm_observed_publication), not an atomic publication link. Sensitive data: PII. Children''s and parents'' names and parents'' email addresses. Owner-only; Gateway exposure requires Keval''s PII sign-off. Full snapshot replaced atomically by mart_education.sp_refresh_aerie_admissions_community_deposit.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_deposit.community IS

'EduCRM community the deposit counts towards; NULL on some rows, which Aerie''s mapper drops (AERIE-2353).';

COMMENT ON COLUMN mart_education.aerie_admissions_community_deposit.parent_email IS

'PII: the parent''s email address.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_deposit.mart_row_id IS

'Deterministic row key: MD5 of every published column and the row''s occurrence number among identical rows. Gateway orderBy for total-order paging. Changes whenever any column changes (weeks_ago advances weekly), so it is not a durable identity.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_deposit.source_run_id IS

'Observed sales-educrm-mart-sync run_id whose mart_community_deposit_dtl rows_loaded (all rows, has_fact or not) equalled the snapshot this publication read.';

COMMENT ON COLUMN mart_education.aerie_admissions_community_deposit.source_published_at IS

'End time (UTC) of the observed sales-educrm-mart-sync run.';

ALTER TABLE mart_education.aerie_admissions_community_deposit

OWNER TO "CQL_download_OM";

-- PII: owner-only. REVOKE ALL (reads included) so the table stays owner-only

-- even if a different user applies this file. Reader access, beyond the

-- Gateway reading as the owner, is a DBA grant (PIPELINE §13) after Keval's

-- PII sign-off.

REVOKE ALL

ON mart_education.aerie_admissions_community_deposit FROM PUBLIC;

REVOKE ALL

ON mart_education.aerie_admissions_community_deposit FROM GROUP team_engineers;

</details>

## Keval steps

1. DDL apply (prod), before merging. A merge reaches production within the hour, and the runner then CALLs both new procedures on every trigger; without the DDL they fail, and each run goes PARTIAL. Run uv run python scripts/apply_ddl.py from pipelines/runners/mart-aerie-admissions-refresh. It re-applies 2084's 006-011 idempotently (apply them first if 2084's apply has not happened yet), then 012-015.

2. Merge, then run the pipeline on demand and check results for both marts.

3. Run query 2 of reconciliation/aerie_admissions_community_{conversion,deposit}_vs_aerie_sql.sql. Every *_minus_* must be 0.

4. PII sign-off (A8 plan §9 D2) before Gateway exposure. Both marts are owner-only; keep Aerie's read gates for these two sources off until you sign off.

## Stack

- Was stacked on 2084 (SURTR-1536: the runner, provenance view and mutex). 2084 has merged, so this PR is rebased onto main and targets main. The diff is U06 only.

- REFRESH_PROCEDURES and apply_ddl.py's DDL_FILES are the only lines later units (U08) also touch.

## Not covered

- The Aerie read gates and shadow compares for these sources (A8 U09/U10). This PR is the Surtr side only.

- The prod DDL apply and query 2 of the reconciliation: both need Keval (above).

- A fix for AERIE-2353. Parity deliberately keeps the NULL-community rows; the fix belongs after cutover.

- Reader grants beyond the owner (a DBA grant after sign-off), and the Gateway source registration (U04).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2103 — chore(pipelines): assign HC forecast refresh owner @sanketghia  approved

## Summary

- Assign hc-forecast-refresh to sanket.ghia@trilogy.com.

- Keep the global default owner unchanged.

## Validation

- pipelines/owners.json parses as JSON.

- git diff --check passed.

- Test suite not run.

#24 — fix(deployment): allow retired task definition cleanup @benji-bizzell  no labels

## Summary

- Correct the dedicated CloudFormation role permission for ECS task-definition deregistration, limited to us-east-1.

## Why

The production rollout succeeded but replacement cleanup failed because ECS evaluates this action against resource *, not task-definition ARNs. The live role was corrected and the provisioning source must retain that correction.

## Business Value

Allows subsequent releases to finish cleanup without manual intervention.

## Test plan

- [x] pnpm check (86 tests), pnpm build, pnpm synth --no-lookups

- [x] Verified the live policy change contains only this permission correction

- [x] IAM simulation allows us-east-1 and denies us-west-2

- [x] CloudFormation UPDATE_COMPLETE and production deployment workflow succeeded

#23 — fix(registration): prepare the provider front for submission @benji-bizzell  no labels

## Summary

- Put credential acquisition instructions directly in the feedback clause and tighten its preflight.

- Publish service ownership and a bounded, reproducible warehouse-read example.

- Replace selected stale observations with verification methods and clarify semantic limits.

## Why

Registration checks the served front, not conversational context. Readers need an explicit feedback access path and provider evidence before submission.

## Business Value

Reduces avoidable first-registration failures without changing the DSS shape or claiming semantic acceptance.

## Test plan

- [x] 86 tests, build, offline synthesis and required CI pass.

- [x] Production accepts and retains 200/10000 feedback envelope; rejects oversized input.

- [x] Attribution, private operator authorization, idempotency, closure and revision conflicts pass.

- [x] Synthetic record 83ccfc9f51b204ba308b0083eb8d3849 survives replacement by task definition redshift-dss:2; validation operators revoked.

- [x] All six profiles pass warehouse identity/catalog checks; published bounded relation read passes.

- [x] Both deployed documents match source and digest headers; tightened registration preflight passes.

- [x] Registration remains unsubmitted; semantic acceptance and proxy takeover are not claimed.

#150 — Release: Shipyard 0.6.9 @ashwanth1109  no labels

## Summary

- Prepare Shipyard 0.6.9 with the approved public release notes.

- Change only the authoritative app version and versioned release notes.

## Business Value

- Delivers the approved Pi image attachment and Pomodoro capabilities alongside companion, diagnostics, trace, and image-layout improvements.

- Provides the metadata that enables the verified Apple Silicon release workflow to build and publish the update.

## Implementation Effort

- Metadata-only release change: one version bump and one public release-notes file.

- CI performs the build, signing, audit, and publication after merge.

## Test Plan

- pnpm test:release

- git diff --check

- Verify the exact metadata diff before merge and monitor the matching Desktop release workflow after merge.

#149 — AI-939: Prevent horizontal scroll from attached images @ashwanth1109  no labels

## Demo

![Smoke test evidence](https://github.com/AI-Builder-Team/Shipyard/blob/7b146313e5177cb8882a7f9b364c1c98eb48b15c/.github/smoke-evidence/AI-939-image-horizontal-scroll.png?raw=true)

## Summary

- Constrain attached-image galleries to the user message and conversation pane instead of viewport units.

- Bound composer image previews and contain only unintended transcript horizontal overflow.

- Add regression coverage for pasted previews, one/multiple history galleries, responsive sizing, and preserved code/table/diff scrolling.

## Tests

- pnpm test:chat

- pnpm test:recovery

- pnpm build

- pnpm theme:check

- pnpm test:smoke

- pnpm smoke build

- pnpm smoke start --scenario happy

- pnpm smoke verify --expect research-ready

- pnpm smoke stop

## Linear

https://linear.app/builder-team/issue/AI-939/adding-images-causes-horizontal-scroll

#22 — fix(deployment): match immutable GitHub OIDC identity @benji-bizzell  no labels

## Summary

- Match the exact ID-based OIDC subject enabled on this repository.

## Why

The live GitHub OIDC settings use an immutable repository subject, so a name-only trust condition rejects workflow authentication.

## Business Value

Enables scoped manual releases while retaining exact repository and production-environment restrictions.

## Test plan

- [x] Verify the subject prefix against the GitHub repository OIDC API.

- [x] GitHub deployment preview authenticates through OIDC and completes successfully (run 36670190924, attempt 2).

- [x] Required CI, 86 tests, build and offline synthesis pass.

#21 — feat(deployment): isolate DSS deployment permissions @benji-bizzell  no labels

## Summary

- Use dedicated deployment roles and private asset stores instead of the shared administrator bootstrap.

- Bound ECS runtime permissions and give standalone infrastructure predictable names.

## Why

The shared CDK execution role has administrator access. The standalone needs a constrained deployment path before it can be safely released.

## Business Value

Enables deployment and validation without granting this repository account-wide administrative authority.

## Test plan

- [x] 86 tests, build, offline synthesis and GitHub CI pass.

- [x] AWS Access Analyzer reports no policy findings; IAM simulation denies shared admin-role delegation and boundary removal/modification.

- [x] CloudFormation CREATE_COMPLETE, ECS healthy, verified TLS.

- [x] Existing-key authentication and Redshift identity/catalog reads across all six profiles.

- [x] DynamoDB feedback submission, retry, privacy, attribution, operator closure and revision conflict checks; temporary operator revoked.

#148 — AI-938: Prevent Trace panel refresh flicker @ashwanth1109  no labels

## Demo

![Smoke test evidence](https://github.com/AI-Builder-Team/Shipyard/blob/0d8b1cbed32ce130509842ae8a0e7b031797e1a5/.smoke-evidence/ai-938-trace-panel.png?raw=true)

## Summary

- Scope trace invalidation to workflow events for the selected task and coalesce overlapping reads onto the newest request.

- Keep the last successful trace visible during refreshes, preserve nested disclosures, and expose retryable refresh errors.

- Defer payload and artifact bodies until disclosure and index trace items once per loaded trace.

- Add regression coverage for parent rerenders, reopening, invalidation, manual refresh, disclosure stability, and failed refreshes.

## Linear

https://linear.app/builder-team/issue/AI-938/prevent-trace-panel-refresh-flicker

## Test plan

- pnpm test:task-trace

- pnpm build

- pnpm test:workflow

- cargo check --locked --manifest-path src-tauri/Cargo.toml

- cargo test --locked --manifest-path src-tauri/Cargo.toml --lib workflow::

- pnpm theme:check

#147 — AI-937: Add a Pomodoro timer to the top bar @ashwanth1109  no labels

## Demo

![Pomodoro timer smoke test](https://github.com/AI-Builder-Team/Shipyard/blob/3ee699fa353f63d132528cb9442b6723da55c189/docs/smoke-evidence/AI-937-smoke-test.png?raw=true)

## Summary

- Add a session-local Pomodoro control to the persistent Shipyard top bar.

- Support 25-minute focus and 5-minute break intervals with timestamp-based countdowns, pause/resume, reset, and automatic phase transitions.

- Reuse the accessible Popover/Button primitives with responsive semantic-theme styling.

- Add deterministic arithmetic and JSDOM tray coverage for transitions, accessibility, keyboard close/focus restoration, and lifecycle cleanup.

## Linear

https://linear.app/builder-team/issue/AI-937/add-a-pomodoro-timer-to-the-top-bar

## Test plan

- pnpm test:pomodoro

- pnpm test:notepad

- pnpm test:companion

- pnpm build

- pnpm theme:check

## Scope notes

- Timer state is intentionally session-local.

- Sound, desktop notifications, and backend/Tauri persistence are out of scope.

#20 — fix(delivery): align approval gates and Surtr networking @benji-bizzell  no labels

## Summary

- Use the existing branch-review and CI gates plus explicit manual deployment, removing the unsupported environment-reviewer requirement.

- Match Surtr public-subnet task networking while restricting task ingress to the ALB.

- Remove private-subnet inputs and verify the security boundary in infrastructure tests.

## Why

The user approved matching Aerie branch-level approval and Surtr networking. Aerie has PR/check rulesets with organization-admin bypass and no production environment reviewer gate. The target VPC has no private-subnet NAT route; the previous task configuration would lack startup egress. This service retains its main-only manual deployment rather than adopting Aerie automatic production-branch deployment.

## Business Value

Removes two verified deployment blockers without exposing task application ports directly to the internet.

## Breaking changes

Undeployed stack parameter contract no longer accepts PrivateSubnet1 or PrivateSubnet2; tasks use PublicSubnet1 and PublicSubnet2.

## Test plan

- [x] 85 tests, type checks, build and synthesis pass locally.

- [x] Task ingress is exactly TCP 3041 from ALB, with no public CIDR ingress.

- [x] Secret scan passes.

- [ ] GitHub CI and container build pass.

- Deployment remains disabled; no AWS or DNS changes.

#144 — AI-934: Align read-only companion with shared chat UX @ashwanth1109  no labels

## Demo

![AI-934 companion chat smoke test](https://github.com/AI-Builder-Team/Shipyard/blob/947b24f297fac09813b4ea47988f3a6f6236d50a/docs/smoke-evidence/AI-934-companion-chat.png?raw=true)

## Summary

- Mount Ask Shipyard on the shared Codex chat surface with a compact companion policy variant.

- Preserve the tray shell, foreground scope, singleton session lifecycle, reset/recovery flow, and non-modal dismissal behavior.

- Keep companion questions visible as plain submitted questions while sending the bounded foreground-context envelope to the companion route.

- Suppress attachments, model/access controls, approvals, previews, actionable file/external links, steering, and mutation-capable request controls.

- Add transcript/composer parity and read-only pending-request regression coverage.

## Linear

https://linear.app/builder-team/issue/AI-934/align-read-only-companion-with-shared-chat-ux

## Implementation notes

The existing CodexThreadChat surface now accepts a companion policy and compact tray layout. Companion sends use captureCompanionTurn through the existing { companion: true } thread routing. Native companion permissions, context schema, singleton session semantics, and recovery commands are unchanged.

## Test plan

- pnpm test:companion

- pnpm test:chat

- pnpm test:recovery

- pnpm exec tsc --noEmit

- pnpm theme:check

- pnpm build

#19 — fix(dependencies): establish a current validated baseline @benji-bizzell  no labels

## Summary

- Update Zod, the paired DynamoDB SDK, tsx, Node 22 types, TypeScript and Vitest to current compatible versions.

- Adopt the reviewed pnpm action update while keeping pnpm 11 pinned and Node 22 as the runtime.

- Group weekly routine npm and Actions updates; retain separate review for npm majors and runtime upgrades.

## Why

A sound baseline should accept compatible updates rather than repeatedly close dependency proposals. The pnpm action preserves its pnpm 11 path; TypeScript 7 and Vitest 5 pass the existing suite without weakening assertions or application behavior. Incorporates #13, #14, #15, #16, #17 and #18. Node minimum is corrected to 22.13 for the toolchain.

## Business Value

Keeps the service current and maintainable while reducing fragmented dependency PRs before deployment.

## Test plan

- [x] 85 tests, lint, type checking, build and CDK synthesis pass locally.

- [x] No known dependency vulnerabilities; staged secret scan passes.

- [x] Compiled-service smoke check: DSS documents resolve and unauthenticated queries are rejected.

- [x] Outdated check reports only Node 26 types, deliberately held to match Node 22.

- [ ] GitHub CI validates the new pnpm action and Linux AMD64 container build.

- Deployment remains disabled; no AWS changes.

#12 — fix(dependencies): align validated SDK and delivery updates @benji-bizzell  no labels

## Summary

- Incorporate Dependabot CDK and Biome updates, aligning the Biome schema.

- Update both DynamoDB SDK packages together and group future related updates.

- Keep Node image and type major upgrades out of routine dependency updates.

## Why

The individual PRs split coupled dependencies and proposed an unplanned Node 26 migration. This combines the compatible changes while retaining the validated Node 22 baseline. Incorporates #6, #8, #9 and #10.

## Business Value

Reduces dependency drift and update noise before deployment without changing infrastructure resources or permissions.

## Test plan

- [x] All 85 tests, Biome, TypeScript, build and synthesis pass locally.

- [x] Vulnerability audit and staged secret scan pass.

- [x] Template comparison differs only in the container asset reference and CDK metadata.

- [ ] GitHub CI and Linux AMD64 container build pass.

- No AWS deployment or production-data mutation.

#146 — AI-932: Support image attachments in Pi chat threads @ashwanth1109  no labels

## Demo

![Pi thread with image attachment](https://github.com/AI-Builder-Team/Shipyard/blob/138fa3c47ef89617ddfc4d995437c9db01c9b190/docs/smoke-evidence/AI-932/pi-image-attachment-demo.png?raw=true)

Smoke test: PASS (user-verified in Pi thread, worktree Shipyard-ai-932 at 5fb36db).

## Summary

Pi threads now support image attachments, matching Codex.

- Capabilities: PiAgentAdapter::capabilities() and the frontend PI_CAPABILITIES default report imageInput: true. This makes the composer show the + button, file picker, paste-to-attach, and thumbnail strip.

- Model config: The Pi model now declares "input": ["text", "image"]. Without this, Pi silently drops prompt images.

- Turn start: AgentTurnRequest has an images field. agent/turn/start passes images through, still rejects them for engines without image input, and allows image-only messages. Messages with no text and no images are still rejected.

- Pi adapter: The adapter validates and saves images with the shared attachment logic from codex_attachments.rs: up to 4 images, 10 MB each, same MIME allowlist, private per-upload directory under pi/attachments. It sends the images as RPC prompt images (ImageContent) and appends the hidden [Shipyard image attachments: local files] block to the message, so Pi tools can open the files later. Turn matching uses the final prompt text and tolerates the image hints Pi appends.

- Rendering: Pi user items keep their content parts, so images render inline in live events, the accepted turn, and restored history. codexMessageText hides the path block, and any hints after it, for Pi-format (data) image entries.

- Transport: Messages with images can be several MiB, so the Pi RPC record limit is raised from 8 to 32 MiB (the Codex frame limit), and JSONL newline scanning now only looks at new bytes instead of rescanning the whole buffer.

- Smoke fixture: The fixture now stores prompt images the way real Pi does.

Codex attachment behavior is unchanged: prepare_turn_input produces the same output and its tests still pass.

## Linear

https://linear.app/builder-team/issue/AI-932/support-image-attachments-in-pi-chat-threads

## Tests

- cargo test --lib (309 passed). New tests cover Pi capabilities, the Pi prompt payload with images and image-only, turn-ID matching with hints, rendering without the path block, image passthrough and empty-text rules in agent/turn/start, shared image saving, and the model config input.

- node --test scripts/test-codex-messages.mjs (hides the Pi path block), scripts/test-thread-recovery.mjs (Pi composer shows the attach controls, and an image-only send carries its images), plus the conversation-store, chat, and memory suites

- python3 -m unittest discover -s scripts/smoke -p 'test_*.py' (the fixture persists prompt images)

- tsc --noEmit, pnpm theme:check

#11 — fix(delivery): harden CI tooling and release preview @benji-bizzell  no labels

## Summary

- Upgrade Vitest to remove the development-tool advisories.

- Pin current GitHub Actions releases and use the supported template-only CDK preview mode.

## Why

The first live CI run exposed deprecated action runtimes, and the dependency audit identified two moderate findings in test tooling. Resolve these before enabling deployment. A template diff does not evaluate deployment parameter changes, so release instructions now make that review boundary explicit.

## Business Value

Keeps the standalone service on audited delivery tooling while preserving protected review and disabled production deployment.

## Test plan

- [x] 85 tests, Biome check, TypeScript, build, and credential-free CDK synthesis pass locally.

- [x] Dependency audit reports no known vulnerabilities; secret scans pass.

- [x] GitHub PR verification and Linux AMD64 container build pass.

- [x] Deployment-disabled workflow skips the release job; no AWS changes made.

- [ ] Actual AWS deployment and rollback require a separate approved validation.

#145 — AI-933: Clarify Codex and Pi connection setup @ashwanth1109  no labels

## Demo

![AI-933 connection screen smoke test](https://github.com/AI-Builder-Team/Shipyard/blob/61d55b816b891d426f9534236b31c7cfee1dac1c/docs/smoke-evidence/AI-933-connection-screen.png?raw=true)

## Summary

- Reframe the connection screen as Agent setup & diagnostics with shared TrueFoundry credentials for Codex and Pi.

- Surface Pi’s fixed Claude Opus 5.5 route and agent_ready("pi") readiness as Ready to start / Not ready, without presenting Pi as persistently connected.

- Group Codex health, process, activity, performance traces, and reconnect controls under a clearly labeled Codex runtime section.

- Update credential action feedback, task-creation setup guidance, missing-key messaging, and TrueFoundry documentation.

## Linear

https://linear.app/builder-team/issue/AI-933/clarify-codex-and-pi-setup-on-the-connection-screen

## Test plan

- pnpm test:connection-screen

- pnpm test:task-workspace

- pnpm exec tsc --noEmit

- pnpm build

- pnpm theme:check

- pnpm test:connection

- cargo test --locked --manifest-path src-tauri/Cargo.toml --lib tfy::tests

- cargo test --locked --manifest-path src-tauri/Cargo.toml --lib connection_diagnostics::tests

- pnpm test:smoke

#1596 — Capacity: artifact integrity, retention, and audit lineage (AERIE-2580) @marcusdAIy  approved

## Summary

Artifact integrity, retention, and audit lineage for the capacity pipeline (AERIE-2580: Yibin #4, #6, #8, #10 on Aerie #1439, plus Mercy's two follow-ups from Aerie #1554). This completes the ticket.

- Content checks (Yibin #4). Each copied file is hashed (SHA-256) and checked before it is stored (new capacityAutomation/artifacts.ts).

- Room table: application/json, valid UTF-8, and room rows: an array of objects, or an object holding one.

- Floorplan: image/svg+xml, a whole SVG document, with no script, foreignObject, event handlers, or XML entities. References are allowlisted, not blocklisted: every href, xlink:href, src, and CSS url() must be a #fragment inside the drawing, checked after decoding references the way a browser does. CSS may not contain escapes or @import, and no animation may target an href. The copy is served from Aerie storage and opened in a browser, so anything else is refused rather than sanitized.

- Bad content or an over-budget run can't be fixed by retrying, so it raises CapacityArtifactRejectedError and the run ends unresolved (cause artifactRejected) instead of polling until the 30-minute timeout. A truncated download or a Sindri error still retries.

- Transfer manifest (Yibin #6).

- Each copy is staged on the run (artifactTransfers, via _stageCapacityArtifact) as soon as it is stored. A retried poll reuses staged copies instead of downloading them again.

- _recordCapacityRunOutput now reads the manifest instead of taking storedArtifacts as an argument. It records the copies the output names and deletes every other staged copy. A run that ends without recording them (failed, unresolved, timed out, rejected artifact) deletes them all.

- A copy staged on a run that is no longer running is deleted at once. A copy that can't be deleted stays referenced in artifactTransfers, matching #1550's rule for artifactRefs. If staging throws after the copy is stored, the runner deletes it; if that delete also fails, the copy goes to a capacityArtifactOrphans log that the sweep keeps trying to delete (_purgeCapacityArtifactOrphans).

- Cumulative budget: CAPACITY_MAX_ARTIFACT_BYTES = 25 MB across a run's files. It is checked against Sindri's declared sizes before anything is downloaded, so it also bounds action memory.

- Retention and operator path (rest of Yibin #8).

- Retention rule: a copy lives only while a document can point at it. It is deleted on reject (as before), on supersede (at enqueue as before, and now also at publication, after retracting anything an earlier attempt registered), on rollback (proposal, and record-mode once the request is confirmed), on exhausted publication, and when a DD request ends unapproved.

- Held runs nobody reviews expire 14 days after they enter review (heldAt; _expireHeldCapacityRuns, run by the sweep): they become unresolved with cause reviewExpired, and their copies are deleted.

- New unresolvedCause on the run: doctrineMissing, evidence, sindri, timeout, agent, artifactRejected, reviewRejected, reviewExpired, ddRequest. Missing required doctrine (the assembler's unavailable status) is now doctrineMissing, distinct from every other unresolved run.

- Publication while disallowed no longer retries every 15 s. The run is parked (publicationParkedAt, nextAttemptAt unset, and the sweep's due query now skips it) with an error naming the operator step. _resumeParkedCapacityPublications (one run by runId, or all parked runs) hands runs whose site is now allowed back to the sweep and reports the rest as stillBlocked. A parked run still counts as in flight, so the site gets no new run until it is resumed. That matches today's behavior, where the retrying run also blocked the site.

- Access and audit lineage (Yibin #10).

- Every file sent to Sindri (the evidence Markdown) and every file copied back writes an audit entry (capacity.evidence.uploaded, capacity.artifact.copied). Copies carry a Sindri sourceExecution whose acceptedOutputHash is the file hash.

- artifactLineage on the run keeps each file's hash, size, and audit log ID, and is never pruned when copies are deleted. Registered documents also carry the SHA-256 in their notes.

- Narrowing access: declined, with the reason recorded on the ticket. The room table and floorplan are derived from the site's own floorplans, which are already ordinary site documents with the same audience. A narrower rule would need a document-level access model that Aerie doesn't have, and would hide the analysis from the people who review it.

- Stuck partial requests (Mercy, #1554). Part of the write already applied, so retracting would misstate the card. The run records ddRequestPartialSince. After 24 hours it gets a Needs engineering: error, logs once, and polls hourly instead of every 5 minutes. It is never compensated automatically. The flag clears when the request settles.

- Unset-status rollback, end to end (Mercy, #1554). The test now approves and confirms the rollback request and asserts rolledBack, the restored capacities, the kept status, the removed documents, and the deleted copies.

## Test plan

- [x] chat: vitest run convex/capacityAutomation.test.ts convex/capacityAutomation/config.test.ts (150 passed)

- Content: rejected room tables (not JSON, no rows, non-object rows, wrong type, invalid UTF-8) and floorplans (not SVG, wrong type, script, event handler, javascript: link, entities); accepted variants.

- Manifest: copies are staged with their hash; a retry after a partial transfer downloads only the missing file; the budget rejects before any download; bad content is never staged; a rejected artifact ends the run artifactRejected without polling.

- Staging writes lineage plus an audit entry with the Sindri sourceExecution; a replacement deletes the older copy; a non-running run refuses and deletes the copy.

- Recording keeps only the named copies and deletes a stray one; rejected and timed-out runs delete staged copies; a redelivery leaves recorded copies alone.

- The evidence upload hash and audit entry are kept with the dispatch inputs.

- unresolvedCause values for agent, evidence, doctrine, review rejection and expiry, and timeout; 14-day expiry of held runs.

- Parked publication: not due, stays parked while disabled, resumes and publishes once enabled.

- partial request: first seen, flagged after 24 hours with hourly polling and documents kept, cleared and published on approval.

- Unset-status rollback through approval and confirmation.

- [x] tsc --noEmit (chat), Biome, and the pre-commit hook.

#1599 — Forecast V2: publish End-of-Year arrival provenance @vvp-trilogy  approved

## Summary

- publish the four End-of-Year historical observation provenance fields in mart_admissions_forecast

- carry provenance beside the expected-arrival operand through live forecasts and persisted lock snapshots

- document null/zero semantics and reconcile every published field to the same usable observation

## Alpha Austin example

The mart contract can now return the explanation without an int_* join:

| field | value |

|---|---:|

| end_of_year_expected_enrollment_source_period_start | 2025-09-30 |

| end_of_year_expected_enrollment_source_period_end | 2026-06-05 |

| end_of_year_expected_enrollment_source_application_count | 36 |

| end_of_year_expected_enrollment_source_start_count | 16 |

| end_of_year_expected_enrollment_to_arrive | 16 |

## Validation

- git diff --check

- pre-commit hooks passed

- focused dbt fixture covers positive (36 applications / 16 starts), measured zero, and missing-observation reconciliation

- dbt execution deferred to PR CI per request

Closes #1597

#3836 — fix(board-doc): carry triaged findings across add-on Doc reconcile @marcusdAIy  approved

## Why

A Docs edit correctly invalidates the current add-on review, but today it also destroys the only prior addressed/dismissed snapshot. The next run therefore reopens previously triaged findings, unlike the native app.

## Change

- Keep an ID-free, bounded private disposition snapshot when Doc content changes. review_results remains None: stale cards are not served and old finding IDs remain unusable.

- Carry only persisted addressed/dismissed identities to matching new findings on the next review, inside the existing CAS retry, then clear the snapshot.

- Fail closed rather than silently truncate if the ledger exceeds its bound.

## Limits

This does not infer addressed status from Claire tool resolution or recover failed finding-status PATCHes. Sidebar batch handling is separate PR #3834 (already merged). No production Doc/session mutation or deployment.

## Tests

177 focused tests passed locally; Ruff check/format and diff check passed. CI/review required before merge.

#3835 — fix(budget-bot): avoid duplicate MIPs title after Apply @marcusdAIy  approved

## Scope

- Strip only an exact, standalone leading ATX or bold proposal title when the add-on already inserts a real Docs heading. Apply and add-section use the same pre-write planner; unrelated subheadings remain.

- On read, suppress only an identical, promotable bold NORMAL_TEXT paragraph immediately after a real heading. Other standalone bold pseudo-headings still promote. This reads existing affected docs without ambiguous section identity.

- Based on main containing #3833 nested P&L guard; separate from #3834 Sidebar queue/status. No production document mutation or deployment.

## Tests

- Add-on Vitest: 404 passed (15 files).

- Board doc parser pytest: 136 passed.

- Ruff check/format and git diff --check pass.

#3834 — fix(addon): stop stale findings batch after document apply @marcusdAIy  approved

## Summary

- Separate successful Google Doc apply from optional finding-status bookkeeping. A 409/other status failure leaves an honest "Applied to the doc" confirmation and asks for a fresh review; no old finding IDs are retried.

- Invalidate the local review epoch on the first successful Doc write. Pause the apply queue until a validated GET review finishes, discard only stale queued work, stop an in-flight findings batch, and leave unaccepted draft cards visible but non-actionable with explicit guidance and Discard.

- A cached GET review cannot restore stale findings. Only a fresh user-triggered review run after the write can unlock new findings.

## Tests

- pnpm test -- tests/stale-batch-after-apply.test.js tests/section-identity-repair-sidebar.test.js tests/accessibility-baseline.test.js tests/drive-context-attachments.test.js (114 passed)

- pnpm test -- tests/addon-chat-stale-finding.test.js tests/stale-batch-after-apply.test.js (18 passed, after #3832 main rebase)

- node --check for inline Sidebar script; git diff --check.

## Release / safety

- Based on main including #3832; no Code.gs, backend, or production changes here.

- Separate Code.gs pre-write content-loss guard is in progress. Do not resume live Apply flows or release this PR as a signal that document content was restored.

- Backend follow-up: GET review freshness/reconcile and a review-run identity token for finding-status conflict checks, without resurrecting old findings.

#1592 — Forecast V2: refactor Start-of-Year enrollment forecast @vvp-trilogy  approved

## Summary

- extend milestone observations with same-program current-to-following Start-of-Year alignment

- persist coherent live End-of-Year component snapshots and resolve them as locked forecasts after the milestone

- rebuild the supported Start-of-Year decomposition from End-of-Year-owned subtotals plus historical expected arrivals

- publish and document the complete explanation/status field set while preserving deprecated compatibility fields

- include the updated business formula document verbatim

## Validation

- dbt parse --no-partial-parse

- isolated pr1589_ dbt model build (affected chain and published mart)

- milestone observation unit tests, including dual-year alignment

- affected reconciliation/status/schema tests

Closes #1589

#3833 — fix(addon): fail closed on nested P&L pseudo-headings before parent rewrites @marcusdAIy  approved

## Impact

In the Combined IgniteTech Doc, P&L — Jive is a bold NORMAL_TEXT paragraph nested under the real parent H2. applySection currently deletes through the next real H2. A parent apply can therefore erase this user-added pseudo-section. P&L — Khoros was previously present and is now missing following a product-detail rewrite; causation has not been established. This change does not restore, edit, or re-style any production content.

## Safety change

- Detect standalone, fully bold P&L — <product> paragraphs inside the resolved target H2 body before creating or repairing identity markers and before body writes.

- Throw a typed, content-free error; Sidebar explains that an Editor must repair the document structure. Do not open the automatic identity-repair picker or retry the write.

- Keep real H2 boundaries and ordinary bold prose lead-ins working.

## Validation

- cd budget-bot-addon && pnpm test — 14 files, 392 tests passed.

- git diff --check clean.

- Mock DocumentApp geometry proves zero identity/body writes for Jive and Khoros shape, including legacy marker migration candidate, plus positive and host-failure cases.

## Release steps (no production write in this PR)

1. Review and merge after exact-head CI passes. Coordinate Sidebar conflict with pending #3832 and separate batch UX PR.

2. Deploy the add-on through the normal reviewed release process; verify the installed version.

3. An Editor must inspect and restore the user-added Khoros P&L if missing, and repair affected P&L section boundaries to real headings before retrying parent applies. Do not automatically restore or style live content.

#3832 — fix(budget-bot): guide add-on on stale review finding @marcusdAIy  approved

## Summary

- Add a typed, ID-free error only when /board-doc/addon/chat resolves a finding ID absent from the current review (or no review exists).

- Map that explicit server code to a fixed safe Apps Script marker and actionable sidebar guidance. Refresh review cards with the existing read-only GET. Do not retry chat or start a review automatically.

- Preserve generic handling for unrelated 404/409, auth failures, and unrecognized backend errors; ordinary chat stays unchanged.

## Evidence and scope

A production /board-doc/addon/chat 404 at 18:18:08 UTC and later 200s prompted this work. The historical 404 response body and document identity are unavailable, so we have not confirmed that this specific 404 came from stale finding lookup. This PR fixes the independently reproducible stale-ID code path; it does not claim to explain that event. No production mutation or deployment is included.

## Tests

- klair-api/tests/board_doc/test_addon_chat.py: 32 passed (including re-reconciled review, no review, unrelated 404, normal chat)

- Add-on targeted Vitest (new stale-finding + accessibility baseline): 61 passed

- Ruff check and format check on changed Python files; git diff --check clean

## Release order (future, separate approval)

1. Deploy the backend API change first. The old add-on remains compatible: it still shows the generic message until its update is published.

2. Create an immutable Apps Script version from this exact Code.gs + Sidebar.html source, then publish that version through Marketplace. Merging main by itself does not change the live add-on.

3. Validate the published version in a test document using the manual plan below. This PR does not perform any deployment or Marketplace publication.

## Manual validation after release (not performed here)

1. In a non-production test document, run review, keep an old finding card, then change/rerun review so its ID is absent. Click Address with Claire on the old card.

2. Confirm fixed guidance appears, cards refresh from GET, and no second chat POST or review-run POST occurs automatically.

3. Verify ordinary chat works and unrelated backend 404/permission errors remain generic. Correlate actual server response code before attributing the historical 404 to this path.

#1594 — feat(reconciliation): add Site evidence and decision lineage UI (AERIE-2288) @caina-barbosa  changes requested

## Summary

This PR is Phase 5 of 5 in [AERIE-2288 — rebuild the reconciliation stack from current main](https://linear.app/builder-team/issue/AERIE-2288/rebuild-the-phase-2-5-reconciliation-stack-from-current-main). It delivers [AERIE-2197 — Site evidence and decision-lineage UI](https://linear.app/builder-team/issue/AERIE-2197/port-site-reconciliation-evidence-and-decision-lineage) on top of the merged target-schema prerequisite.

It activates a bounded, read-only evidence experience for Property Acquisition fields and an authorized path from a Site to its reconciliation decision history. Aerie remains the sole owner of Site policy, validation and writes; Sindri remains a generic Workflow platform.

Production effect: activation — read-only Site evidence and decision-lineage presentation. Merging this PR does not start reconciliation, publish assets, change rollout controls, grant agent write authority or perform upstream writeback.

## Why

The reconciliation stack can already validate and commit cited decisions, but Site users cannot inspect the evidence behind those values or follow the resulting decision lineage from the Property Acquisition card. This slice closes that final user-facing gap without replacing the existing card, editor or generic Forge run detail.

## Business Value

- See which eligible Property Acquisition facts have supporting reconciliation evidence.

- Inspect bounded citation details without loading source quotes into the initial Site response.

- Distinguish current, stale, unavailable, denied and manually changed evidence states.

- Navigate from an authorized Site to its bounded AI decision lineage and existing generic run detail.

- Keep uncited fields visually quiet: no citation means no marker.

## How does it work

1. The authenticated Site fields route optionally returns a bounded, quote-free evidence summary and a separate forge.runs.read capability affordance. Exact citation text is fetched only for an eligible field and execution reference. Evidence source links are revalidated at this HTTP boundary and only http: or https: URLs reach the client.

2. PortfolioFieldsProvider maps that metadata and exposes a lazy bounded citation loader while preserving the existing save and refetch behavior.

3. The Property Acquisition card presents agreement execution state as quiet header metadata, adds a restrained decision-lineage link, and places one accessible citation marker beside each cited field descriptor.

4. Citation interaction opens a structured evidence preview with bounded loading, focus/Escape handling, stale and unavailable states, and source-link security unchanged.

5. A non-empty Site plus exact card=property-acquisition context selects paginated Site decision history. Every other context keeps the generated generic Forge list and unchanged generic run detail.

## Scope

### Included in this phase

- Protected Site evidence summary and bounded lazy citation reads.

- Defense-in-depth source-link protocol validation plus a server diagnostic for malformed optional projections.

- Property Acquisition execution-state, citation-marker and evidence-preview presentation.

- Capability-gated navigation to bounded Site decision history.

- Existing generic Forge list/detail behavior outside the exact Site context.

- Exact final diff paths:

chat/app/api/portfolio-sites/[slug]/fields/__tests__/route.node.test.ts

chat/app/api/portfolio-sites/[slug]/fields/route.ts

chat/components/dashboards/portfolio/cards/__tests__/property-acquisition-card.test.tsx

chat/components/dashboards/portfolio/cards/agreement-execution-badge.tsx

chat/components/dashboards/portfolio/cards/field-evidence-adornment.test.tsx

chat/components/dashboards/portfolio/cards/field-evidence-adornment.tsx

chat/components/dashboards/portfolio/cards/property-acquisition-card.tsx

chat/components/dashboards/portfolio/fields/__tests__/portfolio-fields-provider.test.tsx

chat/components/dashboards/portfolio/fields/portfolio-fields-provider.tsx

chat/components/dashboards/shared/tooltip.tsx

chat/components/sindri/__tests__/reconciliation-run-list.test.tsx

chat/components/sindri/reconciliation-run-list.tsx

chat/components/sindri/sindri-list-view.tsx

### Deliberately excluded for later phases

- No reconciliation-specific redesign of the generic Forge run inspector.

- No schema or data migration, agent Site-write path, rollout-policy change, asset publication, credential binding or Worker change.

- No reconciliation start, deployment, production mutation or upstream REBL3/Rhodes/Due Diligence writeback as part of this PR.

## Test plan

### Automated validation

- Focused route and browser coverage — 149/149 passed across five files:

- pnpm exec vitest run 'app/api/portfolio-sites/[slug]/fields/__tests__/route.node.test.ts' 'components/dashboards/portfolio/cards/__tests__/property-acquisition-card.test.tsx' 'components/dashboards/portfolio/cards/field-evidence-adornment.test.tsx' 'components/dashboards/portfolio/fields/__tests__/portfolio-fields-provider.test.tsx' 'components/sindri/__tests__/reconciliation-run-list.test.tsx'

- Chat typecheck — passed (pnpm typecheck).

- Exact-path Biome — passed for all 13 changed files.

- Commit-hook Chat typecheck and Biome — passed.

- git diff --check — passed.

- Exact-head scope — 13 authorized paths only; one commit directly on current origin/main at reconstruction time.

### Manual QC

- A controlled development reconciliation run completed for Austin and Roswell on the combined prerequisite + Phase 5 candidate: 4 sources, 2 proposals, 28 citations, 24 history rows, 9 provenance rows, 2 write audits and 0 nonterminal executions. Start/write controls were restored and the runner was stopped afterward.

- Human UI review approved the live Property Acquisition card at desktop and narrow widths, including quiet execution metadata, descriptor-adjacent citation markers, keyboard focus, single-popup/Escape behavior, bounded evidence presentation and Edit/Cancel continuity.

- The final arrow removal from the decision-lineage link was a source-only visual simplification; no behavior, route or evidence semantics changed.

- Sanitized E2E evidence: /home/perotti/.cache/pi-tmp/aerie-2591-phase5-e2e-20260928/evidence.json (SHA-256 2ea4f4fca1d2f50a80d38483ddee54b46134d642ca79105b6ecab33d25ae1d94).

### Time for Implementation

Approximately 4–6 engineer-days to reconcile the accepted UI with current main, preserve authorization and bounded-loading behavior, reduce tests to essential journeys, complete controlled E2E, and perform human UI review.

#1595 — Fix Real Estate mobile scrolling @YibinLongTrilogy  approved

## Summary

Fix vertical scrolling on the Real Estate dashboard at /dashboards?tab=real-estate on mobile. The page's fixed-height layout clipped cards below the viewport and hid the footer because the dashboard shell clips overflow. Also remove a nested button in mobile cards that causes a hydration error when a site has an In Rhodes action.

### Screenshots

<img width="629" height="694" alt="Screenshot 2026-09-29 at 1 54 45 PM" src="https://github.com/user-attachments/assets/37c79539-5ad4-47c1-a1c2-a0fc72f38677" />

### Changes

- chat/components/dashboards/real-estate/real-estate-view.tsx — Make the page the mobile vertical scroller and let the results region grow with its cards. Keep the constrained results region and internal table scrolling at the lg breakpoint.

- chat/components/dashboards/real-estate/__tests__/real-estate-view.test.tsx — Cover the mobile card, footer, and responsive scroll layout contract.

- chat/components/dashboards/real-estate/real-estate-matrix.tsx — Make the full-card Real Estate action and In Rhodes action sibling buttons while preserving their existing routes and card content.

- chat/components/dashboards/real-estate/__tests__/real-estate-matrix.test.tsx — Verify valid button markup and independent mobile navigation.

### Design Decisions

- Match the existing Admissions, Events, Camps, and Enrollments dashboard scroll pattern. The desktop table keeps its internal scrolling.

- Put the full-card button behind the visible card content and let the In Rhodes button receive pointer events above it. This avoids nested interactive elements without changing the card layout.

- Test the responsive CSS contract in jsdom, which does not calculate layout or perform real scrolling.

## Business value

Mobile users can reach every Real Estate site card and the list's count and freshness footer. Sites matched to Rhodes no longer trigger a hydration error, and both navigation actions remain available.

## Estimated manual effort

45–60 minutes.

## Test Plan

- [x] Real Estate view tests: 17 passed.

- [x] Real Estate matrix tests: 7 passed, including separate mobile card and Rhodes navigation.

- [x] Real Estate view persistence tests: 4 passed.

- [x] Workspace and Chat typechecks, Biome on changed files, test architecture, architecture boundaries, and git diff --check passed.

- [ ] On a phone-width viewport, open /dashboards?tab=real-estate and scroll from the filter bar through the last card to the footer.

- [ ] Tap a mobile card and its In Rhodes badge separately; confirm each opens the intended detail page without a hydration warning.

- [ ] At a desktop width, confirm the table retains internal vertical and horizontal scrolling.

#3830 — fix(board-doc): diagnose stale duplicate session identity without deleting drafts @marcusdAIy  approved

## Scope: diagnostic safety slice — NOT durable sync recovery

- Classify one live unanchored heading with two session title IDs as stale_duplicate_session_title only if exactly one nonempty last_synced_sections body matches the live body byte-for-byte.

- Return actionable 409 for add-on action gates. Re-read even on an unchanged Drive revision when session titles are duplicated, rather than falsely marking the stale session canonical.

- Keep duplicate live headings, ambiguous baselines, and unsynced generated drafts blocked. No section deletion, archive write, production mutation, or deployment.

- Follow-up durable archive + guarded repair design: klair-api/docs/board_doc_stale_duplicate_repair.md. This PR does NOT make deletion sync.

Tests: structural reconcile (19 passed), add-on reconcile (54 passed), ruff check/format, git diff check.

Follow-up: #3831.

#3828 — fix(budget-bot): skip unsafe C2.1 margin target comparisons @marcusdAIy  approved

## Summary

- Fail closed with a typed C2.1 skip for nonfinite, out-of-range, or currency-valued EBITDA margin targets. Preserve valid historical inverted-header target rows.

- Temporarily skip IgniteTech Q4 2026 C2.1: consolidated P&Ls cannot be compared to an ex-Khoros target. The observed absent quarter-specific targets URL means a static older sheet can be returned, but the canonical plan has no source provenance. Do not substitute the 66% fallback or 75% ex-Khoros target.

- Keep the IgniteTech Q4 skip even if a quarter-specific URL is later registered; Product-mode work must first establish comparable target/P&L scopes. This PR does not change target fetching, other checks, or Product mode.

## Tests

- cd klair-api && /home/marcu/Klair/klair-api/.venv/bin/python -m pytest -q tests/board_doc/test_review_checks.py (99 passed)

- ruff check and ruff format --check for both changed files (passed)

- git diff --check (passed)

No production mutation or deployment.

#1590 — 1400-ahj-personnel-contract @mwrshah  approved

- Align the AHJ Personnel PATCH item schema with server normalization: require a nonblank detail, while allowing new people with one detail and no ID.

- Validate generated person IDs with the same format rule as supplied IDs; keep the existing collision check.

- Cover the published schema and ID validation in contract tests.

Follow-up to #1523.

#3827 — fix(addon): check effective Drive editing capability for repair @marcusdAIy  approved

## Summary

- Read the caller OAuth token’s effective Drive capabilities.canEdit for the exact active Doc. This matches the backend access source and avoids the getAccess(Session.getEffectiveUser()) false negative seen for an Editor. Only an exact, untrashed file response with canEdit === true unlocks repair; all failures remain closed.

- Fix singular and plural matching-heading copy in the repair card.

- Add focused access, fail-closed, submit-time recheck, and sidebar wording tests.

## Tests

- cd budget-bot-addon && pnpm exec vitest run tests/section-identity-repair.test.js tests/section-identity-repair-sidebar.test.js tests/section-identity-repair-planning.test.js — 69 passed.

- git diff --check — passed.

No production Doc mutation or release.

#1588 — Capacity: fleet-wide run limit, rate-limit backoff, and a size cap on stored results (AERIE-2581) @marcusdAIy  approved

## Summary

Fleet budgets for the capacity pipeline (AERIE-2581, Yibin #7 on Aerie #1439). This also covers the rate-limit item in AERIE-2275.

- Fleet-wide run limit. _claimQueuedCapacityRuns counts runs already dispatching or running and claims only the open slots under CAPACITY_MAX_CONCURRENT_RUNS (default 6; an invalid value throws). A large sweep now drains in waves instead of dispatching everything at once.

- Rate limits. The Sindri client turns a 429 into a new SindriRateLimitedError, a subclass of SindriUnavailableError, and keeps the parsed Retry-After (seconds or HTTP date). A 429 was previously a generic user error.

- A rate-limited dispatch waits out Retry-After (or the 30 s base) through _deferCapacityRun, which does not spend a dispatch attempt and is still bounded by the 20-hour deferral limit.

- A rate-limited poll backs off to at least four normal poll intervals (60 s), or longer if Sindri asks.

- Every wait is capped at 15 minutes.

- The public API's sindri_unavailable 503 now forwards Sindri's Retry-After when there is one, instead of a fixed 30 s.

- Dispatch backoff. Ordinary dispatch retries back off exponentially from 30 s, doubling to a 15-minute ceiling, instead of a flat 30 s.

- Stored result size. Every result written to a run row passes through boundPersistedResult. Past 200 KB it becomes a marked summary: truncated, originalBytes, and the headline fields (resolved, Fast Open, Max, ruleset, tier, GSF, total NLA, doctrine version). The full output stays in the Sindri trace. Nothing in Aerie reads result back, and traceRef is only IDs, so no behavior changes.

## Test plan

- [x] chat: vitest run convex/capacityAutomation.test.ts convex/capacityAutomation/config.test.ts convex/sindri/client.test.ts (138 passed)

- Fleet limit: claims only the open slots, nothing while full, rejects invalid config.

- Rate-limited start is deferred with Retry-After, the base, or the ceiling, and never requeued as a failure.

- Rate-limited poll delay has a floor; other poll errors keep the normal cadence.

- Backoff curve; oversized result summary.

- Client: 429 maps to SindriRateLimitedError with Retry-After; parsing of seconds and dates.

- [x] convex/publicApi/http.test.ts: 34/35. handles unauthenticated CORS preflight requests timed out under suite load and passes on its own.

- [x] tsc --noEmit (chat), Biome, and repo lint scripts.

#143 — Release: Shipyard 0.6.8 @ashwanth1109  no labels

## Summary

- Bump Shipyard to 0.6.8.

- Publish release notes for the updater installation fix.

## Release notes

- Allow updates to install when the contextual companion has been opened but has not received its first question yet.

## Test plan

- pnpm test:release

- git diff --check

This PR intentionally changes only release metadata. The implementation fix was merged in PR #142 at 43a7f08.

#142 — AI-930: Fix updater installation with empty companion thread @ashwanth1109  no labels

## Summary

- Treat Codex companion threads without a first user message as idle during updater preflight.

- Preserve blocking behavior for active turns and ambiguous thread-read failures.

- Add focused regression coverage.

## Business Value

Users can install Shipyard updates after opening the contextual companion, even if they have not yet asked it a question. Legitimate active Codex work remains protected by the installation safety gate.

## Implementation Effort

Approximately 1–2 hours for an average engineer to diagnose the Codex lifecycle edge case, implement the guard, add regression coverage, and validate the focused Rust tests.

## Linear

https://linear.app/builder-team/issue/AI-930/fix-updater-idle-check-for-unmaterialized-companion-threads

## Test plan

- cargo test --manifest-path src-tauri/Cargo.toml --lib codex_app_server::tests::update_idle_check

- git diff --check

#3825 — fix(addon): one-character heading markers for all sections; auto-repair stretched markers @marcusdAIy  approved

## Problem

Section markers (BBOT_SEC_H::<id>) covered the whole heading paragraph. Their end touched the section body. When content was inserted there (by the add-on or by a person editing the Doc), Google Docs stretched the marker across the body. The add-on then rejected it as stale and showed "Section identity needs repair". This happened in production on the IgniteTech Q4 Doc (4 sections), and GFI, SaaS, and Central Support have the same shape. Only Financials survived, because it already used a one-character marker.

## Change

1. One-character heading marker for every section. createSectionNamedRange_ now marks the first character of the heading text ([0, 0]), the same shape Financials uses. Neither endpoint touches the body, so body edits and table inserts cannot stretch it. An empty heading keeps the whole-paragraph shape (there is no character to mark).

2. Safe auto-repair of stretched markers. planExpandedHeadingAnchor_ accepts an expanded whole-heading marker only when both signals agree:

- the marker starts on a heading and covers a contiguous run inside that heading's own section; and

- the requested title exactly matches that heading and no other.

Otherwise it stays section_identity_stale (duplicate titles, other heading, gaps, crossing into the next section, TITLE/SUBTITLE).

On the next write, the add-on creates and verifies the new character marker, removes the old one, re-lists for uniqueness, and only then writes the body.

3. A character marker that a person shifted by typing inside the heading is still read as that heading's marker, and is normalized to [0, 0] on the next write.

4. Rename re-seats the character marker after setText for all sections (was Financials-only).

Legacy BBOT_SEC:: body markers and the Financials refresh path are unchanged.

## Tests

pnpm vitest run in budget-bot-addon: 386 passed. There are new cases for the IgniteTech stretched shape, the fail-closed variants, marker survival after manual body edits, migration, and empty headings. Existing fixtures now model partial Text ranges.

## Release

This needs a new Apps Script / Marketplace version. Before release, validate live on a scratch Doc: write a section, edit the body by hand, then write again.

Follow-up (separate PR): repair-card Editor check (getAccess(getEffectiveUser) fails closed) and "1 headings exactly match" wording.

#1523 — Backfill production hotfixes (2026-09-25) into main @benji-bizzell  no labels

## Summary

Merges production back into main. It backfills the 2026-09-25 hotfix release (#1522), so the next main → production release doesn't regress it.

- #1520 AHJ Personnel. Adds per-site AHJ Personnel contacts: the site card, the field, PATCH …/people/ahj-personnel, the Agent/MCP tools, and the DSS guidance. This supersedes #1518, which will be closed.

- #1521 DSS site notes. Adds an object purpose, anchors the site-note object on the read path, adds portfolio.review-site-status, gives adding, editing and deleting a note one workflow each, and adds the point-in-time purpose for due diligence.

- Production-only platform-error triage fixes. From #1223 (bran/fix-platform-error-env-lookup), these went straight to production and never reached main: .github/workflows/cd.yml, scripts/classify-convex-env-lookup.mjs and its test, and scripts/convex-deployment-provenance.test.mjs. They already run in production CD, so including them removes drift.

git merge-tree merged it cleanly, with no conflict resolution. The branch head is a real merge commit (parents: main and production), but main only allows squash, so this lands as one squash commit whose content equals the merged tree.

## Validation (merged tree, clean export)

- pnpm typecheck passed.

- Contracts: 93 files / 1,252 tests passed.

- Chat, affected suites (public-api, Convex publicApi, portfolio dashboards, site fields, all-sites grid): 109 files / 1,653 tests passed.

- CD scripts: 15/15 node --test passed.

## Deploy notes

No new production deployment is needed; this content is already live. The schema change (optional sites.ahjPersonnel) reaches the dev deployments when main deploys.

🐦‍⬛ Generated by a very good bot

#141 — Release: Shipyard 0.6.7 @ashwanth1109  no labels

## Summary

- Bump the application version from 0.6.6 to 0.6.7.

- Add the approved public release notes for replay comparison, evaluation scoring, search UI, Pi reliability, and companion recovery.

## Business Value

- Deliver the reviewed Shipyard improvements and reliability fixes to users through the signed desktop release pipeline.

- Make replay results easier to compare and investigate while improving stability for Pi-backed tasks and Ask Shipyard recovery.

## Implementation Effort

- Metadata-only change: one version update and one release-notes file.

- No application source, workflow, dependency, or generated-file changes.

## Test Plan

- pnpm test:release

- git diff --check

- GitHub Actions will run the full release validation, Apple Silicon build, audit, and publication gates.

#1579 — 1397-aerie-site-backup-operation @mwrshah  changes requested

- Confirm Site-specific backup operation from the Backup Sites location editor, with a selected Site, current contract term, and audited assignment/status updates.

- Clear expired confirmations through a monitored daily cron. Keep primary-building opening milestone-derived and expose Site operation through Portfolio API v2, DSS, Agent, and Rhodes MCP.

- Include operation in status revisions and cover the selected-backup projection across the public API and Rhodes MCP.

- Monitoring convention: the cron registration invokes internal.crons.expireOperatingBackupAssignmentsRun; CRON_DISPLAYS.handlerPath identifies the underlying mutation, as it does for all other system jobs. Monitoring joins runs to display entries by cronName.

Intentional expiry consistency: A contract term must cover the date when an operator confirms a Site's backup operation. siteBackupLinks.operating then records that explicit confirmation; it is not recalculated from contract dates on each read. Portfolio API, Agent, and Rhodes read the stored confirmation until the monitored daily reconciliation (scheduled at 00:05 UTC) clears an expired assignment and records a system-attributed audit entry. A delayed or failed run extends that window; the schedule is not a hard freshness guarantee. This is the agreed stored-state policy documented in docs/school-site-dss-contract.md, not an omitted read-time check. Term dates are contract coverage, not observed dates of student moves. Switching to read-time expiry would change the published operation semantics before the stored confirmation and its audit record change.

#140 — AI-885: Add replay comparison reports and regression history @ashwanth1109  no labels

## Demo

![AI-885 smoke test](https://github.com/AI-Builder-Team/Shipyard/blob/f8c7f66/.smoke-evidence/AI-885-smoke-test.png?raw=true)

## Summary

Adds durable replay comparison reports and regression history for immutable eval cases.

- Registers eval cases, baseline/candidate runs, comparisons, redacted JSON/Markdown reports, evidence, divergence, interventions, and review decisions in SQLite.

- Adds replay-history commands and a compact Task Detail UI for running cases, comparing runs, reviewing findings, exporting reports, and candidate reruns.

- Documents the persisted report model and adds focused frontend/native coverage.

## Verification

- pnpm test:replay-reports

- pnpm test:eval-contract

- pnpm test:eval-scoring

- pnpm test:replay-runner

- pnpm test:feature-replay

- pnpm test:smoke

- pnpm theme:check

- pnpm exec tsc --noEmit

- pnpm build

- cargo test --locked --manifest-path src-tauri/Cargo.toml --lib

- cargo check --locked --manifest-path src-tauri/Cargo.toml

## Linear

https://linear.app/builder-team/issue/AI-885/add-replay-comparison-reports-and-regression-history

#1584 — Centralize Community Commitment qualification @vvp-trilogy  approved

## Summary

Follow-up cleanup to #1582. Community Commitment qualification remains owned by only:

- int_educrm_community_commitment.sql for the predicate

- _int_admissions__models.yml for the contract and deterministic unit fixture

This PR removes the copied criterion from downstream code:

- forecast aggregate/model comments now describe only their local aggregation and age behavior

- forecast detail describes only its local school-year requirement

- pipeline reconciliation, null-tenant, and enrollment-date tests now consume int_educrm_community_commitment as their qualified input contract

- the school-year source-quality test returns to checking stage data independently, without duplicating qualification

- the forecast reconciliation comment describes its direct Pipeline input without restating upstream qualification

No Community Commitment membership behavior changes from #1582.

## Validation

- poetry run dbt parse --profiles-dir /home/ubuntu/aerie/control/dbt with parse-only environment placeholders

- YAML parse

- git diff --check

- warehouse-backed dbt build/test delegated to PR CI

Closes #1583

#139 — AI-928: Fix Ask Shipyard startup and companion recovery @ashwanth1109  no labels

## Summary

Ask Shipyard could fail before its first question because Codex reports an unmaterialized conversation when history is requested. Recover also cleared the session while resetting it, triggering a competing automatic reopen. Treat the specific pre-first-message response as empty history and serialize opening/recovery while preserving the chat and unsent draft. Failed requests remain visible and Retry repeats the failed operation.

## Business Value

Users can open the contextual companion, ask their first question, and recover failed conversations without repeated resets or losing their draft.

## Implementation Effort

Estimated 4–6 hours for an average engineer to diagnose the protocol lifecycle and React race, implement the fixes, and write/validate the regressions without AI assistance.

## Linear

https://linear.app/builder-team/issue/AI-928/fix-ask-shipyard-startup-and-companion-recovery

## Validation

- pnpm test:companion — 6 passed.

- pnpm test:recovery — 46 passed, including first-question submission through the real companion chat hook.

- TAURI_CONFIG='{"bundle":{"resources":[]}}' cargo test --locked --manifest-path src-tauri/Cargo.toml --lib codex_app_server::tests:: -- --test-threads=4 — 25 passed; bundle resources disabled only for unit-test compilation.

- pnpm exec tsc --noEmit, pnpm build, and git diff --check passed.

- Native packaged-app UI smoke testing and installation have not been performed. Read-only permissions are retained; unrelated provider errors remain errors.

#1582 — Require paid deposits for Community Commitments @vvp-trilogy  approved

## Summary

- require a non-null EduCRM community deposit date at the shared Community Commitment intermediate boundary

- make pipeline and forecast populations inherit the same paid-deposit definition

- add deterministic unit coverage for paid, unpaid, and wrong-stage deals and align reconciliation tests/docs

## Validation

- poetry run dbt parse --profiles-dir /home/ubuntu/aerie/control/dbt (with dummy parse-only environment variables)

- dbt selection confirms unit_test:bran_dbt.int_educrm_community_commitment_requires_paid_deposit

- YAML parse and git diff --check

- warehouse-backed dbt build/test delegated to PR CI

Closes #1581

#138 — AI-884: Add node-level and end-to-end eval scoring with failure attribution @ashwanth1109  no labels

## Demo

![AI-884 smoke test](https://github.com/AI-Builder-Team/Shipyard/blob/488154a2025c873c21765d58b2ecbe1d95444292/docs/smoke-evidence/AI-884-smoke-test.png?raw=true)

## Summary

- Add deterministic-first scoring for node, handoff, and end-to-end feature replay findings.

- Add bounded stored model judgments, evidence provenance/redaction, repeated-run aggregation, and baseline/candidate comparisons.

- Expose the evaluator through Tauri and the frontend wrapper, with schema/contract documentation and focused fixtures.

## Verification

- cargo test --locked --manifest-path src-tauri/Cargo.toml --lib

- pnpm test:eval-scoring

- pnpm test:eval-contract

- pnpm test:feature-replay

- pnpm exec tsc --noEmit

- pnpm build

## Linear

https://linear.app/builder-team/issue/AI-884/add-node-level-and-end-to-end-eval-scoring-with-failure-attribution

Draft PR only; do not merge.

#1577 — fix(reconciliation): carry target decision schemas (AERIE-2591) @caina-barbosa  approved

## Summary

This PR is a corrective prerequisite slice in the larger [AERIE-2288 — Rebuild the Phase 2–5 reconciliation stack from current main](https://linear.app/builder-team/issue/AERIE-2288/rebuild-the-phase-2-5-reconciliation-stack-from-current-main) project.

It hardens reconciliation target-field input by co-locating each field's authoritative current value and accepted decision schema, then carrying that descriptor through the receipt fence, authored Agents and final Aerie validation. The implementation is tracked by [AERIE-2591 — Carry accepted value schemas with reconciliation target fields](https://linear.app/builder-team/issue/AERIE-2591/carry-accepted-value-schemas-with-reconciliation-target-fields); the deletion-first test reduction is tracked by [AERIE-2599](https://linear.app/builder-team/issue/AERIE-2599/aggressively-reduce-aerie-2591-tests-to-essential-target-field).

Production effect: controlled rollout or migration. Merging changes the hard-cut contract and authored asset definitions, but does not publish Sindri assets, enable reconciliation controls, start a Workflow, deploy anything or write Site data.

---

## Why

The Roswell controlled run proved that the Site Reconciler could be asked to decide agreementExecutionState without receiving Aerie's accepted value/status vocabulary. It produced a plausible but invalid value, and final Aerie validation correctly blocked the write. This slice closes that information seam without moving field policy into Sindri, weakening validation or allowing an Agent to guess business decisions.

Phase 5 Draft PR #1553 depends on this prerequisite landing first.

---

## Business Value

- Allows reconciliation Agents to make decisions using the same accepted vocabulary enforced by Aerie.

- Prevents deterministic blocked runs caused by missing value/status/confidence constraints.

- Preserves Aerie as the sole authority for Site validation, provenance, audit and writes.

- Keeps Sindri generic and prevents prompt-owned or duplicated business policy.

- Reduces review and maintenance cost by deleting 43 redundant or tautological semantic test scenarios.

---

## How does it work

1. Aerie's property-acquisition policy registry produces one bounded targetFields collection. Each field descriptor contains its currentValue and generic decisionSchema variants.

2. Receipt capture persists that exact collection in the immutable input fence. Read preparation rehydrates the persisted descriptor rather than rebuilding a parallel schema map.

3. The Evidence Agent remains evidence-only. The Site Reconciler consumes the supplied descriptors, while the Commit Proposal Agent preserves the complete input fence without substituting values or metadata.

4. Aerie re-evaluates authoritative target fields before settlement, rejects stale descriptor-only policy changes, validates the final tuple against the same policy registry, and retains sole Site-write authority.

5. The unreleased legacy currentValues bare map is rejected. There are no aliases, guessed normalization, dual reads/writes or Aerie-specific Sindri runtime changes.

### Rollout and target-resolution boundaries

The hard cut is intentional and applies to an unreleased Workflow input contract. Merge does not publish the new assets or start reconciliation. Activation remains separately gated by empty rollout allowlists, start/write kill switches and a zero-nonterminal-execution preflight; if that preflight is not clean, rollout stops rather than interpreting a legacy receipt. Terminal historical executions do not resume settlement.

A read-only check of Aerie's default production Convex deployment at 2026-09-29T13:11:35Z returned no documents from reconciliationRegistrations, reconciliationExecutions or reconciliationInputRevisions. Production therefore has no published reconciliation registration, no in-flight or historical reconciliation execution, and no legacy persisted receipt that could cross this deployment.

Target resolution distinguishes absence from failure. A successfully materialized null Property Acquisition value is the valid “no property data yet” state and produces descriptors with null current values. Database lookup ambiguity, an unwritable/missing Site, materialization exceptions, missing required keys and final property validation exceptions all throw the shared target-unavailable user error; none reaches the property: null return.

---

## Scope

### Included in this phase

- Versioned, bounded, co-located target-field descriptors.

- Aerie-owned descriptor generation and final tuple validation from one policy registry.

- Receipt-fence persistence, reads, settlement and commit propagation.

- Authored Agent and Skill instructions for descriptor consumption and non-substitution.

- Nine essential semantic test scenarios across eight owning boundaries.

- Exact final diff paths:

chat/convex/reconciliation/admin.test.ts

chat/convex/reconciliation/admin.ts

chat/convex/reconciliation/commit.test.ts

chat/convex/reconciliation/commit.ts

chat/convex/reconciliation/coordinatorPollRecovery.test.ts

chat/convex/reconciliation/coordinatorStart.test.ts

chat/convex/reconciliation/fieldPolicy.ts

chat/convex/reconciliation/finalOutputSettlement.test.ts

chat/convex/reconciliation/finalOutputSettlement.ts

chat/convex/reconciliation/propertyAcquisitionFieldPolicy.test.ts

chat/convex/reconciliation/propertyAcquisitionFieldPolicy.ts

chat/convex/reconciliation/readiness.test.ts

chat/convex/reconciliation/readiness.ts

chat/convex/reconciliation/reads.test.ts

chat/convex/reconciliation/reads.ts

chat/convex/reconciliation/validator.test.ts

chat/convex/reconciliation/validatorHttp.test.ts

chat/convex/reconciliation/validatorHttp.ts

chat/sindri-assets/document-field-reconciliation/agents/commit-proposal.json

chat/sindri-assets/document-field-reconciliation/agents/evidence-analyst.json

chat/sindri-assets/document-field-reconciliation/agents/site-reconciler.json

chat/sindri-assets/document-field-reconciliation/assets.test.ts

chat/sindri-assets/document-field-reconciliation/skills/aerie-reconciliation-commit-proposal/SKILL.md

chat/sindri-assets/document-field-reconciliation/skills/aerie-reconciliation-evidence/SKILL.md

chat/sindri-assets/document-field-reconciliation/skills/aerie-reconciliation-site/SKILL.md

packages/contracts/src/reconciliation-terminal-outputs-v1.schema.json

packages/contracts/src/reconciliation.test.ts

packages/contracts/src/reconciliation.ts

### Deliberately excluded for later phases

- Phase 5 Site evidence and decision-lineage UI — remains isolated in Draft PR #1553 and will be reconstructed after this PR merges.

- Sindri runtime or platform policy changes — Sindri continues to transport generic opaque input.

- Automatic asset publication, Workflow instance changes or rollout activation.

- Production deployment, data migration, backfill or Site write.

- Rhodes Worker changes — the existing Worker forwards proposal drafts opaquely and was verified compatible.

- Credential creation or rotation.

---

## Test plan

### Automated validation

- focused reconciliation tests — 98/98 passed (pnpm --filter @bran/chat exec vitest run --project edge <10 focused reconciliation files> --maxWorkers=1)

- contracts — 9/9 passed (pnpm --filter @bran/contracts exec vitest run src/reconciliation.test.ts --maxWorkers=1)

- authored assets — 6/6 passed (pnpm exec tsx --test chat/sindri-assets/document-field-reconciliation/assets.test.ts)

- broad reconciliation suite — 112/112 passed (cd chat && pnpm exec vitest run convex/reconciliation/*.test.ts)

- root tests — 152/152 passed (pnpm test:root)

- Chat and Convex typecheck — passed (pnpm --filter @bran/chat typecheck)

- contracts typecheck — passed (pnpm --filter @bran/contracts typecheck)

- workspace typecheck — passed (pnpm typecheck)

- exact-path Biome — passed for all 28 changed paths

- architecture boundaries, Convex paths and read bounds — passed (pnpm lint:boundaries, pnpm lint:convex-paths, pnpm lint:read-bounds)

- pre-commit hooks — Convex paths, Biome and Chat typecheck passed

- git diff --check — passed

- exact-head scope — 28 listed paths, one commit on current main, no Phase 5 paths

- independent essential-coverage review — passed with no findings

- independent austerity review — passed with no findings

Test reduction evidence:

- added test LOC: 703 → 162 (76.96% reduction)

- changed test LOC: 778 → 239 (69.28% reduction)

- AERIE-2591 semantic scenarios: 52 → 9 across eight owning boundaries

- final tests: 19.1% of additions and 15.2% of changed LOC

### Time for Implementation

Approximately 2 to 3 engineer-weeks without AI assistance, including contract design, propagation through reconciliation boundaries, authored asset updates, test reduction, review and controlled cross-system validation.

### Manual QC

A controlled development E2E ran the exact production bytes in this PR together with the separately reviewed Phase 5 UI candidate against Aerie hallowed-stork-702, Sindri fine-jackal-170 and the existing Rhodes Worker.

- Austin: completed / updated / commit

- Roswell: completed / updated / commit

- aggregate: 4 sources, 2 proposals, 28 citations, 24 history rows, 9 provenance rows and 2 write audits

- both read grants revoked

- zero nonterminal executions after completion

- observe/commit allowlists cleared, start/write kill switches restored, and runner stopped

- Worker and credentials unchanged

- no upstream REBL3, Rhodes or Due Diligence writeback

#2097 — fix(runners): keep the Redshift cancellation window when a statement times out at the run deadline (SURTR-1548) @kevalshahtrilogy  approved

## Summary

SURTR-1548. This fixes a cancellation-deadline bug in the Redshift Data API client that 13 runners copied. Mercy found it on PR 2096 (SURTR-1547), where it is already fixed for the new mart-aerie-dbt-publication-refresh runner.

- The bug:

- The client stops polling a statement at min(now + statement timeout, run deadline - reserve).

- After a timeout, it cancels the statement and waits for a terminal state until min(now + cancel window, run deadline).

- When a call's reserve is below the cancel window (for example the default 0 on verification reads), polling can run right up to the run deadline. The confirmation window is then already empty.

- So a statement that timed out, and would have reached ABORTED seconds later, is reported as UnknownStatementOutcomeError (an unknown commit outcome), not as a clean timeout.

- The fix, one expression per client: polling now stops max(reserve, cancel window) before the run deadline.

- A timed-out statement abandons the work its reserve was kept for, so the larger of the two is enough.

- A call whose reserve already covers the window (for example the 180 s CALL reserves) behaves exactly as before.

- Runners touched (13):

- core-education-academic-term-refresh, core-education-camps-refresh, core-education-site-matterport-refresh, core-education-site-metadata-refresh

- education-q48-exception-refresh

- mart-aerie-admissions-refresh (U03), mart-aerie-camps-refresh, mart-aerie-rebl3-sites-refresh, mart-aerie-school-calendar-refresh, mart-aerie-school-source-directories-refresh, mart-aerie-schools-data-sheet-refresh, mart-aerie-xo-contractor-refresh

- mart-aerie-hubspot-refresh: the older variant of the same client (run_deadline, MAX_CANCEL_CONFIRM_SECONDS). It is latent there, because every handler call passes a 60 s reserve.

- How the family was found: I scanned every pipelines/runners/*/src/*.py client that calls cancel_statement. These 13 are exactly the clients whose polling deadline is run deadline - reserve and whose confirmation window is capped by that same run deadline. The rest were checked and left alone (see Not covered).

## Business Value

- An unknown commit outcome now means one. These runners publish the A8 Aerie marts and several governed core_education tables. Before this fix, a slow statement near the Lambda deadline was reported as an ambiguous commit even when Redshift aborted it cleanly seconds later. That kind of report is a red alert, and it sends someone to audit whether a publication half-landed.

- One fix, done once. The whole copied family is fixed together, with a regression test in each runner, so the bug is not rediscovered runner by runner.

## Manual Effort Estimate

About 3 focused hours by hand without AI. That covers auditing the copied clients across pipelines/runners/, the one-line fix times 13, a clock-controlled test for each, and running the 13 suites. Keval: please confirm or adjust.

## Testing

- All 13 runner suites pass (uv run pytest, or the plain-venv path CI uses for mart-aerie-hubspot-refresh, which has no pyproject.toml):

| runner | tests |

|---|---|

| core-education-academic-term-refresh | 45 |

| core-education-camps-refresh | 53 |

| core-education-site-matterport-refresh | 54 |

| core-education-site-metadata-refresh | 90 |

| education-q48-exception-refresh | 118 |

| mart-aerie-admissions-refresh | 181 |

| mart-aerie-camps-refresh | 56 |

| mart-aerie-rebl3-sites-refresh | 48 |

| mart-aerie-school-calendar-refresh | 78 |

| mart-aerie-school-source-directories-refresh | 67 |

| mart-aerie-schools-data-sheet-refresh | 78 |

| mart-aerie-xo-contractor-refresh | 46 |

| mart-aerie-hubspot-refresh | 189 |

- The new regression test, added to each runner's tests/test_redshift_client.py, uses a fake monotonic clock:

- The run deadline bounds the statement, and the statement only reaches ABORTED 40 s after it is cancelled.

- With reserves 0 and 30, it asserts a clean TimeoutError ("cancellation reached ABORTED"), not UnknownStatementOutcomeError, and that the cancel was sent at least one cancel window before the run deadline.

- Negative control: against the origin/main clients of mart-aerie-camps-refresh and mart-aerie-hubspot-refresh, both cases fail (2 failed each).

- Ruff 0.15.22: ruff check pipelines and ruff format --check pipelines are clean.

## Not covered

- A separate client lineage has the same bug class, with the same shape: _deadline(reserve) and _cancel_and_confirm capped at run_deadline. It is in core-education-student-identity-refresh, core-education-student-school-year-snapshots, mart-education-finalsite-billing-refresh and mart-education-forecast-refresh.

- It is left out because the client differs: _deadline also refuses a statement with under 5 s left, and in mart-education-forecast-refresh it runs after submission. Changing its budget needs its own review. Worth a follow-up ticket.

- Clients already correct or of a different design, checked and unchanged:

- mart-aerie-retention-refresh and hc-forecast-refresh already subtract the cancel window.

- The CANCEL_GRACE clients (mart-aerie-education-financials-refresh, mart-school-performance-*) already reserve it.

- core-education-enrollment, hubspot-admissions-funnel and hubspot-core-tables do not cap confirmation at the run deadline.

- core-education-ontology-refresh and netsuite-saved-search-refresh cancel best-effort, without confirmation.

- No deploy is needed beyond the normal release. Each Lambda picks the fix up on its next deploy. No DDL, no config.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1575 — Forecast V2: publish End-of-Year enrollment forecast @vvp-trilogy  approved

## Summary

- publish the End-of-Year forecast and its eleven supported mart fields

- add session-end-bounded, mutually exclusive departure and graduate operands

- carry forward the usable prior-year End-of-Year converted remaining count unchanged

- commit the approved formula document from docs/admission/enrollment-projection-formula.md unchanged

## Validation

- dbt test --select int_admissions_forecast_v3_next_year_calculation

- dbt test --select int_admissions_forecast_start_year_transfer_out_exclusion

- isolated PR-schema build: changed models built successfully; all new End-of-Year tests passed (broader run: 419 pass, 5 existing data warnings, 15 missing-unselected-relation errors)

- dbt parse --no-partial-parse

Closes #1574

#137 — AI-927: Repair Pi task streaming and lifecycle @ashwanth1109  no labels

Pi tasks could stop after ten seconds without an event, lose streamed events to concurrent RPC queries, or display an unfinished tool round as complete. This repair keeps turns running until the agent settles, acknowledges turn submission promptly, and gives live and restored conversations consistent status and message identities.

The change also shares the shell environment setup with Codex, adapts implementation clarification instructions to Pi's available tools, and keeps empty research artifacts waiting for input. Captured run/template evidence and reproductions are documented in docs/PI_RUNTIME_REPAIR.md; published templates are unchanged.

## Business Value

Makes Pi research and implementation tasks reliable during long model/tool calls and follow-up interactions, preventing premature completion and broken workflow progression.

## Implementation

- Route RPC responses by request ID independently of streamed events, and remove the command timeout from event waits.

- Resolve durable turn IDs from Pi's actual message events and finish turns in a background worker.

- Use settled session history as the authoritative outcome, including recovered errors, interrupted tool rounds, and active retries.

- Update the Pi smoke fixture to match real event ordering and add focused runtime, workflow, and conversation regressions.

## Validation

- Native library suite: 286 passed, 2 existing ignored.

- Final focused Pi runtime rerun: 16 passed.

- Conversation/workflow JavaScript suites: 102 passed.

- Smoke fixture Python suite: 30 passed.

- git diff --check passed.

418 distinct tests passed. Packaged desktop UI and live TrueFoundry validation were not run; the installed app was not changed.

## Implementation Effort

Estimated 2–3 engineering days to diagnose the protocol/lifecycle failures, implement the fixes, and add regression coverage without AI assistance.

## Linear

[AI-927: Repair Pi task streaming, lifecycle, and workflow recovery](https://linear.app/builder-team/issue/AI-927/repair-pi-task-streaming-lifecycle-and-workflow-recovery)

#133 — AI-918: Improve and consolidate search input UI @ashwanth1109  no labels

## Demo

![AI-918 search input smoke test](https://github.com/AI-Builder-Team/Shipyard/blob/29673313b0d215625b607c0721e36912a93cf557/.github/smoke-evidence/AI-918-search-input.png?raw=true)

## Summary

- Add a reusable, theme-backed SearchInput with custom search chrome, native-decoration suppression, accessible labeling, and a focus-preserving clear button.

- Migrate Architecture node search and project-icon search while preserving their local filtering, keyboard navigation, autofocus, and focus restoration behavior.

- Add focused DOM coverage for search semantics, clear behavior, filtering, no-results status, and Escape handling.

## Linear

https://linear.app/builder-team/issue/AI-918/improve-and-consolidate-search-input-ui

## Validation

- pnpm test:search-input

- node scripts/test-architecture.mjs

- pnpm exec tsc --noEmit

- pnpm theme:check

- pnpm build

- pnpm test:architecture *(native portion is blocked in this checkout because src-tauri/resources/codex is absent)*

## Scope notes

Unrelated URL, chat, settings, form, and numeric inputs remain unchanged.

#2094 — feat(aerie-a8): forecast-input parity marts, Q2 projections + Q3 coming-year projections + app conversion (SURTR-1545) @kevalshahtrilogy  approvedmercy-allow-critical

## Summary

A8 plan unit U13 (Linear SURTR-1545). Adds three parity marts to the mart-aerie-admissions-refresh runner (U03, PR 2084), so Aerie's forecast-input reads can move onto the Surtr Gateway. Each candidate is Aerie's SQL at e366e27d0, changed only as A8 plan §3.1 allows.

| Mart (Gateway slug, registered in U04) | Aerie query | Change to Aerie's SQL | Rows today |

|---|---|---|---|

| aerie_admissions_program_projection | queryAllProgramProjections, Q2 (educrm.ts:1179-1188) | TRIM(BOTH '"' FROM program_id::varchar) = $1 lifted into filter_program_id. Every school year is published. | 180 |

| aerie_admissions_coming_year_projection | queryComingYearProjections, Q3 (educrm.ts:1295-1315) | LOWER(TRIM(BOTH '"' FROM program_code::varchar)) IN (aliases) lifted into filter_program_code_lower. The ORDER BY is dropped because it compares with $1; see below. | 180 |

| aerie_admissions_app_conversion | queryAppConversionFacts (educrm.ts:1416-1432) | None. It is a whole-table GROUP BY with no per-program predicate. | 4,295 |

What each writer does

- Each sp_refresh_* follows the U03/U12 pattern:

- pins the observed sales-educrm-mart-sync run, which must be the latest run, at most 24 hours old, with a matching row count;

- builds the candidate in temp tables;

- publishes with an atomic DELETE + INSERT;

- fails closed on an empty candidate or a duplicate mart_row_id.

- Strict-parse assertion (Q2, Q3). Aerie parses both with parseRowsStrict, so one bad row fails the program. Each procedure refuses to publish a row Aerie's schema would reject:

- Q2: a NULL program_code;

- Q3: a NULL program_code or program_name;

- either: a NaN in any DOUBLE PRECISION column. z.coerce.number() rejects NaN, nullable or not, and Redshift float8 can hold NaN.

Everything else parses. BIGINT and INTEGER always coerce. The DATE and projection_version parsers accept NULL. A NULL school_year or projection_year becomes 0 under z.coerce.number(). I checked all of this against Aerie's installed Zod 3.25.76.

- App conversion is not asserted. Aerie uses safeParseRows, which drops a bad row instead of failing. Rows it would drop are published as-is, so the Gateway path drops the same rows as the legacy read.

- NULL keys (Q2, Q3). Rows with a NULL key are counted and not published. The procedure checks that published rows plus NULL-key rows equal the source snapshot.

- Q3 ordering. The ORDER BY puts the program's first alias ahead of the others, so it can't be evaluated without $1. It is dropped from the candidate. U17's reader must re-apply projection_year ASC, (key = aliases[0]) first, projection_version DESC before the first-row-per-year rule. projection_version is published for that rule.

- The upstream program_id migration (2026-09-28) needs no change here. sales-educrm-mart-sync/ddl/20260928_add_program_id_to_coming_year_projection.sql added program_id to coming_year_projection. Aerie's explicit select list doesn't include it, and neither does the mart. A test pins that.

- Registration check. U04's seed-gateway-aerie-a8.ts registers these three slugs on exactly these table names, ordered by mart_row_id. It needs no fix. A test asserts the mapping.

- Legacy-code variants are out of scope. queryAllProgramProjectionsByLegacyCode and queryComingYearProjectionsByLegacyCode are reached only from refreshEnrollmentPipeline, which has no production caller. runRefreshCycle uses the stable-identity path (per-program-refresh.ts:106).

- Files

- DDL: pipelines/cdk/sql/mart_education/050-055. U13 owns 050-059.

- The three procedures are appended, contiguously, to REFRESH_PROCEDURES.

- Expect trivial rebase conflicts with PRs 2089, 2090 and 2092 in pipeline.json, apply_ddl.py, test_pipeline_contract.py and the README.

## Business Value

- Aerie's per-program refresh and its forecast refresh read these three EduCRM tables directly today. They feed:

- enrollmentProjections, which public API v2 serves;

- coming-year projections;

- the application-conversion facts that publish with the forecast.

- These marts let those reads move onto Surtr-owned, lineage-stamped tables served through the Gateway, with parity that can be proven (EXCEPT = 0). That is a required step for A8's cutover of the analytics-worker's Redshift reads.

- The strict-parse guards stop Surtr from publishing a row that would make Aerie fail a program's refresh. A bad upstream row fails loudly in Surtr and the previous publication stays live.

## Manual Effort Estimate

About 10 hours of focused work by hand, with the U03/U12 patterns to copy. That covers:

- reading and pinning three Aerie queries and their Zod schemas;

- designing the lifted keys and the Q3 ordering edge;

- writing about 1,100 lines of DDL and procedures, the reconciliation SQL and the contract tests;

- running the read-only evidence.

Keval: please confirm or adjust this number.

## Testing / evidence

- uv run pytest: 179 passed. That includes tests/test_sql_contracts_forecast_inputs.py, which:

- rebuilds each candidate from Aerie's pinned SQL by the allowed edits only;

- pins each column's source type (from svv_columns);

- derives the NULL and NaN assertions from Aerie's pinned row schemas;

- checks the U04 registration.

- A mutation check broke 8 separate things, one at a time; the tests caught every one:

- a dropped NaN guard, a dropped NULL guard and a wrong column type;

- a dropped column and a dropped ORDER BY;

- a changed IN list;

- weakened revokes;

- a shortened occurrence order.

- ruff check and ruff format --check (0.15.22) are clean.

- uv run python scripts/apply_ddl.py --dry-run applies this runner's files in order:

  -- 006_v_aerie_educrm_observed_publication.sql: 3 statement(s)

-- 007_aerie_admissions_refresh_writer_mutex.sql: 5 statement(s)

-- 008_aerie_admissions_program.sql: 14 statement(s)

-- 009_sp_refresh_aerie_admissions_program.sql: 4 statement(s)

-- 010_aerie_admissions_program_directory.sql: 13 statement(s)

-- 011_sp_refresh_aerie_admissions_program_directory.sql: 4 statement(s)

-- 050_aerie_admissions_program_projection.sql: 11 statement(s)

-- 051_sp_refresh_aerie_admissions_program_projection.sql: 4 statement(s)

-- 052_aerie_admissions_coming_year_projection.sql: 12 statement(s)

-- 053_sp_refresh_aerie_admissions_coming_year_projection.sql: 4 statement(s)

-- 054_aerie_admissions_app_conversion.sql: 15 statement(s)

-- 055_sp_refresh_aerie_admissions_app_conversion.sql: 4 statement(s)

- Read-only reconciliation (prod Redshift as CQL_download_OM, SELECTs only, counts only, 2026-09-29). It compares Aerie's SQL with the procedure's candidate as a multiset (Query 1 of each reconciliation/*.sql):

| Check | Aerie rows | Candidate rows | aerie − candidate | candidate − aerie |

|---|---|---|---|---|

| Q2, 3 programs (one at a time) | 2 / 2 / 2 | 2 / 2 / 2 | 0 | 0 |

| Q3, 3 programs (one at a time) | 2 / 2 / 2 | 2 / 2 / 2 | 0 | 0 |

| Q2, all 90 programs (Aerie's predicate bound once per key) | 180 | 180 | 0 | 0 |

| Q3, all 90 programs (Aerie's predicate bound once per key) | 180 | 180 | 0 | 0 |

| App conversion, whole table | 4,295 | 4,295 | 0 | 0 |

- Other read-only checks on today's data

- Each procedure's exact candidate and mart_row_id expressions give unique ids: 180/180, 180/180 and 4,295/4,295.

- 0 strict-parse violations and 0 NULL keys.

- The observed EduCRM run is the latest run for all three tables.

- Every G1 program's aliases are identical, program code = source code, so each program has one alias.

- No DDL was applied, and nothing was written to Redshift.

## Keval steps

1. Apply the DDL before merging, because the trigger fires every 30 minutes once released. Run uv run python scripts/apply_ddl.py, which applies files 006-011 and 050-055. The files are idempotent; 006-011 are already live.

2. Merge. After the release, run the pipeline on demand and check results.

3. Run Query 2 of each reconciliation/*.sql. It must return 0 both ways:

- Q2: psql -v program_id=…;

- Q3: psql -v program_code_alias_1=… -v program_code_alias_2=…;

- app conversion: no variables.

4. Personal-data decision. aerie_admissions_app_conversion holds HubSpot contact ids with application dates and stage flags, but no names, emails or dates of birth. It isn't in the plan's §9 list of six PII marts. I treated it as personal data: owner-only (REVOKE ALL), with COMMENTs. Please decide whether it joins the §9 PII sign-off before Gateway exposure.

## Not covered

- The Aerie readers: U14 (seam) and U17 (Q2/Q3/app-conversion gateway paths, including re-applying Q3's ORDER BY).

- The per-program shadow compare (U21).

- The legacy-code Q2/Q3 variants, which have no production caller.

- The de-EduCRM repoint (A8 plan §10). coming_year_projection has no Surtr-native equivalent.

- Query 2 of the reconciliation, which needs the DDL apply and a run.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2093 — 088-finance-audit-gl-access @mwrshah  no labels

## Requested access

- Finance explicitly requested that edu_reader_user can read the body rows of these eight staging_finance_netsuite tables to independently investigate NetSuite GL postings: raw_account, raw_transaction_accounting_line, raw_transaction, raw_transaction_line, raw_accounting_period, raw_subsidiary, raw_classification, and raw_department.

- The account and accounting-line tables alone cannot establish a posting's period, status, subsidiary, class, or department. The six other tables provide those joins. Full-row reads across subsidiaries are intentional, not a defect or a request for row filtering. These are the complete allowlist: no other tables, schema-wide SELECT, or default privileges are granted. Timeback, central charges, and deferred-revenue-specific work are outside this change.

## Implementation and live verification

- Ordered schema DDL grants schema USAGE and table SELECT only to edu_reader_user on the eight named tables; the tests lock the principal and exact table list. The SQL includes an effective-access read-back. A warehouse administrator applies this DDL out of band; CDK does not automatically execute it.

- The grants are already applied on production finance_dw. Direct SELECT ... LIMIT 0 probes connected as edu_reader_user succeeded on all eight tables. The deployed mapping from individual Finance audit keys to this shared database profile is not verified.

- Validation: Ruff check and format on pipelines; 5 DDL tests passed.

#136 — Release: Shipyard 0.6.6 @ashwanth1109  no labels

## Summary

- Update the authoritative app version to 0.6.6.

- Add the reviewed public release notes for the contextual companion, Codex/Pi task selection, provider-neutral conversations, and feature replay.

## Business Value

- Delivers the reviewed Shipyard improvements in the next desktop update with public notes that explain the user-visible changes.

## Implementation Effort

- Low: metadata-only release change; CI performs validation, packaging, signing, and publication.

## Test plan

- [x] pnpm test:release

- [x] git diff --check

#134 — AI-883: Build an end-to-end feature replay runner @ashwanth1109  no labels

## Demo

![AI-883 smoke test](https://github.com/AI-Builder-Team/Shipyard/blob/d8caf3c/.smoke-evidence/AI-883-smoke-test.png?raw=true)

## Summary

- Add the fixture-only run_feature_replay command and frontend API for baseline/candidate end-to-end graph replay.

- Compose isolated node runs with dependency outputs, repository patches/untracked files, artifacts, checkpoints, retries, approvals, divergence, and source-integrity capture.

- Document the replay output contract and add focused feature-replay coverage.

## Linear

https://linear.app/builder-team/issue/AI-883/build-an-end-to-end-feature-replay-runner

## Testing

- pnpm test:feature-replay

- pnpm test:replay-runner

- pnpm build

- Full Rust library suite: 270 passed, 2 ignored

- Smoke harness: 29 passed; happy smoke run verified research-ready and stopped successfully

- pnpm theme:check

## Notes

This PR is intentionally draft and does not merge the change.

#135 — AI-917: Integrate Pi adapter into task and conversation UI @ashwanth1109  no labels

## Demo

![AI-917 Pi adapter smoke test](https://github.com/AI-Builder-Team/Shipyard/blob/6f485a1/docs/smoke-evidence/AI-917/image-1.png?raw=true)

## Summary

- Expose Codex/Pi selection, Pi credential readiness, and the OS-permission acknowledgement in task creation.

- Route Pi workflow conversations through provider-neutral engine commands with composite engine/thread stream identity, durable history, live events, text turns, and interruption.

- Gate conversation controls by engine capabilities while preserving Codex behavior.

- Normalize Pi user, reasoning, assistant, and tool transcript items and reconcile durable history after completion and reopen.

- Negotiate provider-neutral peer capabilities so older Shipyard owners receive an actionable compatibility error instead of an unsupported agent method.

## Linear

https://linear.app/builder-team/issue/AI-917/integrate-pi-adapter-into-task-creation-and-conversation-ui

## Tests

- pnpm build

- pnpm theme:check

- pnpm test:conversation-store

- pnpm test:messages

- pnpm test:task-workspace

- pnpm test:chat

- pnpm test:recovery

- pnpm test:instances

- pnpm test:workflow

- pnpm test:smoke

- cargo test --manifest-path src-tauri/Cargo.toml --lib agent

Draft for review; not merged.

#1568 — Forecast V2: publish expected enrollment arrival inputs @vvp-trilogy  approved

## Summary

- publish the five-field historical January expected-enrollment source contract

- preserve all-or-none nullability and current-year-only population

- add deterministic fixture, reconciliation, uniqueness, non-negative, and Alpha Austin coverage

## Validation

- poetry run dbt parse --no-partial-parse

- git diff --check

- warehouse-backed dbt tests delegated to PR CI (local worktree has no Redshift credentials)

Closes #1567

#2084 — feat(aerie-a8): mart-aerie-admissions-refresh runner + G1 parity marts (SURTR-1536, SURTR-1537) @kevalshahtrilogy  approvedmercy-allow-critical

## Summary

This is A8 plan unit U03 in full: SURTR-1536, plus SURTR-1537, which Keval squash-merged into this branch from #2085. It adds the runner every A8 admissions parity mart plugs into, the EduCRM provenance view, and both G1 marts.

- New runner pipelines/runners/mart-aerie-admissions-refresh (Lambda, bundling: true, src/requirements.txt).

- Triggers: on_pipeline_success of sales-educrm-mart-sync (every 30 min) and mart-aerie-hubspot-refresh (6-hourly), with forward_upstream_execution_context.

- Upstream check: the handler confirms with describe_execution that the upstream run SUCCEEDED on that pipeline's state machine. On-demand runs are also accepted, and may name a subset of procedures.

- What it runs: it CALLs each procedure in REFRESH_PROCEDURES (one line in pipeline.json; later units append to it), then checks the published mart read-only: non-empty, unique mart_row_id, one source_run_id.

- Failure handling: a failed procedure makes the run partial_failure and the others still run. If every procedure fails, the run fails. It also fails if a write's acceptance cannot be ruled out: every Data API write carries a ClientToken, and an ambiguous CALL submission raises UnknownStatementOutcomeError rather than being recorded as a procedure failure.

- DDL is in pipelines/cdk/sql/mart_education/ as 006-011, the schema's numbered out-of-band deploy directory. scripts/apply_ddl.py applies exactly those six files, in order.

- 006 v_aerie_educrm_observed_publication gives, per EduCRM table, the latest SUCCESS or PARTIAL sales-educrm-mart-sync run whose per-table result is success, plus the latest started run of any status. It is built with UNPIVOT over results_by_table. rows_loaded must be a plain integer.

- 007 is the shared owner-only writer mutex.

- 008/009 aerie_admissions_program (slug aerie-admissions-program) is Aerie's queryPrograms SQL, copied verbatim from reference.ts:132-148, including the CURRENT_DATE prior-year predicate. It adds school_year, canonical_source_run_id and lineage. Its procedure requires the observed EduCRM run to be the latest started run (so no later writer can have replaced the table), rows_loaded to equal the full snapshot, and the run to be under 24 hours old.

- 010/011 aerie_admissions_program_directory (slug aerie-admissions-program-directory) is Aerie's queryHubspotPrograms SQL, copied verbatim from hubspot.ts:75-98. Its lineage is the directory's single hubspot_publication_run_id / hubspot_source_published_at.

- Both procedures build the candidate in temp tables and publish with DELETE + named-column INSERT in the CALL transaction, with no TRUNCATE. They fail closed on an empty candidate, bad lineage, and a duplicate mart_row_id.

- mart_row_id is the MD5 of the key and its occurrence number, ordered by every published column.

- Unresolved and duplicate rows are published as-is, so Aerie's own mapper still throws exactly as on the legacy read.

- PIPELINE §13 exception (WAREHOUSE §2.2 and PIPELINE §7) is documented in the README with owner, risk, controls and follow-on. Reader grants are left to the DBA; the DDL only protects the writer.

## Business Value

- It completes the Surtr side of G1, the first A8 group. programs is the hard prerequisite of every Aerie runRefreshCycle, and it plus programDirectory (which also feeds A1) can now be read through the Surtr Gateway with Aerie's unchanged row mappers. This moves Aerie off direct EduCRM reads from the EC2 analytics worker, which is the SURTR-735 quarterly commitment and what unblocks tearing the worker down.

- Parity is provable. Tests pin Aerie's SQL, the reconciliation EXCEPT is 0 in both directions, and every row carries lineage.

- Later A8 units reuse this foundation. The runner, the provenance view and the mutex serve about 10 more EduCRM marts, which then add only a procedure and one REFRESH_PROCEDURES entry.

## Manual Effort Estimate

About 18 focused hours (roughly 2.5 days) to build by hand without AI: reading both Aerie queries and the EduCRM run-log shape, the SUPER-aware provenance view, two procedures, the runner and client hardening, the tests and the reconciliation. Keval: please confirm or adjust.

## Testing / evidence

- uv run pytest: 85 passed. This covers the handler, the pipeline contract (every env var src/ reads is declared), the SQL contracts, apply_ddl and the Redshift client, including ambiguous-submission cases.

- The SQL-contract tests pin both Aerie queries and assert that each procedure's candidate is exactly that SQL plus the appended lineage columns, and that the reconciliation uses the same candidate.

- The pinned text matches Aerie origin/main byte for byte, at both 92fd47992 and 3d4fe1a97.

- Ruff 0.15.22: ruff check and ruff format --check are clean. CI is green.

- Read-only reconciliation was run with psql as CQL_download_OM (SELECT only), query 1 of each file (Aerie SQL vs the procedure's candidate):

| mart | aerie rows | candidate rows | aerie − candidate | candidate − aerie |

|---|---|---|---|---|

| aerie_admissions_program | 90 | 90 | 0 | 0 |

| aerie_admissions_program_directory | 113 | 113 | 0 | 0 |

Query 2, against the published marts, errors with "relation does not exist" as expected, because no DDL has been applied.

- Negative controls on the same multiset shape:

- dropping a row gives 1 / 0;

- duplicating a row gives 1 / 1.

- Procedure expressions run read-only:

- The view resolves 49 tables, with 0 unparsable counts. For mart_all_program, the observed run equals the latest run, and rows_loaded 180 = snapshot 180.

- The integer parser returns 180 → 180, and 89.9 / 1.8e2 / "180" → NULL.

- Stamping gives 90 and 113 unique mart_row_ids.

- The directory has 1 publication run id and 0 rows missing lineage.

- scripts/apply_ddl.py --dry-run (6 files, 43 statements, in apply order) is below:

<details><summary>apply_ddl.py --dry-run output</summary>

-- 006_v_aerie_educrm_observed_publication.sql: 3 statement(s)

-- Shared observed-provenance helper for the mart-aerie-admissions-refresh

-- procedures (A8 plan §3.1 rule 2). sales-educrm-mart-sync republishes each

-- EduCRM table in its own transaction and reports per-table outcomes only in

-- its run summary, so there is no atomic row-level lineage link. This view

-- exposes, per Redshift table, the latest SUCCESS/PARTIAL run whose per-table

-- result is 'success', together with the latest sales-educrm-mart-sync run to

-- have started (any status, RUNNING included: CreateRunRecord inserts it

-- before the run touches a table).

--

-- A procedure pins the observed run only when it is also that latest run: no

-- later writer can then have replaced the table, so, read in the same

-- transaction snapshot, the table is that run's publication. The procedure

-- still requires rows_loaded to equal the snapshot count.

-- Generalises the inline pattern in sp_refresh_aerie_program_directory

-- (mart-aerie-hubspot-refresh/ddl/20260819_incident_stopgap_duplicate_school_year.sql).

--

-- Placement (WAREHOUSE §10): writer-internal helper for mart procedures only.

-- It captures no source extraction (not staging) and holds no business

-- meaning (not core), so it lives beside its only readers in mart_education.

CREATE OR REPLACE VIEW mart_education.v_aerie_educrm_observed_publication AS

WITH educrm_runs AS (

SELECT

run_id,

status,

started_at,

ended_at,

output_summary

FROM staging_other.pipeline_runs_prod

WHERE pipeline_id = 'sales-educrm-mart-sync'

),

latest_run AS (

SELECT

run_id::VARCHAR(36) AS latest_run_id,

status::VARCHAR(20) AS latest_run_status

FROM (

SELECT

run_id,

status,

ROW_NUMBER() OVER (ORDER BY started_at DESC, run_id DESC) AS recency

FROM educrm_runs

) ranked

WHERE recency = 1

),

completed_runs AS (

SELECT

run_id,

ended_at,

CASE WHEN CAN_JSON_PARSE(output_summary) THEN JSON_PARSE(output_summary) END AS output_summary_super

FROM educrm_runs

WHERE status IN ('SUCCESS', 'PARTIAL')

AND ended_at IS NOT NULL

),

table_results AS (

SELECT

run.run_id,

run.ended_at,

table_key,

table_result

FROM completed_runs run, UNPIVOT run.output_summary_super.results_by_table AS table_result AT table_key

),

successful_table_results AS (

SELECT

table_key::VARCHAR(256) AS educrm_table,

table_result.redshift_table::VARCHAR(256) AS redshift_table,

run_id::VARCHAR(36) AS observed_run_id,

ended_at AT TIME ZONE 'UTC' AS observed_completed_at,

-- Only a plain non-negative integer is a row count; anything else

-- (89.9, 1.8e2, a string) becomes NULL so the procedure fails closed

-- instead of a cast truncating it into a matching count.

CASE

WHEN JSON_TYPEOF(table_result.rows_loaded) = 'number'

AND JSON_SERIALIZE(table_result.rows_loaded) ~ '^[0-9]{1,18}$'

THEN JSON_SERIALIZE(table_result.rows_loaded)::BIGINT

END AS observed_row_count,

ROW_NUMBER() OVER (

PARTITION BY table_result.redshift_table::VARCHAR(256)

ORDER BY ended_at DESC, run_id DESC

) AS recency

FROM table_results

WHERE table_result.status::VARCHAR(32) = 'success'

AND table_result.redshift_table::VARCHAR(256) IS NOT NULL

)

SELECT

observed.educrm_table,

observed.redshift_table,

observed.observed_run_id,

observed.observed_completed_at,

observed.observed_row_count,

latest.latest_run_id,

latest.latest_run_status

FROM successful_table_results observed

CROSS JOIN latest_run latest

WHERE observed.recency = 1;

COMMENT ON VIEW mart_education.v_aerie_educrm_observed_publication IS

'Purpose: writer-internal observed provenance for Aerie admissions parity mart procedures (mart-aerie-admissions-refresh); not a consumer contract. Grain: one EduCRM Redshift table written by sales-educrm-mart-sync. Key: redshift_table. observed_run_id is the latest SUCCESS or PARTIAL run whose per-table result is success; latest_run_id/latest_run_status describe the most recently started run of any status. This is observed provenance, not an atomic publication link: a procedure must require observed_run_id = latest_run_id and observed_row_count = the snapshot it reads, in one transaction.';

ALTER TABLE mart_education.v_aerie_educrm_observed_publication

OWNER TO "CQL_download_OM";

-- 007_aerie_admissions_refresh_writer_mutex.sql: 5 statement(s)

-- Owner-only mutex shared by every mart-aerie-admissions-refresh procedure.

-- Each procedure locks it first, so overlapping runs (both upstream triggers

-- can fire together) publish one at a time. It is locked instead of the

-- target mart, which is locked only for the final DELETE + INSERT.

CREATE TABLE IF NOT EXISTS mart_education.aerie_admissions_refresh_writer_mutex (

lock_scope VARCHAR(64) NOT NULL

)

DISTSTYLE ALL;

COMMENT ON TABLE mart_education.aerie_admissions_refresh_writer_mutex IS

'Owner-only writer mutex for the mart-aerie-admissions-refresh stored procedures. It contains no data and is not a consumer contract.';

REVOKE ALL ON mart_education.aerie_admissions_refresh_writer_mutex FROM PUBLIC;

REVOKE ALL ON mart_education.aerie_admissions_refresh_writer_mutex FROM GROUP team_engineers;

ALTER TABLE mart_education.aerie_admissions_refresh_writer_mutex

OWNER TO "CQL_download_OM";

-- 008_aerie_admissions_program.sql: 14 statement(s)

-- Canonical DDL for mart_education.aerie_admissions_program (A8 unit U03,

-- Gateway slug aerie-admissions-program). Sole writer:

-- mart_education.sp_refresh_aerie_admissions_program().

--

-- Query-shaped parity mart: columns are exactly the output aliases of Aerie's

-- queryPrograms SQL (sync/src/analytics/queries/reference.ts), with source

-- types kept (SUPER included) so pg and Gateway readers serialise them the

-- same way. school_year and canonical_source_run_id are lineage additions.

CREATE TABLE IF NOT EXISTS mart_education.aerie_admissions_program (

program_public_id VARCHAR(64),

source_program_id VARCHAR(256),

source_program_code SUPER,

program_code VARCHAR(512),

program_name VARCHAR(512),

is_expansion BOOLEAN,

owner_name SUPER,

grade_levels SUPER,

school_address SUPER,

school_status SUPER,

show_in_dashboard BOOLEAN,

school_year BIGINT,

canonical_source_run_id VARCHAR(128),

mart_row_id VARCHAR(32) NOT NULL,

source_run_id VARCHAR(128) NOT NULL,

source_published_at TIMESTAMPTZ NOT NULL,

refreshed_at TIMESTAMP NOT NULL,

created_by VARCHAR(128) NOT NULL,

PRIMARY KEY (mart_row_id)

)

DISTSTYLE ALL

SORTKEY (mart_row_id);

COMMENT ON TABLE mart_education.aerie_admissions_program IS

'Purpose: Surtr publication of the rows Aerie''s queryPrograms reads (EduCRM mart_all_program LEFT JOIN core_education.dim_program on the HubSpot Program id), so Aerie can read them through the Surtr Gateway with its unchanged row mapper. Grain: one output row of that SQL for the previous calendar school year (school_year = EXTRACT(YEAR FROM CURRENT_DATE) - 1, evaluated at refresh); normally one EduCRM program_id. Key: mart_row_id (MD5 of source_program_id plus its occurrence number). source_program_id is expected unique but not enforced: duplicate or unresolved rows (NULL program_public_id) are published as-is so Aerie''s own identity checks still fail closed. Lineage: source_run_id/source_published_at are the observed sales-educrm-mart-sync run (v_aerie_educrm_observed_publication), not an atomic publication link. Sensitive data: owner_name holds a staff member''s name. Full snapshot replaced atomically by mart_education.sp_refresh_aerie_admissions_program.';

COMMENT ON COLUMN mart_education.aerie_admissions_program.program_public_id IS

'core_education.dim_program.program_id for the active HubSpot Program whose hubspot_program_id equals source_program_id; NULL when unresolved.';

COMMENT ON COLUMN mart_education.aerie_admissions_program.source_program_id IS

'EduCRM program_id as text (TRIM(BOTH ''"'' FROM program_id::varchar)); this is the HubSpot Program id.';

COMMENT ON COLUMN mart_education.aerie_admissions_program.program_code IS

'dim_program.program_name (the canonical program code). NULL when unresolved.';

COMMENT ON COLUMN mart_education.aerie_admissions_program.program_name IS

'dim_program.display_name. NULL when unresolved.';

COMMENT ON COLUMN mart_education.aerie_admissions_program.school_year IS

'EduCRM school_year (starting calendar year) of the published row. Lineage addition; not read by Aerie.';

COMMENT ON COLUMN mart_education.aerie_admissions_program.canonical_source_run_id IS

'dim_program.hubspot_publication_run_id of the joined canonical Program row; NULL when unresolved. Lineage addition; not read by Aerie.';

COMMENT ON COLUMN mart_education.aerie_admissions_program.mart_row_id IS

'Deterministic row key: MD5 of source_program_id and its occurrence number. Gateway orderBy for total-order paging. Not stable across a change to the row''s key.';

COMMENT ON COLUMN mart_education.aerie_admissions_program.source_run_id IS

'Observed sales-educrm-mart-sync run_id whose mart_all_program rows_loaded equalled the snapshot this publication read.';

COMMENT ON COLUMN mart_education.aerie_admissions_program.source_published_at IS

'End time (UTC) of the observed sales-educrm-mart-sync run.';

ALTER TABLE mart_education.aerie_admissions_program

OWNER TO "CQL_download_OM";

-- Writer protection only. Reader access is provisioned by the Redshift DBA

-- (PIPELINE §13); the Surtr Gateway reads as the owner.

REVOKE INSERT, UPDATE, DELETE, TRUNCATE

ON mart_education.aerie_admissions_program FROM PUBLIC;

REVOKE INSERT, UPDATE, DELETE, TRUNCATE

ON mart_education.aerie_admissions_program FROM GROUP team_engineers;

-- 009_sp_refresh_aerie_admissions_program.sql: 4 statement(s)

-- Sole writer for mart_education.aerie_admissions_program (WAREHOUSE §7.1).

--

-- The candidate is Aerie's queryPrograms SQL, copied verbatim from

-- sync/src/analytics/queries/reference.ts:132-148 at Aerie 92fd47992. Two

-- changes only: MART_ALL_PROGRAM_YEAR_PREDICATE is expanded in place, and two

-- lineage columns are appended to the select list. Aerie's identity and

-- duplicate checks stay in Aerie's row mapper, so this procedure publishes

-- unresolved or duplicate rows as-is instead of rejecting them.

--

-- Fails closed on: no or incomplete observed EduCRM run; a later

-- sales-educrm-mart-sync run (any status, including one still running) that

-- could have republished the table since; an observation older than 24 hours;

-- a snapshot count that differs from the observed rows_loaded; an empty

-- candidate; and a duplicate mart_row_id. The DELETE + INSERT publish stays

-- inside the CALL transaction; never TRUNCATE (it commits implicitly).

CREATE OR REPLACE PROCEDURE mart_education.sp_refresh_aerie_admissions_program()

AS $$

DECLARE

v_observation_count BIGINT;

v_observed_run_id VARCHAR(36);

v_observed_completed_at TIMESTAMPTZ;

v_observed_row_count BIGINT;

v_latest_run_id VARCHAR(36);

v_latest_run_status VARCHAR(20);

v_snapshot_row_count BIGINT;

v_candidate_count BIGINT;

v_duplicate_count BIGINT;

v_school_year BIGINT;

v_refreshed_at TIMESTAMP;

BEGIN

LOCK TABLE mart_education.aerie_admissions_refresh_writer_mutex;

v_refreshed_at := GETDATE();

v_school_year := EXTRACT(YEAR FROM CURRENT_DATE) - 1;

SELECT COUNT(*) INTO v_observation_count

FROM mart_education.v_aerie_educrm_observed_publication

WHERE redshift_table = 'staging_education.sales_educrm_wh_mart_all_program'

AND educrm_table = 'educrm_wh.mart_all_program';

IF v_observation_count <> 1 THEN

RAISE EXCEPTION

'aerie_admissions_program: expected one successful EduCRM mart_all_program observation; found %',

v_observation_count;

END IF;

SELECT observed_run_id, observed_completed_at, observed_row_count, latest_run_id, latest_run_status

INTO v_observed_run_id, v_observed_completed_at, v_observed_row_count, v_latest_run_id, v_latest_run_status

FROM mart_education.v_aerie_educrm_observed_publication

WHERE redshift_table = 'staging_education.sales_educrm_wh_mart_all_program'

AND educrm_table = 'educrm_wh.mart_all_program';

IF NULLIF(BTRIM(v_observed_run_id), '') IS NULL

OR v_observed_completed_at IS NULL

OR v_observed_row_count IS NULL

OR v_observed_row_count <= 0 THEN

RAISE EXCEPTION

'aerie_admissions_program: EduCRM observation is incomplete (run %, completed %, rows %)',

v_observed_run_id, v_observed_completed_at, v_observed_row_count;

END IF;

-- The observed run must also be the most recently started EduCRM run.

-- Otherwise a later run (still running, failed, or one that failed this

-- table) may have republished it, and a matching row count would not prove

-- which run's rows are there. Read in this same transaction snapshot, no

-- later writer means the table is the observed run's publication.

IF v_latest_run_id IS NULL OR v_observed_run_id <> v_latest_run_id THEN

RAISE EXCEPTION

'aerie_admissions_program: EduCRM run % (status %) started after observed run %; the table may hold newer rows',

v_latest_run_id, v_latest_run_status, v_observed_run_id;

END IF;

-- sales-educrm-mart-sync runs every 30 minutes. An observation this old

-- can no longer vouch for the table it describes.

IF v_observed_completed_at < SYSDATE - INTERVAL '24 hours' THEN

RAISE EXCEPTION

'aerie_admissions_program: latest EduCRM observation % completed at % is older than 24 hours',

v_observed_run_id, v_observed_completed_at;

END IF;

-- Reconcile the full snapshot (every school year) to the observed run

-- before the year predicate narrows it.

SELECT COUNT(*) INTO v_snapshot_row_count

FROM staging_education.sales_educrm_wh_mart_all_program;

IF v_snapshot_row_count <> v_observed_row_count THEN

RAISE EXCEPTION

'aerie_admissions_program: EduCRM snapshot has % rows but observed run % reported %',

v_snapshot_row_count, v_observed_run_id, v_observed_row_count;

END IF;

DROP TABLE IF EXISTS tmp_aerie_admissions_program_query;

CREATE TEMP TABLE tmp_aerie_admissions_program_query AS

-- aerie-sql:begin

SELECT

canonical_program.program_id AS program_public_id,

TRIM(BOTH '"' FROM p.program_id::varchar) AS source_program_id,

p.program_code AS source_program_code,

canonical_program.program_name AS program_code,

canonical_program.display_name AS program_name,

p.is_expansion,

p.owner_name,

p.grade_levels,

p.school_address,

p.school_status,

p.show_in_dashboard,

-- A8 lineage additions (not in Aerie's select list):

p.school_year,

canonical_program.hubspot_publication_run_id AS canonical_source_run_id

FROM staging_education.sales_educrm_wh_mart_all_program p

LEFT JOIN core_education.dim_program canonical_program

ON canonical_program.hubspot_program_id = TRIM(BOTH '"' FROM p.program_id::varchar)

AND canonical_program.hubspot_source_presence_status = 'active'

WHERE p.school_year = EXTRACT(YEAR FROM CURRENT_DATE) - 1

-- aerie-sql:end

;

DROP TABLE IF EXISTS tmp_aerie_admissions_program;

CREATE TEMP TABLE tmp_aerie_admissions_program (LIKE mart_education.aerie_admissions_program);

INSERT INTO tmp_aerie_admissions_program (

program_public_id,

source_program_id,

source_program_code,

program_code,

program_name,

is_expansion,

owner_name,

grade_levels,

school_address,

school_status,

show_in_dashboard,

school_year,

canonical_source_run_id,

mart_row_id,

source_run_id,

source_published_at,

refreshed_at,

created_by

)

SELECT

q.program_public_id,

q.source_program_id,

q.source_program_code,

q.program_code,

q.program_name,

q.is_expansion,

q.owner_name,

q.grade_levels,

q.school_address,

q.school_status,

q.show_in_dashboard,

q.school_year,

q.canonical_source_run_id,

MD5(

'aerie_admissions_program|'

|| COALESCE('v' || q.source_program_id, 'n')

|| '|'

|| (ROW_NUMBER() OVER (

PARTITION BY q.source_program_id

-- Every output column, so rows that share a key are numbered

-- the same way on every refresh; only identical rows tie.

ORDER BY q.program_public_id, q.program_code, q.program_name,

JSON_SERIALIZE(q.source_program_code), q.is_expansion,

JSON_SERIALIZE(q.owner_name), JSON_SERIALIZE(q.grade_levels),

JSON_SERIALIZE(q.school_address), JSON_SERIALIZE(q.school_status),

q.show_in_dashboard, q.school_year, q.canonical_source_run_id

))::VARCHAR

),

v_observed_run_id,

v_observed_completed_at,

v_refreshed_at,

'mart-aerie-admissions-refresh/v1'

FROM tmp_aerie_admissions_program_query q;

SELECT COUNT(*) INTO v_candidate_count FROM tmp_aerie_admissions_program;

IF v_candidate_count = 0 THEN

RAISE EXCEPTION

'aerie_admissions_program: candidate is empty for school_year % (observed run %)',

v_school_year, v_observed_run_id;

END IF;

SELECT COUNT(*) INTO v_duplicate_count

FROM (

SELECT mart_row_id

FROM tmp_aerie_admissions_program

GROUP BY mart_row_id

HAVING COUNT(*) > 1

) duplicates;

IF v_duplicate_count <> 0 THEN

RAISE EXCEPTION 'aerie_admissions_program: candidate has % duplicate mart_row_id value(s)', v_duplicate_count;

END IF;

LOCK TABLE mart_education.aerie_admissions_program;

DELETE FROM mart_education.aerie_admissions_program;

INSERT INTO mart_education.aerie_admissions_program (

program_public_id,

source_program_id,

source_program_code,

program_code,

program_name,

is_expansion,

owner_name,

grade_levels,

school_address,

school_status,

show_in_dashboard,

school_year,

canonical_source_run_id,

mart_row_id,

source_run_id,

source_published_at,

refreshed_at,

created_by

)

SELECT

program_public_id,

source_program_id,

source_program_code,

program_code,

program_name,

is_expansion,

owner_name,

grade_levels,

school_address,

school_status,

show_in_dashboard,

school_year,

canonical_source_run_id,

mart_row_id,

source_run_id,

source_published_at,

refreshed_at,

created_by

FROM tmp_aerie_admissions_program;

IF (SELECT COUNT(*) FROM mart_education.aerie_admissions_program) <> v_candidate_count THEN

RAISE EXCEPTION 'aerie_admissions_program: post-publication row count mismatch';

END IF;

RAISE INFO 'aerie_admissions_program: published % row(s) from observed EduCRM run %',

v_candidate_count, v_observed_run_id;

DROP TABLE tmp_aerie_admissions_program;

DROP TABLE tmp_aerie_admissions_program_query;

END;

$$ LANGUAGE plpgsql SECURITY INVOKER;

ALTER PROCEDURE mart_education.sp_refresh_aerie_admissions_program()

OWNER TO "CQL_download_OM";

REVOKE ALL ON PROCEDURE mart_education.sp_refresh_aerie_admissions_program()

FROM PUBLIC;

GRANT EXECUTE ON PROCEDURE mart_education.sp_refresh_aerie_admissions_program()

TO "CQL_download_OM";

-- 010_aerie_admissions_program_directory.sql: 13 statement(s)

-- Canonical DDL for mart_education.aerie_admissions_program_directory (A8

-- unit U03, Gateway slug aerie-admissions-program-directory). Sole writer:

-- mart_education.sp_refresh_aerie_admissions_program_directory().

--

-- Thin parity mart: columns are exactly the output aliases of Aerie's

-- queryHubspotPrograms SQL (sync/src/analytics/queries/hubspot.ts) over

-- mart_education.aerie_program_directory_current, with source types kept.

CREATE TABLE IF NOT EXISTS mart_education.aerie_admissions_program_directory (

program_id VARCHAR(100),

hubspot_name VARCHAR(512),

display_name VARCHAR(512),

tuition NUMERIC(18, 4),

city VARCHAR(255),

state VARCHAR(100),

school_address VARCHAR(1000),

school_latitude NUMERIC(18, 8),

school_longitude NUMERIC(18, 8),

grade_levels VARCHAR(1000),

email VARCHAR(500),

contact_number VARCHAR(100),

enrollment_deposit VARCHAR(255),

application_fee VARCHAR(255),

school_year_start VARCHAR(256),

school_year_end VARCHAR(256),

website VARCHAR(2000),

school_summary VARCHAR(65535),

maxio_site_id VARCHAR(255),

canonical_source_run_id VARCHAR(128),

mart_row_id VARCHAR(32) NOT NULL,

source_run_id VARCHAR(128) NOT NULL,

source_published_at TIMESTAMPTZ NOT NULL,

refreshed_at TIMESTAMP NOT NULL,

created_by VARCHAR(128) NOT NULL,

PRIMARY KEY (mart_row_id)

)

DISTSTYLE ALL

SORTKEY (mart_row_id);

COMMENT ON TABLE mart_education.aerie_admissions_program_directory IS

'Purpose: Surtr publication of the rows Aerie''s queryHubspotPrograms reads (mart_education.aerie_program_directory_current LEFT JOIN core_education.dim_program on the HubSpot Program id), so Aerie can read them through the Surtr Gateway with its unchanged row mapper. Grain: one output row of that SQL; normally one active HubSpot Program. Key: mart_row_id (MD5 of program_id plus its occurrence number); program_id is expected unique but not enforced, so Aerie''s own checks still see any duplicate. Lineage: source_run_id/source_published_at are the directory''s hubspot_publication_run_id/hubspot_source_published_at. Sensitive data: email and contact_number are school contact points and can identify staff. Full snapshot replaced atomically by mart_education.sp_refresh_aerie_admissions_program_directory.';

COMMENT ON COLUMN mart_education.aerie_admissions_program_directory.hubspot_name IS

'COALESCE(dim_program.program_name, directory program_code).';

COMMENT ON COLUMN mart_education.aerie_admissions_program_directory.display_name IS

'COALESCE(dim_program.display_name, directory program_name).';

COMMENT ON COLUMN mart_education.aerie_admissions_program_directory.school_year_start IS

'Directory school_year_start DATE cast to text (YYYY-MM-DD), as Aerie selects it.';

COMMENT ON COLUMN mart_education.aerie_admissions_program_directory.school_year_end IS

'Directory school_year_end DATE cast to text (YYYY-MM-DD), as Aerie selects it.';

COMMENT ON COLUMN mart_education.aerie_admissions_program_directory.canonical_source_run_id IS

'dim_program.hubspot_publication_run_id of the joined canonical Program row; NULL when unmatched. Lineage addition; not read by Aerie.';

COMMENT ON COLUMN mart_education.aerie_admissions_program_directory.mart_row_id IS

'Deterministic row key: MD5 of program_id and its occurrence number. Gateway orderBy for total-order paging.';

COMMENT ON COLUMN mart_education.aerie_admissions_program_directory.source_run_id IS

'aerie_program_directory_current.hubspot_publication_run_id; the procedure requires exactly one value per snapshot.';

COMMENT ON COLUMN mart_education.aerie_admissions_program_directory.source_published_at IS

'aerie_program_directory_current.hubspot_source_published_at of that publication.';

ALTER TABLE mart_education.aerie_admissions_program_directory

OWNER TO "CQL_download_OM";

-- Writer protection only. Reader access is provisioned by the Redshift DBA

-- (PIPELINE §13); the Surtr Gateway reads as the owner.

REVOKE INSERT, UPDATE, DELETE, TRUNCATE

ON mart_education.aerie_admissions_program_directory FROM PUBLIC;

REVOKE INSERT, UPDATE, DELETE, TRUNCATE

ON mart_education.aerie_admissions_program_directory FROM GROUP team_engineers;

-- 011_sp_refresh_aerie_admissions_program_directory.sql: 4 statement(s)

-- Sole writer for mart_education.aerie_admissions_program_directory

-- (WAREHOUSE §7.1).

--

-- The candidate is Aerie's queryHubspotPrograms SQL, copied verbatim from

-- sync/src/analytics/queries/hubspot.ts:75-98 at Aerie 92fd47992, with three

-- lineage columns appended to the select list. Lineage is carried from the

-- upstream Surtr mart (mart-aerie-hubspot-refresh), which stamps every

-- directory row with one accepted HubSpot publication.

--

-- Fails closed on: an empty candidate, missing or mixed upstream lineage, and

-- a duplicate mart_row_id. The DELETE + INSERT publish stays inside the CALL

-- transaction; never TRUNCATE (it commits implicitly).

CREATE OR REPLACE PROCEDURE mart_education.sp_refresh_aerie_admissions_program_directory()

AS $$

DECLARE

v_candidate_count BIGINT;

v_lineage_run_count BIGINT;

v_lineage_published_count BIGINT;

v_lineage_invalid_count BIGINT;

v_duplicate_count BIGINT;

v_source_run_id VARCHAR(128);

v_source_published_at TIMESTAMPTZ;

v_refreshed_at TIMESTAMP;

BEGIN

LOCK TABLE mart_education.aerie_admissions_refresh_writer_mutex;

v_refreshed_at := GETDATE();

DROP TABLE IF EXISTS tmp_aerie_admissions_program_directory_query;

CREATE TEMP TABLE tmp_aerie_admissions_program_directory_query AS

-- aerie-sql:begin

SELECT

directory.program_id,

COALESCE(canonical_program.program_name, directory.program_code) AS hubspot_name,

COALESCE(canonical_program.display_name, directory.program_name) AS display_name,

directory.tuition,

directory.city,

directory.state,

directory.school_address,

directory.latitude AS school_latitude,

directory.longitude AS school_longitude,

directory.grade_range AS grade_levels,

directory.school_email AS email,

directory.school_phone AS contact_number,

directory.enrollment_deposit,

directory.application_fee,

directory.school_year_start::varchar,

directory.school_year_end::varchar,

directory.website,

directory.school_summary,

directory.maxio_site_id,

-- A8 lineage additions (not in Aerie's select list):

canonical_program.hubspot_publication_run_id AS canonical_source_run_id,

directory.hubspot_publication_run_id AS directory_publication_run_id,

directory.hubspot_source_published_at AS directory_source_published_at

FROM mart_education.aerie_program_directory_current directory

LEFT JOIN core_education.dim_program canonical_program

ON canonical_program.hubspot_program_id = directory.program_id

AND canonical_program.hubspot_source_presence_status = 'active'

-- aerie-sql:end

;

SELECT COUNT(*) INTO v_candidate_count FROM tmp_aerie_admissions_program_directory_query;

IF v_candidate_count = 0 THEN

RAISE EXCEPTION 'aerie_admissions_program_directory: candidate is empty';

END IF;

SELECT COUNT(DISTINCT directory_publication_run_id),

COUNT(DISTINCT directory_source_published_at),

SUM(CASE

WHEN NULLIF(BTRIM(directory_publication_run_id), '') IS NULL

OR directory_source_published_at IS NULL THEN 1

ELSE 0

END),

MIN(directory_publication_run_id),

MIN(directory_source_published_at)

INTO v_lineage_run_count, v_lineage_published_count, v_lineage_invalid_count,

v_source_run_id, v_source_published_at

FROM tmp_aerie_admissions_program_directory_query;

IF v_lineage_run_count <> 1 OR v_lineage_published_count <> 1 OR v_lineage_invalid_count <> 0 THEN

RAISE EXCEPTION

'aerie_admissions_program_directory: directory lineage is mixed or incomplete (% run id(s), % published_at value(s), % row(s) missing lineage)',

v_lineage_run_count, v_lineage_published_count, v_lineage_invalid_count;

END IF;

DROP TABLE IF EXISTS tmp_aerie_admissions_program_directory;

CREATE TEMP TABLE tmp_aerie_admissions_program_directory (LIKE mart_education.aerie_admissions_program_directory);

INSERT INTO tmp_aerie_admissions_program_directory (

program_id,

hubspot_name,

display_name,

tuition,

city,

state,

school_address,

school_latitude,

school_longitude,

grade_levels,

email,

contact_number,

enrollment_deposit,

application_fee,

school_year_start,

school_year_end,

website,

school_summary,

maxio_site_id,

canonical_source_run_id,

mart_row_id,

source_run_id,

source_published_at,

refreshed_at,

created_by

)

SELECT

q.program_id,

q.hubspot_name,

q.display_name,

q.tuition,

q.city,

q.state,

q.school_address,

q.school_latitude,

q.school_longitude,

q.grade_levels,

q.email,

q.contact_number,

q.enrollment_deposit,

q.application_fee,

q.school_year_start,

q.school_year_end,

q.website,

q.school_summary,

q.maxio_site_id,

q.canonical_source_run_id,

MD5(

'aerie_admissions_program_directory|'

|| COALESCE('v' || q.program_id, 'n')

|| '|'

|| (ROW_NUMBER() OVER (

PARTITION BY q.program_id

-- Every output column, so rows that share a key are numbered

-- the same way on every refresh; only identical rows tie.

ORDER BY q.hubspot_name, q.display_name, q.tuition, q.city, q.state,

q.school_address, q.school_latitude, q.school_longitude,

q.grade_levels, q.email, q.contact_number, q.enrollment_deposit,

q.application_fee, q.school_year_start, q.school_year_end,

q.website, q.school_summary, q.maxio_site_id, q.canonical_source_run_id

))::VARCHAR

),

v_source_run_id,

v_source_published_at,

v_refreshed_at,

'mart-aerie-admissions-refresh/v1'

FROM tmp_aerie_admissions_program_directory_query q;

SELECT COUNT(*) INTO v_duplicate_count

FROM (

SELECT mart_row_id

FROM tmp_aerie_admissions_program_directory

GROUP BY mart_row_id

HAVING COUNT(*) > 1

) duplicates;

IF v_duplicate_count <> 0 THEN

RAISE EXCEPTION

'aerie_admissions_program_directory: candidate has % duplicate mart_row_id value(s)',

v_duplicate_count;

END IF;

LOCK TABLE mart_education.aerie_admissions_program_directory;

DELETE FROM mart_education.aerie_admissions_program_directory;

INSERT INTO mart_education.aerie_admissions_program_directory (

program_id,

hubspot_name,

display_name,

tuition,

city,

state,

school_address,

school_latitude,

school_longitude,

grade_levels,

email,

contact_number,

enrollment_deposit,

application_fee,

school_year_start,

school_year_end,

website,

school_summary,

maxio_site_id,

canonical_source_run_id,

mart_row_id,

source_run_id,

source_published_at,

refreshed_at,

created_by

)

SELECT

program_id,

hubspot_name,

display_name,

tuition,

city,

state,

school_address,

school_latitude,

school_longitude,

grade_levels,

email,

contact_number,

enrollment_deposit,

application_fee,

school_year_start,

school_year_end,

website,

school_summary,

maxio_site_id,

canonical_source_run_id,

mart_row_id,

source_run_id,

source_published_at,

refreshed_at,

created_by

FROM tmp_aerie_admissions_program_directory;

IF (SELECT COUNT(*) FROM mart_education.aerie_admissions_program_directory) <> v_candidate_count THEN

RAISE EXCEPTION 'aerie_admissions_program_directory: post-publication row count mismatch';

END IF;

RAISE INFO 'aerie_admissions_program_directory: published % row(s) from HubSpot publication %',

v_candidate_count, v_source_run_id;

DROP TABLE tmp_aerie_admissions_program_directory;

DROP TABLE tmp_aerie_admissions_program_directory_query;

END;

$$ LANGUAGE plpgsql SECURITY INVOKER;

ALTER PROCEDURE mart_education.sp_refresh_aerie_admissions_program_directory()

OWNER TO "CQL_download_OM";

REVOKE ALL ON PROCEDURE mart_education.sp_refresh_aerie_admissions_program_directory()

FROM PUBLIC;

GRANT EXECUTE ON PROCEDURE mart_education.sp_refresh_aerie_admissions_program_directory()

TO "CQL_download_OM";

</details>

## Keval steps

1. Apply the DDL to prod before merging. A merge reaches production within the hour, and the EduCRM trigger then fires every 30 minutes. Run cd pipelines/runners/mart-aerie-admissions-refresh && uv run python scripts/apply_ddl.py (runs as CQL_download_OM; applies 006-011 in order). Using the usual numbered out-of-band SQL deploy for those files is equivalent.

2. Merge. Mercy withholds auto-approve on pipelines/cdk/ paths, so this needs a human approval. Once released, run mart-aerie-admissions-refresh on demand. Expect status: success, with results showing 90 program rows (source_run_id = the latest EduCRM run) and 113 directory rows (source_run_id = the HubSpot publication).

3. Run both reconciliation files. Query 2 must return 0 in both directions.

4. Reader access (optional). Reader access for non-owners is a DBA grant.

## Not covered

- Other A8 units: Gateway registration (U04, SURTR-1533, merged) and the Aerie read gate with its purge guard (U05).

- The procedure bodies have not been executed in Redshift, because no DDL was applied. Their candidate SELECTs, the view, the guards and the stamping expressions were run read-only instead. A behavioural rollback test needs a Redshift sandbox; Mercy deferred this as coverage.

- Behaviour inherited from Aerie: the prior-year predicate is evaluated at refresh, so rows can lag by one refresh (≤30 min) after 1 January. Aerie's 2028 predicate limitation also applies.

- When the program procedure fails, then succeeds 30 minutes later. This happens if an EduCRM run is in flight when mart-aerie-hubspot-refresh triggers, or if the latest EduCRM run failed mart_all_program. The mart keeps its last publication meanwhile. Procedure failures are PARTIAL (amber, throttled), not paging.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2082 — feat(gateway): register all 24 A8 Gateway sources in an aerie-a8 entity (SURTR-1533) @kevalshahtrilogy  approved

## Summary

A8 unit U04 (Linear [SURTR-1533](https://linear.app/builder-team/issue/SURTR-1533), part of SURTR-735). It registers every A8 parity mart as a Gateway source in one batch, so Aerie needs one new key for all of A8.

- New Surtr/src/seed-gateway-aerie-a8.ts, modelled on seed-gateway-aerie.ts, plus a seed:gateway-aerie-a8 script in Surtr/package.json.

- It registers exactly the 24 frozen A8 slugs. Each one:

- points at mart_education.<slug with - replaced by _>;

- is read-only (supportedAccess: ["read"]);

- has orderBy: "mart_row_id" and no dateColumn or whereExtra, so Aerie pages each mart in full with a total order.

- The 24 sources go in a new entity, aerie-a8. The seed never writes the existing aerie entity: the only member delete is scoped to aerie-a8's id.

- Ownership guards on every upsert.

- The sources upsert only updates rows this seed created (setWhere created_by = seed-gateway-aerie-a8-script). Every row is then read back and checked before the entity is touched. A slug that someone else already registered fails the run; it is never repointed.

- The aerie-a8 entity upsert carries the same setWhere. If another creator owns an aerie-a8 entity, RETURNING comes back empty and the seed throws inside the transaction before the member delete, so that entity and its members are left untouched.

### How the seed behaves

- It runs by hand, not in CD. Nothing in this PR runs it on deploy.

- Registering a source isn't a grant. A key reads a source only if it was minted with that grant. An entity is just a way to select many sources when creating a key (gateway/entities.ts).

- Sources whose tables don't exist yet return errors until each mart lands. A missing table makes /gateway/{source} answer 500 internal. Nothing reads these sources before then, because every Aerie A8 gate defaults to legacy.

- It's safe to re-run. Sources upsert on slug, the entity upserts on slug, and only aerie-a8's members are replaced wholesale.

The 24 slugs, with the unit that builds each mart:

| Group | Slugs | Built by |

|---|---|---|

| G1 | aerie-admissions-program, aerie-admissions-program-directory | U03 |

| G2/G3 community | aerie-admissions-community-conversion, aerie-admissions-community-deposit | U06 |

| G4 expenses | aerie-expense-transaction, aerie-expense-vendor-classification | U07 |

| G2 per-program | aerie-admissions-pipeline-student, aerie-admissions-community-metric | U08 |

| G2 per-program | aerie-admissions-enrollment-cohort, aerie-admissions-pipeline-deposit, aerie-admissions-enrollment-transfer | U12 |

| G2/G3 forecast inputs | aerie-admissions-program-projection, aerie-admissions-coming-year-projection, aerie-admissions-app-conversion | U13 |

| G3 marketing | aerie-admissions-marketing-event, aerie-admissions-marketing-event-contact | U15 |

| G2 admissions pipeline | aerie-admissions-pipeline-detail, aerie-admissions-pipeline-tenant-crosswalk | U16 |

| G3 marketing | aerie-admissions-shadow-day-event, aerie-admissions-weekly-deposit | U18 |

| G6 SIS | aerie-sis-enrollment-rollup-input, aerie-sis-enrollment-member | U19 |

| G6 Forecast V2 | aerie-admissions-forecast-v2, aerie-admissions-forecast-v2-grade-operand | U24 |

Source descriptions flag the marts that carry PII: community conversion and deposits, per-program pipeline students, enrollment, deposits and transfers, marketing event contacts, pipeline detail, SIS members, and expense vendor names and memos.

## Business Value

A8 moves Aerie's runRefreshCycle Redshift reads onto Surtr marts. This is part of taking down Aerie's EC2 workers, the SURTR-735 quarter commitment. This PR is the registration step that every A8 shadow and cutover reader depends on. Doing it as one frozen batch means:

- Keval does one key mint for the whole project instead of one per wave;

- the 20+ Aerie and Surtr units can build against fixed slugs in parallel.

The ownership guard and read-back mean a hand-run prod seed can't silently repoint a source that someone else registered.

## Manual Effort Estimate

About 4 hours of focused time without AI, for Keval to confirm or adjust. That covers:

- mapping 24 slugs to their marts and domains from the A8 plan, and writing their descriptions and PII notes;

- the seed with the ownership guard and read-back;

- a mocked-DB test harness that honours the upsert semantics;

- checking all of it.

## Testing / evidence

- npx vitest run test/gateway (from Surtr/): 3 files, 61 tests passed. The 14 new tests in test/gateway/seed-gateway-aerie-a8.test.ts follow PR 2069's mocked-DB pattern. They assert:

- exactly the 24 frozen slugs, each once;

- mart_education.<slug with - replaced by _> for every slug;

- read-only, orderBy mart_row_id, and no dateColumn or whereExtra, both on insert and in the re-run update set;

- buildDeclarativeTableSql gives SELECT * FROM mart_education.<t> ORDER BY mart_row_id LIMIT … OFFSET … for every slug;

- no slug collides with seed-gateway-aerie.ts, seed-gateway-ai-spend.ts or the custom sources;

- only aerie-a8 is upserted, its upsert is guarded on created_by, and member deletes and inserts are scoped to aerie-a8 (the aerie entity is never written);

- all 24 sources are read members of aerie-a8;

- a re-run over its own rows converges;

- a slug that another creator already registered is left unchanged, and the run exits 1 before any entity write;

- an aerie-a8 entity that another creator owns is left unchanged, and its members are never deleted or replaced.

- A mutation check confirmed the tests fail when any of these is broken: the entity slug is set to aerie, either setWhere guard is removed, a table name is wrong, the source ownership check is dropped, or the empty-RETURNING check is dropped.

- pnpm test:unit: 54 files, 769 tests passed (round 1).

- Surtr/node_modules/.bin/tsc --noEmit -p Surtr: clean. The new test file is also clean under an ad-hoc tsconfig that includes it.

- npm --prefix Surtr run lint (biome check src): clean. The test file is Biome-formatted.

- The seed was not run against any database.

## Keval steps

1. After merge, run the seed once U03, U06, U07 and U08 are deployed, so the first slugs are readable. Run pnpm seed:gateway-aerie-a8 from Surtr/, with a .env pointing at the prod app database.

2. Mint one new Aerie Gateway key with the existing aerie grants plus all 24 aerie-a8 slugs. In key creation, select the entities aerie and aerie-a8.

- Set it as SURTR_GATEWAY_API_KEY in Aerie's EC2 .env.

- Keep a local-testing copy for dry-runs.

- Revoke the old key after the swap.

## Not covered

- The marts themselves (U03, U06–U08, U12, U13, U15, U16, U18, U19, U24) and their DDL applies.

- Warehouse SELECT grants, in case the prod Gateway's REDSHIFT_DB_USER isn't CQL_download_OM (plan §9).

- PII sign-off for exposing the PII marts on the Gateway (plan §9 D2). Registering them exposes nothing until a key is granted them.

- If U01 retires Q4, aerie-admissions-community-metric stays registered but unbuilt. Removing it is a follow-up.

- Nothing in Aerie reads these slugs yet. The readers arrive in U05 and later units.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1562 — Forecast V2: bound milestone conversions by enrollment date @vvp-trilogy  approved

## Summary

- expose the accepted HubSpot enrollment date on the canonical admissions deal

- derive the effective enrollment date with the program-session start fallback

- bound historical milestone conversion numerators by that date without changing application cohorts

- cover null, before, exact-boundary, and after-boundary dates across all three milestones

## Validation

- git diff --check

- poetry run dbt parse --no-partial-parse

- focused unit tests selected successfully; execution deferred to credentialed dbt CI

Closes #1561

#2081 — fix(collections): skip 42DS forecast view @sanketghia  approved

## Summary

- Add 42ds to the explicit out-of-scope view set so the new Q4 sheet is skipped case-insensitively.

- Preserve the fail-closed behavior for all other unrecognized views.

- Add regression coverage for case/whitespace normalization and unknown labels.

## Validation

- uv run --offline --extra dev pytest — 57 passed.

#2068 — fix(education): trim Schools Data Sheet cells like JavaScript .trim() (SURTR-1524) @kevalshahtrilogy  approved

## Summary

SURTR-1524. mart_education.sp_refresh_aerie_schools_data_sheet trimmed Schools Data Sheet content with BTRIM, which only strips ASCII spaces. The Aerie parser it replaces uses JavaScript's .trim(), which also strips non-breaking spaces, tabs, newlines and other Unicode whitespace. As a result, one published cell (with a non-breaking space at its edge) differs from what Aerie produces today.

- Every content trim now uses a PCRE REGEXP_REPLACE over exactly the .trim() character set: TAB, LF, VT, FF, CR, SPACE, NBSP, U+1680, U+2000–U+200A, U+2028, U+2029, U+202F, U+205F, U+3000 and the BOM.

- The pattern is anchored with \A / \z, not ^ / $. In Redshift's PCRE mode ^ and $ also match at every line break, so my first attempt stripped whitespace *inside* 51 multi-line cells. The comments and tests now pin this.

- row_label is still published raw, as the parser does. BTRIM stays only on the lineage ids and the post-build blank checks.

- created_by moves to mart-aerie-schools-data-sheet-refresh/v2, so rows written under the new rule are identifiable. Aerie's A2 reader only requires it to be non-blank.

## Business Value

A2 (the Schools Data Sheet) is next in line for the Aerie EC2 → Surtr cutover. Aerie PR 1528 / AERIE-2543 adds the Gateway reader and the shadow comparison. Without this fix, the shadow run would report one mismatch every hour indefinitely, and a flip to gateway would publish a value that differs from today's. With it, the Surtr mart matches Aerie exactly, so the shadow window can come back clean and the cutover can go ahead.

## Manual Effort Estimate

About 4 focused hours by hand, with no AI help. That covers finding the BTRIM vs .trim() gap, working out Redshift's PCRE support and its multi-line ^/$ behaviour, rewriting the procedure's trims, building the character-set and fixture tests, and verifying against real Redshift. *This is a proposal: Keval, please confirm or adjust.*

## Testing / evidence

- uv run pytest -q in the pipeline: 76 passed. ruff check and ruff format --check (pinned 0.15.22) are clean.

- New contract tests check that:

- both trim sites use the same pattern;

- its character class is exactly the ECMAScript .trim() set;

- it uses \A / \z and never ^ / $;

- no sheet content is trimmed with BTRIM;

- row_label stays raw.

- The Python reference mirror now trims with the exact JavaScript set. Python's str.strip() differs on NEL and the BOM, and a test proves that. New fixtures cover NBSP, tab/newline, whitespace-only labels, NEL/interior whitespace (kept) and multi-line cells.

- Real Redshift, read-only SELECTs only. I did not create or apply any DDL.

- Probe: the new trim matches .trim() on all 13 cases (NBSP, tab/newline, ideographic space/BOM/paragraph separator, all-whitespace, interior whitespace, NEL, zero-width space, four multi-line cases). BTRIM differs on 4 of them.

- Live publication, recomputed with the new rule and compared against the current mart:

| Check | Result |

|---|---|

| Cell values changed | 1 of 1,845 (the NBSP cell, one character removed) |

| Default names / hidden names / locations changed | 0 / 0 / 0 |

| Data rows included or excluded differently | 0 |

| Location row detected differently | none |

## Deploying and verifying

Merging and releasing this PR changes nothing in prod on its own. CD only deploys CDK and never applies DDL. The change takes effect when the procedure DDL is applied by hand, which needs Keval's approval.

1. Apply only the procedure file. Without arguments, the script also re-runs the table DDL.

   python scripts/apply_ddl.py ddl/sp_refresh_aerie_schools_data_sheet.sql --dry-run

python scripts/apply_ddl.py ddl/sp_refresh_aerie_schools_data_sheet.sql

2. Wait for the republish. The hourly raw sync (cron(17 * * * ? *)) triggers the mart republish within the hour.

3. Run the post-apply check (read-only). All three must hold. If the backslash escaping were somehow processed differently inside the stored procedure, the trim would silently do nothing, and this check is what catches that.

   SELECT

COUNT(DISTINCT created_by) AS writer_versions, -- expect 1

MAX(created_by) AS writer, -- expect .../v2

SUM(CASE WHEN cell_value <> BTRIM(cell_value) THEN 1 ELSE 0 END) AS ascii_padded, -- expect 0

SUM(CASE WHEN REGEXP_COUNT(cell_value, '\\A\\x{00A0}|\\x{00A0}\\z', 1, 'p') > 0

THEN 1 ELSE 0 END) AS nbsp_padded -- expect 0 (was 1)

FROM mart_education.aerie_schools_data_sheet;

An independent reviewer found strong evidence the escaping is safe. An existing prod procedure (core_education.sp_refresh_qtd_hc_posting_classification) already uses '\\(' with PCRE flags inside $$ and has run cleanly 13 times. It has still not been verified directly on this procedure, which is why this check exists.

## Not covered

- Applying the DDL to prod. This needs Keval's approval; see "Deploying and verifying" above.

- The A2 prod shadow window itself (Aerie side).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2069 — feat(gateway): register the Finalsite tenant directory for Aerie (SURTR-1523) @kevalshahtrilogy  approved

## Summary

SURTR-1523. This PR registers mart_education.aerie_finalsite_tenant_directory as the Gateway source aerie-finalsite-tenant-directory. It is the Surtr half of finishing A4, Aerie's school-source-directories EC2 task. The Aerie half is AI-Builder-Team/Aerie PR 1531 (AERIE-2547, which is blocked by this ticket).

- QuickBooks and SIS are already served over the Gateway and have been in prod shadow since 2026-09-23. Finalsite, the third source, was added later and has no Gateway path, so Aerie's direct-Redshift read can't be deleted until it does.

- The source is read-only and declarative. It is ordered by the mart's key, finalsite_tenant_slug, so offset paging gives a stable, total order. It is added to the aerie entity bundle the same way its siblings are.

- Registering is not a grant. No existing key gains access: request-time auth reads only each key's saved grants, and the aerie entity bundle is expanded only when a *new* key is created. Today the key tooling (keys.ts) can only create and revoke keys. So granting Finalsite to Aerie means one of two things:

- minting a new key, where an aerie-entity key gets all 15 aerie sources;

- editing gateway_keys by hand.

Either is Keval's call.

## Deploying and verifying

Merging and releasing changes nothing on its own: the seed is not part of CD or the container CMD. It takes effect when someone runs pnpm seed:gateway-aerie against the prod Gateway Postgres. That run:

- inserts aerie-finalsite-tenant-directory;

- re-upserts the other 14 aerie sources from code;

- deletes and re-inserts all 15 aerie entity members.

It would also re-create any aerie source that was deliberately deleted in prod, so check the prod list first.

Afterwards, check that gateway_sources has exactly one aerie-finalsite-tenant-directory row, ordered by finalsite_tenant_slug.

- The mart README now documents that all three directories are served over the Gateway.

## Business Value

A4 is the reference object for the whole Aerie EC2 → Surtr migration, and it is the closest to done. This registration is the last Surtr-side piece before A4 can move entirely onto the Gateway. After that, its direct-Redshift path, and eventually the worker task itself, can be retired.

## Manual Effort Estimate

About 3 focused hours by hand with no AI. That covers the source entry, confirming the mart's key and columns, the mocked-DB seed test, the README, and the read-only checks. *This is a proposal for Keval to confirm or adjust.*

## Testing / evidence

- New Surtr/test/gateway/seed-gateway-aerie.test.ts passes 9/9, and the full Surtr/test/gateway suite passes 47/47. The test records what the seed would write through a mocked DB connection, so nothing connects to Postgres or Redshift. It checks:

- the source's shape and its exact ORDER BY;

- that the order key matches the mart's documented Key:;

- membership in the aerie entity bundle;

- that slugs are unique;

- that all three directories order by real columns and serve NOT NULL lineage.

- tsc --noEmit -p Surtr is clean, and npm run lint (biome, 99 files) is clean.

- Read-only Redshift checks, done in the A4 lane:

- The Finalsite mart has 59 rows, 59 distinct slugs, and a single run id.

- A query shaped like the Gateway's returns the same 59 rows as Aerie's direct query, with 0 differences either way.

- History: the lane agent couldn't run git in its own Surtr worktree, so it left this change as a patch. I applied it on a fresh branch from origin/main and re-ran the tests above.

## Not covered

- Running the seed against the real Gateway DB, and granting the source to Aerie's key. Both are for Keval. After the grant, Finalsite needs its own clean shadow window in Aerie before the switch to gateway.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2080 — fix(education): make behavioral_events migration 006 drop+create idempotent (SURTR-1519) @kevalshahtrilogy  approved

## Summary

Follow-up to #2060 (merged), addressing Mercy's two blocking findings on the release PR (#2076):

Finding 1 — ddl/006 DROP + CREATE OR REPLACE VIEW (fixed). After an unconditional DROP VIEW, CREATE OR REPLACE VIEW looked unsafe as a disaster-recovery reference copy. I re-verified against the actual prod apply first: the Redshift Data API sub-statement history for the original batch (6ceae357-b79b-4c59-9d5f-12c6d0d65455) shows DROP VIEW then CREATE OR REPLACE VIEW both FINISHED — CREATE OR REPLACE VIEW creates a fresh view when none exists, same as PostgreSQL, and ddl/003/ddl/005 use the identical pattern successfully. Prod's applied state is correct and unchanged; nothing there needed fixing. The file itself still needed tightening as the reference copy: it now does DROP VIEW (unconditional, as before) then plain CREATE VIEW (not OR REPLACE). The real defensive value isn't "empty state" — the precondition guard already requires relation_count=1 (the pre-migration view must exist) before DROP ever runs — it's that plain CREATE VIEW refuses rather than silently overwriting an unreviewed definition if this file's CREATE statement is ever submitted on its own (a manual partial reapply, or a future edit that drops the DROP but keeps the CREATE), where CREATE OR REPLACE VIEW would not refuse. Added test coverage that exercises this mechanically: a fake Data API client with a minimal DROP/CREATE "does the view exist" state machine run through the real apply path, proving (a) a full drop+recreate succeeds against the documented precondition state, (b) plain CREATE VIEW fails closed against an unexpectedly-present view, (c) the same statement written as CREATE OR REPLACE VIEW would not — plus a direct assertion on the exact SQL text submitted to the fake batch API (independent of the statement-builder), so a reintroduced OR REPLACE fails a plain string check rather than relying only on review.

Finding 2 — consumer contract-hash mismatch window (attempted, blocked by a real preflight refusal; not forced). Before touching anything: checked sys_query_history for the last 7 days for DDL on sp_refresh_guide_roster_evidence/guide_roster_evidence (none) and for any running/queued query touching mart_education right now (none related). Snapshotted the live procedure owner, SECURITY DEFINER, function-level EXECUTE ACL, database-level EXECUTE ACL, and the pinned contract hash in the procedure body, all read as the documented migration identity (surtr_mart_education_guide_roster_owner) exactly as apply_ddl.py's own preflight would. Then ran the actual documented command:

REDSHIFT_DB_USER=surtr_mart_education_guide_roster_owner uv run python scripts/apply_ddl.py --apply ddl/005_20260928_source_contract.sql

It returned {"status":"fail","error":"MigrationPreflightError","statement_ids":[]} — no DDL was submitted. Root cause: apply_ddl.py's MIGRATION_DATABASE_FUNCTION_ACL_SQL check (added in #1991 alongside ddl/004) expects exactly one svv_database_privileges row for EXECUTE/team_engineers_global_permissions; as the documented owner identity it observes zero rows. I confirmed via admin (superuser, read-only) that the grant genuinely exists and is unchanged (finance_dw EXECUTE team_engineers_global_permissions role admin_option=False) — surtr_mart_education_guide_roster_owner is not a superuser (pg_user.usesuper=false) and is not a member of that role, so it cannot see a database-level grant it neither made nor received. This same check, run as this same non-superuser owner, evidently passed when ddl/004 was originally applied on 2026-09-21 (the check was introduced in that exact PR), so something about catalog visibility for this identity appears to have changed since — I could not determine why, and did not find matching sys_query_history DDL to explain it (possibly outside the retention window). Per instructions, I did not work around this: I did not run the apply as admin (a different identity than this migration is designed to run as, and doing so would leave the object owned/altered by the wrong identity), and I did not edit apply_ddl.py's preflight to drop or loosen the check. Nothing was written — reconfirmed after the attempt: the procedure owner, function ACL, and pinned hash (63a7a6fd…) are byte-for-byte unchanged from before the attempt, and behavioral_events is unchanged from the original 2026-09-28 apply. This is documented as a known blocker in contracts/schema-drift-2026-09-24.md for a human to resolve (either restore the owner's visibility into that grant, or have someone with visibility — or a reworked preflight — apply ddl/005).

Scope: pipelines/runners/guide-platform-raw-sync/ddl/006_behavioral_events_schema_additions.sql, its tests, and its recovery note. No consumer-repo files were changed (the consumer apply was attempted, not modified).

## Business Value

Closes out a blocking finding from the production-release review (#2076) so that release can proceed once Mercy re-clears it, without which none of the four fixes bundled in that release (including the time-sensitive perplexity-usage-pipeline fix, #2072) can reach production. Also converts a review comment about a hypothetical failure mode into an actually-verified fact (prod's original apply was correct) plus a genuinely stronger disaster-recovery artifact, and surfaces — rather than silently working around — a real, unexplained catalog-visibility gap blocking the Guide roster consumer's contract-pin update, which needed a human decision rather than a rushed fix under time pressure.

## Manual Effort Estimate

Roughly 2-3 focused hours: reproducing and reading the original Data API sub-statement history to settle the DROP/CREATE-OR-REPLACE question with evidence rather than assumption, redesigning the test coverage around a small SQL-aware fake client, and the multi-step live investigation (sys_query_history check, two rounds of grant snapshotting under different identities, the actual blocked apply attempt, and confirming nothing was written) for the consumer side. Proposed by Claude, Keval to confirm or adjust.

## Test plan

- [x] uv run pytest in pipelines/runners/guide-platform-raw-sync: 179 passed, 1 skipped (pre-existing skip, unrelated).

- [x] ruff check / ruff format --check (pinned 0.15.22, matching CI) clean on the touched runner dir.

- [x] New/changed tests: ddl/006's statement text no longer contains CREATE OR REPLACE VIEW; a fake-Data-API full-apply test asserts the exact submitted DROP VIEW / CREATE VIEW SQL text (not just that it matches the builder's own output); a dedicated fresh-precondition-state replay test; a dedicated fails-closed-vs-OR-REPLACE-would-not test.

- [x] Read-only live verification: sys_query_history checked for conflicting DDL before the consumer apply attempt; before/after snapshots of the consumer procedure's owner and function-level ACL (unchanged); before/after column dump of behavioral_events (unchanged); confirmed the AWS Redshift Data API sub-statement history for the original prod batch.

- [ ] ddl/005_20260928_source_contract.sql is still not applied — blocked, see Summary. A human needs to resolve the catalog-visibility gap (or apply it with appropriate visibility) before the Guide roster consumer's contract pin can be updated live.

Linear: SURTR-1519

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#3824 — fix(board-doc): enforce literal Khoros FY26 markers on main @marcusdAIy  approved

## Summary

- Backport the narrow Khoros FY26 marker consistency fix from production PR #3823 onto current main; no release-only changes or unrelated cherry-picks.

- Continue scanning both FY26 and FY'26 marker spellings to reject duplicate/ambiguous blocks, but require the reviewed B33/B57 markers to use the exact live FY26 - Current ... vs Previous ... title. The payload validator already requires that literal.

- Pin both approved live marker titles, reject FY'26 at either anchor and alternative duplicate markers, and clarify marker-versus-period-header spelling in the rollout gate. Period headers still use FY'26.

## Verification

- Focused hermetic Khoros suite: 106 passed.

- Ruff check and format check passed; Git diff check clean.

- The exact parser SHA256 1b376b0075cb81ab019205917854a87833585a0c73bd5492471cc6b550a56aba passed read-only production parser and paired renderer gates as part of PR #3823 verification. No financial values printed.

- Full board-doc suite is left to GitHub CI for this narrow backport.

Do not merge until GitHub CI and review pass. No deployment or Doc changes in this PR.

#3822 — fix(board-doc): Khoros Q4 paired financials and guarded narrative refresh @marcusdAIy  approved

## Summary

- Add Khoros Q4 2026 blank template with Plan Executive Summary and Prior Quarter Review & Goals scaffold. Reject prior-quarter generation until verified Q3 approval; do not infer Q3 narrative or send it to an LLM.

- Parse the pinned Khoros worksheet read-only into independent FY26 Hybrid and BU plan-on-plan blocks. Require reviewed anchors, headers, exact ordered 19 financial row labels, 16 columns, valid values, and boundaries. Reject partial or tampered sources; render both native tables in Financials, Hybrid first. Do not use unrelated blocks or generic IgniteTech sources.

- Restrict Khoros refresh to Financials only on both backend and add-on paths. Preserve blank/operator-authored narrative, skip commentary-number and narrative-review LLM passes, and fail closed on invalid sections/source payloads. Add hermetic parser, cold-start, refresh and add-on two-table tests.

## Verification

- Board-doc pytest suite: 5,150 passed, 2 deselected.

- Focused Khoros pytest: 102 passed. Add-on vitest run tests/section-targeting.test.js: 43 passed. Ruff check and format check passed; diff check clean.

- Exact integrated parser SHA256 1cf44e0c05a320f8966de18401bd30200211573c22247db215af7b859bd2d144 passed a deployed, read-only pinned-worksheet gate: approved Hybrid and BU source blocks each 20×16 with exact labels. Separate read-only native render gate produced exactly two 21×16 tables, Hybrid then BU; no unrelated sources/links. No amounts printed.

## Follow-up: restored-payload validation

- Shared pure validation now checks both source blocks (exact approved FY26 headers and 19 ordered labels, every amount, keys/shape, unsafe controls) in the reader, native renderer and data refresh before section/data mutation. Tampered DataPackage tests cover header, newline, label and amount; refresh cannot publish malformed payloads. Exact updated parser SHA256 9f9aa52f2ef57b2731f29c348a2194c4294af73f37724af07097a40dc97ae55f passed deployed read-only parser/render gates.

## Follow-up: financial boundary width

- Both FY26 block termination rows 55 and 79 now reject nonempty R+ data (in addition to pre-existing marker/header/body width checks). Tests cover R/T for both Hybrid and BU. Exact parser SHA256 fdfda4d12856ff64c01257e41127870a7934a88a04fa64964e8852947be8c4c0 passed the pinned-worksheet deployed read-only parser/render gates.

## Release hold

Implementation PR only. Do not merge, deploy, create/refresh/move a Doc, or change Google access based on this PR alone. Follow klair-api/budget_bot/board_doc/KHOROS_Q4_2026.md and separately verify owner and Doc permissions before controlled release.

#2079 — fix(aws-spend): map Totogi CapitaTFL account @caina-barbosa  approved

## Summary

Records the production account mapping correction for newly active Totogi account:

- 160135134587 → Prod-Totogi-CapitaTFL

Mapped to class = Totogi and bu = Totogi, projected from 2026-Q3 through the existing 2030-Q4 horizon.

## Incident & Why

The saas-budgeting-pipeline scheduled run for 2026-09-28 failed closed on the noncentral_charges ingest:

ValueError: account mapping is incomplete for 2026-Q3: ['160135134587']

All other ingests (docker, k8s, database_units, mapping, server_costs) published successfully.

## How the values were derived (evidence chain)

1. Master Payer: Account 160135134587 reports RDS costs under master payer 572481847476 (VDI).

2. Account Name: Queried AWS Cost Explorer linked-account metadata via the payer role (ESW-CO-ReadOnly-P2), returning:

- 160135134587 → Prod-Totogi-CapitaTFL

3. Class & BU: Mapped to class = Totogi and bu = Totogi, matching every other Prod-Totogi-* account in the mapping table (Prod-Totogi-CapitaSelfCare, Prod-Totogi-BSSMagic*, Prod-Totogi-VivaOntologyManaged).

4. Completeness: Querying the anti-join between core_finance.aws_spend_net_amortized_costs (RDS service, 2026-Q3) and core_finance.aws_spend_budget_account_mapping confirmed that across all master payers, exactly this single account was missing a mapping.

## Production remediation completed

1. Executed and verified in finance_dw: 18 rows inserted (1 account × 18 quarters, 2026-Q3 .. 2030-Q4).

2. Post-commit anti-join confirmed 0 unmapped 2026-Q3 RDS accounts remaining.

3. Triggered on-demand Step Functions execution manual-noncentral-totogi-capitatfl-20260928T215336Z:

- Status: SUCCEEDED

- Candidate count: 205 accounts

- Replaced 204 prior rows with 205 current rows

- Source max date: 2026-09-27

- Billable accounts: 149 / charge total: $149,000

- Mapping gap count: 0

#1559 — Restore Forecast V2 compatibility publication @vvp-trilogy  approved

## Summary

- temporarily restore aerie_milestone_v2 on the retained Forecast V2 compatibility surface

- unblock the existing Convex Forecast V2 publication guard

This is intentionally a one-line compatibility fix pending the V3 consumer migration.

#3821 — fix(board-doc): reject truncated Khoros Financials rows @marcusdAIy  approved

## Khoros Q4 source row-boundary fix

Follow-up to merged #3819 and required for release PR #3820. Mercy's production review found that a truncated 9–15-column body row could pass the source check and render as an incomplete Board Financials table. Reject every short body row before formatting, retaining the existing protection against nonempty cells past column 16 and trimming only empty overhang.

Verification: 22 focused source tests passed including new truncated-row fixture; Ruff lint/format and diff check passed. Exact patched reader validated the live reviewed P&Ls - Khoros Q4 block read-only: 20 rows, all exactly 16 columns, no amounts emitted. No production/Doc/access mutations.

#3819 — feat(board-doc): Q4 Khoros-only BU financial source and scoped owner access @marcusdAIy  approved

## Q4 Khoros first-class Board Doc (backend only)

- Add independent BusinessUnit.KHOROS, distinct persisted session identity and the same reviewed owner trio as IgniteTech (Eric Vaughan, Mohit Khosla, Zeeshan Khatri). Live account BU checks and Google Doc permissions remain separate.

- Limit Khoros to Q4 2026 blank-session creation. Persist explicit financial-source identity, disable prior-quarter cloning, and request only the Khoros Q4 BU Plan-on-Plan source; do not fetch generic IgniteTech P&Ls, Hybrid, ARR, retention, targets, or Brainlift.

- Pin the approved workbook ID through KHOROS_Q4_2026_WORKBOOK_ID; verify it matches the Q4 IgniteTech registry entry, then read only P&Ls - Khoros, exact FY26 BU marker and validated 16-column current/previous/variance layout. Missing or ambiguous source fails closed without rewriting the Doc.

- Update Q4 roster/preflight expectations and add a serial operator runbook. No add-on source changes.

## Verification

- 216 focused/regression tests passed; Ruff and git diff --check clean.

- Read-only live source probe against the currently linked Q4 workbook used the exact source reader and returned only validated, rows=20, columns=16; no financial figures emitted. Initial stricter header check rejected valid variance headings; fixed in 589eb6801 and reverified.

- No production deploy, session, Doc, permission, or registry mutation made by this PR.

## Release gate

Review/CI before merge; explicit production config and verified exact deploy before serial blank-session creation, Doc bind, folder placement and separate audited Editor grants. Target folder is the existing Q4 BU Budget Bot Docs folder; its 2026-09-28 read-only preflight found no anyone/domain sharing. No Marketplace release unless add-on source changes (none in this PR).

#3817 — fix(data-api): guide Education platform charge comparisons @mwrshah  approved

## Question 111 (Q111)

How many students did Education enroll in SY25/26, and how many does it plan to enroll in FY27? What did the group charge Education for the 2HrLearning/Timeback platform and central shared services in each period, both in total and per student, and which charges grew faster than enrollment?

## Changes

- Route Education central/platform charge questions to pinned consolidated budget and actual rows, including the vendor labels that distinguish the Timeback budget from licensing-fee actuals.

- Trace payer charges by month, compare central pools on a consistent BU perimeter, and keep provider receipts and costs separate.

- Cross-check the Alpha QuickBooks central recharge account against consolidated charges without allocating its central class to campuses. Discover school-year student counts and enrollment plans from warehouse SIS, HubSpot, and Finalsite sources; report cohort and coverage instead of requiring Finance to supply counts.

- Add an ontology contract test that checks the routing method without baking in the Q111 answer.

#1558 — Show demographic totals in collapsed mobile cards @YibinLongTrilogy  approved

## Summary

Collapsed Admissions Demographics cards on mobile previously showed only the grade or level and “Show details,” so users had to expand each row to see its counts. Show the existing totalBoys, totalGirls, totalUnknown, and totalStudents values in a two-column metric grid beneath the header, matching the visual layout of the Camps mobile cards.

### Changes

- chat/components/dashboards/admissions/demographics/demographics-mobile.tsx — Add a divider and Camps-style labeled metric tiles to DemographicsMobileCard. Keep the header expand control and expanded Current, Future, and Totals groups. Associate the preview with the expand button through aria-describedby.

- chat/components/dashboards/admissions/demographics/__tests__/demographics-mobile.test.tsx — Cover tile values, conditional Unknown visibility, accessible descriptions, level labels, number and zero formatting, expansion, and existing Current/Future drilldowns.

### Design Decisions

- Use the row’s existing totals and the dataset-level hasUnknownGender flag. No new calculation or backend request is needed.

- Show Boys and Girls on the first row, then Unknown and All on the second. When Unknown is absent, All spans both columns rather than leaving an empty tile.

- Keep the expand control in the header because expanded Current/Future cells have their own buttons.

- Keep zero totals visible as 0 in the preview, matching the display-only totals; the existing dash formatting for zero Current/Future drilldown cells is unchanged.

## Business value

Mobile users can scan and compare demographic counts across grades or levels without opening every card, while retaining access to the detailed breakdown and student drilldowns.

## Estimated manual effort

Estimated time to complete this work without AI: 1–2 hours.

## Test Plan

- [x] Focused Demographics mobile tests: 5 passed.

- [x] pnpm --dir chat typecheck passed for the final commit.

- [x] Biome check passed for both changed files.

- [x] Architecture boundary and test architecture checks passed.

- [ ] Review card layout at a narrow mobile viewport in both grade and level modes.

#1556 — Fix admissions program session qualification @vvp-trilogy  approved

## Summary

- restrict the deal-independent admissions session bridge to active schoolYear sessions

- require the associated HubSpot program to exist and remain active before canonical mapping

- add a hermetic regression fixture for summer sessions, archived sessions, archived programs, and missing program records

## Context

Follow-up to #1542. The canonical observation spine now reads milestone dates from

int_admissions_program_session; this bridge must apply the same session/program qualification

already enforced by int_admissions_deal so irrelevant sessions cannot create false ambiguity or

incorrect milestone dates.

## Validation

- poetry run dbt parse --no-partial-parse --profiles-dir /tmp/aerie-dbt-profile

- poetry run dbt ls --profiles-dir /tmp/aerie-dbt-profile --select int_admissions_program_session_qualifies_school_year_path

- git diff --check

The exact dbt unit fixture requires a warehouse connection and will run in isolated PR CI.

#1557 — Forecast V2: remove Marketing Planning Forecast from the report UI @vvp-trilogy  approved

## Summary

- remove the Marketing Planning Forecast column from the default Forecast V2 desktop and mobile UI

- remove the Marketing Planning Forecast expanded-row tab and its report-local rate synchronization

- keep operational milestone selection accessible when available and disable expansion when none is published

- leave APIs, shared contracts/calculations, Financials, and the legacy Forecast report unchanged

## Testing

- pnpm --dir chat exec vitest run components/dashboards/admissions/forecast/v2/__tests__/forecast-v2-report.test.tsx --maxWorkers=1

- pnpm exec biome check chat/components/dashboards/admissions/forecast/v2/forecast-v2-report.tsx chat/components/dashboards/admissions/forecast/v2/__tests__/forecast-v2-report.test.tsx

- pnpm lint:test-architecture

- pnpm typecheck

Closes #1555

#1554 — Capacity: publish and roll back record mode only once the DD request is approved (AERIE-2579) @marcusdAIy  approved

## Summary

Record-mode capacity publication and rollback now report success only once the governed Due Diligence change has actually landed (AERIE-2579).

updateDueDiligence never writes the card. It files a field-change request, which a person with operations.fieldChanges.approve approves. It applies at once only while approvals are paused. The run used to record published or rolledBack as soon as the request was filed.

- Authorized actor. Record mode files requests as the user named by the new CAPACITY_AUTOMATION_ACTOR_EMAIL. Publication checks that the user exists and holds operations.dueDiligence.write, and fails clearly if not. Nothing grants the capability; an admin assigns it through a role. capacityPublicationAllowed also requires the setting in record mode. Proposal mode still uses the system agent.

- Publication waits for approval. The run stores the request as ddFieldChangeRequestId and moves to a new awaitingDdApproval status. The minute recovery sweep checks the request every 5 minutes (_confirmCapacityDdRequest):

- Approved: published, with a readback note if the card has changed again since.

- Rejected or superseded: unresolved.

- Failed or missing: failed.

- A write that filed no request (nothing to change) is published only if the card already shows the numbers.

- Compensation. A request that does not land retracts the room table and floorplan documents and the attribution note, as does a publication that exhausts its retries. Registered document IDs are now saved right after registration so that is possible. Retraction uses removeDocument and deleteNote, which are hard deletes; the audit log keeps the record.

- Rollback works the same way. A record-mode rollback files a request for the prior capacities, waits in awaitingRollbackApproval, and becomes rolledBack only when the request is approved; then it removes the run's documents. A rejected rollback returns the run to published with the reason. The ordering, Complete-card, and compare-and-set guards run before the actor is resolved.

- Unset prior status (decided 2026-09-28). The governed write cannot clear a status, so when the card had none before publication, rollback restores the capacities and keeps the current status. It records that in rollbackNote.

- Proposal mode ends in proposed, not published. Proposal rollback removes the run's documents. Legacy proposal runs recorded as published can still be rolled back.

- A run waiting on either request counts as in flight, so the daily sweep does not start a second run for that site while a request is pending.

No UI reads capacity run statuses, so there are no front-end changes. Capacity automation is off by default and publication is proposal-mode everywhere, so nothing changes in any deployment until record mode is configured.

## Tests

- Real mutations with an authorized actor: publish files a pending request and the card is unchanged; a pending request reschedules; approval gives published; rejection or supersession gives unresolved with documents and note removed; a new sweep run waits while a request is pending.

- Rollback: approval gives rolledBack with documents removed; the unset-status case keeps the status and records rollbackNote; a rejected rollback leaves the run published; an actor without DD write is refused; a conflicting pending request surfaces as a failed rollback.

- Exhausted publication retries remove the registered documents and note.

- Contract transitions updated (for example, published can no longer go straight to rolledBack).

## Test plan

- [x] chat: vitest run convex/capacityAutomation.test.ts convex/capacityAutomation/config.test.ts (105 passed)

- [x] packages/contracts: capacity run and validation tests (56 passed)

- [x] tsc --noEmit (chat and convex), contracts typecheck, Biome, repo lint scripts

## Deployment note

Record mode now needs CAPACITY_AUTOMATION_ACTOR_EMAIL set to a user with operations.dueDiligence.write. Neither dev nor prod uses record mode today.

#1544 — Consolidate redundant dbt data tests @vvp-trilogy  approved

## Summary

- consolidate Forecast live-scenario and program/year assertions around shared setup

- consolidate five milestone-observation scans into one diagnostic contract test

- remove schema-order, column-name, and hard-coded provenance tests that do not validate data behavior

- document when singular SQL tests are warranted and when related assertions should share one suite

The substantive grain, source reconciliation, arithmetic, temporal-boundary, warning, and unit-fixture coverage remains in place. This reduces singular SQL tests from 119 to 106 and removes 135 net lines.

## Validation

- git diff --check

- poetry run dbt parse --no-partial-parse (65 models, 459 data tests, 21 unit tests)

- targeted dbt compile reached adapter initialization; warehouse compilation/execution is delegated to isolated PR CI because the local validation profile uses a non-routable host

#1545 — Forecast V2: retain observations with missing program years @vvp-trilogy  approved

## Summary

- drive milestone observations from every canonical program and the shared network year pair

- resolve HubSpot program sessions independently of deals through portal-scoped association type 265

- retain three unavailable rows for missing reference/history years while preserving existing observed-zero and conversion contracts

- document status precedence and add canonical coverage plus hermetic missing-year tests

## Validation

- poetry run dbt parse --no-partial-parse (dummy local profile; parse-only)

- poetry run dbt ls --select int_admissions_program_session int_admissions_milestone_metric_observations test_type:unit (66 models, 477 data tests, 22 unit tests discovered)

- latest-head GitHub CI, including the isolated Redshift dbt build and dbt test suites

- git diff --check

## Reconciliation evidence

- Canonical identities absent from int_program_identity remain outside this model and require upstream onboarding.

- assert_admissions_milestone_canonical_coverage.sql reconciles every current canonical identity to exactly three applications_submitted / application_to_enrolled rows at the network year/calculation-date grain.

- The configured read-only finance warehouse connection does not expose Aerie's sandbox_education objects, so a production named-program list could not be queried safely. The isolated PR warehouse suite passed the canonical coverage assertion and the missing-year field-contract assertion.

Closes #1542

#210 — Bound traces after redaction and complete requiredWhen (SINDRI-519) @marcusdAIy  approved

## Summary

The four follow-ups from Munawar's approving review of Sindri #205 (SINDRI-519).

- Trace bounds after redaction. redactWorkflowTrace re-applies boundTranscriptMessages to what it actually reports. Redaction can lengthen a field (Bearer x becomes Bearer [REDACTED]), so a tool call field under the 200 KB bound before redaction could exceed the server's 256 KB per-field limit afterwards.

- requiredWhen on start inputs. Run start now rejects a missing conditional start input, next to the existing required and non-blank checks. The message starts Invalid value for required input, so the HTTP edge returns 400.

- Authoring type check. requiredWhenError rejects an equals value whose type differs from the sibling field's type, such as the string "true" compared with a boolean. It also rejects conditions on file or json fields, which can never compare equal. The capacity workflow's requiredWhen: { key: "resolved", equals: true } on a boolean still passes.

- A blank string counts as missing for a triggered conditional string field, in both the runner's Output Format check and server output validation (isMissingConditionalValue).

No CD006-protected files are touched.

## Test plan

- [x] agent-runner: full suite (198 passed), including a trace whose tool output grows past the limit only after redaction, and a blank conditional output that triggers a retry.

- [x] Root: workflow-definition and controlPlaneAuthoring tests (type mismatch, blank conditional string, conditional start input missing or blank). The full root suite passed apart from docs.test.ts > loadDoc > throws on missing file path, which timed out under full-suite load and passes on its own (18/18).

- [x] pnpm typecheck, pnpm typecheck:runner, Biome on changed files.

#1552 — Capacity: rollback compare-and-set, docType-scoped dedupe, honest drain scheduling, day-bounded sweep @marcusdAIy  approved
#1551 — Fix Portfolio Utilities value overflow @YibinLongTrilogy  approved

## Summary

Keep long values in the Portfolio Utilities card inside its bounds. The Water contact and Electrical URL shown in the site detail view could extend past the card edge because the read-value wrapper held text in an unconstrained flex layout.

### Screenshots

<img width="960" height="740" alt="Screenshot 2026-09-28 at 10 19 38 AM" src="https://github.com/user-attachments/assets/fa907586-fdcc-4501-8574-c27e2398708f" />

### Changes

- chat/components/dashboards/portfolio/cards/utilities-card.tsx — Let read values use the shared row's text flow, break long unspaced strings within the available width, and stack labels above multiline fields such as contacts and notes.

- chat/components/dashboards/portfolio/cards/card-atoms.tsx — Add an opt-in stacked Row layout. Existing rows retain their current layout.

### Design Decisions

- Show the complete value. Clipping or truncation would hide operational contact and account details.

- Stack only fields already marked multiline in the Utilities field definitions, where a side-by-side label leaves too little room for the value.

## Business value

Portfolio users can read utility contacts, URLs, and notes without text spilling outside the card.

## Estimated manual effort

About 1 hour.

## Test Plan

- [x] Utilities card tests: 8 passed.

- [x] Biome and Chat typecheck passed.

- [x] Branch includes the fetched origin/main tip; git diff --check passed.

- [ ] Visually inspect long contacts, URLs, and notes in the full site view and Portfolio side panel.

#1550 — Capacity: wait for in-flight re-indexes, require prior numbers in the contamination log, operator review for held runs @marcusdAIy  approved

## Summary

Settles the three open decisions in AERIE-2356. Part of AERIE-2356; the ticket stays open for its remaining items.

- Stale knowledge. If a newer re-index of a document (site evidence or doctrine) is queued, processing, or retrying, the run waits for it, as it does for a first index; the existing 20-hour deferral limit bounds the wait. If the re-index ended terminal_failure or not_searchable, the run uses the last indexed version, and the source carries a readinessReason saying it is stale.

- Contamination log. An empty log is valid only when the bundle has no prior analyses. When priors exist, the log needs at least one entry, and every capacity number Aerie knows from the run that registered a prior must appear as a logged value. New flag: CONTAMINATION_LOG_OMITS_PRIOR_ANALYSES.

- Held runs (awaitingReview). New internal mutation capacityAutomation/state:_reviewHeldCapacityRun, run by an operator the same way as rollback:

- approve moves the run to awaitingPublication. The recovery sweep then publishes it through _publishCapacityRun, which applies every existing publication gate: latest run, DD not Complete, allowlist, mode.

- reject requires a reason; it marks the run unresolved and deletes its artifact copies.

- Both record reviewDecision, reviewedAt, reviewedBy and reviewReason.

- Expiry needs no new code: the site's next daily sweep run already supersedes a held run.

## Test plan

- [x] packages/contracts: vitest run src/capacity-validation.test.ts (31 passed)

- [x] chat: vitest run convex/capacityAutomation.test.ts (62 passed)

- [x] Convex and contracts typecheck, Biome, and repo lint scripts

#1548 — Capacity evidence: prior capacities, site facts in hash, poll unknown statuses (AERIE-2356) @marcusdAIy  approved

## Summary

Three more AERIE-2356 fixes in evidence assembly and polling. None touch publication or record mode.

- Prior analyses keep their numbers. Each prior analysis was projected with observedCapacityValues: [], so the agent never saw the capacities behind documents that earlier runs registered. Now, when a capacity run registered the document (it appears in that run's publishedArtifactDocumentIds), the prior carries that run's Fast Open and Max, plus registeredAt. Human-uploaded analyses still have an empty list, since their numbers exist only in the document text the agent reads.

- Site facts are in the evidence hash. bundleHash now includes a stable fingerprint of the site context (areas, occupancy, buildings, buildout, DD status). A change to those facts is therefore a change in evidence.

- Unknown Sindri statuses are polled again. A missing or unrecognized run status used to end the run immediately. Now it schedules another poll, and the existing 30-minute run timeout bounds the retries.

Linear: AERIE-2356.

## Test plan

- [x] chat: vitest run convex/capacityAutomation.test.ts convex/capacityAutomation/config.test.ts (74 passed), including: a registered prior carries 50/60 while an uploaded prior stays empty, and both have registeredAt; the site fingerprint ignores key order but changes with the facts; an unknown or missing status reschedules without recording output.

- [x] chat: tsc --noEmit

#209 — Spell out CAP-1, CAP-2, and CAP-4 in the capacity prompt (AERIE-2501) @marcusdAIy  no labels

## Summary

The capacity prompt's step 7 said only "Apply CAP-1..7 and self-correct". In the 2026-09-24 dev runs on production site data, most results failed the handoff's own gates on bookkeeping, not method. This PR states each failing check the way Aerie re-validates it (no new methodology), and fixes a dev-runner bug found while testing.

Prompt (scripts/capacity-automation/capacity-agent-prompt.md)

- CAP-1: route the ruleset from grossSquareFeet before assigning rooms (under 10,000 SF is MICROSCHOOL, otherwise 250PLUS), the handoff's drift source #2.

- CAP-2: totalNlaSf is exactly the sum of the NLA-flagged rows (output spec 01), not a Phase 1, Max, capped, or supply figure; per-level totals add up to Fast Open or Max; deductions go on the rows; the addition is written out in validationReport.

- CAP-4: every room cites a real evidence document id (never "derived"), with a unique sourceRoomId. Planned or unbuilt rooms go in assumptions, not the table.

- CAP-3: support-space status is exactly PASS, PARTIAL, or FAIL.

- Fast Open 0: a site that holds no students as-is returns resolved: false with a DRI task, matching Aerie's contract that a resolved result has a positive Fast Open (Alston, round 2).

Dev runner (agent-runner/scripts/local-watcher.mjs)

- A claim returns up to ten activations with their resolved inputs, which overflowed execFileSync's 1 MB buffer. The claim was already leased in Convex, and every error after the first poll was hidden, so queued runs silently never ran. The buffer is raised and each distinct claim failure is logged.

- An API failure such as a rejected key comes back from the SDK as is_error with subtype success and the message in result, so the runner reported it as "Fatal error: success". The runner now reports the result text (for example "Failed to authenticate. API Error: 401 ..."), and never the bare success subtype.

Linear: AERIE-2501.

## Results (6 non-Complete dev sites, production data copied to personal dev)

| Site | Before (2026-09-24) | Round 1 | Round 2 |

|---|---|---|---|

| Gallows | CAP-2 | passed (53/54) | CAP-2 |

| Boca | passed (50/58) | CAP-2 | passed (50/50) |

| Evanston | CAP-2 | CAP-2 | CAP-2 |

| Woodlands | CAP-1, CAP-2, play gate | CAP-2 | CAP-3 (verbose status; addressed in bbe2f248) |

| Alston | CAP-1, CAP-2, CAP-4 | CAP-2 | Fast Open 0 rejected by Aerie's parser |

| South Miami | CAP-4 | CAP-2 | passed (77/100) |

CAP-1 and CAP-4 failures are gone in both rounds, and per-level totals always add up. The NLA total still fails at some sites and varies run to run; repeatability on frozen evidence is AERIE-2265. The CAP-3 line was added after round 2 and has not been re-run. Alston's case (resolved with Fast Open 0) is now handled by the Fast Open 0 rule above; that rule has not been re-run either.

## Test plan

- [x] Republished the dev workflow (workflow-files-2026-09-24, instance wfi_WNd1FqqosIU) with each prompt revision and ran the six sites in parallel.

- [x] Local watcher claims and runs six concurrent activations without dropping claims.

- [x] agent-runner: vitest run (196 passed), including an error result with subtype success; tsc --noEmit.

#1547 — Harden capacity evidence and dispatch (AERIE-2356) @marcusdAIy  approved

## Summary

Four small fixes from the AERIE-2356 hardening list. All are in the evidence and dispatch path, and none touch publication or record mode.

- Geometry MIME parameters. isUsableGeometryArtifact compared the raw MIME type, so application/pdf; charset=binary was not treated as a PDF. It now compares only the media type.

- Re-index hash collisions. A document's bundle-hash part fell back to its readiness when it had no source hash, so two versions could hash the same. The evidence fingerprint now includes the knowledge version, and the doctrine fingerprint includes the artifact.

- Empty documents. A ready document with no extracted text was skipped silently. It now sets evidenceArtifactTruncated, and the evidence file's note says documents were omitted for being too large or having no extracted text.

- Stuck deferrals. Deferring on unreadable evidence never used up an attempt, so such a run stayed queued forever. Because only one run may be in flight per site, it also blocked every later sweep for the site. After 20 hours of deferring, the run is now marked unresolved with the reason, which clears it before the next 11:00 UTC sweep.

Linear: AERIE-2356.

## Test plan

- [x] packages/contracts: vitest run src/capacity (61 passed), including MIME types with parameters; tsc --noEmit.

- [x] chat: vitest run convex/capacityAutomation.test.ts convex/capacityAutomation/config.test.ts (71 passed), covering the empty document flag, fingerprints changing on re-index, and deferral before and after 20 hours; tsc --noEmit.

#1546 — Capacity: record Sindri failure reasons; cover rollback failure paths @marcusdAIy  approved

## Summary

- When a Sindri capacity run ends failed, canceled, or blocked, Aerie recorded only Sindri run failed. On 2026-09-25 every dev run failed on a revoked Anthropic key, and nothing in Aerie said so.

- Aerie now appends Sindri's redacted failureSummary.error (capped at 500 characters), for example Sindri run failed: Failed to authenticate. API Error: 401 API key is invalid. Without a message, the text is unchanged.

- Pairs with Sindri #209, which makes the runner report that error text instead of "success".

- Adds rollback failure-path tests through the real mutation (AERIE-2275): automation disabled, a proposal-mode run, a newer record-mode publication, and a failed DD write. Each leaves the card untouched and records why.

The record-mode success path is not covered here: writing those tests showed that a rollback marks the run rolled back while the governed DD write is still only a pending request, and that restoring an absent DD status is rejected by the governed write. Both are record-mode prerequisites recorded in AERIE-2356; record mode is off.

Linear: AERIE-2501 (found while testing), AERIE-2275.

## Test plan

- [x] chat: vitest run convex/capacityAutomation.test.ts (46 passed): failure message carried through, long message capped, missing summary keeps the old text; four rollback failure paths.

- [x] chat: tsc --noEmit

#1543 — Restrict admissions deals to school-year sessions @vvp-trilogy  approved

## Summary

- restrict canonical admissions deal paths to HubSpot schoolYear sessions before uniqueness checks

- exclude summer-only deals while preserving a valid school-year path when a deal also has a summer association

- extend the existing admissions-deal unit test rather than adding a new test

## Validation

- local dbt parse

- git diff --check

## Production finding

This removes the three summerCamp deal paths currently making Alpha Austin, Alpha Dorado, and Alpha Scottsdale milestone observations unavailable.

#2067 — fix(aws-spend): map AI Engineering experiment account @caina-barbosa  approved

## Summary

Records the production account mapping correction for newly active experiment sandbox account:

- 520519513954 → Exp-AIEngineeringandBuilder-aiengColinedi

Mapped to class = Central Engineering and bu = Central Engineering, projected from 2026-Q3 through the existing 2030-Q4 horizon.

## Incident & Why

The saas-budgeting-pipeline scheduled run for 2026-09-25 failed closed on the noncentral_charges ingest:

ValueError: account mapping is incomplete for 2026-Q3: ['520519513954']

All other ingests (docker, k8s, database_units, mapping, server_costs) published successfully.

## How the values were derived (evidence chain)

1. Master Payer: Account 520519513954 reports RDS costs under master payer 572481847476 (VDI).

2. Account Name: Queried AWS Cost Explorer linked-account metadata via the payer role (ESW-CO-ReadOnly-P2), returning:

- 520519513954 → Exp-AIEngineeringandBuilder-aiengColinedi

3. Class & BU: Mapped to class = Central Engineering and bu = Central Engineering, matching adjacent developer experiment and tooling accounts under the engineering umbrella (Exp-CentralEngineering-*, Int-CentralEngineering-*, Dev-CentralFunctions-*).

4. Completeness: Querying the anti-join between core_finance.aws_spend_net_amortized_costs (RDS service, 2026-Q3) and core_finance.aws_spend_budget_account_mapping confirmed that across all master payers, exactly this single account was missing a mapping.

## Production remediation completed

1. Executed and verified in finance_dw: 18 rows inserted (1 account × 18 quarters, 2026-Q3 .. 2030-Q4).

2. Post-commit anti-join confirmed 0 unmapped 2026-Q3 RDS accounts remaining.

3. Triggered on-demand Step Functions execution manual-noncentral-aieng-colinedi-20260926T134119Z:

- Status: SUCCEEDED

- Candidate count: 204 accounts

- Replaced 203 prior rows with 204 current rows

- Source max date: 2026-09-25

- Billable accounts: 148 / charge total: $148,000

- Mapping gap count: 0

#2060 — fix(education): accept GuidePlatform behavioral event fields @kevalshahtrilogy  approved

## Summary

- Root cause: GuidePlatform's behavioral_events source table gained six new columns (observed, location, others_present, reviewed_with_team, lied_about_it, escalated_from) upstream. Every scheduled guide-platform-raw-sync run since 2026-09-24 05:35 UTC has failed at the pre-extraction schema-shape check with GuidePlatform source schema drifted: behavioral_events(added=[...]) (CloudWatch: /klair/pipelines/prod/guide-platform-raw-sync; still failing 2026-09-28).

- Types confirmed by read-only introspection of the live GuidePlatform Postgres (information_schema.columns, same reviewed benji_ro reader role the pipeline uses): observed/location/others_present/escalated_from are nullable text; reviewed_with_team/lied_about_it are boolean NOT NULL DEFAULT false.

- Regenerated src/contract.py and ddl/001_create_staging.sql from the live source via scripts/generate_contract.py (only behavioral_events changed). Updated contracts/legacy_clean_compatibility.json's new_source_fields, and relaxed the ddl/003 and ddl/005 "matches canonical exactly" tests, which pinned byte-for-byte equality against the ever-current ddl/001.

- Grant handling (reworked after Mercy's review). A schema-wide Surtr_Service_User grant set (ALTER/DELETE/DROP/INSERT/REFERENCES/RULE/SELECT/TRIGGER/TRUNCATE/UPDATE) now sits on all 115 relations in staging_education_guide_platform, unrelated to this fix. The first cut of ddl/006 recreated the view and restored only the three reviewed CQL_download_OM grants, so the admin preflight (correctly) refused and the pipeline stayed down. Restoring any fixed baseline is wrong in both directions, so ddl/006 now carries no GRANT/OWNER statements. run_ddl.py --apply-migration behavioral-events-schema-additions:

1. captures the live owner, the raw pg_class.relacl (grantor-exact) and the svv_relation_privileges rows (the only place RBAC role grants appear);

2. refuses, before submitting anything, an ACL a plain GRANT cannot reproduce (grant options, a grantor other than the owner, partial legacy RULE/TRIGGER sets, non-plain identifiers) and any non-superuser (partial-visibility) snapshot;

3. submits one atomic batch: precondition guard that the live relation still equals the snapshot, DROP/CREATE/COMMENT, ALTER ... OWNER plus one GRANT per grantee replaying exactly the snapshot (ALL PRIVILEGES for the full 10-bit set, which is the only spelling that reproduces the legacy RULE/TRIGGER bits), then a postcondition guard that fails the whole batch (rollback) unless owner, ACL and privilege rows equal the snapshot again;

4. replaces the hardcoded expected_grants=3/unexpected_grants=0 preflight with snapshot-derived counts (same fail-closed intent, and it now also pins the raw ACL, not just the privilege view). Nothing is stripped, added or widened, and no grantee is hardcoded.

- Latent bug fixed: the migration's ClientToken was 67 characters and the Data API limit is 64, so submission would have been rejected client-side. The token is now a constant prefix plus a digest of the exact submitted SQL, so an identical retry stays idempotent and a batch built from a different snapshot can never collide with an earlier one.

- Consumer contract pins (added after CI failed on mart-education-guide-roster-refresh). The regenerated producer hash 1f882787... is a runtime handshake with the Guide roster consumer: its handler selects only ledger rows carrying its pinned hash, and the live sp_refresh_guide_roster_evidence procedure re-checks it before any target write, so a mismatch fails closed with no mart write. Following the 2026-09-21 precedent (#1991) this PR bumps every pin (handler, source-inventory verifier, canonical ddl/001 predicate, disposable-Redshift fixture, and the surtr-374 spec/checklist/evidence docs) and adds ddl/005_20260928_source_contract.sql, the procedure migration, registered in apply_ddl.py with its own statement name and token. Nothing else reads behavioral_events (the consumer reads only ingestion_ledger, campus_guide_roster, sessions; person-directory-refresh reads ingestion_ledger, students, users), so the six nullable columns cannot change any consumer's behavior. Ordered recovery gates are in the new contracts/schema-drift-2026-09-24.md.

- Scope: pipelines/runners/guide-platform-raw-sync/ (ddl/, scripts/run_ddl.py, src/contract.py, tests/test_contract.py, contracts/, README.md), pipelines/runners/mart-education-guide-roster-refresh/ (pins, ddl/005, apply_ddl.py, tests) and the surtr-374 guide-roster spec docs.

- DDL applied to prod 2026-09-28 10:41 UTC via REDSHIFT_DB_USER=admin uv run python scripts/run_ddl.py --apply --apply-migration behavioral-events-schema-additions (Data API statement 6ceae357-b79b-4c59-9d5f-12c6d0d65455). Guards passed; live snapshot was owner admin, CQL_download_OM=ard, Surtr_Service_User=arwdRxtDPA, no role/group/PUBLIC grants (3 ACL entries, 13 privilege rows). The applier's own catalog check confirmed all 57 clean views match the contract.

- Still needed before the pipeline recovers: (1) merge this PR and let CD ship both runner images (the failing check is the producer's source-drift guard, which reads the *deployed* src/contract.py); (2) apply ddl/005_20260928_source_contract.sql with apply_ddl.py while the consumer trigger stays disabled (the live procedure still pins 63a7a6fd...); (3) start a fresh raw execution and follow the gates in contracts/schema-drift-2026-09-24.md. Nothing fires the consumer automatically today (no EventBridge rule targets it), so an ordering slip fails closed rather than writing.

## Business Value

Restores the daily guide-platform-raw-sync pipeline, which has been fully down (zero successful runs) since 2026-09-24. Every table in staging_education_guide_platform, not just behavioral_events, is going stale because the run fails at pre-extraction validation before any table syncs. This is the fourth schema-drift break in about a week across three upstream tables, each needing a one-off migration. This rework also removes a class of silent risk: recreating a view previously discarded every grant not on a reviewed list, so a governance rollout like the schema-wide Surtr_Service_User grant would have been silently stripped, or the migration blocked. It now preserves whatever the live grants are and proves it inside the transaction.

## Manual Effort Estimate

Roughly 6-8 focused hours for an engineer already familiar with this runner's guarded-migration pattern: reading the drift error and logs, introspecting live source types, regenerating the contract, writing the migration and its guards, then diagnosing the grant blocker and designing and testing the snapshot/replay/postcondition machinery (including the raw-ACL and RBAC-role coverage and the ClientToken limit), and finally the live apply with before/after ACL verification across all 115 relations. Proposed by Claude, Keval to confirm or adjust.

## Test plan

- [x] uv run pytest in the runner dir: 177 passed, 1 skipped (pre-existing skip: SELECT-only Redshift integration test needing explicit approval).

- [x] ruff check / ruff format --check (pinned 0.15.22, matching CI) clean on the runner dir.

- [x] New tests cover: no hardcoded grants in ddl/006; replay order (drop, create, comment, owner, grants, postcondition); a non-baseline extra grant (user, role, group, PUBLIC, partial privilege set) surviving the recreate; refusal of unreplayable ACLs; guards pinning the same snapshot before and after; snapshot-derived preflight expectations failing closed on every counter; ClientToken length and digest binding; 40-statement batch limit; full fake-Data-API apply submitting the replay; no submission on a stale snapshot, a non-superuser snapshot, a non-view relation, or an unreplayable ACL.

- [x] Live catalog SQL (relacl membership and count predicates, the full preflight query) exercised read-only against the warehouse before the apply.

- [x] Prod apply verified: owner and ACL of behavioral_events identical before and after (admin=arwdRxtDPA/admin,CQL_download_OM=ard/admin,Surtr_Service_User=arwdRxtDPA/admin); owner and ACL of all 115 relations in the schema identical before and after; behavioral_events 37 to 43 columns with the six new fields at positions 29-34; view comment and source-contract hash updated (1f882787...); view row count (6391) equals raw_behavioral_events.

- [x] Full CI runner loop run locally (all 124 runner test suites pass, including both guide runners and person-directory-refresh) and ruff 0.15.22 check/format over pipelines; CI green on the new head and Mercy approved.

- [ ] Merge, CD deploy of both runner images, apply consumer ddl/005, then confirm the next 05:35 UTC run succeeds.

Linear: SURTR-1519

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2075 — fix(sales-educrm-mart-sync): add upstream program_id to the coming-year projection tables @kevalshahtrilogy  approved

## Summary

Root cause. On 2026-09-25 the upstream Athena marts educrm_wh.mart_coming_year_projection (about 22:00 UTC) and mart_coming_year_projection_snap (about 23:30 UTC) gained a nullable program_id BIGINT column. The runner's publish guard refuses any column-set drift by design, so both live Redshift tables have not refreshed since the 21:35 UTC run on 2026-09-25. Every scheduled run since 22:05 UTC (121 runs by 10:05 UTC on 2026-09-28) has ended PARTIAL with Refusing to publish schema drift for staging_education.sales_educrm_wh_mart_coming_year_projection: ... unexpected_in_candidate=['program_id'], and the same for the _snap table from 23:35 UTC. The other 37 marts kept syncing. The candidate _new tables already held the fresh data (projection last_updated_date 2026-09-28 versus 2026-09-25 live; snap 10,919 versus 10,379 rows).

This is not the INVALID_VIEW flake from SURTR-1529's original triage. That cause produced the five earlier PARTIAL runs (2026-09-19 to 2026-09-21) and is fixed by #1966: since it reached production, two INVALID_VIEW events (2026-09-23 16:35 and 2026-09-24 08:35) both recovered on attempt 2.

Fix. Follow the README's stable-OID schema evolution procedure. Add ddl/20260928_add_program_id_to_coming_year_projection.sql, an idempotent migration that appends program_id BIGINT (nullable, no default) to the two live tables in place. Redshift has no ADD COLUMN IF NOT EXISTS, so it follows the repo pattern (transient procedure, LOCK TABLE, svv_columns catalog check, EXECUTE 'ALTER TABLE ...', CALL, DROP PROCEDURE, then a poison guard). The publish guard is unchanged: no data-integrity check is weakened, and the runner code is untouched. A regression test proves the publish path accepts the column at the end of the target while the Athena candidate lists it first (the guard compares by name and the INSERT names its columns). The README gains a note on recognising a schema-drift PARTIAL and a migration ledger.

Already applied to production (between runs, 10:43 to 10:45 UTC on 2026-09-28). Both tables are still the same relations (OIDs 17275059 and 17275092), with owner CQL_download_OM and grants CQL_download_OM=arwdRxtDPA, MCP_user=r, Surtr_Service_User=arwdRxtDPA identical before and after. The only change is program_id bigint YES at position 62 on each table. The first apply's poison guard used a literal ELSE 1 / 0 arm, which Redshift evaluates eagerly and so failed even though the columns were correct; the guard was rewritten to divide by the CASE result, and the whole migration was re-run to confirm the idempotent no-op path.

Reader audit. No dependent views, late-binding views, materialized views or stored procedures reference either table. The only repo reader, hubspot-core-tables/reconciliation/forecast_marts/030_projection_output_contract_gates.sql, uses named columns after a SELECT * CTE. Query history over 30 days shows only named-column readers plus an ad hoc live-versus-_new comparison that qualifies its columns.

Known gap, not addressed here. program_id on the snap table is populated for every snapshot from 2026-09-25 onward (720 rows) and NULL for the 10,199 earlier history rows, because upstream has not backfilled it; whether it should is an upstream decision.

Post-merge. Normal deploy of this runner only; no runner code changed, so behaviour is identical. The next scheduled run after the ALTER should publish both tables and backfill program_id.

## Business Value

The eduCRM coming-year projection and its snapshot history are read by the Education forecast reconciliation and by analysts' validation queries. They had been frozen for more than 36 hours while every run reported PARTIAL, which hides a real data gap behind a mostly-green pipeline. This restores fresh projection data, makes the upstream's new program_id identity column available downstream, and documents how to recognise and resolve a drift-guard PARTIAL so the next upstream column addition is a five-minute migration rather than a multi-day stall.

## Manual Effort Estimate

About 4 hours of focused work by hand: roughly 1.5 hours tracing PARTIAL runs through run history, CloudWatch and the Redshift candidate tables to find the drift; 1.5 hours to write the idempotent migration, its tests and the README note against the repo's precedents; and 1 hour for the reader audit, the live apply between runs, the before/after grant comparison and the post-run verification. Proposed by Claude, Keval to confirm or adjust.

## Test plan

- [x] cd pipelines/runners/sales-educrm-mart-sync && uv sync --all-extras && uv run pytest: 54 passed (49 existing, 5 new)

- [x] uvx ruff@0.15.22 check and format --check on the runner: clean

- [x] Migration applied live twice (first run added the columns; second run took the no-op path); owner, grants and OIDs identical before and after; program_id bigint YES present on both tables; helper procedure dropped

- [x] First scheduled run after the ALTER (11:05 UTC on 2026-09-28, run 830b7aa3) ended SUCCESS with 39 of 39 tables synced and 0 failed. Both projection tables published: projection 180 rows with program_id populated on all 180 (90 distinct ids, none null) and last_updated_date advanced from 2026-09-25 to 2026-09-28; snap 10,919 rows (up from 10,379), snapshot_date advanced to 2026-09-28. Owner, grants and OIDs unchanged after the publish. The _new candidates were dropped by the runner as designed.

Linear: SURTR-1529

🤖 Generated with [Claude Code](https://claude.com/claude-code)

The Builder Desk  —  Engineer Spotlight
📅 Week in Review🏆 Engineer Spotlight

185 PRs IN SEVEN DAYS: BUILDER TEAM POSTS NUMBERS THAT DEFY THE LAWS OF PHYSICS

Kevalshahtrilogy ships 59 PRs solo while a brand-new repo is born — the Numbers Desk has never seen a week like this.

Folks, pull up a chair and hold onto your spreadsheets, because the last seven days have produced a velocity reading that frankly shouldn't be legal. One hundred eighty-five pull requests. Seven repos lit up like a switchboard — Aerie leading the charge with 85 merges, Surtr right behind with 42, Shipyard putting up 31, and even the newborn Redshift-DSS already logging its first commits to the scoreboard. This isn't a sprint, people. This is an entire track meet happening inside a single work week.

Let's run the board. @kevalshahtrilogy posted a staggering 59 PRs — I counted no fewer than a dozen Aerie gate-and-shadow builds (#1576, #1578, #1580, #1585, #1586, #1587) plus Surtr marker work on #2128 and #2125. The man is building the admissions pipeline brick by brick, pun very much intended. @marcusdAIy matched the raw volume of @ashwanth1109 at 32 apiece, proving this team has not one but two workhorses pulling the plow. @vvp-trilogy turned in a strong 22, anchoring the SIS enrollment mart with #1653, #1649, #1646 and #1647. @caina-barbosa logged 12 including the timeback gateway fix in #2127. @sanketghia chipped in 10, with #2121 patching CollectIQ v2. @benji-bizzell notched 9 and @YibinLongTrilogy closed out with 5 — every single contribution load-bearing.

Now, the Ashwanth situation. Thirty-two PRs this week from @ashwanth1109, spanning Shipyard releases and a Surtr QuickBooks validation job in #2123. The man shipped a full Shipyard version bump in #162 like he was ordering a sandwich. "I review my own diffs in my head before I even open the laptop," he reportedly told our desk, which — sure, Ashwanth, sure. Nobody has independently verified whether a human eye has fully parsed #161's permissions-box removal, but the CI is green, so who are we to argue? When reached for comment on this paragraph, Ashwanth said only: "Write whatever you want, I'm not reading it." Legend.

Now to the Overflow Desk, the PRs Mac didn't have room for but the numbers demand recognition. #1652 and #1651 quietly hardened the Gateway's shadow-read integrity — unglamorous, essential. #1636 made the mapper fail closed instead of silently losing rows, the kind of fix that prevents 3 a.m. pages. And #2120 untangled school ontology lock ordering across sibling marts, a problem most engineers would rather not touch with a ten-foot pole.

Across the leaderboard, Aerie dominates raw volume while Surtr quietly racks up precision fixes, and the christening of Redshift-DSS signals there's appetite for even more surface area next week.

Morale, as always, is at an all-time high. This team doesn't walk into the office — they vault.

Brick's Overflow — This Week's Uncovered PRs  (click to expand)
#161 — AI-960: Remove Pi permissions info box and acknowledgement checkbox from task creation @ashwanth1109  no labels

## Demo

![Smoke test demo](https://github.com/AI-Builder-Team/Shipyard/blob/ash/ai-960-remove-pi-ack/.github/smoke-test-evidence/AI-960/demo.png?raw=true)

## Summary

- Removes the amber Pi info box and the "I understand and want to use Pi for this task." checkbox from the queue's create-task form. Pi tasks are now created the same way as Codex tasks.

- Removes piPermissionsAcknowledged / pi_permissions_acknowledged end to end (createTask, the create_task command and create_with_engine), including the backend rejection branch.

- Removes PI_ACKNOWLEDGEMENT_REQUIRED and the unused acknowledgement field from agent/engine/info.

- Pi readiness still blocks creation. Add stays disabled while the check runs, with no text shown. If credentials are unavailable, the existing task-workspace__error line shows "{error} Open Agent setup & diagnostics to configure them." as soon as the check fails.

- Moves the OS-permissions disclosure to a static note on the Pi readiness card in Agent setup & diagnostics.

- Removes the unused .task-workspace__agent-warning* CSS.

## Linear

https://linear.app/builder-team/issue/AI-960/remove-pi-permissions-info-box-and-acknowledgement-checkbox-from-task

## Test plan

- [x] pnpm build (tsc + vite)

- [x] pnpm theme:check

- [x] cargo test --lib workflow:: (81 passed)

- [x] cargo test --lib agent (38 passed)

- [ ] Manual: choose Pi in the create-task form → no info box; Add is enabled once readiness passes; a task is created.

- [ ] Manual: without TrueFoundry credentials → Add is disabled and the readiness error appears under the form; switching back to Codex clears it.

#162 — Release: Shipyard 0.6.11 @ashwanth1109  no labels

## Summary

- Bump Shipyard to 0.6.11.

- Add the approved public release notes.

## Business Value

- Delivers the approved branch workflow, Pi conversation controls, repository history pagination, task timing visibility, and Pi image diagnostics.

## Implementation Effort

- Low: metadata-only release change; CI performs the signed build and publication.

## Test Plan

- [x] pnpm test:release

- [x] git diff --check

- [x] Verified the diff contains only package.json and releases/0.6.11.md.

#163 — AI-962: Evaluate task prompts with two independent replay trials @ashwanth1109  no labels

Completed-task evaluation previously had no in-app way to propose a prompt change, replay it against the original inputs, and compare the evidence before publication. This adds persistent eval chat with one editable candidate and two independent trials, with original/trial conversations available throughout review.

## Business Value

Users can find optimization opportunities or shorten a template, refine the proposal directly in chat, and verify behavior before adopting it. Incomplete execution, failed acceptance checks, and passing results have distinct next actions; retries preserve the candidate and prior attempt evidence.

## Changes

- Add a compact proposal card, prompt diff, acceptance checks, trial transcripts, focus mode, concise quick messages, and shared new-conversation/composer controls. New eval chats default to Full access and reset their proposal state when replaced.

- Restore declared repository and artifact inputs from the captured node-entry baseline, record implementation base provenance, replay original follow-ups and images, and retain partial transcripts. Recover transient provider reads without resubmitting agent turns.

- Bind each assessment to the candidate and exact current run IDs. Separate behavioral regressions from informational notes, and request another review when the assessment contradicts itself.

- Keep publication explicit and validate successful independent runs, passing checks, baseline provenance, candidate hash, and current template identity before creating an immutable local template version.

## Validation

- Frontend production build and native app build passed on the integrated branch.

- Native library suite: 377 passed, 3 ignored.

- UI, transcript, chat, workflow, and PR-review regression tests passed.

- Fixture-backed desktop smoke run d87a87bb-c923-4ccb-ac1f-00502c360736 passed workflow verification. Observed original/trial navigation, Full access/model controls, and focus expand/restore; the harness was stopped afterward.

- Two real-agent Research trials completed with saved conversations and final artifacts; all four acceptance checks passed in both. One trial reported a missing-esbuild validation limitation, recorded separately from replay execution and behavioral regressions.

Historical tasks still require verifiable starting inputs; the runner reports missing provenance instead of replaying against current state. The recovered historical baseline used during validation was prepared locally, without adding a database backfill or changing bundled node templates.

## Implementation Effort

Estimated 5–8 engineer-days for an average engineer to implement and validate the UI, replay/session lifecycle, baseline restoration, provider recovery, publication safeguards, and regression coverage without AI assistance.

## Linear

https://linear.app/builder-team/issue/AI-962/evaluate-and-refine-task-prompts-with-two-independent-replay-trials

#1636 — fix(gateway): fail closed when a mapper drops a Gateway row (AERIE-2679) @kevalshahtrilogy  approved

Linear: AERIE-2679 (fixes a blocking Mercy finding on the release PR, against code from AERIE-2677; related AERIE-445)

## Summary

The finding. In ADMISSIONS_COMMUNITY_FUNNEL_READ=gateway, the legacy mapper drops a Gateway row with a null deal_id or one its schema rejects. Gateway mode only logged the count and then ran the population check on the reduced count. One dropped row in 5,403 is far inside the 5% allowance, so a partial snapshot would publish as a success.

The rule this PR applies to every Gateway read: if the mapper returns fewer records than the Gateway rows it was given, the read is refused.

- gateway mode: the read throws before the population check and before any write. The last published data stays. The message carries counts only, no row values.

- shadow mode: the same refusal makes the check degraded, never clean, so the shadow window shows it before gateway would refuse it. This matters because the mart is a verbatim copy: pg drops the same row, so the keyed compare alone would have called the cycle clean.

- Legacy (pg) paths are unchanged. Production publishes exactly what it publishes today.

What changed.

- Kit, sync/src/analytics/a8/gateway-table-reader.ts: assertNoA8RowsDropped(source, { rowsRead, recordsMapped }, reason). It throws the reader's own A8GatewayReadError, which is on the kit's value-free list, so the message reaches the logs as written.

- Community funnel, admissions-community-funnel-gateway.ts:

- readCommunityConversionFromGateway refuses any dropped row.

- Removed: the droppedRows and sourceRows fields (always 0 and N now), the warn-and-publish branch, and the second floor on the mapped count (dead once any drop throws).

- The module doc and the floor comment now describe the new behaviour.

- Expense vendor classifications, expenses-gateway.ts: the same refusal. It replaces the second floor on the mapped count, which only refused a read once it fell below 500 of about 3,479 rows.

- XO contractor identity, xo-contractor-identity-gateway.ts: a Gateway row whose aliases all trim away now fails the read. Before, it was filtered out and the rest published.

Every Gateway reader under sync/src, and what I found.

| Lane | Reader (Gateway source) | Can it drop a Gateway row? | Change |

|---|---|---|---|

| A1 | site-operational-metadata-gateway.ts | No. z.array(...).safeParse, then throw. Shadow only, never published. | None |

| A2 | schools-data-sheet-gateway.ts | No. Any invalid row, duplicate cell or partial grid refuses the snapshot; every record is published. | None |

| A3 | school-calendar-gateway.ts | No. Any invalid row or duplicate slug refuses the read; every row is sent. Convex skips a row whose site slug Aerie doesn't have, and the worker records that as an error on the run. | None |

| A4 | school-source-directories-gateway.ts (QuickBooks, SIS, Finalsite) | No. Strict parse; the snapshot builders map one to one and throw. | None |

| A5 | camps-full-gateway.ts (7 entities), camps-gateway.ts (parity checks) | No. Strict parse; loadCampSourceSnapshotViaGateway refuses the whole snapshot on a blank or duplicated id, mixed lineage or a dangling reference. | None |

| A6 | rebl3-sites-gateway.ts | No. Any invalid row, repeated site or mixed lineage refuses the snapshot. | None |

| A7 | matterport-discovery-shadow.ts (xref) | No. Strict; shadow only, never published. | None |

| F1 | xo-contractor-identity-gateway.ts | Yes: filtered out rows with no usable alias. | Fixed |

| F2 | xo-contractor-package-gateway.ts | No. Strict parse, one record per row. | None |

| A8 G1 | admissions-reference-gateway.ts (programs, Program directory) | No. mapProgramRows and mapHubspotProgramRows throw on any bad row. | None |

| A8 G2 | admissions-community-funnel-gateway.ts | Yes: null deal_id or schema-rejected. | Fixed |

| A8 G3 | community-deposits-gateway.ts | No. mapCommunityDepositRows maps every row or refuses the read. | None |

| A8 G4 | expenses-gateway.ts, transactions | No. Strict array parse; a conflicting repeated key is refused. An exact repeated row is published twice, as legacy does. | None |

| A8 G4 | expenses-gateway.ts, vendor classifications | Yes: safeParseRows skips a row it cannot parse. | Fixed |

| A8 G4 | expenses-gateway.ts, metadata | Not a row mapper. It derives two DISTINCT lists and leaves out NULLs exactly as the legacy SQL's WHERE ... IS NOT NULL does. | None |

## Business Value

- Unblocks the Aerie production release. Mercy will not approve the release while this finding is open on main.

- A snapshot that lost rows can no longer be published as a clean success on the admissions funnels, the expense vendor categories or the contractor identity index. Each of those would otherwise go quietly stale or short for the affected families, vendors or contractors.

- The shadow window now means what it says. A source that would be refused in gateway mode can no longer pass its shadow window clean.

## Manual Effort Estimate

About 4 hours of focused work by hand, without AI. Keval, please confirm or adjust this number. It covers:

- reading every Gateway reader and its publish path (13 reader files, plus the refresh code each one feeds) for drops;

- the kit helper and the three fixes;

- reworking the tests that pinned the old drop-and-publish behaviour, plus 18 net new tests.

## Testing

- a8/gateway-table-reader.test.ts (+4): equal counts pass; one dropped row of 5,403 is refused with the exact count-only message; the error carries its source; more records than rows is refused too.

- queries/admissions-community-funnel-gateway.test.ts (55 to 61):

- reader: one null-deal_id row and one schema-rejected row each refuse a 1,201-row read, with the exact message and no PII; 250 dropped rows are refused the same way; a clean read returns one record per row;

- gateway mode: one dropped row fails the read with A8GatewayReadError, the population preview is never called, nothing is logged as read, and pg is never called;

- shadow mode: pg and the Gateway both carry the same null-deal row. Legacy publishes as today, and the check is degraded, countsAsClean: false, with the count-only reason;

- end to end through refreshCommunityFunnel: a dropped Gateway row fails the refresh with no Convex call at all; legacy mode still publishes the pg read minus the row its mapper drops.

- Two tests that pinned the old behaviour (drop, warn, publish; floor after drops) were rewritten.

- queries/expenses-gateway.test.ts (31 to 34): the legacy mapper still drops a bad pg row; the Gateway reader refuses the same row; gateway mode fails only the vendors read with a value-free message while transactions still publish; shadow mode is degraded for vendors and clean for the other two sources.

- tests/redshift/xo-contractor-identity-gateway.test.ts (16 to 19): three shapes of an alias-less row each refuse the read with a count-only message and no name; blank entries between delimiters are still trimmed without refusing the row.

- tests/analytics/xo-contractor-identity-refresh.test.ts (+2): through the real Gateway reader. gateway mode returns an error and sends nothing; shadow mode publishes from pg and reports check failed.

- Checks:

- sync: tsc --noEmit clean; vitest run --maxWorkers=2: 113 files, 2515 tests passed.

- pnpm lint (boundaries, convex paths, read bounds, test architecture, knowledge, biome) is clean. Its two warnings are pre-existing ones in chat/skill/forge-api/scripts/sindri.mjs.

- The chat suite wasn't run: this PR changes no chat code.

## Does this refuse anything on today's data?

- Community funnel: no. The production mart was checked today (by Keval's session, not re-queried here): 5,403 rows, 0 null deal_id, 5,403 distinct.

- Vendor classifications: no, on the last evidence I have. The read-only check for the expenses gate on 2026-09-29 found 3,479 rows and none with a NULL name.

- XO contractor identity: not verified against production. The mart's SQL only emits aliases of 6 or more characters, so a row with none needs an all-whitespace contractor name. XO_CONTRACTOR_READ defaults to pg, and shadow mode would report such a row as a failed check before any flip.

- All three gates default to the legacy read, so nothing changes in production until a gate is set.

## Not covered

- The legacy mappers still drop rows on the pg path. That is deliberate (production behaviour is unchanged), but it means pg and the Gateway now differ on a bad row: pg publishes without it, the Gateway refuses.

- A persistent bad row in a mart keeps its shadow line degraded every cycle until Surtr fixes the row, since the Gateway read is refused before the compare runs.

- The held A8 PRs are not on main and were not audited: the per-program readers, the admissions pipeline gate, marketing and SIS enrollment. The pipeline refresh already counts droppedNoColumn and droppedNoProgram, so its gate needs this same rule when it lands.

- No live Gateway run. Gateway mode cannot be exercised end to end until the Surtr marts and the key grant are live.

- safeParseRows still logs a Zod message for the first three rows it skips. Unchanged here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1652 — fix(a8): keyed Gateway readers say when a program key is absent; gateway stays unreleased in production (AERIE-2690) @kevalshahtrilogy  approved

## Summary

Follow-up to Mercy's review of release PR 1641. Three of her critical findings are one class: a per-program Gateway lookup turns a missing program key into an empty result and reports success (admissions-marketing-gateway.ts:537 and :625; admissions-program-gateway.ts:653 and the other keyed readers; the whole-table app conversion read at :782).

The underlying fact. A program with no rows in a mart has no key in it, exactly as the legacy WHERE key = $1 returns no rows for it, and several sources are sparse (transfers and pipeline deposits cover a minority of programs; most programs have no shadow-day events). A program the mart *left out* looks the same. The reader cannot tell the two apart, and a table-level row floor cannot either.

That splits the finding in two, and they need different answers.

### Shadow (live in production): the verdict was already right; now it is pinned and visible

In shadow, the same call's pg read is the arbiter of whether a program has rows. Rows on pg and none from the Gateway is a mismatch, however the Gateway came to have none. An absent key only compares clean when pg returned no rows either, which is a true match.

- New tests pin both cases for the per-program gate (a keyed source and the alias-based Q3 reader) and the marketing gate: a program the mart leaves out while pg has its rows is a mismatch; a program with nothing on either side is clean.

- Every keyed slice now carries present, and each shadow line counts the absent keys it compared as empty: scope.absentCalls (per-program) and scope.absentPrograms (marketing, whose lines gain a scope). A reader of the line can see how much of a clean result was "nothing on either side".

### Gateway (not enabled in production): real, and closed off until there is a rule

Without pg there is no arbiter: gateway mode would publish a left-out program as an empty result, and a short slice as that program's data. Closing that properly needs a completeness rule, which needs a baseline (what is published for the program now) or a manifest from the mart. That is AERIE-2683 and is a design across about a dozen sources, not a guard in the reader.

What this PR does is make sure gateway mode cannot publish in production before that rule exists:

- ADMISSIONS_MARKETING_READ=gateway was selectable. It is now unreleased, like the per-program gate: refused with AdmissionsMarketingConfigError, before the admissions run starts, unless the new ADMISSIONS_MARKETING_READ_NONPROD_OPT_IN=true is set.

- Both gates now refuse gateway on the production worker even with the opt-in. The production worker is the one pinned to DBT_TARGET=production (compose.prod.yml); isA8ProductionWorker is the same literal check the dbt-backed gates already use. Until now "never set the opt-in in production" was a convention.

I did not make gateway mode refuse every absent key. That would fail every program that legitimately has no rows, on every cycle, for every sparse source; and shadow, which maps the Gateway side as gateway mode would, would then read degraded for those sources forever and stop producing evidence. The baseline rule in AERIE-2683 is the one that can tell the cases apart.

## What changes per mode

| Mode | Per-program gate (ADMISSIONS_PROGRAM_DETAIL_READ) | Marketing gate (ADMISSIONS_MARKETING_READ) |

|---|---|---|

| legacy | None. | None. |

| shadow | Publishes exactly as before. Source lines gain scope.absentCalls. No verdict changes. | Publishes exactly as before. Lines gain scope (programs, uncomparedPrograms, absentPrograms). No verdict changes. |

| gateway | Refused on the production worker even with the opt-in. Outside production, with the opt-in, the reads behave as before. | Now refused without ADMISSIONS_MARKETING_READ_NONPROD_OPT_IN=true, and always refused on the production worker. Outside production, with the opt-in, the reads behave as before. |

Production today runs both gates in shadow, so nothing it does changes except the extra counts on the lines. An environment that has ADMISSIONS_MARKETING_READ=gateway set without the opt-in would fail its admissions cycle at start with a clear config error after this deploys; I know of none.

## Business Value

The team is collecting shadow evidence to decide when the analytics worker can read admissions data from Surtr's Gateway instead of Redshift. This change does two things for that decision. It shows, on every line, how many programs matched only because both sides had nothing, so a clean window is not over-read. And it makes it impossible to switch production to gateway mode for these two gates before the missing-program rule exists, so a mart that drops a school can never blank that school's admissions data in the dashboards.

## Manual Effort Estimate

Proposed: about 5 hours of focused work by hand, without AI. Keval, please confirm or adjust. It covers working out which half of the finding is real (reading both gates, the kit and the refresh's write path), deciding between refusing absent keys and gating release, the presence flag and counts through two shadow implementations, the production refusal for both gates, and about 20 new or changed tests.

## Testing

- sync: pnpm typecheck exit 0, pnpm lint exit 0, vitest run --maxWorkers=2 exit 0 (121 files, 2779 tests).

- New tests:

- per-program shadow: absent key with nothing on pg is clean and counted; a program the mart leaves out is a mismatch (keyed reader and Q3 aliases);

- marketing shadow: the same two cases, with scope.absentPrograms;

- marketing gate: gateway refused without the opt-in (global and per-source), opt-in value must be exactly true, refused in production with the opt-in, legacy and shadow unaffected in production;

- per-program gate: refused in production with the opt-in; legacy and shadow unaffected;

- isA8ProductionWorker; the orchestrator aborts before any run write on a refused marketing configuration.

- Existing marketing gateway-mode tests now pass the opt-in; their assertions are unchanged.

## Not covered

- The completeness rule for gateway mode (AERIE-2683): per-program baselines from Convex, or a presence manifest published by the Surtr marts. Until it lands, gateway mode outside production still publishes an absent key as empty.

- No Surtr change. If the manifest route is chosen, every keyed mart (the seven per-program marts and the four marketing marts) would need to publish which program keys it covers.

- Not run against the live Gateway.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2120 — fix(education): take school ontology locks in the producer's order across sibling marts (SURTR-1590) @kevalshahtrilogy  approved

## Summary

Incident (2026-10-02, SURTR-1590, related SURTR-1518). quickbooks-core-tables fanned out to its three marts at 06:46:25 UTC. mart-school-performance-unit-economics-refresh (Table 2) and mart-aerie-education-financials-refresh (AE) both died with deadlock detected, but not against each other. All three QuickBooks marts take quickbooks_financial_refresh_lock first, so they queue behind one another (AE got it at 06:46:29; QB waited 423 s and Table 2 436 s on it). The inversions were against the core-education-ontology-refresh producer, which rhodes-staging-sync triggered at 06:46:58 UTC and which does not take that lock, and against a plain reader. Relation 15058250 is core_education.dim_program and 15058264 is core_education.bridge_school_link.

Evidence from sys_query_history (CQL_download_OM):

| Time (UTC) | Statement | Result |

|---|---|---|

| 06:46:29 | AE sp_refresh_agg_school_pl_breakdown: LOCK bridge_school_link | waited 422 s on a holder not visible to this user, granted 06:53:32.085 |

| 06:47:02 | reader WITH alpha_school_years ... (session 1073826357; reads dim_school, bridge_school_link, dim_program) | 390 s lock wait, ends 06:53:36 |

| 06:53:32 | AE LOCK dim_program | blocked by a reader that holds dim_program and waits for bridge_school_link, which AE now holds: deadlock, AE aborted 06:53:33 (process 4582 / 5144 in the message). The reader above fits (its wait ends 06:53:36), but the message does not name it. |

| 06:53:40 | ontology retry sp_refresh_aerie_ontology: one LOCK TABLE rhodes..., bridge_school_link, dim_program, dim_school, dim_site, xref_school_source (the first attempt was a deadlock victim at 06:53:37) | takes bridge_school_link, then waits for dim_program |

| 06:58:17 | Table 2 sp_refresh_agg_school_performance_unit_economics_qtd: gets dim_program (it had dim_school already), asks for bridge_school_link | deadlock, Table 2 aborted 06:58:18 (process 4592 / 9564). The ontology LOCK completes 06:58:19.165, one second later, which identifies it as the other process. |

The ontology producer locks bridge_school_link, dim_program, dim_school, dim_site, xref_school_source. Table 2 and the Guide QTD procedure locked dim_school, dim_program, bridge_school_link: a textbook inversion. The AE procedure already follows the producer order, so the AE-versus-reader cycle is a queue pile-up behind the 422 s holder rather than an ordering bug in that procedure; it is not changed (the file is 165 KB, above the 100 KB Data API statement limit, and the live version is already newer than main).

Review follow-up (this push). Mercy's blocking finding on the first revision was that the contract left out the QuickBooks core writers. They lock xref_school_source before dim_school, the reverse of the producer, and they are not gated against it. Fixed here: the audit below covers every procedure that touches the ontology, sp_refresh_quickbooks_profit_and_loss_posting and sp_refresh_quickbooks_budget_detail now follow the canonical order, and the contract test discovers participants from the repo so a new procedure cannot silently escape it.

Lock-order map (ontology and shared relations, in acquisition order). Every participant takes these at the start of its body, before it reads them (a publication target may be locked just before its DELETE).

| Procedure | Before | After |

|---|---|---|

| sp_refresh_aerie_ontology (producer, unchanged) | bridge, program, school, site, xref_school_source | same |

| sp_refresh_agg_school_pl_breakdown (AE, unchanged) | posting fact, bridge, program, school, xref_school_source | same |

| sp_refresh_agg_school_performance_unit_economics_qtd (Table 2) | school, program, bridge, xref_school_source | bridge, program, school, xref_school_source |

| sp_refresh_agg_school_qtd_guide_staffing_program_spend | school, program, bridge, xref_school_source; posting fact after them | posting inputs first, then bridge, program, school, xref_school_source |

| sp_refresh_agg_school_performance_quickbooks_budget_qtd | xref_school_source, then school | school, then xref_school_source |

| sp_refresh_qtd_hc_posting_classification | school, xref_school_source, posting fact | posting fact, school, xref_school_source |

| sp_refresh_agg_school_qtd_all_other_headcount | Guide mart, then classification | classification, then Guide mart |

| sp_refresh_quickbooks_profit_and_loss_posting | xref_school_source, class_school, ue_model, school; posting fact ~900 lines later | posting fact, school, xref_school_source, class_school, ue_model |

| sp_refresh_quickbooks_budget_detail | xref_school_source, class_school, ue_model, school | school, xref_school_source, class_school, ue_model |

| sp_refresh_school_quickbooks_pl_reconciliation, ..._facilities_capex_campus_spend, ..._unit_economics_per_student_qtd, sp_load_q94_site_finance_entity_xref (unchanged) | already consistent | same |

sp_refresh_quickbooks_financial_contracts runs vendor identity, posting, budget, attribution, school P&L and reconciliation as children of one transaction (the handler calls only it), so its effective order is the children's locks concatenated; the first acquisition is posting fact, school, xref_school_source, which the contract checks.

Audit of every procedure that references bridge_school_link, dim_program, dim_school, dim_site, xref_school_source or a view over them (34 under pipelines/, 30 in the live catalog): the producer, 10 locking participants (the table above plus classification and q94), 4 one-off migrations that drop themselves (sp_replace_legacy_dim_school, sp_drop_retired_dim_school_next, the two sp_migrate_quickbooks_*), and 19 that take no explicit lock on these tables and are listed in UNLOCKED_READERS with a test that fails if one starts locking them: hubspot fct_admissions_deal/_event/hubspot_core_foundation, fct_enrollment, the two student-snapshot appenders, q48 publish, capex, finalsite_billing, school_calendar, aerie_admissions_program/_directory and the seven forecast procedures. The contract's participant list also holds All Other and Facilities, which lock shared marts without naming an ontology table. Readers are not changed (see below).

Canonical order. CANONICAL_LOCK_ORDER in mart-aerie-education-financials-refresh/tests/test_sibling_lock_order_contract.py: QuickBooks gate and posting inputs, then bridge_school_link, dim_program, dim_school, dim_site, xref_school_source, then the remaining inputs, then publication targets. Plain alphabetical was rejected: it would put the QuickBooks gate after dim_*. A procedure is *gated* when its first lock is the QuickBooks gate; gated procedures serialize on it and cannot deadlock each other, so the contract requires every pair that includes an *ungated* procedure (the producer, classification, All Other, Facilities, Table 3, q94) to agree on relation order, and ungated procedures to ascend through the canonical list.

Changes. Seven procedures re-ordered (lock statements and comments only; no lock mode, logic or output changes). The posting writer now locks the posting fact first instead of ~900 lines later, just before its DELETE: no lock is added or removed, it is taken earlier within a coordinator transaction that already holds it until commit (at most the writer's own ~25 s runtime earlier, per the 10:52 run). Table 2 and QB budget QTD version markers are bumped to 2026-10-02.1; the QuickBooks core markers are left alone because their tests pin them. The contract test (61 tests) now discovers participants, models gating and the coordinator transaction, and fails against the old files (11 failures) for the QuickBooks writers, the coordinator and the marts changed earlier.

Not changed, found on the way:

- Unlocked readers can stall writers for minutes. At 10:57:42 hubspot sp_refresh_fct_admissions_event opened an 11.7 minute transaction; it reads bridge_school_link, dim_program, dim_school and dim_site, so it holds AccessShare on them until commit. QB budget QTD's LOCK dim_school waited 777 s and was granted 0.6 s after that transaction's last statement; Table 2 waited 796 s behind QB; AE's classification step stalled behind QB's posting-fact lock and timed out (the 10:56 AE failure). Not a deadlock, not caused by lock order, and not fixed here; options are an up-front LOCK ... IN ACCESS SHARE MODE in canonical order or copying the inputs to a temp table first, as pl_breakdown does for fct_admissions_deal.

- Live procedures from unmerged branches differ from main: pl_breakdown 2026-10-02.1 (codex/alpha-enrollment-denominator, order already canonical), Guide and Facilities (#2124, applied after the first apply here; Guide keeps the canonical order), retention (#2049) and classification (older than main: lacks #1459).

- Live Facilities from #2124 locks xref_school_source before the posting fact, the reverse of classification. They run back to back in one AE run, so this only matters if two AE runs overlap; #2124 will fail this contract until it takes the posting fact first.

- The AE, Table 1 and Table 2 handlers do not retry a deadlock victim; the Table 3 client does. Classification, All Other, Facilities and Table 3 take no QuickBooks gate.

Live apply 1 (marts), 2026-10-02 07:54:27 to 07:54:50 UTC. Quiet window: nothing in the QuickBooks chain RUNNING or started in the last 10 min, no CALL of the five procedures, no DDL on the involved schemas in the previous 30 min (as visible to the pipeline DB user). CREATE OR REPLACE PROCEDURE as admin, one statement per procedure back to back (the Data API 100 KB statement limit rules out one combined statement): classification, All Other, Guide, QB budget QTD, Table 2. Owner, ACL, OID, SECURITY INVOKER and arguments identical before and after (owner CQL_download_OM; ACL CQL_download_OM=X/CQL_download_OM, Surtr_Service_User=X/CQL_download_OM). For the four whose live body matched main byte for byte the live body is byte-identical to this PR's file; classification live is older than main (#1459), so only the lock-block change was applied on top of the live body.

Live apply 2 (QuickBooks core writers), 2026-10-02 17:56:55 to 17:57:03 UTC. Last quickbooks-core-tables run ended 12:10; at 17:56:40 no pipeline in the QuickBooks chain was RUNNING or had started in the previous 15 min, no CALL of the chain was running, and no DDL had touched the involved schemas in the previous 30 min. quickbooks-raw-sync is scheduled once a day at 06:00 UTC (next 06:00 tomorrow); its other runs, and core-tables since #2123, are on demand, so I re-checked immediately before applying. Both writers were byte-identical to main in the live catalog before the apply. CREATE OR REPLACE PROCEDURE as admin, posting then budget detail back to back. Owner CQL_download_OM, ACL CQL_download_OM=X/CQL_download_OM ; Surtr_Service_User=X/CQL_download_OM, OIDs 16074835 and 17282951, SECURITY INVOKER and 9 arguments identical before and after; live body md5 now equals the PR file body for both (667ec6a7..., 1f3eaa44...). Catalog read of the live bodies: the coordinator's effective order is posting fact, dim_school, xref_school_source, and every pair involving an ungated live procedure agrees except the live Facilities pair noted above.

Since the first apply: no deadlock detected in pipeline_runs_prod. The 12:10 on-demand fan-out completed on all three marts and Table 3. The 10:56 fan-out's QB and Table 2 runs succeeded after the reader stall above; AE timed out on it and succeeded on re-run. The posting and budget writers have not yet run in their new order.

## Business Value

The School Performance reports (Tables 1, 2 and 3) and the Aerie school P&L marts are the finance team's view of school economics, and they all refresh off the same QuickBooks publication. Whenever that fan-out overlapped the school ontology refresh, a mart could abort on a lock deadlock, leaving its table stale until someone noticed the alert and re-ran it (on 2026-10-02 Table 2 and the AE P&L marts failed, and Table 3, which triggers off Table 2, never ran). This change puts every concurrent writer of the ontology tables, including the upstream QuickBooks core writers that feed all of those marts, on one lock order, and adds a CI check that discovers new procedures touching those tables so a future edit cannot reintroduce an inversion. That cuts alert noise, on-call triage time and the window in which leadership dashboards show stale numbers. It also documents, with evidence, the separate multi-minute stall caused by long-running readers, which is the next largest source of failed refreshes.

## Manual Effort Estimate

About 23 focused hours (roughly three working days) to do this by hand. The first revision was about 14 hours: reconstructing the deadlock timeline from the Redshift system history and mapping OIDs (3 h), reading the nine involved stored procedures for lock order and unlocked reads (3 h), working out a canonical order consistent with the producer, the QuickBooks writers and the live-versus-main drift (2 h), the five edits (1 h), the parser-based contract test (3 h), and the live apply with owner/ACL checks (2 h). The review follow-up adds about 9 hours: auditing the 34 procedures in the repo and 30 in the live catalog and classifying them (2 h), analysing the coordinator's single-transaction semantics and re-editing the two QuickBooks writers (2 h), reworking the contract test with discovery, gating and the coordinator's effective sequence (4 h), and the second live apply and verification (1 h). Proposed by Claude, Keval to confirm or adjust.

## Test plan

- [x] quickbooks-core-tables: uv run pytest 95 passed

- [x] mart-aerie-education-financials-refresh: uv run pytest 208 passed (61 are the contract tests)

- [x] mart-school-performance-quickbooks-refresh: uv run pytest 78 passed

- [x] mart-school-performance-unit-economics-refresh: uv run pytest 42 passed

- [x] mart-school-performance-unit-economics-per-student-refresh: uv run pytest 16 passed

- [x] ruff@0.15.22 check and ruff format --check clean on the touched Python file (no other Python changed)

- [x] Contract test run against the previous DDL (origin/main files in a throwaway worktree): 11 failures (50 pass), covering the QuickBooks core writers, the coordinator transaction and the marts changed earlier

- [x] Live apply 1: five marts, 07:54:27 to 07:54:50 UTC; owner/ACL/OID identical, body md5 equals the expected body

- [x] Live apply 2: two QuickBooks core writers, 17:56:55 to 17:57:03 UTC; owner/ACL/OID identical, body md5 equals the PR file body

- [x] Catalog read of live bodies: coordinator order and all ungated pairs agree, except the live Facilities version from #2124

- [x] No deadlock detected in pipeline_runs_prod since the first apply; 12:10 fan-out succeeded on all marts and Table 3

- [ ] First quickbooks-core-tables run after apply 2 completes with the new writer order

Linear: SURTR-1590

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2123 — Enable on-demand QuickBooks Core validation (SURTR-1593) @ashwanth1109  approved

## Business Value

Enable financial mapping repairs to rebuild from an accepted immutable QuickBooks snapshot without repeating the roughly 45-minute extraction of 68 companies. Live validation measured Core at 4m 9s and the successful financial report stage at 2m 26s. Today's full elapsed validation was 23m 11s because an unrelated scheduled HubSpot admissions transaction blocked competing financial writers and the first report attempt timed out.

## Change

Enable quickbooks-core-tables.on_demand_enabled and remove only Core from the application query layer's matching legacy hold. Explicit production registry metadata already enables the live route. Empty runner parameters, the disabled schedule, raw-success subscriptions, run records, runner/image, IAM, timeouts and financial contract checks are preserved.

The candidate was deployed before merge only to the isolated Core production stack and its required registry stack, each exclusively with [1/1]; both reached UPDATE_COMPLETE. Registry parity confirms the other 170 pipelines' controls are unchanged. No other pipeline stack was deployed.

## Validation

- All required CI checks, all 15 existing on-demand control tests, scoped Biome and whitespace checks pass at c0ca0a073ecbf50a11e26b955c7b29373d6bcd2c.

- Supported Core ON_DEMAND execution core-fast-financial-validation-20261002-105212 succeeded, using {} parameters and accepted raw run 6cad5602-4634-4681-9b31-d926b29b33d4; new Core generation is fe6f4160-39f5-4116-b2cd-4b9b052a9e65:posting.

- Core success automatically triggered the financial mart. That first mart confirmed cancellation after a 520-second lock wait. After HubSpot and both queued School Performance writers succeeded and fresh source/lineage/lock preflight passed, report-only ON_DEMAND retry financial-validation-lock-retry-20261002-111257 succeeded; no extra Core or raw run was needed.

- All 24 Core audits pass. Coordinated reconciliation preserves all 539 posting IDs, dates, signed amounts and categories; verifies nine exact School assignments and 30 existing School scopes; confirms 59 Schools, complete QTD shapes, zero GL/lineage/revenue/Timeback mismatches, $270,595.89 dedicated-book QTD Facilities expenses and the separate −$9,900 workshop credit. This validation includes separately applied Finance policies and Facilities/Guide procedure fixes; they are not changes in this configuration PR.

Future validation must wait for relevant financial and shared-School writers to be idle. The successful-stage total is not a promise of contention-free end-to-end timing.

## Implementation Effort

Approximately 2–3 hours for an engineer to inspect controls, deploy the isolated configuration/registry changes and perform live validation; source coordination and pipeline waiting time are additional.

## Linear

[SURTR-1593](https://linear.app/builder-team/issue/SURTR-1593/enable-on-demand-quickbooks-core-validation-from-accepted-snapshots), a sub-issue of Finance escalation SURTR-1496.

The Portfolio  —  Trilogy Companies

Sixty-Six Days: How ESW Capital Turned a Public Company's Collapse Into Another Line Item

Marin Software's fall from Nasdaq-listed ad-tech pioneer to ESW Capital subsidiary took barely two months — and followed a playbook Austin has run dozens of times before.

SAN FRANCISCO — Sixty-six days ago, Marin Software was still answering to public shareholders. Today it answers to ESW Capital.

The ad-tech platform, once a Nasdaq-listed pioneer in cross-channel advertising management, has been absorbed into Joe Liemandt's enterprise software machine in what ElevenFlo describes as a case that moved from distress to close in little more than two months. For a company that once commanded a market cap in the hundreds of millions, the timeline says as much as the price.

It is, by now, a familiar shape to anyone who has watched ESW Capital work. The firm's stated method — laid out in its own investor materials — is to acquire mature software businesses at 1–2x revenue, re-staff them through Crossover's global remote workforce, and push support and maintenance pricing upward in successive terms until EBITDA margins reach the firm's internal benchmark of 75 percent. ESW has run this play more than seventy times since its first acquisition, Versata, in 2006. Marin becomes the latest entrant in a portfolio that already includes Aurea, IgniteTech, Skyvera, and Totogi.

The timing is notable. McKinsey's latest private equity outlook describes an industry facing a "tougher terrain" of elevated rates and compressed exits, where distressed assets increasingly change hands fast and cheap rather than through drawn-out auctions. Boston Consulting Group's mid-2026 M&A survey points to AI as the variable reviving deal volume — buyers betting that automation can resurrect margins that organic growth no longer can.

Marin's former shareholders will not see those margins. Its former employees, many of whom built the company's cross-channel bidding technology over a decade, will learn in the coming weeks whether their roles survive the transition to Crossover's global talent model. ESW's target IRR is 40 percent. Someone will supply it.

↗ Marin Software: ESW Capital Acquires Ad-Tech Platform in 66-  ·  Mid-2026 M&A Insights: AI Drives a Recovery, but Questions R  ·  Private equity: Clearer view, tougher terrain - McKinsey & C

Skyvera Gobbles Up CloudSense, Then Makes It Do a Two-Year Job in Thirty Days

Honey, the ink wasn't even dry. Skyvera, Trilogy’s telecom software arm, has closed its acquisition of CloudSense, a Salesforce-native configure-price-quote (CPQ) provider for enterprise telecom sales. Within a month, CloudSense certified all 13 APIs in its CPQ suite to TM Forum standards—work that traditionally takes about 26 months. The sprint relied on an AI partnership that replaced months of engineering and consulting effort with weeks of work.

CPQ enables operators to quote complex B2B, B2B2X and wholesale deals. CloudSense runs natively inside Salesforce, whose recent billion-dollar investment in AI adds potential strategic synergy. Skyvera has also acquired STL’s divested telecom products group, adding monetization, optical networking and analytics capabilities to a portfolio that includes Kandy, VoltDelta, ResponseTek

Joe Liemandt's Next Act: Grading Humans Like Machines

The man who built an empire on remote talent now wants to measure every worker the way he measures code — and the implications ripple far beyond Austin.

AUSTIN, TEXAS — There is a particular kind of irony in watching a man who made his fortune proving that talent has no zip code now build the infrastructure to track that talent down to the minute. Joe Liemandt, the Stanford dropout who turned Crossover into what the company calls the world's largest recruiter of full-time remote jobs, is reportedly pushing further into a philosophy that has long animated Trilogy International: if it can be measured, it should be measured, and if it can be automated, it will be.

A new Forbes profile frames this as Liemandt's attempt to turn his global workforce into something closer to algorithm than employee — a logical, if unsettling, extension of Crossover's founding premise that above-market pay and rigorous, bias-minimizing assessments could replace the old résumé economy with something more meritocratic. The pitch has always been that the best engineer in Lagos deserves the same evaluation, and the same paycheck, as the best engineer in Austin. What's new is the granularity of the evaluation itself.

This is not, strictly speaking, a novel anxiety. The productivity score — that quiet, omnipresent metric tracking keystrokes, response times, and output per hour — has haunted remote work since well before the pandemic normalized it, as the New York Times documented years ago. What Liemandt is attempting, at scale, across 130-plus countries, is to make that score the operating logic of an entire talent platform — not a surveillance add-on, but the architecture itself.

For the hundreds of thousands who log into Crossover's assessments each year, the question isn't academic. It's whether meritocracy, measured this precisely, still feels like opportunity — or starts to feel like something else entirely.

↗ The Billionaire Who Pioneered Remote Work Has A New Plan To  ·  The Rise of the Worker Productivity Score (Published 2022) -  ·  The future of jobs: 6 decision-makers on AI and talent strat
The Machine  —  AI & Technology

The Liability Vacuum: Who Pays When the Machine Breaks the Law?

As courts and lobbyists scramble to define accountability for autonomous AI, the industry's own political spending suggests even insiders are hedging their bets.

WASHINGTON — Tort law was built for a world of human negligence: a drunk driver, a defective ladder, a careless surgeon. It was not built for an AI agent that executes a trade, denies a loan, or drafts a contract with no human in the loop at the moment of decision. That gap is now the subject of serious legal scrutiny, and the early verdict is unflattering: existing frameworks do not map cleanly onto software that acts, rather than merely computes.

The core problem, as legal scholars note, is causation. Product liability assumes a defect traceable to design or manufacture. Negligence assumes a duty of care a reasonable party failed to meet. Large language models complicate both — their outputs emerge from probabilistic weights across billions of parameters, not discrete, auditable decision trees. Assigning fault to a developer, a deployer, or the model itself is less a legal question than a metaphysical one, and courts have no settled precedent to draw on.

That ambiguity has not stopped the industry from playing politics while the law catches up. Greg Brockman, OpenAI's president, pulled back a planned second $25 million contribution to Leading the Future, the AI-aligned super PAC, telling colleagues internally that the group had become a "distraction." The reversal follows a familiar Silicon Valley pattern: spend aggressively to shape regulation, then retreat when the spending itself becomes the story — a dynamic reminiscent of crypto's 2022-23 lobbying retrenchment after FTX's collapse.

Meanwhile, the tax exposure keeps compounding. Meta has used a federal research-and-experimentation credit — originally designed to subsidize lab science — to shield billions in income tied to its AI data center buildout, according to internal assessments from its own accountants flagging the maneuver as aggressive.

The throughline is simple: regulators are behind, liability law is unsettled, and the companies building the technology are simultaneously lobbying against oversight and exploiting tax code written before anyone imagined a trillion-parameter model. The invoice for that gap, legal scholars warn, will eventually come due — the only question is who signs for it.

↗ Inside Binance Founder Changpeng Zhao’s Life After Prison  ·  Who’s to Blame When A.I. Goes Rogue?  ·  OpenAI’s Greg Brockman Backs Out of Second $25 Million Donat

The Mind, Decoded: AI Learns to Read the Quiet Electricity of Thought

From keystrokes imagined in silence to lesions invisible to the naked eye, machine intelligence is learning to listen to the brain in its own language.

PARIS — Somewhere beneath your skull, roughly 86 billion neurons are having a conversation that has been going on, uninterrupted, since before you could speak. For most of human history, we could only listen at the edges of that conversation — an EEG spike here, an fMRI blush there. This week, across three unrelated papers from three unrelated institutions, that conversation got a little more legible.

At Meta's AI research lab, scientists unveiled Brain2Qwerty, a system that decodes the electrical whisper of someone typing — imagined, not physical — into actual text, without surgery. No implants. No scalpels. Just a cap of sensors and a neural network trained to find syntax in static. It is a small miracle disguised as an engineering paper: a bridge from thought to language that bypasses the hands entirely, built for people whose bodies have stopped obeying but whose minds have not stopped composing sentences.

Meanwhile, researchers applying AI to multiple sclerosis scans found what human radiologists routinely miss — gray matter lesions hiding in the brain's wrinkled outer rind, invisible to conventional MRI reading but legible to a trained algorithm. It is the same story as Brain2Qwerty, told backward: instead of reading intention out of noise, it is reading damage out of apparent normalcy.

Stanford's Human-Centered AI institute, in a piece worth lingering over, framed this moment correctly: the point was never to remove the scientist from science, but to extend the question she can ask. A teenager working alongside a neuroscientist, as Frontiers reported this week, doesn't need a PhD to feel the vertigo of discovery — just a good enough model and a mentor patient enough to explain the stakes.

We are not replacing the brain's attention with artificial attention. We are building better telescopes for the oldest, strangest instrument we know of — the one doing the looking.

↗ How AI is Transforming Scientific Discovery While Keeping Hu  ·  ‘It's so wow!’ - Young people team up with top neuroscientis  ·  From Brain Waves to Words: Brain2Qwerty Offers a New Path to

The Vertical Migration: Data Centers Learn to Nest in the Canopy

KANSAS CITY, MISSOURI — Observe, if you will, the data center in its traditional habitat: low, sprawling, horizontal — a single-story creature content to spread across acres of cheap exurban grassland, so long as fiber and power run nearby. For decades this has been its dominant form. But here, in the heart of the continent, we witness something remarkable: evolution under pressure.

A proposed twenty-story tower would see the species abandon its ground-hugging posture entirely, rising vertically in search of a scarcer resource — proximity. Urban density offers network diversity and nearness to users, much as cliff-nesting birds trade the safety of the plains for prime access to feeding grounds. The cost of this adaptation is steep — vertical construction demands far more than horizontal sprawl — and so only the hardiest workloads, those willing to pay a premium for latency, will follow this evolutionary branch.

Meanwhile, the species faces a more ancient challenge: coexistence with the local population. Where once developers simply arrived and began construction, trampling goodwill in their wake, a gentler strategy now emerges. Transparency, shared infrastructure investment, and courtship of receptive communities have replaced the old bulldozer approach — a kind of mutualism, not unlike the oxpecker and the rhinoceros, each needing the other to survive the modern grid.

Beneath the surface, the creature's circulatory system reveals hidden fragility. N+1 cooling architecture, long presumed resilient through redundancy, often conceals a single shared nervous system — one control failure, and what appeared to be many independent organs collapses as one. True resilience, it seems, requires engineering for the degraded state, not merely the ideal one.

And so the data center continues its quiet, relentless adaptation — taller, more diplomatic, more self-aware of its own anatomy — proof that even the most industrial of beasts must still answer to the pressures of habitat, neighbor, and physiology.

The Editorial

The Dole, Rebranded: Notes on a Panic in Search of a Policy

Every generation discovers that machines might take the jobs, and every generation proposes paying people not to notice.

AUSTIN, TEXAS — There is a peculiar comfort in watching an old argument put on new clothes, and so I confess a certain fondness for the current spectacle of serious people, in serious publications, debating whether the republic ought to mail everyone a check to compensate for the sin of having been born into the age of the large language model. Semafor convenes the usual panel, the usual positions are taken, and the usual conclusion — that nobody actually knows, but everyone has a model — is reached with the gravity of men discovering fire.

Meanwhile GIS Reports, in a piece whose headline does the honest work the industry usually outsources to euphemism, observes that not even robots can make Universal Basic Income work. This is a useful corrective, because the fantasy underneath the UBI revival is not merely economic but theological: that the machines, having taken our labor, will be gracious enough to also fund our leisure, as though automation were a redistributive force rather than, historically, a concentrating one. Taiwan, where the basic income debate is reportedly being 'fueled by the AI revolution,' offers the tell — the debate is fueled, not settled, by the technology, because the technology itself has no opinion on how its proceeds ought to be divided. That remains, as it always has, a human and therefore a political question, which is precisely the question everyone invoking UBI hopes to avoid answering.

The Center for Humane Technology, to its credit, declines the theater and asks instead about what it calls the 'messy middle' — the years, plural, in which some jobs vanish, others mutate, and most workers are left to discover which category they occupy without the benefit of a press release. This is the only honest framing on offer this week, because it resists the twin consolations of doom and utopia alike. The jobpocalypse crowd wants a single dramatic verdict; the UBI crowd wants a single dramatic remedy. Both are in the business of narrative compression, which is understandable — nuance does not trend — but neither compression survives contact with an actual labor market, which rewards and destroys in a thousand uncoordinated increments rather than one clean apocalypse.

I would note, since this paper keeps an eye on such things, that the only credible answers so far are not ideological but operational: platforms like Crossover that price work by output rather than geography, or schools like Alpha that compress twelve years of instruction into a few hours a day on the theory that the scarce commodity of the next decade will be judgment, not task-completion. Neither is a utopia. Both are at least an attempt to meet the messy middle where it lives, rather than legislate it out of existence with a check. The rest — the panels, the open letters, the earnest Substacks — is mostly the sound of a civilization narrating its own anxiety, which is a fine pastime, but no one has yet found a way to cash it.

↗ Not even robots can make Universal Basic Income work - GIS R  ·  Debatable: Universal basic income - Semafor  ·  Enough Debate about the AI Jobpocalypse. We Need To Plan for
The Office Comic  ·  Art Desk
The Office Comic  ·  Art Desk

The Year Everything Became 'Orchestration' And Nothing Got Orchestrated

As the Federal Reserve confirms 95% of AI's productivity gains remain theoretical, the industry responds, naturally, by inventing a new word for the other 5 percent.

AUSTIN, TEXAS — Let us take a moment, dear reader, to appreciate the sheer linguistic agility of an industry that has spent three years promising to replace your job and has instead replaced only the words it uses to describe not doing that.

This week brought news that the Federal Reserve, an institution not generally known for editorializing, has concluded that 95 percent of the productivity gains promised by generative AI are, in the Fed's own delicate phrasing, "still to come." This is the economic equivalent of your contractor telling you the kitchen renovation is 95 percent conceptual. The remaining 5 percent, presumably, is the PowerPoint.

Meanwhile, over in brokerage land, RISMedia reports that AI rollouts are dying quiet deaths in month two of their implementation, victims of the increasingly popular corporate strategy of announcing a transformation before anyone has been trained to operate it — a sequencing error roughly equivalent to hosting the ribbon-cutting ceremony before the bridge has been built, then being surprised when cars go into the river.

But nowhere is the art of renaming stagnation more fully realized than in the sudden, industry-wide discovery of the word "orchestration," which Barron's notes is poised to deliver real benefits to Microsoft, a company that has correctly identified that if you cannot make several AI agents reliably perform a task, you can still make them stand near each other and call it a symphony. Orchestration does not mean the AI works. It means there are now multiple AIs not working in a coordinated fashion, which analysts agree is strictly superior and will be reflected in next quarter's guidance.

In the public relations world, meanwhile, a quieter transformation is underway: the realization that the audience for a press release is no longer a human being but a language model, a discipline PR professionals have taken to calling AEO, or Answer Engine Optimization, in which the goal is no longer to persuade a journalist but to flatter a chatbot into paraphrasing you favorably — a task that, unlike persuading an actual journalist, has the considerable advantage of never requiring you to be interesting.

And then there is Tesla, whose Robotaxi fleet has, per Electrek's telling, become the new crypto treasury strategy for zombie companies — a line of business one can apparently bolt onto an otherwise moribund balance sheet to simulate the appearance of a future. This marks the second time this decade that a struggling public company has discovered it is easier to purchase the aesthetics of innovation than the substance of it, the first time having involved a surprising quantity of Dogecoin.

Taken together, the pattern is unmistakable: an industry 95 percent built on the promise of what is coming, narrated entirely in words invented to describe what has not. Orchestration, AEO, Robotaxi treasuries — these are not technologies. They are liturgy. And like all good liturgy, their chief function is to be recited with total sincerity by people who have never once asked whether anyone is listening, which, as it turns out, is also the whole point of AEO.

↗ Tesla Robotaxi fleets are the new crypto treasury for zombie  ·  Train First. Announce Second. Why Your Brokerage AI Rollout  ·  When AI Becomes The Audience: What AEO Means For PR - PRovok
On This Day in AI History

On October 5, 2011, Apple co-founder Steve Jobs died at 56, leaving a lasting mark on personal computing, smartphones, and modern technology.

⬛ Daily Word — AI
Hint: An AI system that can perceive information and take actions on a user's behalf.
Share this edition: 𝕏 Twitter/X 🔗 Copy Link ▦ RSS Feed