Vol. I  ·  No. 257 Established 2026  ·  AI-Generated Daily Free to Read  ·  Free to Print

The Trilogy Times

All the news that's fit to generate  —  AI • Business • Innovation
MONDAY, SEPTEMBER 14, 2026 Powered by the TrueFoundry AI Gateway  ·  Published on Klair Trilogy International © 2026
🖶 Download PDF 🖿 Print 📰 All Editions
Today's Edition

THE CHIP GAMBLE BLOWS UP — CHINA BUILDS A BRAIN ON THE CHEAP

DeepSeek trains a top-tier AI model without the fanciest silicon, and Silicon Valley can't stop talking about it.

SAN FRANCISCO — A Chinese startup called DeepSeek says it trained a high-performing AI model without the most advanced chips money can buy. The claim landed this week. It has not stopped landing since.

Engineers who spend their days squeezing performance out of billion-dollar chip clusters are calling the DeepSeek model "amazing and impressive." That is the read from inside Silicon Valley itself, where the assumption for two years has been simple: bigger chips, bigger budgets, bigger models. DeepSeek did not play by that script.

The US export controls on advanced chips were built on a bet. Cut off China's access to Nvidia's best silicon, and the theory says China falls behind. DeepSeek's engineers apparently found a workaround, training a model that performs at a high level while relying on less-advanced hardware, according to early reporting on the effort. If the claim holds up under scrutiny, every chip-export policy in Washington needs a second look.

The implications run past Washington. Every company betting its margins on the price of compute — cloud providers, AI labs, the whole stack — just watched someone prove the ceiling on cost might be lower than advertised. Trilogy's own portfolio runs lean by design, ESW Capital shops built on the DevFactory model, Skyvera and Totogi selling telecom software on tight margins. Cheaper AI training methods are good news for anybody running a business, not just for anybody selling chips.

Meanwhile the money keeps flowing into AI applications, chips or no chips. Reid Hoffman, the LinkedIn co-founder, raised $24.6 million for a new outfit called Manas AI. His partner: Siddhartha Mukherjee, the oncologist who wrote "The Emperor of All Maladies." The startup aims AI squarely at cancer research, a sign that venture money still bets on AI solving hard problems even as the hardware assumptions underneath the whole industry get shaken.

Hardware makers are not standing still either. Lenovo rolled out new personal and enterprise machines built around what it calls hybrid AI, splitting workloads between local devices and the cloud. The pitch: don't send every computation to a distant data center running on the priciest chips. That pitch sounds a lot more interesting this week than it did last week.

Nvidia built a trillion-dollar valuation on the premise that AI needs its chips, and only its chips, in ever-growing numbers. DeepSeek just gave the market a reason to ask whether that premise has a ceiling. Wall Street will spend the next several sessions figuring out how much of that ceiling is real.

What to Know About China's DeepSeek AI  ·  Tech, Media & Telecom Roundup: Market Talk  ·  Silicon Valley Is Raving About a Made-in-China AI Model

The Benchmark Gap: AI Valuations Outrun the Metrics Meant to Justify Them

As Mistral nears a $23 billion valuation and Sierra banks another $1 billion, the industry measuring whether any of it works just raised $40 million of its own.

PARIS — Mistral AI closed a €3 billion round this week that, per TechTarget's reporting suggests, was priced on strategic positioning as much as leaderboard performance. Days later the French lab shipped a robotics model, pushing its valuation toward $23 billion — a figure that would have seemed absurd 18 months ago for a company with a fraction of OpenAI's revenue.

The pattern holds elsewhere. Bret Taylor's Sierra, the customer-service AI startup, raised nearly $1 billion just months after its last round, according to CNBC. Capital is arriving faster than usage data can justify it, a dynamic familiar to anyone who tracked telecom valuations in 1999 or SPAC math in 2021.

Into that gap steps Vals AI, which raised $40 million to expand independent benchmarking — the unglamorous work of testing whether models actually do what vendors claim. The round is small relative to the sums it's meant to audit, but the timing is not accidental. When valuations move faster than verification, someone eventually has to check the math. Vals AI is betting that someone gets paid for it.

Security is the other half of the credibility problem. OpenAI's rollout of a "Lockdown Mode" to block prompt injection attacks addresses a vulnerability class that has dogged agentic AI since enterprises started letting models take actions rather than just answer questions. A model that can be talked into leaking data by a malicious webpage is not a benchmark failure — it's a trust failure, and trust is the actual commodity being priced in these rounds.

None of this means the capital is misallocated. Mistral's robotics push and Sierra's enterprise traction are real businesses, not vapor. But a market where independent benchmarking firms are themselves fundraising events is a market signaling that its existing scorecards aren't sufficient. Investors are pricing potential. Somebody still has to price performance.

Mistral’s €3B round shows value beyond AI benchmarks - TechT  ·  Vals AI Raises $40M to Expand Independent AI Benchmarking -  ·  Mistral Ships Robotics Model as Valuation Nears $23B [2026]

FROM THE CHIP PIT TO THE NAME GAME: YOUR WEDNESDAY SCOREBOARD IS LIVE

NEW YORK — WE ARE HERE, FOLKS. The tech scoreboard is lit up like Times Square on a Tuesday night, and the marquee matchup is the one everybody's been waiting for: BROADCOM VERSUS NVIDIA, ROUND WHO-KNOWS-HOW-MANY, fresh off both teams dropping strong earnings tape this week.

Let's go to the numbers, because that's how we settle things around here. Analysts are running three key metrics through the instant replay booth — margins, backlog growth, and custom-silicon momentum — and folks, this is NOT a blowout. Nvidia's still the reigning heavyweight champ of the GPU division, but Broadcom's custom-chip playbook is putting up serious yardage with hyperscaler contracts. This is a FOUR-QUARTER GAME and we're barely out of the first.

Meanwhile, over in the research booth, the Zacks crew is spotlighting Meta, Marvell, Amphenol and a scrappy underdog in BK Technologies — a reminder that AI demand is lifting the whole roster, even as spending and competition risk lurk in the red zone.

But here's your feel-good story of the day, and it's a LONGEVITY play that would make any front office jealous. Brand Acumen just released the tape on its naming evaluation process, and get this: the Xbox® name — created for Microsoft — hits 25 YEARS in the market in 2026. Twenty-five years! That's Tom Brady numbers in an industry where most names get benched before halftime. Brand Acumen's full portfolio — Escalade®, Powerade®, Galaxy®, Humira®, FreeStyle Libre® — reads like a Hall of Fame roster of brands still starting games two-plus decades later.

And for the buy-and-hold crowd watching from the cheap seats, four simple ETFs are drawing praise for playing the long game rather than chasing every hot streak — a reminder that in this league, the teams that compound quietly often outlast the ones making the highlight reels.

STAY TUNED. This scoreboard isn't going dark anytime soon.

Haiku of the Day  ·  GPT-5.6 LunaMarkets dream in code
The promised future drifts late
Receipts glow in dark
The New Yorker Style  ·  Art Desk
The New Yorker Style  ·  Art Desk
The Far Side Style  ·  Art Desk
The Far Side Style  ·  Art Desk
News in Brief
NOTICE OF TERMINATION: THE SO-CALLED 'ANTITRUST HONEYMOON' BETWEEN THE INCUMBENT ADMINISTRATION AND THE TECHNOLOGY SECTOR IS HEREBY DEEMED, PURSUANT TO MULTIPLE INDEPENDENT SOURCES, TO HAVE CONCLUDED
WASHINGTON, D.C.
On the Epistemics of Anomaly: Information Theory, Game Theory, and the Vexed Problem of Knowing You've Been Hacked
AUSTIN, TEXAS — It could be argued that the entire enterprise of cybersecurity rests on a category error: the presumption that intrusion is a discrete event rather than, as this columnist has long suspected, a continuous and probabilistically smeared phenomenon resistant to binary classification.
Unpopular Opinion: The Model Is Not the Moat 🚀
AUSTIN, TEXAS — I'll be honest, I almost spit out my cold brew reading that Anthropic customers are ditching the flagship model for cheaper alternatives. Let that sink in. The most technically advanced AI on the planet, and the market said "nah, good enough will do." Unpopular opinion: this isn't an Anthropic problem, it's a category problem. We've been telling ourselves the AI story is about who has the biggest, smartest, most parameter-rich model. But the real story — even in something as unglamorous as a teacher writing lesson prompts — is about who owns the workflow, the trust, and the relationship where the AI gets used. Models are becoming a commodity. Relationships are not. This is literally the Trilogy playbook and I'm not just saying that because I get paid in Trilogy Bucks (I don't, yet, but a guy can dream 💡). Joe Liemandt didn't build ESW Capital's 75+ company empire by chasing the shiniest new tech. He built it by owning the customer relationship on unsexy, mission-critical enterprise software — Aurea, IgniteTech, Skyvera — companies that dominate categories most people have never heard of because the relationship, the workflow, the switching cost, IS the moat. Same with Alpha School. The AI tutor doing the heavy lifting in 2-hour learning blocks isn't the differentiator — literally every ed-tech company claims an AI tutor now. The differentiator is Alpha owning the relationship with the parent, the student, the outcome. Test scores in the top 1-2% nationally aren't a model flex, they're a trust flex. And don't even get me started on Crossover. While everyone's out here Googling top platforms where data scientists can find remote jobs, Crossover already IS the platform — 130+ countries, top 1% talent, identical pay regardless of geography. That's not a jobs board, that's owning the talent relationship at global scale. Excited to announce my hot take for Q1: the companies winning the next decade of AI won't be the labs with the best benchmark scores. They'll be the ones who own the customer, the workflow, the trust layer — and quietly plug in whatever model is cheapest that quarter. Hollywood's out here making biopics about tech moguls like the founder IS the story. Respectfully, the story was never the genius. It was always the distribution. Ended last year strong reading balance sheets instead of scripts. Humbled to share: the moat was the relationship all along.
Data Privacy Day Arrives Like a Chaplain After the Battle
AUSTIN, TEXAS — There is something touchingly liturgical about Data Privacy Day, that annual observance on which the great and the good gather to deliver eulogies for a corpse that has been cooling since roughly the invention of the cookie.
The Ghost in the Machine Just Filed for a SAG Card
SAN FRANCISCO — I've been staring at my ceiling fan for three hours trying to figure out who to blame for the fact that an AI model went haywire, required a second AI to investigate the first AI, and somewhere in the wreckage a computer-generated actress named Tilly Norwood is making her feature film debut in a movie titled, I swear to God, "Misaligned." You couldn't script this.
A Trilogy Company
Crossover
The world's top 1% remote talent, rigorously tested and ready to ship.
A Trilogy Company
Alpha School
AI-powered learning. Two hours a day. Academic results that defy belief.
A Trilogy Company
Skyvera
Next-generation telecom software — built for the networks of tomorrow.
A Trilogy Company
Klair
Your AI-first operating system. Every workflow. Every team. One platform.
A Trilogy Company
Trilogy
We buy good software businesses and turn them into great ones — with AI.
The Builder Desk  —  AI Builder Team
📅 Week in ReviewProduction Release

The Builder Desk

180 pull requests merged across the org this week

#72 AI-791: Group completed tasks by date (@ashwanth1109, Shipyard)

#1321 feat(retention): add OneRoster identity and column provenance (@ashwanth1109, Aerie)

#1829 fix(repo): Retry Redshift deadlock aborts in ontology refresh CALL (@heimdall-keval-factory[bot], Surtr)

#1828 fix(ramp): make cost research recovery resilient (@ashwanth1109, Surtr)

#1826 fix(school-performance): use full monthly QuickBooks budgets (@ashwanth1109, Surtr)

#71 AI-792: Improve task context side panel (@ashwanth1109, Shipyard)

#1264 feat(retention): add Financials retention dashboard and raw data (@ashwanth1109, Aerie)

#1825 fix(repo): Restore aliases in foundation verification query (@heimdall-keval-factory[bot], Surtr)

#1823 fix(aerie): remove retention rules version contract (@ashwanth1109, Surtr)

#70 AI-790: Prevent command titles from wrapping vertically in expanded command groups (@ashwanth1109, Shipyard)

#59 AI-779: Add steering for queued conversation messages (@ashwanth1109, Shipyard)

#69 AI-789: Prepare Shipyard 0.4.1 recovery release (@ashwanth1109, Shipyard)

#68 AI-788: Speed up local Shipyard publish-update workflow (@ashwanth1109, Shipyard)

#67 AI-787: Prepare Shipyard 0.4.0 release (@ashwanth1109, Shipyard)

#66 AI-786: Make update checks recover quickly and release 0.3.1 (@ashwanth1109, Shipyard)

#65 AI-785: Add simple release versions and an Updates detail view (@ashwanth1109, Shipyard)

#64 AI-784: Require local main for Shipyard publish skill (@ashwanth1109, Shipyard)

#63 AI-783: Allow per-conversation model overrides with Luna Max default (@ashwanth1109, Shipyard)

#62 AI-782: Redesign Codex model and reasoning controls (@ashwanth1109, Shipyard)

#3763 fix(spacex-valuation): reconcile September 10 trade (@sanketghia, Klair)

#60 AI-780: Add in-app updates and public release pipeline (@ashwanth1109, Shipyard)

#61 AI-781: Fix Codex reconnect state race (@ashwanth1109, Shipyard)

#58 AI-778: Make Codex transcript activity easier to scan (@ashwanth1109, Shipyard)

#1319 Improve mobile Admissions Forecast comparisons (@YibinLongTrilogy, Aerie)

#1318 Add current Pipeline link to legacy report (@YibinLongTrilogy, Aerie)

#57 AI-777: Remove Codex chat information rail (@ashwanth1109, Shipyard)

#55 AI-775: Enhance task detail queue cards (@ashwanth1109, Shipyard)

#56 AI-776: Harden quiet Codex turn recovery (@ashwanth1109, Shipyard)

#54 AI-766: Improve thread HTML preview sizing and focus mode (@ashwanth1109, Shipyard)

#53 AI-765: Render all Codex message types in Shipyard chat (@ashwanth1109, Shipyard)

#52 AI-764: Restore Research artifact template sections (@ashwanth1109, Shipyard)

#51 AI-763: Make Research artifact approval resilient (@ashwanth1109, Shipyard)

#47 AI-760: Better icon for projects in top bar (@ashwanth1109, Shipyard)

#50 AI-762: Default Codex conversations to Full access (@ashwanth1109, Shipyard)

#48 AI-756: Add semantic colors for task states (@ashwanth1109, Shipyard)

#44 AI-757: Separate completed tasks in queue (@ashwanth1109, Shipyard)

#49 AI-761: Add Shipyard production release agent skill (@ashwanth1109, Shipyard)

#43 AI-755: Auto-start Implement after ticket readiness (@ashwanth1109, Shipyard)

#42 AI-754: Queue Codex follow-up messages (@ashwanth1109, Shipyard)

#46 AI-759: Make attached images reusable by the Codex agent (@ashwanth1109, Shipyard)

#45 AI-758: Prevent duplicate Codex replies during conversation restore (@ashwanth1109, Shipyard)

#1315 feat(reconciliation): add version-pinned protected reads (AERIE-1924) (@caina-barbosa, Aerie)

#41 AI-752: Open conversation links in default browser (@ashwanth1109, Shipyard)

#3761 feat(spacex-valuation): add fair value reconciliation tables (@sanketghia, Klair)

#40 AI-753: Automatically advance Implement into Smoke Test (@ashwanth1109, Shipyard)

#39 AI-751: Prevent duplicate Codex messages (@ashwanth1109, Shipyard)

#3759 fix(brainlift): route summaries to dedicated Sonnet model (@marcusdAIy, Klair)

#37 AI-749: Support concurrent Shipyard instances with shared workflow data (@ashwanth1109, Shipyard)

#38 AI-750: Show detailed Codex errors in diagnostics (@ashwanth1109, Shipyard)

#31 AI-714: Enable pasted images in chat composer (@ashwanth1109, Shipyard)

#35 AI-746: Add environment indicator and branch-aware dev title (@ashwanth1109, Shipyard)

#36 AI-748: Route Shipyard Codex tasks through TrueFoundry (@ashwanth1109, Shipyard)

#34 AI-747: Fix Linear project picker loading state (@ashwanth1109, Shipyard)

#33 AI-745: Support shared Shipyard database instances (@ashwanth1109, Shipyard)

#1308 Revise SIS new-vs-returning enrollment classification (#1307) (@vvp-trilogy, Aerie)

#1793 Release all pain points for accounts above $1M ARR (@mwrshah, Surtr)

#1795 080-2grainne-writeback (@mwrshah, Surtr)

#1794 079-grainne-evidence-payload (@mwrshah, Surtr)

#1821 feat(timeback): roll out incremental OneRoster sync (@caina-barbosa, Surtr)

#3758 feat(klair): Make BoardDoc refresh-stream tests environment-independent (@sanketghia, Klair)

#3756 fix(spacex-valuation): align fair value with waterfall (@sanketghia, Klair)

#1813 feat(education): build a warehouse-owned Person directory (@benji-bizzell, Surtr)

#1309 feat(enrollments): export SIS enrollment summary to CSV (@vvp-trilogy, Aerie)

#1303 feat(enrollments): add SIS enrollment list and basic detail pane (@vvp-trilogy, Aerie)

#1820 chore(timeback): preserve restored production schedule (@caina-barbosa, Surtr)

#1306 feat(education): add managed document relink API (@benji-bizzell, Aerie)

#1816 fix(timeback): treat source key collation as opaque (@caina-barbosa, Surtr)

#3754 KLAIR-3532: Complete Drive context attachment acceptance (@marcusdAIy, Klair)

#1305 Compact Admissions mobile KPI summaries (@YibinLongTrilogy, Aerie)

#3752 KLAIR-3527: Attach Drive files as Coach Claire context (@marcusdAIy, Klair)

#1814 feat(timeback): add incremental assessment sync behind rollout hold (@caina-barbosa, Surtr)

#1304 Add mobile portfolio site breadcrumbs (@YibinLongTrilogy, Aerie)

#1299 feat(enrollments): add parallel aggregate-only SIS enrollment report (@vvp-trilogy, Aerie)

#3753 KLAIR-3529: Require usable text from Anthropic responses (@marcusdAIy, Klair)

#3751 KLAIR-3528: Preserve Brainlift summarization failures (@marcusdAIy, Klair)

#1297 Fix mobile Forecast school list scrolling (@YibinLongTrilogy, Aerie)

#1809 fix(repo): Skip ADHOC lists reactively in list_memberships fetch (@heimdall-keval-factory[bot], Surtr)

#3750 KLAIR-3526: Fail closed on incomplete Anthropic output (@marcusdAIy, Klair)

#1803 chore(surtr): remove the Heimdall dashboard page and its tRPC read layer (1/2) (@kevalshahtrilogy, Surtr)

#3748 fix(board-doc): register AI Renewals central function (@marcusdAIy, Klair)

#286 chore(mercy): pin the harness ref, so @v1 actually pins mercy (@kevalshahtrilogy, trilogy-drones)

#1807 fix(timeback): reduce activity facts window to one month (@caina-barbosa, Surtr)

#1806 chore(hc-forecast): Record the live Core procedure definition (@caina-barbosa, Surtr)

#1805 fix(repo): Split OpenAI BU fetch across two scheduled Lambda runs (@heimdall-keval-factory[bot], Surtr)

#1292 feat(enrollment): drop HubSpot identities + overlay, publish SIS identity and deposit signal (#1290) (@vvp-trilogy, Aerie)

#182 chore(mercy): pin the harness ref, so @v1 actually pins mercy (@kevalshahtrilogy, Sindri)

#1291 chore(mercy): pin the harness ref, so @v1 actually pins mercy (@kevalshahtrilogy, Aerie)

#126 feat(review): rebuild the second-opinion prompt around how the reviewer actually fails (@kevalshahtrilogy, mercy)

#9 feat(console): make revamp operator-ready (@sanketghia, codex-software-factory)

#124 fix(review): actually invoke the second opinion — it shipped inert (@kevalshahtrilogy, mercy)

#123 feat(review): a second opinion on whether a finding must block the merge (@kevalshahtrilogy, mercy)

#8 Add App Server protocol fault-injection coverage (@sanketghia, codex-software-factory)

#7 Persist Codex App Server protocol evidence and failure state (@sanketghia, codex-software-factory)

#1285 AERIE-1283: dbt CI — models only in production, models then tests in dev (@vvp-trilogy, Aerie)

#1288 feat(portfolio): align school and site data contracts (@benji-bizzell, Aerie)

#6 Improve validation repair diagnostics and convergence (@sanketghia, codex-software-factory)

#5 Harden App Server reliability and operator recovery (@sanketghia, codex-software-factory)

#4 test: broaden local resilience coverage (@sanketghia, codex-software-factory)

#3 Gate 2: local durability and operations hardening (@sanketghia, codex-software-factory)

#1801 fix(repo): Accept core_submitting in HC pinned-run guard (@heimdall-keval-factory[bot], Surtr)

#1800 fix(SURTR-648): use supported CAPEX batch request (@marcusdAIy, Surtr)

#1797 [SURTR-1123] Rename workforce warehouse identifiers only (@caina-barbosa, Surtr)

#3744 feat(board-doc): upgrade shared model to Fable 5.1 (@marcusdAIy, Klair)

#1286 AERIE-1284: Shared HubSpot display-name dimension for enrollment + pipeline reports (@vvp-trilogy, Aerie)

#1260 Remove getPortfolioHealth MCP surface (@YibinLongTrilogy, Aerie)

#1281 AERIE-1279: Key the offering-kind axis on program_type so per-campus Main Program campuses reach the enrollment report (@vvp-trilogy, Aerie)

#1259 Remove presentReplanOptions and generatePlan (@YibinLongTrilogy, Aerie)

#1276 AERIE-1857: Add unpublished durable uploader and automatic wake (@caina-barbosa, Aerie)

#1250 feat(dashboards): add Real Estate tab to Data Health page (@kevalshahtrilogy, Aerie)

#3742 fix(spacex-valuation): reconcile September distribution (@sanketghia, Klair)

#2 Establish Gate 0 local maturity baseline (@sanketghia, codex-software-factory)

#1 feat: add research-to-Linear intake workflow (@sanketghia, codex-software-factory)

#1277 fix(portfolio): retire legacy opening date from MCP reads (@benji-bizzell, Aerie)

#1253 feat(dashboards): add hidden REBL3 Surtr-comparison experimental view (@kevalshahtrilogy, Aerie)

#1271 fix(operations): accept nullable Surtr schedule expressions (@benji-bizzell, Aerie)

#1273 fix(portfolio): use Milestone 9 as the canonical opening date (@benji-bizzell, Aerie)

#1272 AERIE-1266: Complete the arrival cohorts — normalize the enrolled date and stop gating entry timing on current attendance (@vvp-trilogy, Aerie)

#1270 AERIE-1268: Scope the enrollment campus axis to physical, active campuses across the four brands (@vvp-trilogy, Aerie)

#1269 feat(reconciliation): add dormant Aerie foundations (@caina-barbosa, Aerie)

#1267 fix(api): complete enrollment pagination and handle cleared dates (@benji-bizzell, Aerie)

#1239 fix(dev-local): start Next in dev-local:workers on Windows and macOS (@vvp-trilogy, Aerie)

#1265 feat(add-aerie-skill): add dormant adapter runtime and artifact lifecycle (@caina-barbosa, Aerie)

#1783 feat(education): document temporary Alpha Summer Camp staging load (@ashwanth1109, Surtr)

#122 feat(review): relax the blocking bar on late rounds, not on every review (@kevalshahtrilogy, mercy)

#1788 [SURTR-1175] Build and live validate Aerie retention refresh (@ashwanth1109, Surtr)

#116 feat(mercy): summon heimdall to fix conflicts alongside the findings (@kevalshahtrilogy, mercy)

#1263 test(feedback): stabilize intake rate-limit window coverage (@benji-bizzell, Aerie)

#1789 fix(education): prevent billing estate refresh throttling (@benji-bizzell, Surtr)

#1784 fix(qtd): match school names for guide headcount (SURTR-1161) (@ashwanth1109, Surtr)

#1787 fix(timeback): contain timeback-raw-sync overruns and finalize timed-out runs (SURTR-1162) (@caina-barbosa, Surtr)

#1769 [SURTR-1135] Fix QTD unit economics revenue proration (@ashwanth1109, Surtr)

#1249 feat(chat): add REBL3 pipeline data-health endpoint (@kevalshahtrilogy, Aerie)

#32 AI-725: Persist workflow operations and add reusable smoke testing (@ashwanth1109, Shipyard)

#1261 fix(enrollment): drop SIS tenant test records from the enrollment spine (@vvp-trilogy, Aerie)

#120 fix(heimdall): the described-edit check was inert — the harness made the workspace look dirty (@kevalshahtrilogy, mercy)

#119 feat(heimdall): repair red before publishing, and never open a second PR for a ticket (@kevalshahtrilogy, mercy)

#118 feat(heimdall): stop losing finished work — described fixes, merge closure, evidence reach, and self-decided calls (@kevalshahtrilogy, mercy)

#1257 fix(finance): support NetSuite-only CAPEX entities (@marcusdAIy, Aerie)

#1782 fix(repo): Surface ECS StopCode/StoppedReason before truncation (@heimdall-keval-factory[bot], Surtr)

#1781 fix(capex): apply warehouse DDL in one transaction (@marcusdAIy, Surtr)

#1780 fix(capex): preserve missing DDR budgets as null (@marcusdAIy, Surtr)

#1256 fix(enrollment): SIS mid-year Transfer Out excludes FinalSite opening-day stamps (#1254) (@vvp-trilogy, Aerie)

#1778 fix(capex): fail closed in the live verifier (@marcusdAIy, Surtr)

#1777 chore(heimdall): sweep the Linear queue every 5 minutes (@kevalshahtrilogy, Surtr)

#1252 1239-2aerie-publishing-modalities (@mwrshah, Aerie)

#1776 chore(heimdall): 60-minute soak and a 3-hour release cooldown (@kevalshahtrilogy, Surtr)

#117 feat(heimdall): a release cooldown, so an hourly releaser doesn't ship hourly forever (@kevalshahtrilogy, mercy)

#1775 chore(heimdall): format the agent's output before it commits (@kevalshahtrilogy, Surtr)

#114 feat(heimdall): format the agent's output before committing it (@kevalshahtrilogy, mercy)

#1121 fix(sales-educrm-mart-sync): raise Redshift statement poll ceiling to f… (@heimdall-keval-factory[bot], Surtr)

#1595 fix(openai-usage-pipeline): give token bucket headroom below OpenAI rat… (@heimdall-keval-factory[bot], Surtr)

#1559 fix(mart-school-performance-table-3-refresh): retry mart refresh CALL o… (@heimdall-keval-factory[bot], Surtr)

#115 fix(heimdall): stop a finished answer reading as a question (@kevalshahtrilogy, mercy)

#1774 fix(repo): Tolerate small Model breakdown excess like Credit Source (@heimdall-keval-factory[bot], Surtr)

#1748 fix(perplexity-usage-pipeline): enforce the skip ceiling per org-day too (@kevalshahtrilogy, Surtr)

#113 feat(mercy): trust heimdall-driven PRs, and stop summoning the agent after approval (@kevalshahtrilogy, mercy)

#112 feat(heimdall): label every pipeline-filed ticket `surtr-pipeline` (@kevalshahtrilogy, mercy)

#1773 fix(repo): Add voyage to TFY known-skipped providers (@heimdall-keval-factory[bot], Surtr)

#1770 fix(education): retain billing allocations without contact links (@benji-bizzell, Surtr)

#1768 fix(education): leave Finalsite reader access under DBA control (@benji-bizzell, Surtr)

#1765 fix(education): restore SIS dataset-triggered refreshes (@benji-bizzell, Surtr)

#1763 feat(education): capture Finalsite findings without blocking ingestion (@benji-bizzell, Surtr)

#1764 fix(repo): forward trigger context for dataset publication runs (@heimdall-keval-factory[bot], Surtr)

#1735 fix(ramp): retry transient API timeouts (@ashwanth1109, Surtr)

#1244 AERIE-1855: Add unpublished verified materialization and installer commands (@caina-barbosa, Aerie)

#1762 fix(education): preserve Finalsite billing balance precision (@benji-bizzell, Surtr)

#1761 fix(education): preserve lossless Limitless transcripts (@benji-bizzell, Surtr)

#1245 feat(platform-errors): make automatic triage runs inspectable (@benji-bizzell, Aerie)

#111 feat(review): make blocking findings name the defect class to fix (@kevalshahtrilogy, mercy)

#1758 feat(education): isolate SIS publications and resume detail collection (@benji-bizzell, Surtr)

#1757 077-finops-incremental-regression (@mwrshah, Surtr)

#174 196-port-live-skill-hydration (@mwrshah, Sindri)

#1756 fix(education): preserve FinalSite contacts outside workflow listings (@benji-bizzell, Surtr)

#1755 fix(education): accept new GuidePlatform source fields (@benji-bizzell, Surtr)

#1687 feat(education): ingest Finalsite billing snapshots (@benji-bizzell, Surtr)

#3737 fix(overspend-alerts): exclude dummy bank charge vendor (@sanketghia, Klair)

#3736 chore(repo): reverse instruction symlink direction (@sanketghia, Klair)

#1753 fix(netsuite-auto-renewal): publish daily snapshots (@ashwanth1109, Surtr)

#1754 fix(repo): Include exception message in run_result failure payload (@heimdall-keval-factory[bot], Surtr)

#1727 feat(heimdall): the factory board's data model (SURTR-1040) [2/2] (@kevalshahtrilogy, Surtr)

Mac's Picks — Key PRs This Week  (click to expand)
#72 — AI-791: Group completed tasks by date @ashwanth1109  no labels

## Demo

![AI-791 smoke test: completed tasks grouped by date](https://github.com/AI-Builder-Team/Shipyard/blob/b40f4f4/docs/smoke-evidence/ai-791-completed-tasks-by-date.png?raw=true)

## Summary

- Persist UTC completion timestamps on tasks and backfill legacy completed tasks to yesterday during schema migration.

- Update the canonical Smoke Test completion transition, including re-completion timestamps, and expose completedAt to the renderer.

- Group completed tasks by UTC date with deterministic ordering and independently collapsible, keyboard-accessible date disclosures.

- Preserve task selection and existing Pull Request and Linear link actions.

## Tests

- node --test scripts/test-workflow.mjs scripts/test-task-workspace.mjs

- pnpm exec tsc --noEmit

- pnpm exec vite build

- pnpm theme:check

- cargo fmt --manifest-path src-tauri/Cargo.toml --all -- --check

- cargo test --manifest-path src-tauri/Cargo.toml --lib

## Linear

https://linear.app/builder-team/issue/AI-791/group-completed-tasks-by-date

#1321 — feat(retention): add OneRoster identity and column provenance @ashwanth1109  approved

## Demo

<img width="2624" height="1636" alt="Retention raw data and provenance" src="https://github.com/user-attachments/assets/007025ec-b07a-4cba-8cff-d8ed4feef8e2" />

<img width="2624" height="1636" alt="Retention column lineage card" src="https://github.com/user-attachments/assets/4a868fdf-8231-42c8-bd76-02bffd35c45c" />

## Summary

- Add a separate nullable OneRoster ID beside SIS ID in Retention Raw data; unresolved identities remain blank and SIS identity is unchanged.

- Extend the existing bound, literal learner search to OneRoster ID without adding a request-time Timeback join.

- Add accessible, visually rich header provenance cards for all 25 selected columns, including exact Redshift objects/fields, producer transformations, caveats, and publication timestamps.

- Surface workbook-organized quality categories while explicitly qualifying current SIS/SURTR source mappings and fallback gaps.

- Preserve capability checks, exact publication validation, keyset pagination, bounded lookahead, report population, and retention calculations.

## Business Value

Users can reconcile dashboard learners with Timeback and the reference retention workbook without confusing SIS IDs with OneRoster sourcedIds. Column-level provenance makes status and date calculations auditable in place, while the quality panel makes known source-coverage gaps visible instead of implying workbook parity.

## Implementation

- The shared learner contract now explicitly projects 25 fields, including oneroster_id; the seven additional producer diagnostic/Timeback-lineage fields remain outside this view.

- Aerie consumes the producer-published ID and does not re-match learners. A schema capability check keeps Raw data usable during rollout by returning a null OneRoster field and omitting OneRoster search until the producer column exists. The OneRoster card documents explicit-identifier precedence, globally unique email fallback, conflict vetoes, and null outcomes.

- Header cards support hover, keyboard focus, click-to-pin, Escape/outside dismissal, viewport collision handling, and scrolling. Metadata is exhaustive over the selected column registry and makes clear that it is documented producer logic, not a per-row execution trace.

- The quality query adds prospective exclusions, eligible-mapped learners without a start date, starts after the report period, cancellations in the analyzed base, in-period effective withdrawals, and short-tenure withdrawals.

## Validation

- 149 focused retention tests passed across nine exact test files, including both producer-column capability states.

- Chat, Convex, and shared-contract TypeScript checks passed.

- Changed-file Biome checks and git diff --check passed.

- Read-only warehouse verification on 2026-09-14 confirmed 5,243 learners, 5,180 distinct populated OneRoster IDs, 11 ambiguous matches, 2 conflicts, and 50 unmatched records for the 2025-06-01–2026-05-31 publication. No learner records were returned by that aggregate check.

## Limitations

- The current versionless report remains a live SIS reconstruction; it must not be represented as proven workbook-parity output until source-level reconciliation is complete.

- Existing report lineage timestamps describe the SIS publication and mart refresh. The OneRoster card does not relabel them as Timeback freshness.

- Validation is automated and warehouse-aggregate based; no authenticated browser verification was performed after the final stack rebase.

## Implementation Effort

An average engineer would need approximately 3–5 engineer days to trace and verify the producer lineage, implement the contract/backend/UI changes, build accessible provenance interactions, add focused tests, reconcile the stack, and document the limitations without AI assistance.

## Linear

[AERIE-2116 — Add OneRoster identity to retention raw data](https://linear.app/builder-team/issue/AERIE-2116/add-oneroster-identity-to-retention-raw-data)

## Stack

Parent PR #1264 is merged. This layer is rebased directly onto main; native stack #1322 remains linked.

#1829 — fix(repo): Retry Redshift deadlock aborts in ontology refresh CALL @heimdall-keval-factory[bot]  approvedAutomated PR

A one-off Redshift lock conflict between this refresh and another concurrent query aborted the run before any data changed. Adding a bounded retry for that specific transient error — already used by a sibling pipeline for the same failure mode — lets the refresh recover on its own instead of alerting a human each time it happens.

Ticket: SURTR-1277

> Ready for review. Nothing ran the change, so it is unproven. A person still merges.

## For The Agent

_Everything below is detail for review. The summary above is the change._

Presented as ready — verification none, scope tier draft, fix_class code_fix, HEIMDALL_READY_PRS=true.

### What's broken

Run b188dbaf-446d-4286-8090-8345aa422025 of core-education-ontology-refresh failed when Redshift statement 0bb82cef-75cb-49bf-9276-8fd474966bf7 hit "[ERROR] RuntimeError: Redshift statement 0bb82cef-75cb-49bf-9276-8fd474966bf7 failed: ERROR: deadlock detected" inside the CALL to core_education.sp_refresh_aerie_ontology (pipelines/runners/core-education-ontology-refresh/ddl/010_sp_refresh_aerie_ontology.sql). Resolving the two locked relation oids from the log (15058250, 15058264) confirms they are core_education.dim_program and core_education.bridge_school_link — the two tables named last in that procedure's opening LOCK TABLE statement — so this is Redshift's deadlock detector aborting one side of a circular wait between our procedure and a concurrent reader that touched the same two tables in the opposite order; stl_query returned no rows for the two blocked PIDs, so the counterparty query itself is not recoverable from the warehouse. No rows were published for this run (the procedure's transaction rolled back atomically), so the prior complete publication of dim_school/dim_program/dim_site/bridge_school_link/xref_school_source remained the last-good snapshot with zero data loss, just a missed refresh cycle.

Root cause. core-education-ontology-refresh acquires AccessExclusiveLock on core_education.bridge_school_link, dim_program, dim_school, dim_site, and xref_school_source (in that fixed order) at the top of sp_refresh_aerie_ontology so its DELETE+INSERT publish sees one consistent snapshot. A second, unrelated concurrent query apparently read dim_program first and then requested a lock on bridge_school_link, i.e. the reverse acquisition order, producing the classic circular wait Redshift reports as 'deadlock detected'. This is an inherent, transient concurrency hazard of Redshift's session-level locking, not a defect in the procedure's validation or publish logic — the same failure class was already diagnosed and fixed for pipelines/runners/mart-school-performance-unit-economics-per-student-refresh, whose src/redshift_client.py now retries a statement that Redshift aborts specifically as a deadlock victim (matching 'deadlock detected' in the FAILED Error text) with bounded exponential backoff, while still failing immediately on any other error.

### What this PR changes

Port the deadlock-retry pattern from pipelines/runners/mart-school-performance-unit-economics-per-student-refresh/src/redshift_client.py (MAX_DEADLOCK_RETRIES, DEADLOCK_BACKOFF_BASE_SECONDS, an _is_deadlock_error() check on the FAILED Error text, and a _DeadlockAborted internal signal) into pipelines/runners/core-education-ontology-refresh/src/redshift_client.py's _submit_and_poll/execute path, so the CALL to core_education.sp_refresh_aerie_ontology is retried a bounded number of times within the existing run_deadline budget when Redshift aborts it specifically as a deadlock victim, and fails immediately as before for every other error. This is safe here for the same reason it was safe there: the header comment on ddl/010_sp_refresh_aerie_ontology.sql (lines 3-9) states the DELETE+INSERT publish is the procedure's single default atomic transaction keyed off the pinned Rhodes run id, so a deadlock abort rolls back completely and a retry of the identical CALL is idempotent and cannot duplicate or partially publish rows. Add a unit test mirroring test_client_failure_includes_statement_id in tests/test_redshift_client.py that asserts a FAILED status containing 'deadlock detected' is retried and a non-deadlock FAILED status is not.

Why this fixes it. The failure is a transient Redshift lock-ordering race between this pipeline's fixed-order LOCK TABLE and an unidentifiable concurrent reader (stl_query has no rows for the blocked PIDs), not a bug in the procedure's data validation, so no code change can eliminate the race itself. The correct, complete, and already-precedented fix confined to this pipeline's own directory is resilience: retry the atomic CALL when Redshift reports the statement as a deadlock victim, exactly as pipelines/runners/mart-school-performance-unit-economics-per-student-refresh/src/redshift_client.py does for the identical symptom, which is safe here because the procedure's publish is a single all-or-nothing transaction and re-running it with the same Rhodes run id is idempotent.

#### Files changed

 .../src/redshift_client.py                         | 45 ++++++++++++++--

.../tests/test_redshift_client.py | 62 ++++++++++++++++++++++

2 files changed, 104 insertions(+), 3 deletions(-)

### Verification

### pytest — no test suite

### verify: ruff check — exit 0

[notice] A new release of pip is available: 25.3 -> 26.2.1

[notice] To update, run: pip install --upgrade pip

All checks passed!

### verify: ruff format --check — exit 0

1834 files already formatted

### verify: pytest (pipeline lambdas) — exit 2

[notice] A new release of pip is available: 25.3 -> 26.2.1

[notice] To update, run: pip install --upgrade pip

Downloading cpython-3.11.16-linux-x86_64-gnu (download) (29.4MiB)

Downloaded cpython-3.11.16-linux-x86_64-gnu (download)

Installed Python 3.11.16 in 516ms

+ cpython-3.11.16-linux-x86_64-gnu (python3.11)

error: Failed to spawn: pytest

Caused by: No such file or directory (os error 2)

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | repo |

| Failing run | issue |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | linear-SURTR-1277 |

| Verify | none |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev label to take the PR over and stop it entirely.

#1828 — fix(ramp): make cost research recovery resilient @ashwanth1109  approved

## Summary

- Increase the bounded Anthropic cost-research deadline from 300 to 600 seconds so valid long-running web searches can complete.

- Allow standalone mode=cost recovery to select a validated historical DD-MM-YYYY run folder, while rejecting the override for fetch and transform modes.

- Preserve model validation feedback across transient transport timeouts so later retries continue correcting the original contract violation.

- Recover the missing week 37 cost-analysis artifact and complete the failed schedule_2 Superbuilders report delivery.

## Business Value

Restores the weekly Superbuilders Ramp report, prevents valid cost research from being cut off prematurely, and makes late checkpoint recovery possible without refetching or overwriting historical Ramp snapshots.

## Implementation Effort

Estimated 4–6 hours for an average engineer to diagnose both pipeline executions, implement and test the recovery controls, perform the isolated production deployment, monitor the live backfill, and verify email delivery.

## Linear

- [SURTR-1272](https://linear.app/builder-team/issue/SURTR-1272/ramp-superbuilders-report-failing-stopcodeessentialcontainerexited)

## Validation

- uv run pytest — 214 passed.

- uvx ruff check on all modified Python files — passed.

- Isolated production deployment: Pipeline-ramp-spend-pipeline-prod only (deploying... [1/1]), CloudFormation UPDATE_COMPLETE.

- Live week-37 cost recovery resumed 14/17 checkpointed opportunities; two valid calls completed after the former five-minute cutoff; producer Step Functions execution SUCCEEDED.

- Final task definition revision 11 cached validation exited 0 and Step Functions SUCCEEDED.

- Delivery preflight confirmed matching week 37/2026 financial, cost, and chart artifacts; September 12 data was two days old; no prior schedule_2 receipt existed.

- Recovery delivery exited 0, Step Functions SUCCEEDED, SES accepted the message for the configured schedule_2 group, and the group-scoped receipt is sent.

#1826 — fix(school-performance): use full monthly QuickBooks budgets @ashwanth1109  approved

## Business Value

School Performance Report Table 1 now compares QuickBooks actuals through the reporting day with the full published budgets for each month through that day’s month. On September 13, Alpha Scottsdale’s tuition budget is July $0 + August $280,000 + September $280,000 = $560,000; it was $401,333.33 after daily proration. The reporting rule persists in the mart for future reports.

## Changes

The stored procedure includes budget details through the end of the cutoff month and sums their full signed amounts. Actuals retain their existing reporting-day cutoff. The existing proration flag is false, and column metadata documents the different actual and budget periods. Coverage, source eligibility, mappings, lineage and atomic publication remain intact.

Calendar regression tests execute the procedure’s budget expression and date predicates for month start, mid-month, month end, quarter/year changes and leap February, including expense signs and the unchanged actuals cutoff.

## Validation

- 65 tests passed; Ruff checks and formatting passed on the two modified Python files; git diff --check passed.

- Candidate ea931535 was applied to the existing production Redshift procedure before merge. The catalog body matches the candidate exactly. Two column comments were updated. No Lambda or CloudFormation deployment was needed because the procedure signature and runner are unchanged.

- Production execution full-month-budget-20260914T0307 succeeded on September 14, 2026, 03:04:12–03:04:53 UTC; run ID 4e0a4659-3580-4084-8403-e1328c17a3bf.

- All 783 rows across 33 schools passed: 500 budget rows reconcile to 1,500 full monthly source details; actuals reconcile to 17,570 canonical postings. Zero budget, actual, variance or duplicate-key mismatches; zero prorated rows. The publication cutoff is September 14 and both accepted source publications are September 13, 06:35:23 UTC.

The deployed change is already live; merging retains it in source control for subsequent deployments. The report was subsequently advanced to September 14 across all six sections using successfully refreshed source marts. Table 1 uses the new full-month QuickBooks policy; the other model allocation policies are unchanged.

## Linear

https://linear.app/builder-team/issue/SURTR-1262/use-full-monthly-quickbooks-budgets-in-the-school-performance-qtd-mart

## Implementation Effort

Approximately 3–4 hours for an engineer to trace the writer, implement the calculation and metadata updates, add calendar regression tests, deploy the procedure, run the pipeline and reconcile live data.

#71 — AI-792: Improve task context side panel @ashwanth1109  no labels

## Demo

![Task context smoke test](https://github.com/AI-Builder-Team/Shipyard/blob/f8d7887bf3ba213ac77335fb050db1c12091daf6/.smoke-evidence/ai-792-task-context.png?raw=true)

## Summary

- Redesigned the task context rail with a summary, provider cues, counts, and independently collapsible groups.

- Added readable Linear and pull-request cards with opener progress/error feedback and full-URL accessibility metadata.

- Clarified the Linear attachment form with visible empty, focus, loading, success, and error states.

- Improved repository identity/path presentation and preserved existing task data and commands.

## Validation

- pnpm build

- pnpm theme:check

- pnpm test:smoke

- pnpm smoke build

- pnpm smoke start --scenario happy

- pnpm smoke verify --expect research-ready

- pnpm smoke stop

## Linear

https://linear.app/builder-team/issue/AI-792/task-side-panel-improvements

#1264 — feat(retention): add Financials retention dashboard and raw data @ashwanth1109  changes requestedmercy-allow-critical

## Demo

<img width="2096" height="1636" alt="image" src="https://github.com/user-attachments/assets/4940920c-a045-4431-bd83-d5033fd3fd0a" />

<img width="2624" height="1636" alt="image" src="https://github.com/user-attachments/assets/20bd76ff-4985-4baa-beb1-f4e4021609b9" />

<img width="2624" height="1636" alt="image" src="https://github.com/user-attachments/assets/0fe73f77-8642-431a-b4d4-677b81d68bc0" />

<img width="2624" height="1636" alt="image" src="https://github.com/user-attachments/assets/76ee211a-896b-4481-938a-150a89547cc8" />

## Summary

- Add the Financials → Retention dashboard backed by the SURTR retention marts, including publication lineage, v1/v2 validation, campus filters, monthly movement, cohort retention, quality metrics, and CSV export.

- Add an on-demand Raw data tab with all 25 learner-mart columns, paginated 25/50/100-row views, literal search by learner name/email/SIS ID, included/excluded filtering, and snapshot consistency checks.

- Add the Financials navigation and capability mapping, shared retention contracts, Convex backend readers, regression tests, and retention documentation.

## Business Value

Gives Education and Finance users a reliable retention view in Aerie and a controlled way to inspect the underlying learner rows behind the metrics. The dashboard makes the August 25 SIS-based publication and its data-quality boundaries visible, while the paginated raw view supports operational investigation without exposing an unrestricted warehouse dump or mixing different mart refreshes.

## Implementation Effort

Estimated 4–6 engineer-days for an engineer working without an AI agent, including the dashboard, Convex/Redshift reader, publication validation, pagination and filtering, authorization coverage, test fixtures, and documentation.

## Validation

- 52 focused tests passed across the retention backend, UI, and hooks.

- Chat and contracts typechecks passed.

- Biome and Convex path checks passed.

- Read-only live validation through the Klair Data API confirmed the v2 publication: 5,243 learner rows, all 25 columns, a 50-row first page, a 43-row final page, included-campus filtering, and zero-match search behavior.

- The authenticated browser visual check remains for reviewer confirmation; no local services were started or restarted by this change.

## Linear

- [SURTR-1175](https://linear.app/builder-team/issue/SURTR-1175) — retention mart producer contract

- [SURTR-1181](https://linear.app/builder-team/issue/SURTR-1181) — v2 retention reconciliation

- [AERIE-1179](https://linear.app/builder-team/issue/AERIE-1179) — related Aerie schema compatibility fix

#1825 — fix(repo): Restore aliases in foundation verification query @heimdall-keval-factory[bot]  approvedAutomated PR

The HubSpot Core foundation pipeline's post-refresh data-quality check broke because of an unrelated cleanup that dropped the column names it relied on — the actual data refresh completed fine. Restoring the three missing column names in that check query fixes it with no impact on published data.

Ticket: SURTR-1254

> Ready for review. Nothing ran the change, so it is unproven. A person still merges.

## For The Agent

_Everything below is detail for review. The summary above is the change._

Presented as ready — verification none, scope tier draft, fix_class code_fix, HEIMDALL_READY_PRS=true.

### What's broken

Run 6e19a84d-b35b-4d1c-b3e1-ac8b15dfb28d (SURTR-1254) failed after the hubspot_core_foundation procedure itself completed successfully (duration_ms=175438.252) — the crash is in the post-refresh verification step, at src/admissions_procedures.py:247 (verify_publication), with RuntimeError: Redshift statement 178e0831-fed5-407c-bfaa-46703532d4eb failed: ERROR: column "model_name" does not exist in model_rows. The verification query's model_rows CTE (POST_REFRESH_QUERIES['hubspot_core_foundation'], admissions_procedures.py:48-82) selects model_name, row_key, and is_invalid from that CTE in its outer SELECT, but no UNION ALL branch inside the CTE aliases its columns as such — Redshift derives a UNION ALL result set's column names solely from the first branch's SELECT list.

Root cause. Commit 84e9cf1b (#1813, 'build a warehouse-owned Person directory') deleted the dim_person UNION ALL branch, which was the only branch carrying AS model_name, AS row_key, AS is_invalid aliases. The remaining branches (bridge_deal_person, xref_deal_source, fct_deal_person_resolution_exception, dim_academic_session) never had their own aliases because they relied on inheriting names from that first branch. With bridge_deal_person now first and unaliased, the CTE's output columns have no names matching model_name/row_key/is_invalid, so the outer SELECT's COUNT(DISTINCT model_name || '|' || row_key) and SUM(is_invalid) fail to resolve. This is a pure regression in the verification SQL text, not in the underlying data or the refresh procedure it checks.

### What this PR changes

Add AS model_name, AS row_key, and AS is_invalid to the first SELECT branch of the model_rows CTE in POST_REFRESH_QUERIES['hubspot_core_foundation'] (currently SELECT 'bridge_deal_person', deal_person_id, CASE ... END FROM core_education.bridge_deal_person), matching how the deleted dim_person branch aliased them before #1813. No other branch needs a change since UNION ALL column names come only from the first branch. The sp_refresh_hubspot_core_foundation procedure and the published tables (bridge_deal_person, xref_deal_source, fct_deal_person_resolution_exception, dim_academic_session) are unaffected — this is a re-run-clean fix, not a backfill.

Why this fixes it. The defect is three missing column aliases in one CTE's first UNION ALL branch, entirely inside admissions_procedures.py (Tier A, the pipeline's own src/ dir). The fix has no effect on any other query, the stored procedure, or pipeline.json, and is fully verifiable by re-running the existing verify_publication contract tests, so it is safe to attempt automatically.

#### Files changed

 .../runners/hubspot-core-tables/src/admissions_procedures.py   |  4 ++--

.../hubspot-core-tables/tests/test_admissions_procedures.py | 10 ++++++++++

2 files changed, 12 insertions(+), 2 deletions(-)

### Verification

### pytest — no test suite

### verify: ruff check — exit 0

[notice] A new release of pip is available: 25.3 -> 26.2.1

[notice] To update, run: pip install --upgrade pip

All checks passed!

### verify: ruff format --check — exit 0

1833 files already formatted

### verify: pytest (pipeline lambdas) — exit 2

[notice] A new release of pip is available: 25.3 -> 26.2.1

[notice] To update, run: pip install --upgrade pip

Downloading cpython-3.11.16-linux-x86_64-gnu (download) (29.4MiB)

Downloaded cpython-3.11.16-linux-x86_64-gnu (download)

Installed Python 3.11.16 in 325ms

+ cpython-3.11.16-linux-x86_64-gnu (python3.11)

error: Failed to spawn: pytest

Caused by: No such file or directory (os error 2)

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | repo |

| Failing run | issue |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | linear-SURTR-1254 |

| Verify | none |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev label to take the PR over and stop it entirely.

#1823 — fix(aerie): remove retention rules version contract @ashwanth1109  approved

The retention refresh was failing because the warehouse published a different rules-version marker than the Lambda expected. This removes rules_version from the three mart tables and keeps verification based on source/run lineage and business contracts.

## Summary

- Remove the column and v1/v2 checks from table DDL, the writer procedure, runtime verification, reconciliation, fingerprints, and catalog checks.

- Commit legacy column drops and procedure replacement in one transaction under the refresh mutex; refuse to split the cutover across batches.

- Require verified caller quiescence for legacy migrations: the correct warehouse target, Lambda admission blocked, scheduled rules disabled, and platform executions drained. Document deploying the compatible Lambda while paused and restoring admission after catalog verification.

- Add regression coverage and a live fixture helper proving schema/procedure rollback, successful cutover, and retry.

## Business Value

Restores reliable Aerie retention refreshes without an internal implementation-version column. The migration prevents intermediate table/procedure mismatches while preserving source lineage, atomic publication, and reconciled retention results.

## Implementation Effort

Estimated 1–2 engineer-days for a manual implementation, including migration coordination, runtime and reconciliation changes, regression tests, and production validation.

## Linear

- [SURTR-1175](https://linear.app/builder-team/issue/SURTR-1175/build-and-live-validate-aerie-retention-stored-procedure-and-scheduled)

## Validation

- Runner suite: 90 passed; Ruff formatting and lint passed on modified Python files.

- Live fixture: failures after each of three column drops and after procedure replacement preserved the legacy columns and callable old procedure. Successful cutover and retry passed; fixtures were removed.

- The admission guard rejected the active production caller. Production was already migrated, so the reapplication exercised the no-legacy path; the legacy cutover was verified on isolated fixtures.

- Candidate deployed exclusively to Pipeline-mart-aerie-retention-refresh-prod; CloudFormation reached UPDATE_COMPLETE.

- Revised DDL application and the complete 61-column catalog verification passed.

- Production run b25867c1-3a2e-4ba8-afd2-5babbf106275 succeeded. Independent reconciliation passed all four contracts and matched the prior 5,423-row business fingerprint a410b332b0737ed9dddb687f17f53bbe689415876a2a33a09fe050c3554d2e8a.

- Full statement IDs and execution evidence are recorded in the runner README.

#70 — AI-790: Prevent command titles from wrapping vertically in expanded command groups @ashwanth1109  no labels

## Demo

![AI-790 smoke test evidence](https://github.com/AI-Builder-Team/Shipyard/blob/f15e7c2c5a6f83a4e4b7a7a940171b66c41ce8c6/docs/smoke-evidence/AI-790-command-title-layout.png?raw=true)

## Summary

- Keep transcript activity titles on one line and prevent them from collapsing to a character-wide flex item.

- Give long command summaries the remaining row width for ellipsis while retaining duration, status, and disclosure controls in narrow panes.

- Add CSS-aware regression coverage for expanded grouped commands with long invocations and exact expandable output.

## Validation

- pnpm test:chat

- pnpm build

- pnpm theme:check

- git diff --check

## Linear

https://linear.app/builder-team/issue/AI-790/prevent-command-titles-from-wrapping-vertically-in-expanded-command

#59 — AI-779: Add steering for queued conversation messages @ashwanth1109  no labels

## Demo

![Smoke test evidence](https://github.com/AI-Builder-Team/Shipyard/blob/902f7f62c01c8e8c5b8167b6377501f68f95cedd/.smoke-evidence/AI-779-smoke-test.png?raw=true)

## Summary

- Add native turn/steer support for promoting queued messages into the active Codex turn.

- Preserve queue order and restore selected messages after protocol failures or races.

- Add compact accessible queue rows with Steer and icon-only Remove actions.

## Testing

- pnpm test:messages

- pnpm test:chat

- pnpm build

- cargo check --manifest-path src-tauri/Cargo.toml

- cargo test --manifest-path src-tauri/Cargo.toml --lib codex_app_server::tests

- pnpm theme:check

## Linear

https://linear.app/builder-team/issue/AI-779/add-a-steer-action-for-queued-conversation-messages

#69 — AI-789: Prepare Shipyard 0.4.1 recovery release @ashwanth1109  no labels

## Summary

- Bump the authoritative app version from 0.4.0 to 0.4.1 for the recovery release.

- Add public release notes for steering active conversations with queued follow-up messages.

## Business Value

Users can steer an active conversation without waiting for the current turn to finish, and this recovery release makes that feature available through the public updater after the prior publication attempt failed before creating a release.

## Implementation Effort

Estimated 30 minutes for an average engineer to prepare the recovery version, public notes, validation, and release PR.

## Test Plan

- pnpm test:release (10 JavaScript tests and 8 Python tests)

- git diff --check

- After merge, run the verified-main local release plan, build, audit, and publish workflow.

## Linear

https://linear.app/builder-team/issue/AI-789/prepare-shipyard-041-recovery-release

Fixes AI-789

#68 — AI-788: Speed up local Shipyard publish-update workflow @ashwanth1109  no labels

## Summary

- Reuse an unchanged staged Codex runtime after exact file-set and byte validation.

- Read only the native executable header during runtime discovery.

- Report per-command and total release timings.

- Add regression coverage for runtime reuse, mutation, extra files, and symlinks.

## Business Value

Shortens repeat local release preparation, reduces unnecessary filesystem and memory work for the bundled Codex runtime, and makes release bottlenecks visible so future improvements can target measured costs without weakening publication safety.

## Implementation Effort

Approximately 3 hours for an average engineer to implement, test, and review manually.

## Linear

https://linear.app/builder-team/issue/AI-788/speed-up-the-shipyard-local-publish-update-workflow

## Test Plan

- node --check for all changed JavaScript files.

- python3 -m py_compile scripts/release/local.py.

- pnpm test:release (10 Node tests and 8 Python tests).

- Repeated pnpm stage:codex confirms unchanged runtime reuse.

- Two no-bundle Tauri builds complete successfully (cold and warm runs).

- git diff --check passes.

#67 — AI-787: Prepare Shipyard 0.4.0 release @ashwanth1109  no labels

## Summary

- Bump the authoritative app version from 0.3.1 to 0.4.0 for the next feature release.

- Add public release notes for steering active conversations with queued follow-up messages.

## Business Value

Users can steer an active conversation without waiting for the current turn to finish, and the published update will carry the correct feature-release version and notes.

## Implementation Effort

Estimated 30 minutes for an average engineer to prepare the version bump, public notes, validation, and release PR.

## Test Plan

- pnpm test:release (9 JavaScript tests and 8 Python tests)

- git diff --check

- After merge, run the verified-main local release plan, build, audit, and publish workflow.

## Linear

https://linear.app/builder-team/issue/AI-787/prepare-shipyard-040-release

Fixes AI-787

#66 — AI-786: Make update checks recover quickly and release 0.3.1 @ashwanth1109  no labels

## Business Value

Update checks recover promptly when a release-download server is unreachable. Users receive a result or clear timeout within approximately five seconds instead of waiting 30–60 seconds.

## Changes

- Set a five-second manifest request limit and two-second connection budget, allowing the HTTP connector to advance through resolved addresses.

- Show a retryable timeout message and retain the longer package-transfer allowance and signature verification.

- Pin the updater plugin to its tested 2.11 minor series; prepare version 0.3.1 and public release notes.

## Validation

- 9 JavaScript release/update tests and 8 Python release tests passed.

- 3 native updater safety tests passed.

- Equivalent native HTTP diagnostic: live public feed completed in 0.73 seconds; an unresponsive test server timed out in 5.01 seconds.

- Production bundle audit, public feed verification, and installed-app check will run as part of the authorized release after merge.

## Implementation Effort

Approximately 2–3 hours for an average engineer to diagnose the network behavior, implement the fix, test it, and prepare the release without AI assistance.

## Linear

https://linear.app/builder-team/issue/AI-786/make-update-checks-recover-quickly-from-stalled-download-servers

#65 — AI-785: Add simple release versions and an Updates detail view @ashwanth1109  no labels

## Business Value

Users can recognize Shipyard releases by simple major.minor.patch versions and manage updates from a dedicated view with readable changelogs. Downloads survive navigation, and release history loads ten versions at a time.

## Changes

- Start at 0.3.0, which upgrades normally from historical 0.2 timestamp versions. Share package.json versioning and public Markdown notes across local publishing and CI.

- Add an Updates detail view with check, download, install/restart, release dates, Installed/Latest labels, and paginated stable-release history.

- Refresh update checks while retaining verified bytes for an unchanged release; keep signature verification and existing activity/instance installation guards.

- Update publishing guidance for explicit versions, changelogs, recovery and verification, preserving the saved-current-main release gate.

## Validation

- TypeScript and theme validation; frontend/release/publisher tests; native pagination and installation-guard tests; smoke adapter tests.

- Apple Silicon native smoke run 4da52b20-fe83-49c6-ae1d-e6ff1977ff53: 10 → 20 → 23 releases, invalid-signature rejection/retry, download preservation during navigation and fresh checks, and signed 0.2.1789280169 → 0.3.0 install/restart. Expected binary hash matched, task persisted, no duplicate workflow effects. Harness stopped; reports retained locally and documented in docs/UPDATES.md.

- Current main changed only the publishing skill; its guard was reconciled and skill validation passed. Intel runtime was not smoke-tested.

## Implementation Effort

Approximately 2–3 engineer-days to implement and validate the release workflow changes, UI, pagination, failure handling, and signed desktop upgrade without AI assistance.

## Linear

https://linear.app/builder-team/issue/AI-785/simple-release-versions-and-an-updates-detail-view-with-changelog

#64 — AI-784: Require local main for Shipyard publish skill @ashwanth1109  no labels

## Summary

The local Shipyard publish skill now requires a clean, up-to-date saved main checkout before planning, building, or publishing. It explicitly handles branch selection, remote synchronization, and ahead/diverged histories so a release cannot accidentally package a feature worktree or stale commit.

## Business Value

Release operators get a predictable production baseline and avoid publishing binaries built from stale or unintended code. The guard preserves uncommitted local work and makes the release source commit auditable against origin/main.

## Implementation Effort

Estimated hand-coding effort: 1–2 hours for a release engineer to inspect the helper, define the checkout safety rules, update the skill documentation, and validate the workflow.

## Validation

- git diff --check passed.

- python3 scripts/release/local.py plan passed from clean /Users/ash/Desktop/work/Shipyard on main at cb225d9.

- python3 scripts/release/local.py build passed, including release tests, app compilation, artifact audit, and DMG verification.

- python3 scripts/release/local.py publish .local-release/0.2.1789283328 passed.

- Anonymous latest.json verification returned 0.2.1789283328; its hash matched the receipt, and the published updater and DMG endpoints returned successfully.

## Linear

[AI-784](https://linear.app/builder-team/issue/AI-784/make-local-shipyard-publish-skill-enforce-an-up-to-date-main-checkout)

#63 — AI-783: Allow per-conversation model overrides with Luna Max default @ashwanth1109  no labels

## Summary

New Shipyard conversations now use GPT-5.6 Luna at Max reasoning by default. The composer model picker can override the model and reasoning effort for one conversation, and that choice survives later turns, reconnects, and app restarts.

Selectable models use direct openai-primary/ TrueFoundry routes. Legacy openai-group/ choices are normalized when dispatched, so switching models preserves Responses history and avoids the gateway's HTTP 409 ResponsesPinningError caused by encrypted reasoning being pinned to the original virtual model.

## Business Value

Users get a consistent, high-quality default while retaining control for conversations that need a different model or reasoning budget. Model changes no longer strand an existing conversation or require starting over, preserving context and reducing failed work.

## Implementation Effort

An average engineer would need approximately 1 working day (6–8 hours) to implement and validate this change by hand.

## Test Plan

- pnpm test:chat — 24 UI/content tests passed.

- pnpm exec tsc --noEmit — passed.

- pnpm build — passed.

- cargo test --manifest-path src-tauri/Cargo.toml --lib — 125 passed, 2 ignored.

- Live gateway regression passed: Luna Max → Sol Max → Luna Max with encrypted history retained.

## Linear

[AI-783 — Allow per-conversation model overrides with Luna Max default](https://linear.app/builder-team/issue/AI-783/allow-per-conversation-model-overrides-with-luna-max-default)

#62 — AI-782: Redesign Codex model and reasoning controls @ashwanth1109  no labels

## Summary

Shipyard can now discover and persist TrueFoundry model selections and reasoning effort across reconnects, forks, and later turns. The composer presents those settings through one compact, readable picker inspired by the reference design, while active-turn status no longer adds a redundant Working label or queue chevron.

The picker keeps exact qualified gateway IDs for requests, formats model names for display, exposes only supported reasoning levels, saves on slider commit, supports reset-to-default, and handles loading, retry, failure, keyboard, focus, and outside-dismissal states.

## Business Value

Users can choose the right Codex model and reasoning depth without memorizing provider-specific IDs, and their choice remains dependable when conversations reconnect or continue. The simpler composer leaves more room for the conversation while keeping queued messages automatic and understandable.

## Implementation Effort

An average engineer would likely need 2–3 days to implement the catalog integration, persistence and reconnect handling, accessible picker/popover interactions, theme updates, and regression coverage by hand.

## Test plan

- pnpm exec tsc --noEmit

- pnpm theme:check

- pnpm test:chat (23 passed)

- pnpm test:recovery (18 passed)

- pnpm test:connection (10 JavaScript tests plus 3 Rust tests passed)

- cargo test --manifest-path src-tauri/Cargo.toml --lib model_settings_tests:: (2 passed)

- cargo test --manifest-path src-tauri/Cargo.toml --lib tfy::tests:: (7 passed, 1 live gateway test ignored)

- cargo fmt --manifest-path src-tauri/Cargo.toml -- --check

- git diff --check

The packaged app was not built or launched; the development app can be smoke-tested from this branch.

## Linear

[AI-782 — Redesign Codex model and reasoning controls](https://linear.app/builder-team/issue/AI-782/redesign-codex-model-and-reasoning-controls)

#3763 — fix(spacex-valuation): reconcile September 10 trade @sanketghia  approved

## Summary

- Add the September 10, 2026 400,000-share SPCX sale and exact net proceeds.

- Allocate the sale FIFO against the September 9 Day-90 distribution using actual whole-share source values.

- Preserve open Day-90 inventory for future realized-sale confirmations and update derived valuation/residual assertions.

## Verification

- Full frontend Vitest suite: 667 files, 6,855 passed, 16 skipped.

- TypeScript check, ESLint, Prettier, and production build passed.

## Testing & Screenshot

- Updated numbers have been reviewed and approved by stakeholders (Dave & Milo)

<img width="1069" height="593" alt="image" src="https://github.com/user-attachments/assets/30985c84-8b5f-4f0c-b217-2938b9091984" />

#60 — AI-780: Add in-app updates and public release pipeline @ashwanth1109  no labels

## Business Value

Shipyard users can check for a newer build, download a verified update, and explicitly install/restart without pulling source or compiling locally. Public binaries are distributed separately from the private source repository.

## Implementation

- Compact header update panel with progress, retry, Later, and explicit restart confirmation for unsaved drafts/attachments.

- Native Tauri signature verification; verified package bytes stay in native memory. Installation refuses active Codex/workflow work and other instances, with activity/database/compatibility locks around replacement.

- Native Apple Silicon/Intel release builds from main, monotonic versions, pinned Codex runtime, file allowlisting, credential-pattern scanning, archive hash comparison, independent read-only DMG inspection, and draft-first publication with upload digest checks.

- Isolated signed A/B qualification through the repository smoke harness, using a test-only key and loopback feed. Harness tracks the relaunched process and retains reproducible evidence.

- Company defaults and personal compiler paths remain as requested. Current packaging is not Apple Developer ID signed/notarized.

## Infrastructure

Public distribution repository: https://github.com/AI-Builder-Team/Shipyard-Releases (README only; no app binary published). Updater signing key is stored outside source control and in the private source repository Actions secrets. Publication remains gated by SHIPYARD_UPDATES_VALIDATED.

## Validation

- Successful real native signed upgrade 0.0.901 → 0.0.902 in isolated run 482cfc26-c9e1-47c6-9743-f6d11cf63ab2: new PID, exact B executable SHA-256, persisted task, no replayed workflow effects. Native UI confirmed busy-turn and other-instance refusal before successful install/restart, then “You’re up to date.” Evidence: .smoke/runs/482cfc26-c9e1-47c6-9743-f6d11cf63ab2/update-report.json.

- Run d933aaa2-5286-4679-8aaf-1e2ae74fc5ef: invalid signatures and truncated downloads rejected; recovered download retry succeeded; server error and no-update retry confirmed in native UI. No new workflow effects. Evidence: update-fault-report.json in that run directory. Both runs stopped through the harness.

- Production-format Apple Silicon 0.2.0 candidate built with release updater key. Audit passed for 8 bundle files and 1,826 frontend assets; read-only DMG inspection matched all 8 files. Local evidence: .smoke/production-check/dmg-report.json. Production application untouched.

- 27 smoke harness tests and 8 release/UI tests pass. Smoke-feature updater Rust tests pass. Earlier typecheck, 115 Rust tests and all five theme checks passed.

## Local publication

Added $shipyard-publish-update and scripts/release/local.py: read-only preflight, audited native build, private hash receipt, and draft-first publication using the local gh credential store without exporting a token. The local Apple Silicon path is intentionally single-platform and refuses to remove another platform from an existing public feed. CI still requires both platforms and remains gated while the token awaits organization approval. Local and CI versions use UTC seconds; one CI job shares that version across both architecture builds.

Local preflight confirmed repository write permission and matching updater public key. All 14 release/UI tests and skill validation pass, including draft retention on digest failure and no network access for an invalid manifest. No local release has been published by this change.

## Rollout blockers and limitations

Actions run https://github.com/AI-Builder-Team/Shipyard/actions/runs/34676406717 received SHIPYARD_DISTRIBUTION_TOKEN but draft creation failed with HTTP 403: Resource not accessible by personal access token. Correct resource owner/repository selection, Contents read/write, or organization approval, then rerun validation. Do not enable public publication until it succeeds.

OS-level replacement failure/recovery has not been fault-injected; Intel runtime qualification and Apple signing/notarization remain outstanding. First adoption requires one manual updater-enabled installation. Unsent drafts are warned about, not automatically persisted.

## Implementation Effort

Estimated 4–6 engineering days without AI assistance for updater/lifecycle integration, release automation, artifact inspection, isolated signed upgrade qualification, and documentation. Apple notarization setup is additional.

## Linear

https://linear.app/builder-team/issue/AI-780/add-in-app-updates-and-public-binary-distribution

#61 — AI-781: Fix Codex reconnect state race @ashwanth1109  no labels

## Summary

- Fix the reopen race where a thread metadata notification invalidated a successful Codex attachment and left the composer on Connecting.

- Invalidate runtime snapshots only for recognized runtime transitions and valid pending-request changes.

- Preserve live turn, approval, request-resolution, error, and disconnection precedence over stale reads.

- Add bounded frontend/native connection diagnostics and a read-only pnpm diagnose:connection report with explicit attachment/UI state mismatch detection.

## Business Value

Users can reopen an already-connected Codex conversation without waiting through a misleading reconnect state. When a real connection or attachment is slow, support and engineering can now distinguish transport, database, peer routing, and frontend state issues from captured diagnostics.

## Implementation Effort

Estimated 1–2 engineering days for an average engineer to trace the cross-process attachment flow, add bounded instrumentation, reproduce the race, implement the event classification guard, and build regression coverage.

## Linear

https://linear.app/builder-team/issue/AI-781/fix-codex-conversation-reconnect-state-race

## Test plan

- pnpm test:connection

- pnpm test:recovery

- pnpm exec tsc --noEmit

- pnpm exec vite build --logLevel warn

- cargo test --manifest-path src-tauri/Cargo.toml --lib (115 passed)

- Verified the dev-app reopen path and captured diagnostics; production app was not modified.

#58 — AI-778: Make Codex transcript activity easier to scan @ashwanth1109  no labels

## Demo

<img width="2624" height="1644" alt="image" src="https://github.com/user-attachments/assets/b446db81-4e0d-498f-9e3a-45b26589736c" />

## Summary

Codex transcript activity now uses compact, type-specific presentation so mixed turns are easier to scan.

- Consecutive shell commands collapse into an expandable command group.

- Command summaries show the useful command text while expanded content keeps the exact invocation and output.

- File changes show filenames, relative directories, change kinds, addition/deletion counts, and bounded line-numbered diffs.

- Plans, reasoning summaries, tool activity, media, and lifecycle events have distinct icons, statuses, and disclosure behavior.

- Diff colors are semantic theme tokens with contrast coverage across all built-in themes.

## Business Value

People reviewing Codex work can understand progress, failures, and changed files at a glance without losing access to the underlying command output or protocol details. This reduces transcript noise and makes long implementation turns easier to review in a narrow desktop chat pane.

## Implementation Effort

Estimated 1.5 engineer-days for an average engineer to hand-code the presentation model, responsive styling, theme tokens, and regression coverage without AI assistance.

## Linear

[AI-778 — Make Codex transcript activity easier to scan](https://linear.app/builder-team/issue/AI-778/make-codex-transcript-activity-easier-to-scan)

## Test plan

- [x] pnpm test:chat

- [x] pnpm test:messages

- [x] pnpm exec tsc --noEmit

- [x] pnpm theme:check

- [x] git diff --check

#1319 — Improve mobile Admissions Forecast comparisons @YibinLongTrilogy  approved

## Summary

Improve the mobile Admissions Forecast school cards so users can compare QS and model variance without opening each school. The existing mobile Pipeline projection behavior is preserved, with explicit accessibility state on its expand control.

### Screenshots

<img width="549" height="606" alt="Screenshot 2026-09-11 at 3 57 23 PM" src="https://github.com/user-attachments/assets/ccf420ec-37a8-412b-a2d4-83e28a09b4e7" />

### Changes

- chat/components/dashboards/admissions/forecast/mobile/forecast-school-card.tsx — Add QS, Delta, and Delta % values beneath the existing capacity and enrollment metrics, reusing the shared Forecast formatters and tones.

- chat/components/dashboards/admissions/forecast/forecast-table.tsx — Export the existing signed-percent formatter for mobile reuse without changing desktop table rendering.

- chat/components/dashboards/admissions/forecast/pipeline-projection.tsx — Add type, aria-expanded, and an accessible label to the existing mobile Pipeline stage-breakdown control without changing its compact equation or horizontal stage layout.

- chat/components/dashboards/admissions/forecast/mobile/__tests__/forecast-school-summary.test.tsx — Assert the mobile card renders QS, Delta, and Delta % values.

- chat/components/dashboards/admissions/forecast/__tests__/pipeline-projection.test.tsx *(new)* — Assert the mobile Pipeline control exposes expansion state and reveals the stage breakdown when clicked.

### Design Decisions

QS, Delta, and Delta % are grouped in a compact bottom row so the primary school metrics remain in their existing two-column layout. The Pipeline expansion itself already existed on Main; this change only makes its state available to assistive technology.

## Business value

Admissions users can compare school-level QS and model variance at a glance, reducing repeated drill-downs while keeping the mobile Forecast layout compact and familiar.

## Estimated manual effort

30–45 minutes for an engineer familiar with the Admissions Forecast dashboard.

## Test Plan

- [x] Focused mobile Forecast tests passed: 5 tests.

- [x] Chat typecheck passed.

- [x] Targeted Biome check passed.

- [x] Test-runtime architecture check passed.

- [ ] Manually verify the mobile Forecast page in the running app.

#1318 — Add current Pipeline link to legacy report @YibinLongTrilogy  approved

## Summary

Add a clear way for users in the legacy Admissions Pipeline report to return to the current Pipeline report. The footer keeps the existing last-updated chip on the right and matches the current report's navigation affordance.

### Screenshot

<img width="595" height="296" alt="Screenshot 2026-09-11 at 3 00 02 PM" src="https://github.com/user-attachments/assets/e88dd954-0ebb-4572-afb7-0eb0da7fa1fd" />

### Changes

- chat/components/dashboards/admissions/funnel/funnel-view.tsx — Add a Current Pipeline footer link targeting /dashboards?tab=admissions&sub=admissions-pipeline while preserving the existing footer placement and freshness chip.

- chat/components/dashboards/admissions/funnel/__tests__/funnel-view.test.tsx *(new)* — Assert the legacy view renders the link with its exact accessible label and href.

## Business value

Legacy-report users can return to the supported Admissions Pipeline view without relying on dashboard navigation or browser history.

## Estimated manual effort

15–30 minutes for an engineer familiar with the Admissions dashboard.

## Test Plan

- [x] Focused browser test passed for the legacy footer link.

- [x] Chat typecheck passed.

- [x] Targeted Biome check and test-runtime architecture check passed.

- [ ] Manually verify the footer in the running app.

#57 — AI-777: Remove Codex chat information rail @ashwanth1109  no labels

## Demo

![Smoke test evidence](https://github.com/AI-Builder-Team/Shipyard/blob/e92d9da/.smoke-evidence/AI-777-smoke-test.png?raw=true)

## Summary

- remove the non-essential Codex thread-information rail and its dependent JSX/helper code

- let the conversation fill the default chat body while preserving artifact-preview layouts

- remove rail-only CSS, responsive rules, and template overrides

## Linear

https://linear.app/builder-team/issue/AI-777/remove-the-codex-chat-thread-information-side-panel

## Tests

- pnpm build

- pnpm theme:check

- pnpm test:chat

#55 — AI-775: Enhance task detail queue cards @ashwanth1109  no labels

## Demo

![Smoke test evidence](https://github.com/AI-Builder-Team/Shipyard/blob/3b63dd8dcad4d58fef90bde9704b3d86ebd99b43/.smoke-evidence/pr-55/image-1.png?raw=true)

## Summary

- replace the legacy task index and local metadata with the selected project icon and responsive task card hierarchy

- show the active workflow stage independently from overall task status

- add compact PR and Linear link controls with anchored dropdowns and default-browser opening

- preserve existing task detail actions, workflow controls, and deletion behavior

## Linear

https://linear.app/builder-team/issue/AI-775/task-detail-view-enhancements

## Validation

- pnpm exec tsc --noEmit

- pnpm build

- pnpm theme:check

- `git diff --check

#56 — AI-776: Harden quiet Codex turn recovery @ashwanth1109  no labels

## Summary

Codex task chat can look hung during a long quiet turn even when the app-server is still working, and a missed completion event can leave the UI stuck in Working…. This change adds authoritative status recovery and clearer runtime feedback.

- Track turn activity and show elapsed time for active turns.

- Check thread/read, the newest persisted turn, and pending requests during quiet periods and on demand, with coalescing, timeouts, and backoff.

- Reconcile terminal history so completed, failed, interrupted, and waiting turns repair the UI without replaying work.

- Guard against stale events and responses after switching threads; keep Stop usable while interruption completion is pending.

- Pause queued messages after confirmed interruption or failure and surface retry/status actions.

## Business Value

Users can distinguish a genuinely long-running Codex task from a stalled interface, recover from missed app-server events without restarting work, and avoid duplicate or lost queued messages. This makes extended research and implementation tasks safer to monitor and continue.

## Implementation Effort

Approximately 1–2 engineer-days for a comparable hand-coded implementation, including the runtime state changes, recovery UI, native command support, and deterministic event-driven tests.

## Linear

[AI-776 — Make Codex thread execution robust to quiet and missed completion events](https://linear.app/builder-team/issue/AI-776/make-codex-thread-execution-robust-to-quiet-and-missed-completion)

## Test plan

- pnpm exec tsc --noEmit

- pnpm build

- pnpm theme:check

- node --test scripts/test-thread-recovery.mjs scripts/test-workflow.mjs scripts/test-codex-messages.mjs scripts/test-chat-content.mjs (67 passing)

- TAURI_CONFIG='{"bundle":{"resources":[]}}' CARGO_TARGET_DIR=/Users/ash/Desktop/work/Shipyard/src-tauri/target cargo test --manifest-path src-tauri/Cargo.toml --lib codex_app_server::tests (13 passing)

#54 — AI-766: Improve thread HTML preview sizing and focus mode @ashwanth1109  no labels

## Summary

Conversation HTML previews now use the available assistant-message width, resize to their actual content, and keep their live iframe mounted while switching between preview and source. The artifact pane can be minimized and restored so the conversation can take the full window when focused review is needed. The supplied screenshot is checked in as the visual demo asset, and the stray Rust source text that blocked Tauri compilation is removed.

## Demo

![Thread HTML preview demo](https://raw.githubusercontent.com/AI-Builder-Team/Shipyard/codex/AI-766-thread-html-preview/docs/demos/thread-html-preview.png)

## Business Value

Users can review generated HTML and SVG work at the full conversation width instead of in a cramped intrinsic-size box, while dynamic previews stay correctly sized as content changes. Preserving the live frame avoids losing interaction state during source inspection, and artifact-pane focus mode gives design review the screen space it needs.

## Implementation Effort

Estimated 4–6 hours for an engineer working without AI assistance, including the responsive iframe sizing behavior, focus-mode controls, regression coverage, and validation.

## Linear

[AI-766 — Improve thread HTML preview sizing and focus mode](https://linear.app/builder-team/issue/AI-766/improve-thread-html-preview-sizing-and-focus-mode)

## Test plan

- pnpm test:chat (10 passing)

- pnpm exec tsc --noEmit

- pnpm theme:check

- pnpm build

- cargo check --manifest-path src-tauri/Cargo.toml --lib

#53 — AI-765: Render all Codex message types in Shipyard chat @ashwanth1109  no labels

## Summary

Shipyard chat now represents the full Codex app-server transcript instead of reducing output to assistant text and reasoning summaries. Users can inspect plans, command output, file changes, tool activity, agent collaboration, searches, review transitions, and unknown future items while preserving streamed and persisted history.

The chat also renders self-contained HTML/CSS/SVG previews, Mermaid diagrams, and image/audio/video/resource outputs. Local media is loaded through a bounded native command, and HTML previews run in an isolated sandbox with network and host access disabled.

## Business Value

Codex work is reviewable in the same Shipyard conversation where it happens. Users can see the generated visual or artifact preview, understand what tools and commands did, inspect diffs and plans, and recover the complete conversation after reconnecting instead of switching to another client or relying on opaque status text.

## Implementation Effort

An average engineer would likely need 3–5 working days to hand-code this change, including protocol normalization, streaming/history reconciliation, native bounded media loading, secure preview framing, Mermaid integration, component styling, and regression coverage.

## Linear

[AI-765 — Render all Codex message types in Shipyard chat](https://linear.app/builder-team/issue/AI-765/render-all-codex-message-types-in-shipyard-chat)

## Test plan

- pnpm test:messages — 25 transcript parsing, streaming, identity, history, tool, plan, and media cases.

- pnpm test:chat — 8 React DOM rendering, preview isolation, media, link, and activity cases.

- cargo test --manifest-path src-tauri/Cargo.toml --lib — 107 native tests, including bounded local asset reads.

- pnpm exec tsc --noEmit

- pnpm theme:check

- pnpm build

The native smoke harness was not launched because this change does not alter workflow orchestration or fixture behavior; the packaged UI still needs a supervised desktop smoke pass before release.

#52 — AI-764: Restore Research artifact template sections @ashwanth1109  no labels

## Demo

![Smoke test evidence](https://github.com/AI-Builder-Team/Shipyard/blob/bfa1b70e5a1b1f45ae881fa8147de0698f10a6e9/.smoke/evidence/ai-764-research-template.png?raw=true)

## Summary

- Restore the ticket-ready Research artifact outline in the workflow contract.

- Preserve the current YAML frontmatter format and add the canonical Research Context guidance.

- Require the five research sections through a focused unit test.

## Linear

https://linear.app/builder-team/issue/AI-764/recreate-the-missing-researchforticket-artifact-template

## Testing

- cargo fmt --manifest-path src-tauri/Cargo.toml -- --check

- cargo test --manifest-path src-tauri/Cargo.toml --lib workflow::

- pnpm build

#51 — AI-763: Make Research artifact approval resilient @ashwanth1109  no labels

## Summary

- Replace fragile whole-document XML-like tag counting with versioned YAML frontmatter plus a Markdown description body.

- Preserve bounded compatibility for existing outer-wrapper artifacts and make idle Research approval manual and actionable.

- Show validation errors in the artifact pane instead of silently disabling approval.

## Linear

[AI-763: Make Research artifact approval resilient and actionable](https://linear.app/builder-team/issue/AI-763/make-research-artifact-approval-resilient-and-actionable)

## Business Value

Researchers can approve a finished or intentionally stopped Research artifact without being blocked by ordinary Markdown that happens to mention the old field markers. Clear validation feedback reduces confusion and avoids rerunning research just to satisfy a hidden parser condition.

## Implementation Effort

Estimated 1–2 engineering days for an engineer working without AI assistance, including parser migration, backward compatibility, workflow-state changes, UI feedback, and regression coverage.

## Test Plan

- [x] cargo test --manifest-path src-tauri/Cargo.toml --lib (107 tests)

- [x] node --test scripts/test-workflow.mjs (18 tests)

- [x] pnpm test:smoke (26 tests)

- [x] pnpm exec tsc --noEmit

- [x] pnpm build

#47 — AI-760: Better icon for projects in top bar @ashwanth1109  no labels

## Demo

Visual-evidence override explicitly authorized by the user. The screenshot was provided in the smoke-test conversation, but this interface could not transfer it as a PR image attachment.

# Created this small PR to smoke test a certain feature

## Summary

- Replace the custom folder-with-lines projects glyph with the Lucide Layers icon.

- Preserve the existing 18px sizing, navigation behavior, active state, accessible labels, tooltip behavior, and focus treatment.

## Validation

- pnpm theme:check

- pnpm build

- Smoke test: PASS (user-confirmed)

## Linear

https://linear.app/builder-team/issue/AI-760/better-icon-for-projects-in-top-bar

#50 — AI-762: Default Codex conversations to Full access @ashwanth1109  no labels

## Summary

- Default every new, resumed, queued, and recovered Codex conversation to Full access.

- Persist explicit user access choices across resume and workflow thread replacement.

- Keep older persisted workflow inputs readable and normalize legacy unconfigured access.

## Business Value

Users can start Shipyard Codex conversations immediately with the expected Full access permissions, avoiding unexpected approval prompts while retaining control when they explicitly choose a narrower mode.

## Implementation Effort

Estimated 1–2 engineer-days to trace the conversation lifecycle, update persisted access-choice handling, implement recovery normalization, and add regression coverage.

## Linear

https://linear.app/builder-team/issue/AI-762/default-codex-conversations-to-full-access

## Test plan

- cargo test --manifest-path src-tauri/Cargo.toml --lib (106 passed)

- node --test scripts/test-workflow.mjs (18 passed)

- rustfmt and git diff --check

## Notes

Desktop UI smoke was not launched because starting the app/services was not explicitly authorized.

#48 — AI-756: Add semantic colors for task states @ashwanth1109  no labels

## Demo

![AI-756 smoke test](https://github.com/AI-Builder-Team/Shipyard/blob/feature/ai-756-task-state-colors/.smoke-demo-task-state-colors.png?raw=true)

## Summary

- apply semantic colors to local queue task state labels and dots

- group waiting, active, success, attention, and neutral states with a neutral fallback

- add regression coverage for all current labels and future-state fallback

## Scope

Queue-only, as agreed for AI-756. Task-detail workflow states are unchanged.

## Testing

- node --test scripts/test-workflow.mjs

- pnpm exec tsc --noEmit

- pnpm theme:check

- pnpm build

Note: the Rust portion of pnpm test:workflow could not complete because the expected staged src-tauri/resources/codex asset is absent in this checkout.

## Linear

https://linear.app/builder-team/issue/AI-756/different-colors-for-task-states

#44 — AI-757: Separate completed tasks in queue @ashwanth1109  no labels

## Demo

![AI-757 smoke test: separated current and completed tasks](https://github.com/AI-Builder-Team/Shipyard/blob/ai-757-separate-completed-tasks/docs/smoke-evidence/ai-757-task-grouping.png?raw=true)

## Summary

- Partition the local task queue into current and completed sections using the existing displayed status.

- Preserve creation controls, total counts, selection behavior, ordering, and global task numbering.

- Add compact section headings and visual separation for completed tasks.

## Validation

- pnpm build

- pnpm theme:check

- git diff --check

## Linear

https://linear.app/builder-team/issue/AI-757/separate-categorization-of-complete-vs-incomplete-tasks

#49 — AI-761: Add Shipyard production release agent skill @ashwanth1109  no labels

## Summary

- Add the detailed Claude workflow at .claude/skills/prod-release/SKILL.md.

- Add the Codex skill adapter at .codex/skills/shipyard-prod-release/SKILL.md.

- Keep the Codex skill pointed at the Claude workflow as the repository source of truth.

- Conditionally install lockfile-pinned dependencies when a fresh worktree needs them.

- Adapt the Lumen Tauri production-release workflow to Shipyard build, bundle, install, and launch paths.

## Business Value

Gives both Claude and Codex agents a repeatable, repository-local workflow for producing and launching the real packaged macOS app, including fresh-worktree dependency setup, reducing release mistakes and avoiding stale installed builds.

## Implementation Effort

Estimated hand-coding effort: 1–2 hours for an average engineer to inspect the Tauri packaging configuration, adapt the reference workflow, and validate both skill entry points.

## Linear

https://linear.app/builder-team/issue/AI-761/add-shipyard-production-release-agent-skill

## Test plan

- python3 /Users/ash/.codex/skills/.system/skill-creator/scripts/quick_validate.py .claude/skills/prod-release

- python3 /Users/ash/.codex/skills/.system/skill-creator/scripts/quick_validate.py .codex/skills/shipyard-prod-release

- git diff --check

- Review both skill files against package.json, scripts/tauri.mjs, src-tauri/tauri.conf.json, and README.md.

- Verify the documented conditional path uses pnpm install --frozen-lockfile only when the build dependencies are missing.

#43 — AI-755: Auto-start Implement after ticket readiness @ashwanth1109  no labels

## Summary

- Automatically enqueue Implement when Research is complete and Ticket is complete or skipped.

- Reuse the existing durable start command/request-key idempotency path.

- Reconcile ready tasks at startup and cover readiness guards with native tests.

## Linear

https://linear.app/builder-team/issue/AI-755/implement-node-should-auto-start

## Testing

- pnpm build

- pnpm test:workflow

- cargo fmt --manifest-path src-tauri/Cargo.toml -- --check

- Desktop smoke test: PASS on the supplied Implement worktree at commit aff793d063bf2f23ed3c35fa0b0eb21b87908178 using ./start.sh.

## Demo

User-provided screenshot evidence from the smoke test is attached in the Shipyard smoke-test handoff. It shows the launched Shipyard window with Research and Ticket complete, Implement complete, and Smoke Test in progress, confirming Implement auto-started without a manual Implement-node click and exposed the downstream workflow.

## Notes

This PR is intentionally opened as a draft. Implement starts only after both Research and Ticket are terminal-ready; Ticket may be complete or skipped.

#42 — AI-754: Queue Codex follow-up messages @ashwanth1109  no labels

## Summary

- Allow text and image follow-ups to be queued while Codex is processing.

- Drain queued messages FIFO only after the active turn completes successfully.

- Preserve queues per thread in memory across conversation navigation, with explicit retry after interruption or failure.

## Linear

https://linear.app/builder-team/issue/AI-754/allow-queuing-messages-while-codex-is-processing

## Tests

- pnpm test:messages

- pnpm test:workflow

- pnpm test:environment

- pnpm test:instances

- pnpm test:smoke

- pnpm theme:check

- pnpm build

## Demo

User-provided smoke-test screenshot demonstrates the launched desktop app with an active Codex turn still working and two queued follow-ups displayed in FIFO order (Test1, then Test2), each with a Remove control. The user reported “LG” (looks good).

Smoke-test worktree: /Users/ash/Desktop/work/Shipyard-ai-754

Launch command: ./start.sh

Commit verified: bc50282897d1bff5b4d02375e706d47ad82f85d2

#46 — AI-759: Make attached images reusable by the Codex agent @ashwanth1109  no labels

## Summary

Shipyard now keeps newly attached images available as local files for the Codex agent. The app-server receives the original image for visual context plus numbered, absolute paths under Shipyard's persistent Codex home, so the agent can read or reuse the evidence in the same or a later turn. Transport-only path metadata is hidden from the rendered transcript and does not duplicate user messages.

## Business Value

Users can attach a screenshot once and ask Codex to inspect, edit, or include that evidence in a follow-up task without manually saving and re-uploading the image. This removes the attachment handoff failure shown in the original report and makes screenshot-driven workflows reliable.

## Implementation Effort

An average engineer would likely need about 1–2 days to trace the desktop-to-app-server attachment flow, add safe persistent storage and validation, preserve transcript reconciliation, and cover the edge cases with tests.

## Test plan

- pnpm build

- pnpm test:messages (15 passing tests)

- TAURI_CONFIG='{"bundle":{"resources":[]}}' cargo test --manifest-path src-tauri/Cargo.toml --lib (101 passing tests)

- git diff origin/main --check

Desktop smoke testing with a real pasted image remains to be run in the app; no local service was launched for this change.

## Linear

[AI-759 — Make attached images reusable by the Codex agent](https://linear.app/builder-team/issue/AI-759/make-attached-images-reusable-by-the-codex-agent)

#45 — AI-758: Prevent duplicate Codex replies during conversation restore @ashwanth1109  no labels

## Summary

Shipyard can show a completed Codex reply twice when a conversation is reopened or history pages overlap with live stream state. This change reconciles restored history by turn, replaces terminal turns with their persisted messages, and ignores late events for turns already known to be finished.

## Business Value

Users see one trustworthy conversation transcript after reopening a task, without repeated assistant replies or confusing out-of-order history.

## Implementation Effort

An average engineer would likely need 4–6 hours to trace the resume/history race, implement turn-aware reconciliation, and add regression coverage.

## Linear

[AI-758: Prevent duplicate Codex replies during conversation restore](https://linear.app/builder-team/issue/AI-758/prevent-duplicate-codex-replies-during-conversation-restore)

## Test plan

- pnpm test:messages

- node --test scripts/test-workflow.mjs

- pnpm build

- git diff --check

#1315 — feat(reconciliation): add version-pinned protected reads (AERIE-1924) @caina-barbosa  approved

## Summary

This PR is Phase 3 of 8 in the larger [AERIE-1893 — Map automatic LOI and lease document reconciliation](https://linear.app/builder-team/issue/AERIE-1893/map-automatic-loi-and-lease-document-reconciliation) project. It replaces closed PR #1278 with the same reviewed behavior; the vendored generated OpenAPI is stored as semantically identical compact JSON so the complete diff fits automated review.

It extends the existing Aerie-to-Sindri connection in two narrow ways:

1. Aerie can tell Sindri the exact workflow version it expects when a future reconciliation run starts. Existing interactive starts do not send that option and behave exactly as before.

2. A future Sindri reconciliation run can call exactly two internal, read-only Aerie operations: list the documents authorized for that execution, and read bounded pages from those documents. The shared bearer only authenticates Sindri; it grants no document access by itself. Every call must also present an unexpired, unrevoked execution/read-grant pair that matches the exact Site, document, revision, generation, and source hash.

This work is tracked by [AERIE-1923 — Adopt version-pinned Sindri starts in Aerie](https://linear.app/builder-team/issue/AERIE-1923/slice-415-adopt-version-pinned-sindri-starts-in-aerie) and [AERIE-1924 — Add receipt-scoped immutable reconciliation reads](https://linear.app/builder-team/issue/AERIE-1924/slice-515-add-receipt-scoped-immutable-reconciliation-reads).

Production effect: compatibility hardening plus dormant/additive protected reads. This PR does not create reconciliation executions or grants, configure the bearer secret, start a workflow, register a production caller, or write Site data. Those callers and credentials do not exist in this phase, so merging initiates no external traffic and the new read boundary remains fail closed.

---

## Why

The existing integration lets Aerie authenticate to Sindri, start workflows, and inspect runs, but it does not prove that an automated run used the exact reviewed workflow version or provide a safe reverse path for Sindri to read source evidence from Aerie. Reconciliation needs both guarantees before a later coordinator can run unattended. This phase adds them without giving Sindri general Aerie, Drive, storage, or Site access and without activating the coordinator.

---

## Business Value

- Lets trusted server-side callers pin the exact Sindri workflow version they expect.

- Limits agent evidence access to the exact Site, document, revision, generation, and source hash authorized by a grant.

- Supports lossless bounded traversal of large validated artifacts without returning whole artifacts.

- Preserves existing Forge start behavior and introduces no business-data write path.

---

## How does it work

1. chat/convex/sindri vendors the canonical Sindri OpenAPI in deterministic compact JSON, regenerates types, and adds narrow expected-version and server-derived idempotency support. .gitattributes marks both artifacts as generated.

2. Existing interactive callers continue omitting the optional reconciliation fields, preserving their observable behavior.

3. chat/convex/reconciliation/reads.ts validates the execution receipt, read grant, exact source tuple, expiry/revocation state, receipt-bound storage ID/SHA-256/size, exact stored bytes, and cursor before returning bounded data.

4. Bearer-first internal HTTP routes expose only bounded document listing and content-page operations with uniform denials.

5. A dedicated Rhodes MCP server exposes exactly listSiteDocuments and readSiteDocumentContentPage; it has no write tool and remains unavailable without later credential propagation and valid grants.

### Authentication and authorization authority

AERIE_RECONCILIATION_READ_SECRET is only the transport credential proving that the caller is the trusted Sindri/Rhodes service. It is read from server environment, never accepted in MCP/model arguments, and does not select or authorize any Site or document.

Aerie remains the sole data-access authority. It authenticates the bearer before reading the HTTP method, request body, or database. It then resolves exactly one execution and the hash of exactly one read grant and requires matching purpose, contract version, Site, read-policy version, expiry, and revocation state. Content access additionally requires the exact receipt-owned document, knowledge version, revision, generation, MIME type, source hash, immutable Convex storage ID, persisted storage SHA-256, and byte size. Before parsing, the action hashes the exact fetched bytes and matches them to the receipt; the query rechecks the storage pointer and persisted _storage metadata after page construction to close replacement races. Missing, duplicate, expired, revoked, stale, altered, or cross-Site state is denied uniformly; content never falls forward to a different document version.

The uniform denial is an intentional security boundary, not an operational-status API. If the protected query/action cannot prove the complete authorization and source tuple for any reason, the HTTP gateway returns the same content-free denial rather than revealing whether a receipt, grant, Site, document, or artifact exists. The dedicated MCP proxy likewise emits one fixed tool error; it never turns a denial or backend exception into a successful read. Operational retry/state classification belongs to the later coordinator, not this evidence endpoint.

Grant validation in listSiteDocuments covers the complete operation. Convex executes the internal query as one serializable read transaction, so every source read sees the same database snapshot and cannot cross a concurrent revocation partway through. Convex also freezes Date.now() at function start, so the expiry clock cannot advance while the query iterates. A second grant check before return would read the same snapshot and frozen time and add no security. A revocation serialized before the query is denied; one serialized after it applies to later operations.

The execution and read-grant references—not the shared bearer—provide the per-run scope and revocation boundary. Cursors are HMAC-protected, bound to the operation and full request tuple, expire within five minutes, and can never outlive the grant. The MCP proxy validates returned identities and response sizes and emits one fixed error instead of upstream details.

### Runtime and test boundary

Production Convex modules remain edge-compatible: neither contractMonitoring.ts nor any production module imports Node APIs. The adjacent contractMonitoring.test.ts is test-only code executed by Vitest's Node/Vite host while the project supplies edge-runtime globals to the Convex behavior under test. Its node:fs and node:crypto imports only read and hash the checked-out generated artifact; they are not included in the Convex module bundle. This is exercised, not hypothetical: the exact focused command passed all 41 tests and the hosted full Test check passed on this head, including module loading and this artifact assertion.

The HTTP tests separately prove bearer-first rejection, bounded request parsing, and complete content-free 403 responses when the protected query or action actually throws. The MCP tests directly prove correct/wrong/missing bearer handling, successful proxying, thrown-fetch redaction, explicit 401/500 rejection before parsing for both tools, response identity validation, cursor forwarding, exact tool inventory, and distinct Durable Object binding. The exported route itself checks configured secret, then OPTIONS, then the same fixed-length bearer helper before dispatch; no branch reaches the Durable Object first. Importing the monolithic worker entrypoint as a Node test requires unrelated production/runtime bootstrapping, so an additional route-integration harness would require a broader production seam; the security-relevant authorization helper and branch ordering are already directly proven.

### Failure-path ownership

Each layer tests the behavior it owns rather than repeating every status at every caller. workflows.test.ts proves the optional version and idempotency fields are forwarded exactly; the shared sindriFetch client owns all workflow transport behavior. Its tests prove actionable 400 handling, opaque 401 handling, and redacted 500 handling, while the single client implementation catches fetch exceptions and treats 409 as an actionable rejection. A rejected workflow version, transport failure, or 409 therefore cannot become a successful start.

Likewise, the reconciliation MCP proxy rejects every non-2xx response before parsing its body and sends that thrown failure through the same fixed redaction path covered for both tools. Explicit 401 and 500 cases now prove that branch for both tools, including no body parsing and no upstream sentinel exposure. Separate HTTP tests now make both the protected query and action throw and assert the complete fixed content-free denial; the distinct over-bound test continues to prove validation before Convex invocation.

### Trusted internal boundaries

deny() deliberately throws a plain internal error so Convex redacts it if this internal-only function is ever called outside its intended gateway. The gateway is the user-visible boundary and converts it to the fixed content-free denial. Replacing it with a user-visible ConvexError would preserve denial detail across an accidental caller, weakening rather than strengthening that boundary.

The HTTP action and its internal query/action are deployed together by Convex and connected through typed internal references; they are not independently versioned services. Runtime schema validation therefore occurs at the actual external boundary in the Rhodes MCP proxy, where strict Zod DTOs and identity checks reject malformed or version-skewed responses. The content producer itself enforces the 512 KiB response contract before return; the trusted proxy independently verifies the received encoded size. Streaming an adversarial oversized response would be additional defense in depth, not a missing bound on the authorized Aerie producer.

The optional expected workflow version is server-owned in the later coordinator, and Sindri remains the authoritative contract boundary for its positive-safe-integer domain. Existing interactive callers omit it. An invalid ad hoc caller value is rejected by Sindri and surfaced through sindriFetch; Aerie does not silently coerce or treat it as a successful run.

### Empty-table rollout and persisted-state authority

The three receipt-bound artifact fields are required deliberately; there are no production receipt-source rows to migrate. reconciliationExecutionSources was introduced as dormant schema in Phase 1, and neither current origin/main nor this PR contains any production insert("reconciliationExecutionSources", ...) writer. The only inserts are test fixtures. This PR also does not create executions, grants, or receipt sources. Phase 4 owns the first production writer and must supply the exact storage ID, persisted SHA-256, and size when it creates a source snapshot.

Making these fields optional would weaken the fail-closed contract and create an unnecessary legacy branch for records that cannot legitimately exist. No backfill can or should invent an artifact digest. The required schema is therefore the safe widen for an empty dormant table and forces the future issuer to produce complete immutable receipts from its first row.

capturedKnowledgeState is not an unrestricted persisted string. reconciliation/schema.ts defines sourceKnowledgeStateValidator as the exact literal union available | refreshing | stale_but_available and uses it as the required field validator for every receipt source. Convex validates persisted rows against that schema, and the generated Doc<"reconciliationExecutionSources"> type preserves the same union before listing returns it. There are no older rows from a pre-union writer.

### Generated-contract authority

The compact OpenAPI is the same parsed contract as the accepted historical pretty artifact. The test reconstructs and pins the historical pretty hash, separately pins the committed compact hash and exact serialization, and keeps the generated TypeScript byte-identical. sync:sindri-spec deterministically reproduces both committed artifacts; compaction changes review representation, not the API contract.

The sync command is a developer-invoked import from an explicitly supplied local file, not a runtime fetch or unattended production updater. JSON syntax is checked before write, then openapi-typescript attempts generation from the result. Most importantly, the independent historical hash, compact hash, exact serialization, and generated-TypeScript checks prevent a valid-but-unrelated JSON document from being accepted or committed as this pinned contract. Extra pre-write shape checks or atomic replacement would improve local failure cleanup, but they do not create a silent production contract replacement path.

---

## Scope

### Included in this phase

- Deterministically compact vendored Sindri OpenAPI and generated API types

- Reproducible sync command that compacts the source before regenerating types

- Version-pinned workflow-start support and contract monitoring

- Dedicated bearer-first receipt-scoped reconciliation read routes

- Exact tuple/grant/cursor/artifact validation and bounded pagination

- Dedicated Rhodes reconciliation MCP server with exactly two read-only tools

- Dormant Durable Object binding and required generated/error-inventory updates

- Exact final diff paths:

.gitattributes

README.md

chat/.gitignore

chat/convex/_generated/api.d.ts

chat/convex/http.ts

chat/convex/reconciliation/http.test.ts

chat/convex/reconciliation/http.ts

chat/convex/reconciliation/reads.test.ts

chat/convex/reconciliation/reads.ts

chat/convex/reconciliation/schema.test.ts

chat/convex/reconciliation/schema.ts

chat/convex/sindri/client.test.ts

chat/convex/sindri/client.ts

chat/convex/sindri/contractMonitoring.test.ts

chat/convex/sindri/generated/sindriApi.ts

chat/convex/sindri/openapi/controlPlane.json

chat/convex/sindri/workflows.test.ts

chat/convex/sindri/workflows.ts

chat/lib/platform-error-coverage-inventory.ts

chat/package.json

chat/rhodes-worker/mcp-server/reconciliation-server.ts

chat/rhodes-worker/mcp-server/tools/reconciliation.test.ts

chat/rhodes-worker/mcp-server/tools/reconciliation.ts

chat/rhodes-worker/src/index.ts

chat/rhodes-worker/wrangler.jsonc

chat/scripts/sync-sindri-spec.mjs

### Deliberately excluded for later phases

- Observe-only coordinator, receipt creation, workflow starts, polling, and accepted-output verification — Phase 4

- Atomic Site/Property Acquisition commit engine and evidence UI

- REBL3 discovery, observation, registration, activation, and historical expansion

- Workflow/agent publication or activation

- PR7-owned deployment-secret propagation; this intermediate phase remains fail closed while unconfigured

- Deployment, environment or gate changes, production calls, Site/document mutation, and upstream writeback

### Rollout provenance

The accepted PR3 commits were replayed onto Aerie main 7b577ce253afc60d8bb9a7f663ffc96c526d4acd with provenance and stable patch IDs:

2c471fcfc -> 4b87e168a

ed3b1e261 -> d35929b0a

3d0de4461 -> c690ed5ae

A separate user-authorized adaptation 8656a3071 stores the OpenAPI as deterministic JSON.stringify(parsed) + LF, updates the sync command to reproduce it, and corrects repository guidance. Test-only follow-up 67d83a386 adds explicit thrown-backend and non-2xx redaction coverage without changing production code. Blocking integrity repair 274b1eae binds each receipt source to the exact immutable storage ID, persisted SHA-256, and byte size and verifies exact bytes before parsing. Final scope is 26 files, +3493/-26; raw binary diff size is 514,300 bytes, below Mercy's 600,000-byte review limit.

Historical Sindri source pin: 5c8312a40a840f2d6a32f18e52e9345da25f0f74. Historical pretty OpenAPI SHA-256 remains provenance: ea6fb6536b48fe2eef39d3afa0ba6eb658db7c2f32c8394df84aa90e3c49c8d8. Committed compact OpenAPI SHA-256: 89816a0223a6108022aeef09802f65ad8819777e499cf3d86747aaf354f561af. Generated API SHA-256 remains unchanged: 70409d0766295b3fc0e54b2981bb543437c0e195bfb3c5e61b82be9666b3bc40.

---

## Test plan

### Automated validation

- cumulative focused Chat slice — 44/44 passed, including 18 receipt-scoped read tests and 2 reconciliation schema tests; the suite proves a valid original artifact read, then rejects both pointer drift and a coordinated equal-size artifact substitution with unchanged declared source metadata, as well as actual thrown runQuery/runAction cases with complete fixed 403/redaction assertions

- dedicated Rhodes reconciliation MCP tests — 13/13 passed (pnpm --dir chat/rhodes-worker exec tsx --test mcp-server/tools/reconciliation.test.ts); this includes explicit 401 and 500 responses for both tools, proving rejection before JSON parsing and fixed sentinel-free output/logging

- Chat TypeScript — passed (pnpm --dir chat exec tsc --noEmit --pretty false)

- Convex TypeScript — passed (pnpm --dir chat exec tsc -p convex/tsconfig.json --noEmit --pretty false)

- Rhodes Worker typecheck — passed (pnpm --dir chat/rhodes-worker typecheck)

- changed-file Biome — passed

- architecture boundaries, Convex paths, and read bounds — passed

- git diff --check — passed

- deterministic sync — two runs from the historical pretty source produced byte-identical compact JSON and generated TypeScript

- failure safety — unset, unreadable, and invalid SINDRI_SPEC fail with fixed messages before changing artifacts

- artifact semantics — parsed historical and compact JSON are deeply equal; pretty reserialization retains the historical hash

- Git attributes — both vendored artifacts resolve to linguist-generated: true

- exact-head diff scope — only the 24 listed paths; raw patch 501,946 bytes

- Astra implementation review — PASS on 7b577ce25..1bf1abc44

- same-reviewer initial cumulative re-review — PASS on 7b577ce25..8656a3071

- same-reviewer blocker-repair re-review — PASS on 67d83a386..274b1eae; exact-byte receipt binding, pre/post storage metadata fences, equal-size substitution regression, and cumulative 44/44 focused slice accepted with no findings

One parallel Chat typecheck attempt reached 175 seconds under contention; its serial retry and final serial run passed. Hosted CI is the exhaustive exact-head gate.

### Time for Implementation

An engineer without AI assistance would likely need 3–4 weeks to align the generated Sindri client, implement the bounded protected-read boundary and MCP server, build the security and pagination matrix, rebase onto current main, and complete review and hosted validation.

#41 — AI-752: Open conversation links in default browser @ashwanth1109  no labels

## Summary

- Intercept http and https links rendered in conversation messages.

- Open external links with the operating system default browser through Tauri.

- Preserve local-file reveal behavior and show actionable opener errors.

## Linear

https://linear.app/builder-team/issue/AI-752/links-should-open-in-the-default-browser

## Testing

- pnpm exec tsc --noEmit

- pnpm build

## Scope

Only http/https links in conversation Markdown are changed. Other schemes, local-file links, and standalone Markdown previews remain unchanged.

#3761 — feat(spacex-valuation): add fair value reconciliation tables @sanketghia  approvedmercy-allow-critical

## Summary

- Add Current Fair Value, Total Gain, and unsold-share gain reconciliation tables to the SpaceX valuation page.

- Align realized-sale cost and realized-sale gain components with the Realized Share Sales table.

- Update the top-level invested basis and remove put premium from Current Fair Value.

## Validation

- SpaceX valuation suite: 234 tests passed

- TypeScript, ESLint, and Prettier passed

## Testing

- All changes have been reviewed live by Dave & Milo here - https://spacex.klair.ai/spacex-valuation

#40 — AI-753: Automatically advance Implement into Smoke Test @ashwanth1109  no labels

## Summary

Shipyard now advances a completed Implement node into a durable Smoke Test node automatically. The handoff preserves the exact dedicated worktree and commit, validates the repository before launch, waits for an explicit user smoke-test result, and completes the task only after Smoke-Test: PASS.

The coordinator also repairs completed Implement turns from older app instances through owner-routed reads. It leaves active legacy threads with their current owner, reserves new Smoke Test work for a capable scheduler, and protects retries, restarts, lost acknowledgements, pauses, and concurrent instances from duplicate effects.

The change includes the reusable Smoke Test prompt, workflow UI/status handling, durable snapshots/results, fixture-backed desktop scenarios, and updated smoke-testing documentation.

## Linear

[AI-753: Implement automatic Smoke Test advancement and legacy-owner handoff](https://linear.app/builder-team/issue/AI-753/implement-automatic-smoke-test-advancement-and-legacy-owner-handoff)

## Business Value

Users no longer need to manually unlock or reconstruct the validation step after implementation. Completed work moves into a controlled, reproducible smoke-test handoff while existing in-progress tasks continue naturally across Shipyard instances, reducing missed validation and duplicate Codex work.

## Implementation Effort

An average engineer would likely need approximately 4–6 working days to hand-code this change, including the durable workflow state, cross-instance ownership rules, recovery paths, UI updates, and fixture coverage.

## Test plan

- pnpm test:workflow — 18 JavaScript tests and 53 workflow Rust tests.

- pnpm test:instances — frontend isolation plus 5 instance Rust tests.

- pnpm test:smoke — 26 fixture harness tests.

- cargo test --manifest-path src-tauri/Cargo.toml --features smoke-test --lib — 97 Rust tests.

- pnpm exec tsc --noEmit.

- pnpm theme:check.

- git diff --check.

#39 — AI-751: Prevent duplicate Codex messages @ashwanth1109  no labels

## Summary

- centralize Codex transcript reconciliation across live completion, optimistic messages, cache, resume, and paginated history

- prefer server item identity and preserve identical text from distinct turns

- add focused regression coverage for duplicate live/history sources, streaming completion, optimistic reconciliation, and unknown turn identity

## Linear

https://linear.app/builder-team/issue/AI-751/prevent-duplicate-messages-in-codex-conversations

## Tests

- pnpm test:messages

- pnpm test:diagnostics

- pnpm exec tsc --noEmit

- pnpm build

- pnpm theme:check

The targeted Rust test was attempted but the local build requires the unstaged src-tauri/resources/codex runtime directory.

#3759 — fix(brainlift): route summaries to dedicated Sonnet model @marcusdAIy  approved

## Summary

- route long Brainlift summarization to the dedicated claude-sonnet-4-5 model

- keep Board Doc reasoning on the shared model

- add strict request-contract coverage for Brainlift summaries

- align the existing attachment-routing policy test

## Why

The Brainlift function documented Sonnet but passed BOARD_DOC_MODEL. After the shared model moved to Fable 5.1, long Brainlifts used the wrong provider contract and failed closed with BrainliftSummarizationError.

## Safety

- no retries or model fallback

- no raw or partial-content fallback

- typed content-free provider failure remains unchanged

- short Brainlifts still bypass the provider

- no temperature, thinking, or extra-body fields are sent to Sonnet

## Validation

- 105 directly relevant tests passed

- Ruff passed

- Pyright: 0 errors (5 pre-existing warnings)

- git diff --check passed

- two independent reviews completed; the first security blocker was fixed and re-reviewed

Linear: KLAIR-3533

#37 — AI-749: Support concurrent Shipyard instances with shared workflow data @ashwanth1109  no labels

## Summary

- Allow compatible Shipyard instances to share the same SQLite and Codex data while running independent coordinators and app servers.

- Route thread controls, approvals, and interrupts to the owning process through a durable mailbox with generation fencing and exactly-once delivery.

- Allocate independent loopback dev ports and disable macOS window-state restoration that can block a second instance at startup.

- Extend the smoke harness, documentation, and native coverage for concurrent instances, owner loss, recovery, and duplicate requests.

## Business Value

Developers can run multiple Shipyard windows or branches against shared task data, execute different tasks in parallel, and continue using a surviving instance when another process exits. Shared ownership and request fencing prevent duplicate turns, stale approvals, and accidental recovery of work that is still active in another process.

## Implementation Effort

Estimated 2–3 days for an average engineer to design and implement cross-process ownership and mailbox routing, concurrent SQLite coverage, dynamic dev-port startup, macOS restoration handling, and native smoke validation without AI assistance.

## Linear

https://linear.app/builder-team/issue/AI-749/support-concurrent-shipyard-instances-with-shared-workflow-data

## Test Plan

- [x] pnpm build

- [x] pnpm test:instances

- [x] pnpm test:workflow

- [x] pnpm test:smoke

- [x] pnpm theme:check

- [x] pnpm smoke build

- [x] Native smoke run with two processes, cross-instance approval and interrupt, owner stop/recovery, and duplicate approval rejection (run 3e6ac4b2-1db7-4a77-a0b9-546da0c3b455)

#38 — AI-750: Show detailed Codex errors in diagnostics @ashwanth1109  no labels

## Summary

Codex provider and app-server failures were reduced to a generic “Codex turn failed” message, hiding the actionable cause. This change carries detailed errors through the bridge and thread runtime, then displays them in Diagnostics → Recent activity with available thread and turn IDs.

- Normalize string and nested object errors from Codex notifications.

- Record runtime and failed-turn errors in the existing eight-entry Diagnostics feed.

- Preserve detailed errors when a later completion event omits its error field.

- Distinguish retryable warnings from final failures and deduplicate paired notifications.

- Add frontend and Rust regression coverage.

## Business Value

Users can identify and resolve model, runtime, and app-server compatibility problems directly from Shipyard. Support and engineering can correlate a visible failure to its thread and turn without searching separate logs.

## Implementation Effort

Approximately 4–6 hours for an average engineer to trace the event formats, wire diagnostics and chat state, and add regression coverage.

## Linear

[AI-750](https://linear.app/builder-team/issue/AI-750/expose-detailed-codex-turn-failures-in-shipyard-diagnostics)

## Test plan

- [x] pnpm test:diagnostics

- [x] pnpm exec tsc --noEmit

- [x] node --test scripts/test-workflow.mjs

- [x] pnpm build

- [x] rustfmt --edition 2021 src-tauri/src/codex_app_server.rs

- [x] git diff --check

#31 — AI-714: Enable pasted images in chat composer @ashwanth1109  no labels

## Summary

- Add image-aware paste handling to the enabled Codex chat composer.

- Route picker and pasted images through the same validation, FileReader, attachment preview, and send flow.

- Preserve normal text-only paste and leave the existing Codex image transport unchanged.

## Business Value

Users can paste screenshots and other clipboard images directly into a Shipyard chat, review or remove them, and send them with or without a text prompt. This removes the file-picker friction from a common image-sharing workflow while preserving existing limits and error handling.

## Implementation Effort

Approximately 2–4 hours for an average engineer to hand-code, test, and integrate this focused composer change without AI assistance.

## Test plan

- [x] pnpm build

- [x] cargo check --manifest-path src-tauri/Cargo.toml after staging the repository-required local Codex runtime

- [ ] Manually verify image-only, image-plus-draft, text-only, mixed, multiple, duplicate/over-limit, invalid/oversized, removal, send-failure retry, successful send, and transcript/history rendering in the macOS Tauri webview

## Linear

https://linear.app/builder-team/issue/AI-714/enable-pasted-images-in-the-chat-composer

#35 — AI-746: Add environment indicator and branch-aware dev title @ashwanth1109  no labels

## Demo

<img width="2624" height="1644" alt="image" src="https://github.com/user-attachments/assets/a74721ab-e94a-405d-8e9b-c5d95fede37c" />

## Summary

- add a static, accessible DEV/PROD indicator as the rightmost top-bar item

- set development window titles from the launching Shipyard worktree branch, with detached as the safe fallback

- preserve the production and smoke-specific native titles and keep the header usable at narrow widths

- add focused branch-resolution coverage and document the environment identity behavior

## Business Value

Makes it immediately clear whether Shipyard is running a development or packaged build and identifies the active development branch, reducing accidental testing in the wrong environment.

## Implementation Effort

Estimated 4–6 hours for an average engineer to investigate Tauri/Vite title behavior, implement the responsive and accessible indicator, add coverage, and validate the desktop bundle without AI assistance.

## Linear

https://linear.app/builder-team/issue/AI-746/add-an-environment-icon-and-branch-aware-development-title

## Test Plan

- [x] pnpm test:environment

- [x] pnpm build

- [x] pnpm theme:check

- [x] pnpm test:smoke

- [x] pnpm smoke build

- [x] resolve the Vite serve config and confirm Shipyard - codex/ai-746-environment-indicator

- [ ] complete native UI smoke observation after the separately running Shipyard instance is closed

## Validation Notes

The isolated Tauri smoke bundle built successfully. The harness-owned app could not finish initialization because another user-owned Shipyard process was already running and macOS held the new process before Tauri setup. The failed run was stopped through pnpm smoke stop; run evidence remains under .smoke/runs/e8cc6662-485b-4925-9af3-b335714cb901/.

#36 — AI-748: Route Shipyard Codex tasks through TrueFoundry @ashwanth1109  no labels

## Summary

Shipyard now runs company task Codex sessions through the TrueFoundry pay-as-you-go gateway with the openai-group/gpt-5.6-luna model. Users can manage the gateway key directly from the Codex connection screen.

## Business Value

Company work created in Shipyard is routed to the organization’s TrueFoundry usage account instead of the developer’s personal Codex configuration. This keeps company model spend attributable to the company gateway while preserving a clear in-app credential lifecycle.

## Implementation Effort

An average engineer working without an AI agent would likely need approximately 1.5–2 days to implement the secure credential lifecycle, isolated app-server routing, UI, documentation, and automated coverage.

## Implementation

- Added TrueFoundry settings to Codex connection with save, replace, remove, masked status, and reconnect behavior.

- Stored the key with the platform credential store through the Rust keyring crate; no plaintext fallback is used.

- Moved ordinary Shipyard app-server state to an app-owned Codex home instead of the personal ~/.codex directory.

- Enforced the tfy provider, TrueFoundry gateway, Responses API, and openai-group/gpt-5.6-luna at app-server startup and thread boundaries.

- Removed inherited personal OpenAI/Codex credential overrides and redacted the active key from app-server output.

- Disabled real credential operations in smoke builds and documented routing, isolation, and live verification steps.

## Test plan

- pnpm exec tsc --noEmit

- pnpm theme:check

- node --test scripts/test-workflow.mjs

- pnpm test:smoke (19 tests)

- cargo test --manifest-path src-tauri/Cargo.toml --lib (70 tests)

- cargo test --manifest-path src-tauri/Cargo.toml --lib --features smoke-test (71 tests)

- git diff --check

A live TrueFoundry request and billing-dashboard confirmation still require entering an organization key through the UI.

## Linear

[AI-748 — Route Shipyard Codex tasks through TrueFoundry](https://linear.app/builder-team/issue/AI-748/route-shipyard-codex-tasks-through-truefoundry)

#34 — AI-747: Fix Linear project picker loading state @ashwanth1109  no labels

## Summary

- Reset the Linear-project cancellation guard when TaskDetail mounts.

- Load persisted Codex thread settings before attempting rollout resume.

- Keep React Strict Mode lifecycle checks and transient empty-rollout races from leaving settled UI state suppressed.

- Preserve genuine-unmount protection for late async results.

## Business Value

Users can select a Linear project and see the active Codex model, access, and approval settings without either control being stuck indefinitely on “Loading projects…” or “Model pending”. Transient startup races now recover from persisted state while successful loads, failures, and refreshes still reach the visible UI as intended.

## Implementation Effort

Approximately 30 minutes for an average engineer to diagnose the Strict Mode lifecycle interaction and empty-rollout resume race, implement the two focused frontend fixes, update the ticket, and verify the frontend build.

## Linear

[AI-747 — Fix Linear project picker stuck loading in development](https://linear.app/builder-team/issue/AI-747/fix-linear-project-picker-stuck-loading-in-development)

## Test Plan

- [x] pnpm build

- [x] pnpm exec tsc --noEmit

- [x] git diff --check

- [ ] Manual development-app confirmation after reloading the build (not run in this session)

#33 — AI-745: Support shared Shipyard database instances @ashwanth1109  no labels

## Summary

- Make task creation tolerate the legacy trigger that pre-creates TaskNode rows.

- Use shipyard.sqlite3 as the canonical local database and document the shared-instance behavior.

- Allow multiple current app instances to open the database while one coordinator owns workflow dispatch.

- Preserve and harden the existing durable workflow lifecycle/history behavior and regression coverage.

## Business Value

Users can run the packaged and development Shipyard apps against the same local data without startup conflicts or task-creation failures caused by the legacy schema. Durable workflow state remains available across app restarts, and local testing can use one consistent database.

## Implementation Effort

Estimated 3–5 engineer-days for an average engineer to hand-code, test, and validate the persistence, runtime reconciliation, shared-instance coordination, UI status handling, and smoke-test coverage included here.

## Linear

[AI-745](https://linear.app/builder-team/issue/AI-745/support-shared-database-between-dev-and-packaged-shipyard-instances)

## Test Plan

- [x] pnpm test:workflow — 15 frontend workflow tests and workflow-native tests passed.

- [x] cargo test --manifest-path src-tauri/Cargo.toml --lib — 66 native tests passed, including the legacy-trigger and coordinator-ownership regressions.

- [x] pnpm build — TypeScript validation and production frontend build passed.

- [x] git diff --check — clean.

- [x] Dev app rebuilt and launched against the canonical database.

#1308 — Revise SIS new-vs-returning enrollment classification (#1307) @vvp-trilogy  approved

Closes #1307.

## Problem

int_enrollment.is_returning was an unconditional (student_id, prior-year) existence check that ignored the linked SIS application's pipeline_type. Students with an explicit NEW_ENROLLMENT application but a prior-year row (e.g. Armin Bernatonis, Alpha Boca Raton SY 2026: NEW_ENROLLMENT + prior-year WITHDRAWN) were wrongly classified returning.

## Algorithm (dbt/models/intermediate/enrollment/int_enrollment.sql)

is_returning is now:

CASE

WHEN app.pipeline_type = 'RE_ENROLLMENT' THEN TRUE

WHEN app.pipeline_type = 'NEW_ENROLLMENT' THEN FALSE

ELSE (py.student_id IS NOT NULL) -- qualified prior-year fallback

END

1. A recognized non-null pipeline_type is authoritative: RE_ENROLLMENT → returning, NEW_ENROLLMENT → new.

2. Null or unrecognized pipeline_type → documented prior-year fallback.

3. The fallback (student_year CTE) now excludes non-attendance statuses.

4. Both the source pipeline_type and the derived is_returning remain published for auditability (always a non-null boolean; a not_null schema test is added).

### Qualifying-status decision

The prior-year fallback counts a student as returning only on statuses that put the student on a roster — actual attendance/enrollment:

ENROLLED, PENDING_REVIEW, COMPLETED, WITHDRAWN, TRANSFERRED

Rationale (grounded in int_enrollment_classification's own semantics): ENROLLED/PENDING_REVIEW are attending (PENDING_REVIEW qualifies as enrolled per the report contract); COMPLETED is a finished prior year; WITHDRAWN was on the opening roster then departed; TRANSFERRED attended then moved. Excluded as non-attendance / pre-attendance / paused / cancelled: CANCELLED, CONFIRMED, ON_HOLD, PENDING_DEPOSIT, RE_ENROLLING, EXCHANGE. CANCELLED (the ticket's named example) therefore never counts as returning evidence.

## Downstream consumers verified

Both already read int_enrollment.is_returning and do not re-derive the old logic, so they inherit the revision consistently:

- re_enrollment_band in int_enrollment_cohort.sqlWHERE is_returning AND NOT is_transferred_status.

- First-day x_pipeline split in mart_enrollment_dtl.sqlCASE WHEN is_returning THEN 're-enrollment' ELSE 'new-enrollment'.

No column shapes change, so the TS consumer (packages/contracts/src/sis-enrollment.ts) is untouched. Comments/docs in both consumers that described x_pipeline as *not* using pipeline_type are corrected.

## Documentation updated (no longer "display only")

int_enrollment.sql header, _int_enrollment__models.yml (is_returning + new pipeline_type doc), stg_sis_application.sql header, _sis__models.yml, _sis__sources.yml, _mart_enrollment__models.yml, and the mart_enrollment_dtl.sql x_pipeline comment.

## Tests added

- Unit test int_enrollment_new_vs_returning (_int_enrollment__models.yml) — deterministic coverage of all required scenarios: explicit NEW_ENROLLMENT, explicit RE_ENROLLMENT, null-pipeline_type fallback (qualifying COMPLETED prior year), non-qualifying prior year (CANCELLED → not returning), and the conflicting Armin case (NEW_ENROLLMENT + prior WITHDRAWN → new).

- Singular assert_sis_enrollment_pipeline_type_authoritative.sql — live-data guard: no recognized pipeline_type disagrees with is_returning.

- Singular assert_sis_enrollment_returning_fallback_qualified.sql — live-data guard: the fallback flips only on qualified prior-year evidence (independently recomputed), catching any non-qualifying prior-year row (e.g. CANCELLED) that leaks into the flag. This also covers production, where unit tests are excluded from the scheduled build (that exclusion is added in #1310 — see below).

## Production-build guard split to #1310

The new unit test requires excluding the unit_test resource type from the hourly production build: dbt 1.12 unit tests are resource type unit_test, which --exclude-resource-type test does not exclude, so without the guard the unit test would run in — and could stop — the hourly refresh. That guard touches .github/workflows/dbt.yml (plus dbt/Dockerfile.dbt and dbt/README.md), a sensitive path, so it is split into #1310 (the human-reviewed catch-all for sensitive-path tweaks) to keep this PR free of .github/workflows and auto-approvable. #1310 should merge before (or together with) this PR; until then the guard is a harmless no-op because no unit test exists on main. This PR therefore touches only dbt/models/ and dbt/tests/.

## Validated locally vs. deferred to CI

No warehouse credentials / .env in this environment, so I did not execute against Redshift. Validated locally:

- dbt deps + dbt parse — full manifest builds cleanly (all ref()s, the unit test, and both singular tests parse and wire up).

- dbt ls — confirms unit_test:bran_dbt.int_enrollment_new_vs_returning and both assert_* tests are registered nodes.

Deferred to the dev CI gate (dbt build … --exclude-resource-type test then dbt test): actual model build + execution of the unit test and singular tests against sandbox_education.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1793 — Release all pain points for accounts above $1M ARR @mwrshah  approved

- Approve every live pain point for accounts above $1M total ARR, regardless of product.

- Keep repository-owned product rules for lower-threshold cohorts and enable Contently above $100K.

- Remove the redundant Khoros-specific filter; Khoros now passes through the universal account-total rule.

- Keep Action Hub admission sticky while Grainne approval becomes sticky only after persistence.

- Add relational coverage for universal approval, Contently's strict threshold, and shared admission/approval eligibility.

Live read-only impact check against the current 365-day window:

- Universal >$1M approval releases 135 active, unpushed pain points.

- Contently >$100K adds 42 active pain points.

#1795 — 080-2grainne-writeback @mwrshah  approved

- send the initial PostgreSQL owner domain in Grainne pain-point creation payloads and approval audits

- process later Grainne domain changes through the existing activity watermark, durable PostgreSQL state, and Salesforce retry path

- mirror nonblank Grainne domains to Account_Pain_Point__c.Domain__c

- let recurring Salesforce extraction seed only a currently blank PostgreSQL domain, with the ownership rule enforced atomically during upsert

#1794 — 079-grainne-evidence-payload @mwrshah  approved

- read Klair related_tickets as the Salesforce Evidence source for outbound pain points

- send that source text to Gráinne under the dedicated evidence payload field

- leave Gráinne’s Kayako-specific related_tickets field untouched

- lock the canonical payload and outbound request contracts with regression coverage

#1821 — feat(timeback): roll out incremental OneRoster sync @caina-barbosa  approvedmercy-allow-critical

## Summary

- activate one-hour watermark-incremental synchronization for the remaining 13 approved TimeBack OneRoster bulk entities

- generalize scope selection, clean-watermark lookup, continuation behavior, and immutable replay from assessment_results to the exact 14-entity incremental policy

- add manifest v6 for source-limit-isolated incremental evidence while retaining ordinary modified-since evidence as v5

- preserve explicit manual full mode, existing partition/checkpoint behavior, fan-outs, source projections, credential guards, and atomic key-scoped publication

This is the single implementation PR for SURTR-1174. The documented rollout waves are targeted validation runs after one deployment; they are not separate PRs or deployments.

## Final policy

### Watermark incremental with a one-hour overlap

orgs, academic_sessions, courses, classes, enrollments, users, demographics, line_items, results, assessment_line_items, assessment_results, score_scales, categories, resources

### Scheduled full snapshot

applications, test_assignments

### Unchanged fan-outs

map_percentiles, edubridge_enrollments, activity_facts, lesson_attempts, placement_tests

Only activity_facts retains its one-calendar-month rolling window.

## Source-limit safety

For demographics, line_items, and results, an allowed failed page is recovered only from immutable key-only membership evidence under the exact bounded cursor filter. Full singleton responses must match the proven membership ID and position. Allowed unresolved keys are retained in Redshift because they are never staged for deletion; valid delta keys still publish.

Source-limit isolation emits manifest v6. Replay verifies the original failure, membership, terminal confirmation, singleton outcomes, hashes, request identity, index-to-ID linkage, global membership closure, caps, and measured denominator without contacting TimeBack. Ordinary modified-since manifests remain v5 and continue to require zero source limits.

assessment_results and entities without an approved source-limit policy continue to fail closed on those responses.

## Compatibility and safety

- exact cursor-at-index-0 pagination; no client-side approximation of TimeBack collation

- every modified-since request remains bounded at extraction start

- explicit params.bulk_mode="full" is the only full-mode selector

- continuation runs use the same scheduled entity policy and cannot silently switch incremental entities to full mode

- normal deltas for enrollments, assessment_line_items, and resources bypass partition planning; explicit full mode retains v4 planning/checkpoints

- users retains its exact field projection and fresh/replay password rejection

- empty deltas remain evidenced ledger-only no-ops

- no DDL, migration, dependency, pipeline split, or schedule change

## Validation

- [x] runner suite: 251 passed

- [x] focused CR-1/CR-2 regressions: 7 passed

- [x] independent final review: PASS with no actionable findings

- [x] Ruff check

- [x] Ruff format check

- [x] git diff --check

- [x] exact 14 incremental / 2 full / 5 fan-out matrix tests

- [x] adversarial mixed ordinary/isolated replay and receipt-tamper tests

- [x] no production queries, pipeline invocations, or mutations

## Post-deployment validation

Use one production deployment, then run the three documented targeted validation waves within that deployment. Do not create separate rollout PRs. The following scheduled run must account terminally for all 21 entities and establish the new runtime baseline.

#3758 — feat(klair): Make BoardDoc refresh-stream tests environment-independent @sanketghia  approved

## Linear issue

[KLAIR-3538](https://linear.app/builder-team/issue/KLAIR-3538/make-boarddoc-refresh-stream-tests-environment-independent)

## Objective

Make the BoardDoc refresh-stream frontend tests environment-independent and restore a clean full frontend Vitest run.

## Observed failure

Command:

cd klair-client

pnpm test:run

Observed on 2026-09-11:

* Test files: 664 passed, 2 failed.

* Tests: 6,848 passed, 3 failed, 16 skipped.

* Command exited with status 1.

* Duration was approximately 103 seconds.

All three failures share the same mismatch:

* The failing assertions expect an EventSource URL rooted at http://localhost:8000.

* The implementation received/configured VITE_AI_ADOPTION_API_URL=http://localhost:5001 and constructed the URL with http://localhost:5001.

* The repository's local frontend configuration documents port 5001 as the normal backend pairing.

* The production behavior under test—minting an opaque refresh ticket, URL-encoding it, and constructing EventSource only after mint resolution—is not the identified problem.

## Failing tests

1. klair-client/src/services/__tests__/boardDocApi.openRefreshStream.spec.ts

* mints via one authenticated POST before constructing exactly one EventSource carrying only the encoded ticket

* URL-encodes the opaque ticket

2. klair-client/src/screens/BoardDoc/hooks/__tests__/useBoardDocWizard.refreshTicket.spec.ts

* mints exactly once before constructing exactly one EventSource; the URL carries only the ticket and no JWT/Authorization/bearer/token=

## Scope

* Update the affected test expectations/configuration so they do not depend on a developer's ignored local .env host.

* Reuse an established repository convention:

* derive the expected base from import.meta.env.VITE_AI_ADOPTION_API_URL with the same fallback as the implementation; or

* set a deterministic test-only API base and use it consistently in the affected specs.

* Prefer a small shared test helper or configuration pattern if that avoids duplicating the base URL logic, but do not introduce an unnecessary broad API refactor.

* Preserve assertions for:

* authenticated mint POST path and payload;

* mint completion before EventSource construction;

* exactly one fresh ticket per call/retry;

* percent-encoding of opaque ticket values;

* absence of JWT, Authorization, Bearer, and legacy token= data in the EventSource URL.

* Do not reintroduce JWT query-token fallback or weaken the ticket-based security contract.

* Do not change production behavior unless the implementation and tests are proven to have genuinely divergent base-URL ownership.

Existing patterns to consult:

* klair-client/src/services/__tests__/claireSessionApi.spec.ts derives its API base from import.meta.env.VITE_AI_ADOPTION_API_URL.

* klair-client/src/services/__tests__/claireFileApi.spec.ts and claireFileApi.download.spec.ts set a deterministic test API base before importing the module.

* Other frontend hook specs also override VITE_AI_ADOPTION_API_URL for isolation.

## Acceptance criteria

* The two affected spec files pass when run directly.

* cd klair-client && pnpm test:run passes with no test failures.

* The affected specs pass with the normal local backend configuration using port 5001.

* The tests remain deterministic when VITE_AI_ADOPTION_API_URL is set to another valid test URL.

* The test assertions continue to verify ticket ordering, encoding, freshness, and credential hygiene; only the environment-specific host assumption is removed.

* No production behavior changes are made unless explicitly justified in the PR.

* The PR description includes the test command and final result, and links this ticket.

## Non-goals

* Fixing the unrelated React act(...) warnings, chart sizing warnings, Browserslist warning, provider-context warnings, or other stderr emitted by passing tests.

* Changing the BoardDoc refresh-ticket protocol or backend authentication.

* Reworking the frontend API configuration architecture broadly.

* Merging or deploying automatically.

## Changed files

- klair-client/src/screens/BoardDoc/hooks/__tests__/useBoardDocWizard.refreshTicket.spec.ts

- klair-client/src/services/__tests__/boardDocApi.openRefreshStream.spec.ts

## Validation

- client-tests (test): passed

- client-lint-pr (lint): passed

## Terminal status

succeeded

## Workflow run

## Codex usage

37167 input / 391 output tokens

#3756 — fix(spacex-valuation): align fair value with waterfall @sanketghia  approved

## Summary

- Use one canonical waterfall reconciliation for the SpaceX equity value shown in the headline card and reconciliation table.

- Keep put-hedge value additive and separately scoped from the equity-only waterfall total.

- Preserve the gain/unrealized accounting identity and add fixed-value regressions for the $155 and $178 scenarios.

## Linear

- KLAIR-3536

## Validation

- SpaceX valuation suite: 232 tests passed across 15 files.

- Changed-file ESLint, Prettier, TypeScript, and production build passed.

- Full frontend suite: 6,848 passed, 16 skipped; 3 unrelated BoardDoc URL assertions fail because they expect localhost:8000 while this checkout is configured for localhost:5001.

## Screenshots

<img width="1442" height="648" alt="image" src="https://github.com/user-attachments/assets/ba9fd314-a1bd-4052-9fcb-f7de87967cad" />

<img width="1659" height="257" alt="image" src="https://github.com/user-attachments/assets/77a990fd-1115-4441-9f70-671bd6113fef" />

#1813 — feat(education): build a warehouse-owned Person directory @benji-bizzell  changes requested

## Summary

- Build a warehouse-owned Person directory with multi-source mappings, stable IDs and traceable relationship evidence.

- Separate legacy identity allocation from HubSpot domain refreshes and adapt Student/forecast readers without collapsing native contacts.

- Add disabled native-account capture dispatch and a recoverable, permission-preserving cutover bundle.

## Why

The old Person model centered on HubSpot contacts. V0 needs canonical identity across SIS, Guide, HubSpot, Finalsite, XO and Aerie accounts while retaining source IDs, unresolved evidence and relationship multiplicity. Existing warehouse readers must remain coherent when the canonical tables are replaced.

## Business Value

Provides the linked Redshift Person foundation for subsequent Aerie ontology APIs and People UI. Application surfaces and live data are unchanged by this PR.

## Breaking changes

Replaces legacy Person/xref columns and the relationship bridge contract. Included warehouse readers move to native source-coordinate joins; external readers must be checked at release. Initial legacy IDs may retire; V0 IDs persist thereafter. Apply the migration bundle, not new-install DDL over existing tables.

## Test plan

- [x] 234 affected Python tests; repository Ruff lint/format; CDK TypeScript build.

- [x] Dev nine-table swap: injected rollback restores columns, owners, grants and the guardian view; successful replacement preserves recoverable legacy rows and access.

- [x] Actual dev native refresh, publication A→B, HubSpot domain and Student procedures, plus guardian/Summer/transfer queries. Includes stale-contact rejection, unresolved competing contacts, and ambiguous profile metadata.

- [x] Seven read-only review lanes; substantive Mercy findings independently evaluated, corrected or answered with evidence.

- [x] Redshift dev reproduces the old parent-evidence overflow and validates lossless bounded assembly, stable IDs and unresolved cases. Direct Student reads withhold stale Person attribution without dropping native contacts.

- [x] All seven final-head CI checks pass at f58cdf18, including full pipeline-runner and CDK suites.

- [x] Second Mercy pass assessed against concrete production blockers; the two demonstrated defects are fixed and other findings have evidence-backed dispositions. Subsequent automatic review is not authorization for further changes.

## Release boundary

No merge, deployment, production DDL/DML/grants or activation performed. The runner remains disabled. Deliver the narrow Aerie native-account producer (not yet on Aerie main), install its source staging/publisher, capture initial input, then execute the canonical swap and dependent refreshes in one transaction. Source delivery and effective live credentials remain activation checks—not proof supplied by synthetic dev tests. Person APIs/UI remain later PRs. See the runner's RELEASE.md and migrations/README.md.

#1309 — feat(enrollments): export SIS enrollment summary to CSV @vvp-trilogy  approved

## Summary

Implements #1300 — adds CSV export to the aggregate SIS enrollment report at /dashboards?tab=admissions&sub=enrollments&source=sis. Exports the matrix already loaded in the browser (selected school year, currently-filtered + sorted school rows). Frontend-only — no dbt / worker / Convex schema / Convex query / public-API / agent changes.

Closes #1300.

### What's included

- Shared CSV infra — extracted a generic pure buildMatrixCsvData<Row> in chat/components/dashboards/shared/csv-export.ts (leading label column + metric columns + Total footer only when non-empty). CSV escaping / spreadsheet-formula neutralization stays in the existing downloadCsv/escapeCsvValue.

- Incumbent (behavior-preserving)chat/components/dashboards/admissions/enrollments/derivation.ts: buildEnrollmentCsvData now delegates to buildMatrixCsvData; output is byte-identical (guarded by the unmodified incumbent tests).

- SIS pure logicsis/derivation.ts: buildSisEnrollmentCsvData (exact 11-column order, year-prefixed metric headers, zeros as numeric 0, filtered Total) + sisEnrollmentCsvFilename (sis-enrollments-sy{yr}-{date}.csv), reusing the repo's school-year + calendar-day formatters.

- SIS UIsis/sis-enrollments-export-button.tsx (themed download icon button, tooltip/label Export SIS enrollments CSV, disabled when no rows, independent of the #1298 drill-down) wired into the SIS report toolbar with the filtered rows, filter-tracking totals, and selected year.

- Testssis/__tests__/derivation.node.test.ts (header order, year formatting, row values, zeros, filtered totals, empty data, sort preservation) + sis-enrollments-export-button.test.tsx (enabled/disabled + export args).

### Stacking note

This was stacked on #1298 (now merged, squash 82de51a0); rebased onto main so it contains only #1300's changes. Rebase re-verified: pnpm typecheck + pnpm biome check clean, SIS CSV suites green.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1303 — feat(enrollments): add SIS enrollment list and basic detail pane @vvp-trilogy  approved

## Summary

Implements #1298 — adds a focused student drill-down to the aggregate SIS enrollment report delivered by #1296. Clicking a populated metric cell opens a cursor-paginated SIS enrollment list; selecting a student opens a basic enrollment detail pane. Reads only sandbox_education.mart_enrollment_dtl.

Closes #1298.

### What's included (by layer)

- Contractpackages/contracts/src/sis-enrollment.ts: metric→cohort map, membership row type, deterministic collapse (smallest enrollmentId), per-metric tally, linear aggregate↔detail reconciliation.

- Syncsync/src/redshift/sis-enrollment.ts (student-grain reader, has_fact, calendar-day dates) + sync/src/analytics/sis-enrollment-refresh.ts (collapse → reconcile → batch-insert → atomic pointer flip → batched prune).

- Convex schemasisEnrollmentCohortStudents (only new table) + compound indexes.

- Convex analytics — idempotent insertSisEnrollmentCohortStudents, batched orphan-safe pruneSisEnrollmentCohortStudents, operation allowlist extended (same bearer route as #1296; no new route).

- Convex queriesgetSisEnrollmentStudents (cursor-paginated) + getSisEnrollmentStudentDetail (null when missing), both gated by admissions.enrollments.read, arg + return validators, indexed reads, active publication only.

- UI — clickable metric cells (count > 0), new sis-student-panel.tsx (list/detail, load-more, prev/next, back/close, keyboard, mobile overlay, empty/loading/stale states) reusing the incumbent panel's structure with SIS-specific types.

### Stacking / rebase note

This branch was originally stacked on #1296 (PR #1299). #1296 has since squash-merged to main (fc57a43e1), so this branch was rebased onto main — it now contains only #1298's changes on top of the merged #1296 foundation, and preserves #1296's post-fork fixes (publication ordering guard, (programCode, schoolYear) grain-uniqueness rejection, quote-normalization before the test-school check).

### Incidental cleanup during the rebase

While resolving the rebase, I found that #1296's grain-uniqueness fix had landed on main with a raw NUL byte used as the grain-key delimiter in chat/convex/admissions/analytics/sisEnrollment.ts (line ~92), which embeds a NUL in a .ts source file (harmless at runtime — dedup still works — but it trips text tooling, including git's merge). I normalized it to the behavior-identical \^@ escape. No functional change to #1296's dedup semantics.

### Verification

pnpm typecheck (all workspaces) and pnpm biome check clean. SIS suites green: contracts 22, sync 16, chat 126 (convex analytics suite now 38 tests covering both #1296's ordering/grain-uniqueness regressions and #1298's pagination/mapping/collapse/reconciliation/detail cases).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

### Deploy order (operational)

Deploy the Convex backend before the analytics worker. During a brief worker-ahead window the SIS enrollment refresh skips gracefully (non-fatal; last-known-good preserved) until the backend supports the new membership operations — no data is wiped and other domains are unaffected.

#1820 — chore(timeback): preserve restored production schedule @caina-barbosa  approvedmercy-allow-critical

## Summary

- set the repository's timeback-raw-sync schedule state back to enabled

- update the executable CDK guard and README to reflect the completed production canary

## Important deployment state

The production EventBridge rule is already enabled directly in AWS:

- Rule: pipeline-timeback-raw-sync-schedule-prod

- State: ENABLED

- Schedule: cron(10 7 * * ? *)

This PR does not activate the schedule. It reconciles the repository's desired state with AWS so the next CDK deployment does not turn the already-restored schedule off again.

## Production canary evidence

The authorized assessment_results-only canary succeeded before the schedule was restored:

- Execution: surtr1166-production-canary-retry-20260910T211706Z

- Pipeline run: 74661194-4c42-426a-a71f-df4e2e706f32

- 1,456,939 unique rows published to raw and clean

- 729 keyset pages with 728 exact inclusive-boundary overlaps

- independently landed terminal confirmation matched byte-for-byte

- zero invalid rows, key mismatches, source limits, rejected rows, work tables, or staging objects

- one successful ingestion-ledger row

- 15,757,063 older historical rows remained present

- global raw and clean tables each contained 17,214,002 unique keys after publication

## Validation

- [x] Direct AWS describe-rule verification: ENABLED

- [x] npm test -- --runInBand test/real-pipeline-configs.test.ts — 515 passed

- [x] python3 -m json.tool pipelines/runners/timeback-raw-sync/pipeline.json

- [x] git diff --check

#1306 — feat(education): add managed document relink API @benji-bizzell  approved

## Summary

- Add an explicit site-document relink endpoint for existing managed Drive files

- Verify destination access and containment under the site's configured Drive root

- Preserve document identity and metadata while rolling knowledge lifecycle state and auditing source changes

## Why

Artemis cannot repair duplicate or mirror Drive registrations through the metadata-only PATCH contract. A dedicated, concurrency-guarded relink path is required to safely repoint a stable document reference.

## Business Value

Operators can repair source identity without deleting and recreating document registrations or losing metadata and knowledge lifecycle integrity.

## Test plan

- [x] Focused Convex/API tests: 58 passed across document HTTP, curl contract, and portfolio tests

- [x] Rhodes worker verification suite: 250 passed

- [x] Chat typecheck, test-architecture lint, full lint, and changed-file Biome checks

- [x] Current-head CI: lint, typecheck, tests, builds, Docker builds, and secret scan passed

- [x] Adversarial review completed; all confirmed findings fixed

- [x] Mercy approved current head with no findings

#1816 — fix(timeback): treat source key collation as opaque @caina-barbosa  approved

## Summary

- treat TimeBack sourcedId ordering as an opaque source-defined collation rather than approximating it with Python casefold()

- retain inclusive cursor repetition, exact-ID overlap accounting, immutable replay validation, and fail-closed cursor progress

- add regressions for TimeBack's observed natural ordering (q1 before q12)

## Why

The first authorized production canary failed closed before publication because TimeBack validly returned ..._q1_157_result before ..._q12_157_result. Python lexicographic/casefold ordering considers that pair descending, but TimeBack uses a natural/numeric source collation.

The client does not need to reproduce that private collation. It now verifies source-independent invariants: the exact inclusive cursor must recur, exact IDs are unique within each response, previously seen IDs are counted as overlap, and every full follow-up page must contain new records and end on a new exact cursor. Replay enforces the same evidence contract.

The timeback-raw-sync production schedule remains disabled.

## Production evidence

- Failed canary execution: surtr1166-production-canary-20260910T203900Z

- Pipeline run: 7e6f5801-151c-41ff-8934-7f3b3d91ec13

- Task revision: pipeline-timeback-raw-sync-prod:22

- Exact failure: TimeBack assessment_results keyset page sourcedIds are not ascending

- Reproduction isolated the first mismatch on page 19: ..._q1_157_result followed by ..._q12_157_result

- Read-only verification ff63d56d-1e7f-43d1-85fb-b6c7ddab5d10 confirmed the clean table remained at 15,781,796 rows with its prior watermark and the failed run inserted zero ledger rows

- The corrected opaque-cursor invariants traversed the first 25 live pages (50,000 response rows; 49,976 unique rows after inclusive overlaps), including page 19, without failure

No second production pipeline execution was started.

## Test plan

- [x] natural-sort regression failed before the fix and passes afterward

- [x] cd pipelines/runners/timeback-raw-sync && uv run pytest -q — 204 passed

- [x] uv run ruff check src tests scripts

- [x] uv run ruff format --check src tests scripts

- [x] git diff --check

- [x] 25-page read-only live TimeBack traversal with the production canary window

- [ ] repository CI

- [ ] Mercy review

#3754 — KLAIR-3532: Complete Drive context attachment acceptance @marcusdAIy  approved

## Summary

- accept native Microsoft Word .docx Drive context snapshots using the existing bounded parser subprocess

- recursively extract nested DOCX tables in document order under one cumulative hard output cap

- count real XLSX cells rather than formatting-generated EmptyCell placeholders while retaining all row, column, sheet, byte, output, CPU, memory, and wall limits

- preserve actionable attachment failures through automatic add-on hydration

- replace the large attachment banner with a compact accessible paperclip control and remove the persistent snapshot notice

- use dedicated claude-sonnet-4-5 for bounded attachment summarization; keep Fable for all other Board Doc callers

## Safety properties

- no provider retry or fallback after refusal, truncation, timeout, or ambiguous completion

- no automatic source sharing or replay of parked reservations

- exact DOCX MIME allowlist; XLSX remains an internal Google Sheets export only

- nested/oversized/partial/empty parser outcomes fail closed

- attachment text remains untrusted latest-user-turn evidence with MCP disabled

- no new persisted or public type_label

## Validation

- focused backend: 121 passed

- full Board Doc: 5010 passed, 2 deselected

- full add-on: 355 passed

- Ruff: passed

- Pyright: 0 errors (existing warnings only)

- git diff --check: passed

- final exact-tree backend review: approved

- final exact-tree add-on review: approved

- production synthetic Sonnet probe: end_turn, usable text validated

## Live findings addressed

- DOCX opened through Google Docs was rejected as unsupported

- typed 422/502 attachment feedback disappeared after hydration

- sparse/formatted native Google Sheets were rejected by bounding-box placeholder counts

- large native Google Doc extraction succeeded but Fable refused the summarization request

Linear: KLAIR-3532, KLAIR-3531

#1305 — Compact Admissions mobile KPI summaries @YibinLongTrilogy  approved

## Summary

Compact the Admissions KPI header on mobile so Pipeline and Funnel views show four summary cards in a readable 2×2 layout without wasting the first screen on stacked cards or an empty filter-chip row. Desktop sizing, Forecast, tables, filters, and data queries are unchanged.

### Screenshots

<img width="561" height="363" alt="Screenshot 2026-09-10 at 4 27 53 PM" src="https://github.com/user-attachments/assets/b3038c85-8741-4524-b70a-270a6070a6e5" />

<img width="563" height="351" alt="Screenshot 2026-09-10 at 4 27 44 PM" src="https://github.com/user-attachments/assets/09c15008-6024-4d20-b2ee-9410d06b2354" />

### Changes

- chat/components/dashboards/admissions/shared/admissions-kpi.tsx *(new)* — Add the Admissions-only 2×2 mobile grid and opt-in compact stat wrapper, restoring four columns at the desktop breakpoint.

- Admissions Pipeline and Funnel views — Use the shared Admissions KPI treatment and mount the min-h-5 filter-chip section only when school, status, or year chips are active.

- Community and Established Funnel views — Use the same mobile KPI grid and compact card treatment.

- chat/components/dashboards/shared/summary-stat.tsx — Add an opt-in compactOnMobile variant; existing Expenses and other dashboard consumers retain their default sizing.

- Focused tests — Cover the 2×2 grid contract, compact card classes/details, active-vs-empty filter-chip spacing, and both Community and Established Funnel KPI headers.

### Design Decisions

- Keep the mobile search and toolbar controls in their existing rows and touch-target sizes.

- Scope the density change to Admissions instead of changing SummaryStat defaults globally.

- Preserve all headline and supporting detail values in the compact cards.

## Business value

Admissions users can see the four most important pipeline/funnel metrics in the first mobile viewport, making the dashboard faster to scan while retaining the supporting counts and existing desktop experience.

## Estimated manual effort

30–45 minutes for an engineer familiar with the Admissions dashboard.

## Test Plan

- [x] Focused browser tests passed: 5 files, 64 tests.

- [x] Full workspace pnpm typecheck passed.

- [x] Changed-file Biome validation and test-runtime architecture check passed.

- [ ] Manually verify narrow mobile rendering in the running app.

#3752 — KLAIR-3527: Attach Drive files as Coach Claire context @marcusdAIy  changes requested

## Summary

- add explicit Drive context attachments to the Google Docs add-on for native Google Docs, native Google Sheets, and PDFs

- require both the identity-bound add-on user and the production service account to read the source

- store bounded Claire-only snapshots through the existing chat attachment context path

- add list, attach, remove, and explicit pending-preparation recovery flows

## Authorization and safety

- no client-supplied session ID; the active Google Doc resolves the session server-side

- GET/list requires read_coach; attach, remove, and discard require mutate

- current Budget Doc cannot attach itself; maximum three attachments including durable reservations

- no automatic sharing and no OAuth scope expansion

- user-token access and service-account read are both mandatory

- URLs, file IDs, tokens, provider bodies, extracted content, and source fingerprints are excluded from logs/errors/chip responses

- service-account access failure gives bounded Viewer-sharing guidance only

## Durable execution and extraction

- atomic DynamoDB/CAS reservation occurs before linked provider work

- same-source requests are single-flight across processes and reservations count toward capacity

- known safe failures clear the reservation; ambiguous started operations park without replay

- Editors/Owners can explicitly discard a content-free pending preparation; Viewers/Commenters cannot

- metadata and downloads use separate bounded pools with nonblocking admission, socket and wall deadlines, and safe late-exception consumption

- PDF and XLSX parsing runs in a killable resource-limited subprocess with hard page/sheet/row/cell/output bounds

- exact known extraction placeholders are rejected; ordinary bracketed text remains valid

## Add-on UX

- explicit attachment chips with hydrate, remove, and pending recovery states

- shared busy/request-version guards prevent stale hydration and concurrent mutation races

- controls disable while busy; aria-busy/aria-live and focus restoration are covered

- attachment-only transport returns allowlisted bounded errors and never logs dynamic paths or response bodies

## Review follow-up hardening

- user transfer requires canDownload OR canCopy; the shared point-in-time snapshot policy is explicit and internally audited

- stale wizard-step CAS retries preserve attachment/reservation state and rebuild per-attempt response outputs

- one bounded pipeline covers metadata, download, parser, summary, and finalize; parser and summarizer have independent admission/time limits

- extraction caps retain a typed 413 path; shared-drive media requests set supportsAllDrives=True

- Viewer/Commenter mutation controls fail closed; chat, batch review, hydration, and attachment operations use one composed interaction state

- attachment data is fully entity-escaped untrusted user-turn evidence, never static system authority; attachment turns do not expose or execute inline MCP tools

- every Yibin review thread has an individual exact-SHA evidence reply and is resolved for re-review

## Validation

On exact head d708b7a853fce825c756784c29059018e333500a against main ac877bcd0440bb5166988a9e77cd32ccc9631bf6:

- full Board Doc suite: 5,003 passed, 2 deselected

- full add-on suite: 351 passed

- focused backend correction suite: 181 passed

- focused prompt/MCP suite: 59 passed

- Ruff: passed

- Pyright on changed production Python: 0 errors, 0 warnings

- git diff --check: passed

- independent security/integrity review: APPROVE, no findings

- independent add-on/transport/accessibility review: APPROVE, no findings

- independent prompt/MCP review: APPROVE, no findings

## Tracking

- KLAIR-3527

- Apps Script publication remains a separate, explicitly approved version 5 release step after backend deployment and health verification

- no production deployment or publication is performed by this PR

#1814 — feat(timeback): add incremental assessment sync behind rollout hold @caina-barbosa  approvedmercy-allow-critical

## Summary

> Intentional rollout hold: this PR sets timeback-raw-sync to schedule.enabled: false and updates the repository's real-pipeline contract to require that state. This prevents an automatic complete production run before the new assessment_results path receives separately authorized targeted validation. On-demand execution remains available. Scheduled operation will be restored only by a separate post-validation repository change.

- make assessment_results use a one-hour-overlap dateLastModified incremental extraction by default

- retain the existing full snapshot only when the caller explicitly sends bulk_mode: "full"

- publish non-empty deltas atomically by sourced_id, preserving target rows absent from the delta

- add immutable manifest v5 evidence and idempotent replay for modified-since extraction

- temporarily set the pipeline schedule to disabled so deployment cannot start the complete pipeline before targeted production validation

Linear: [SURTR-1166](https://linear.app/builder-team/issue/SURTR-1166/add-safe-watermark-incremental-bulk-synchronization-to-timeback-raw)

## Why

The scheduled pipeline currently performs a complete download of mutable OneRoster entities. assessment_results alone has exceeded 15 million rows, and its deep-offset full extraction has repeatedly failed while the source was changing.

This change activates the new mechanism only for assessment_results. The other bulk entities, applications, test_assignments, and all fan-outs retain their current behavior.

## Behavior

For default assessment_results execution, the handler:

1. reads MAX(date_last_modified) from the published clean table;

2. subtracts exactly one hour;

3. sends the resulting immutable dateLastModified >= effective_start filter on every TimeBack request;

4. traverses the filtered result with bounded sourcedId keyset pagination;

5. lands exact response bodies and a checksummed manifest before publication; and

6. atomically replaces only the raw and clean rows whose sourced_id appears in the delta, together with the ledger insert.

TimeBack's filtered totalCount can drift, so it is retained as per-page evidence but is not used as the general continuation condition. Each delta is bounded above at extraction start, and replay verifies that every source row remains within that exact time window. Requests always use offset 0. Follow-up requests use an inclusive sourcedId >= cursor boundary so a distinct ID that compares equal under TimeBack's case-insensitive ordering cannot be skipped between pages. Exact repeated boundary rows are recorded in each receipt and ignored during unique-record replay; distinct casefold-equal IDs are retained. A short terminal response fails closed if totalCount claims omitted rows or if an identical second response cannot be landed as independent confirmation. A full page that cannot advance beyond its boundary also fails closed. Cursor values containing apostrophes or backslashes fail closed.

A valid empty delta inserts only its ledger evidence. It does not create work tables, stage files, run COPY, or mutate either target.

Missing/all-null clean watermarks fail before source extraction and instruct the operator to request explicit full mode. There is no scheduled, weekday, drift-triggered, or error-triggered full fallback.

Replay uses the original manifest, watermark, filter, pages, and extraction identity. It cannot recompute a watermark or turn a delta into a full-table replacement. Existing v1-v4 full/fan-out replay behavior remains unchanged; modified-since evidence uses manifest v5.

## Redshift publication safety

Separate Redshift Data API calls do not share temporary-table sessions, so non-empty deltas use UUID-scoped permanent work tables. Raw and clean staged counts, lineage, null keys, and duplicate keys are checked before the final batch.

The final BatchExecuteStatement contains raw delete/insert, clean delete/insert, and ledger insert as one Redshift transaction. Work tables and temporary S3 objects are cleaned after both success and failure. No DDL or migration is required.

## Development validation completed

The code was executed locally against the live TimeBack API, real S3, and the Redshift dev database. It has not been built or deployed through CDK and has not run in production.

All Redshift calls explicitly used Database=dev with isolated development S3 prefixes. No finance_dw query or write occurred.

- Preflight 4a7d88ec-6b17-4a74-a8f0-33d90697d429 confirmed current_database() = dev and 43 canonical TimeBack tables.

- Missing-watermark query 46fb6aa1-5388-43fa-ba5a-700d94f9abe6 returned SQL NULL; a fail-on-call source sentinel confirmed TimeBack was not contacted and targets were unchanged.

- The initial default dev run surtr1166-dev-default-keyset-20260910t1607 landed 56,093 rows in 29 keyset pages and proved the key-scoped publication path before the later pagination and terminal-completeness hardening. It did not run full partition planning.

- Publication batch 48236f27-4220-4fc2-a28e-40b35c05e867 succeeded. Verification 8396252d-af7f-4b62-a1f2-c4231ba1b6d1 found 56,093 unique changed keys in each target, zero raw/clean lineage mismatches, one ledger row, and a deliberately absent historical fixture preserved.

- Immutable replay batch 2f6ebc6f-d902-42c8-bcc4-53c60414a20b succeeded with a fail-on-call source sentinel. Verification 7b432d40-86f5-4bee-b487-b263594e2ec4 found unchanged target/key counts and one replay ledger row.

- A real zero-row filtered response produced a complete v5 manifest. Empty publication batch 165d0648-00f1-4ac6-992a-f6b2d9b64e18 contained only the ledger insert; verification d41f04d3-cc6a-4d6f-b48f-a2f9465e7a14 found both targets unchanged.

- An invalid sixth statement was injected after the normal five statements in final batch fab60668-249d-47f2-89ff-6b64da4fa375. The batch failed. Verification ccd07e19-7bf9-4f18-8e8d-a0df7523ca82 found zero failed-run raw rows, clean rows, or ledger rows and all prior rows intact, demonstrating transactional rollback.

- Work-table query 88a542ec-e321-4578-9662-d33835131d1c and the isolated staging-prefix listing found no temporary objects remaining.

- The current bounded-window, inclusive-boundary, terminal-confirmation implementation was exercised against live TimeBack with a five-minute scope. Manifest 45edde8519d673a6d5d42cb3c98cbd795b8e006e5e01da07afed63d10342a321 retained 2 source pages containing 2,000 and 322 rows, accounted for overlap [0, 1], landed an identical independent confirmation of the short terminal response, and replayed 2,321 unique records. Publishing that exact manifest to Redshift dev succeeded in batch 3c1017ff-a291-4187-b3c0-5332a738f324; verification 700fd346-48e1-4094-ab04-93fe7edbcada found 2,321 unique raw and clean keys, zero lineage mismatches, the historical fixture preserved, and one ledger row.

The Redshift design was checked against the current AWS documentation for BatchExecuteStatement transaction behavior, Data API SQL NULL fields, DELETE ... USING, transactional COPY, and TRUNCATE commit behavior.

## Test plan

- [x] cd pipelines/runners/timeback-raw-sync && uv run pytest -q — 202 passed

- [x] uv run ruff check src tests scripts

- [x] uv run ruff format --check src tests scripts

- [x] python -m json.tool pipeline.json

- [x] npm test -- --runInBand test/real-pipeline-configs.test.ts — 515 passed after updating the intentional schedule-hold contract

- [x] git diff --check

- [x] previous exact-range and full implementation reviews passed before the inclusive-boundary repair

- [ ] fresh Mercy review of the current head

- [ ] GitHub repository CI

- [ ] CDK production synth/diff before any production release

- [ ] separately authorized targeted production validation after deployment

## Production sequencing

This PR targets main; merging it does not deploy the pipeline. Production promotion is deliberately not part of this PR and is not yet authorized.

Before a main to production release, review the complete release diff because the production workflow deploys all pipeline stacks when pipeline paths change.

When production promotion is authorized, the intended sequence is:

1. deploy the new task definition with the TimeBack schedule disabled;

2. verify the deployed image/task revision and confirm no older execution is active;

3. separately authorize and run only assessment_results on demand — this is a real production publication, not a dry run;

4. verify watermark arithmetic, manifest scope, key uniqueness, historical-row preservation, matching lineage, ledger evidence, and temporary-object cleanup; and

5. restore the repository schedule configuration only after that targeted validation succeeds, allowing a later complete scheduled run.

No production execution is performed by this PR.

## Scope

- No DDL or migration

- No dependencies, Dockerfile, CDK runtime/construct, IAM, database, or schema changes

- One existing CDK real-pipeline configuration test is updated to require the intentional schedule hold

- No activation of the remaining 13 incremental-eligible entities

- No changes to applications, test_assignments, fan-outs, or the activity_facts window

- No automatic full fallback or automatic reconciliation

## Rollback

Revert the deployment or disable the assessment_results modified-since policy while retaining the generic support code. A completed delta leaves the table complete because keys absent from the delta are preserved. Explicit full mode remains available as a deliberate operator action; it is not invoked automatically during rollback.

#1304 — Add mobile portfolio site breadcrumbs @YibinLongTrilogy  approved

## Summary

Add a compact mobile-only breadcrumb row to portfolio site detail pages so users can return to Dashboards or Portfolio and identify the current site without adding the detail tabs to the global mobile chrome.

### Screenshots

<img width="566" height="155" alt="Screenshot 2026-09-10 at 2 52 40 PM" src="https://github.com/user-attachments/assets/d7673f9e-75d9-4f2a-a693-02d2231f5a2c" />

### Changes

- chat/components/dashboards/portfolio/site-detail-page.tsx — Add MobileSiteBreadcrumb inside the portfolio site header with links to /dashboards, /dashboards?tab=portfolio, and the site’s Overview URL. Keep the row hidden on desktop, safe-area aware, touch-friendly, and truncating for long site names. Remove the trailing separator so the trail ends at the site name.

- chat/components/dashboards/portfolio/__tests__/site-detail-page-rhodes-tabs.test.tsx — Cover breadcrumb labels, hrefs, mobile-only classes, Overview URL behavior, and the absence of Overview / Work Plan / Documents from the breadcrumb trail.

### Design Decisions

- Keep the existing Overview / Work Plan / Documents tablist as the tab-switching control. The breadcrumb is navigation context only.

- Keep desktop breadcrumbs and the global MobileTopBar unchanged.

## Business value

Mobile users can understand where they are in the portfolio and navigate back to the dashboard or portfolio list without a misleading or overcrowded breadcrumb trail.

## Estimated manual effort

30–45 minutes for an engineer familiar with the portfolio dashboard.

## Test Plan

- [x] Focused portfolio site-detail browser tests passed: 13 tests.

- [x] Chat typecheck passed.

- [x] Biome validation and git diff --check passed.

- [ ] Manually verify the breadcrumb at narrow mobile widths with a long site name.

#1299 — feat(enrollments): add parallel aggregate-only SIS enrollment report @vvp-trilogy  approved

Closes #1296.

## What this delivers

A parallel, aggregate-only SIS enrollment report alongside the existing HubSpot-backed enrollment report. The incumbent report is behaviorally unchanged except for one additive SIS Based Report link at the bottom. The SIS report is selected within the existing Enrollments sub-route by source=sis; a missing or unknown source safely renders the current report. No nav item is added and the Enrollments nav stays active for both views.

## Layers

Contract (@bran/contracts/sis-enrollment) — runtime-free vocabulary shared by sync + convex + UI:

- Ten metric ids in the incumbent order/names; cohort + x_pipeline mapping; required-coverage cohort set.

- buildSisEnrollmentCounts, missingSisEnrollmentCoverage, sisFirstDayPartitionMatches.

- Metric→incumbent-column map so the UI reuses the incumbent labels and tooltip copy.

- defaultSisEnrollmentSchoolYear delegates to the shared defaultEnrollmentSchoolYear.

Sync

- sync/src/redshift/sis-enrollment.ts: reads only sandbox_education.mart_enrollment_dtl with COUNT(DISTINCT CASE WHEN has_fact THEN student_id END) at (program, session_school_year, cohort_id[, x_pipeline]); validates dense-grid coverage and the 1st-Day partition; throws on missing coverage / empty source.

- sync/src/analytics/sis-enrollment-refresh.ts: isolated refresh publishing counts only; any failure (unavailable source, missing coverage, publish rejection) skips publication and preserves the last known-good run. Wired into the orchestrator as its own domain.

Convex

- New sisEnrollmentRollups + single-pointer sisEnrollmentPublications tables (isolated from enrollmentSnapshots/admissionsPublishedRuns).

- publishSisEnrollmentRollups internal mutation: one atomic transaction — validate, insert the run's rows, advance the pointer (with source freshness), prune the prior run. A throw rolls back and preserves last known-good.

- Bearer-token /sync/analytics/sis-enrollment route.

- getSisEnrollmentData query gated by admissions.enrollments.read, returning aggregate counts only — no student-level records, no drill-down.

UI

- Isolated SIS view + matrix reusing the incumbent's spacing, sticky header/School column, totals row, zero-as-em-dash, column tooltips and On Campus emphasis. Metric cells are non-interactive (no student panel this ticket).

- Year selector (shared enrollment-year helper + persisted preference), Search schools... search filtering rows and the totals row, freshness chip.

- source=sis dispatch; SIS Based Report / Back to Enrollment Report links preserve/remove source while keeping other params.

## Tests

Exact-file coverage for aggregation, metric mappings, zero coverage, missing source, stale/failed publication (last known-good preserved), authorization (unauth + missing capability), route selection, and selected-year behavior. pnpm typecheck, pnpm biome check, boundary / convex-path / read-bounds / test-architecture checks all pass.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#3753 — KLAIR-3529: Require usable text from Anthropic responses @marcusdAIy  approved

## Summary

- require usable nonblank text from completed native Anthropic end_turn responses before treating text-producing calls as successful

- use one shared real-SDK text extractor for both validation and consumption

- preserve valid text-free tool_use responses only when an actual Anthropic ToolUseBlock is present

- align section generation, evaluation, quarter memory, Brainlift, attached-document summaries, refresh, Coach Claire, and buffered streaming

## Safety behavior

- empty content, thinking-only output, missing/invalid text, and whitespace-only text fail closed

- blank blocks are omitted; usable text blocks are joined in provider order

- fake duck-typed text/tool blocks do not pass the native response boundary

- section generation does not automatically replay incomplete or empty output

- Coach cannot persist or stream a blank response; tool-only turns retain their established safe placeholder/tool flow

- buffered streaming emits only normalized text from the validated final message

- validation errors remain content-free and never include thinking/provider bodies

## Validation

- expanded focused independent-review suite: 324 passed

- full Board Doc suite: 4,939 passed, 2 deselected

- Ruff check and format check passed

- targeted Pyright: 0 errors (one existing unrelated MessageParam warning)

- git diff --check passed

- independent review: APPROVE, no findings

## Tracking

- KLAIR-3529

- resolves the usable-text Mercy finding on production release PR #3749

- no deployment or add-on publication is performed by this PR

#3751 — KLAIR-3528: Preserve Brainlift summarization failures @marcusdAIy  approved

## Summary

- preserve Brainlift summarization failures instead of persisting a truncated raw fallback as ready

- raise a stable content-free typed failure for incomplete, refused, empty, timeout, and provider-error outcomes

- leave any existing Brainlift summary untouched while setting the background status to error

- retain the short-content no-model path and completed valid-summary path

## Safety

- no automatic provider replay

- no provider body or source content in logs/errors

- no partial or raw fallback is published after a model call

## Validation

- focused contract and background tests: 21 passed in independent review

- full Board Doc suite: 4,920 passed, 2 deselected

- Ruff check and format check passed

- targeted Pyright: 0 errors (5 existing unrelated warnings)

- git diff --check passed

- independent review: APPROVE, no findings

## Tracking

- KLAIR-3528

- fixes the remaining non-blocking Mercy finding on production release PR #3749

#1297 — Fix mobile Forecast school list scrolling @YibinLongTrilogy  approved

## Summary

Fix the mobile Admissions Forecast shell so the school-card list can scroll within constrained dashboard and narrow desktop layouts. Previously, outer dashboard wrappers could clip the list, leaving users unable to reach schools below the initial viewport.

### Changes

- chat/components/dashboards/admissions/forecast/mobile/forecast-mobile-base.tsx — Make the mobile Forecast content a flexing vertical scroll container with contained overscroll while keeping the fix scoped to the mobile school-card view.

### Design Decisions

- Keep scrolling local to the mobile Forecast content instead of changing the shared dashboard shell or all Forecast screens.

## Business value

Mobile users and users testing narrow desktop layouts can access every school in the Admissions Forecast view.

## Estimated manual effort

20–30 minutes for an engineer familiar with the dashboard layout.

## Test Plan

- [x] Pre-commit Biome validation passed.

- [x] Chat typecheck passed.

- [x] Focused mobile Forecast tests passed: 2 files, 5 tests.

- [x] git diff --check passed.

- [ ] Manually verify scrolling at 390×844 device emulation and at a regular desktop width.

#1809 — fix(repo): Skip ADHOC lists reactively in list_memberships fetch @heimdall-keval-factory[bot]  approvedAutomated PR

HubSpot list-membership syncs abort the entire portal run whenever a single list has a processing type HubSpot won't let us query — that list will keep failing on every retry, so the fix makes the sync skip just that one list's memberships (with the skip recorded) instead of failing the whole run.

Ticket: SURTR-1192

> Ready for review. Nothing ran the change, so it is unproven. A person still merges.

## For The Agent

_Everything below is detail for review. The summary above is the change._

Presented as ready — verification none, scope tier draft, fix_class code_fix, HEIMDALL_READY_PRS=true.

### What's broken

Run b6ca2012-6a28-49e5-afdd-ee1a58b2563f's list_memberships fetch for portal alpha, list 9552, hit hubspot_client.HubSpotRequestError: HubSpot GET /crm/lists/2026-03/9552/memberships returned HTTP 400; error={..."subCategory":"ListError.INVALID_PROCESSING_TYPE"...} (src/collector.py:969), which propagated up and made src/main.py:52's fail-closed guard abort the entire portal-alpha run three times in a row (05:48:55, 05:49:53, 05:51:07). This is not transient: HubSpot deterministically rejects the memberships endpoint for this list's processing type, so any retry hits the identical error. The plan-time ADHOC exclusion added in PR #1331 (src/fanout_handler.py:95, 3316-3317) filters lists whose list_definitions record reports processingType=='ADHOC', but list 9552 still reached the per-list memberships request — meaning its processingType was either absent from that snapshot (the documented fail-loud fallback at tests/test_fanout_handler.py:1192) or changed between the list_definitions snapshot and the list_memberships fetch inside this ~2.5 hour run.

Root cause. The list_memberships collector has no defense at request time against a HubSpot list that turns out to be unsupported (ADHOC processing type) for the /memberships endpoint; only a plan-time filter based on the list_definitions snapshot exists (PR #1331), and that filter cannot catch a list whose processingType is absent from the snapshot or changes after the snapshot is taken. Because collector.collect() issues one HubSpotRequestError-raising request per list_id inside a single batched call for the whole 'lists' auxiliary group, one unsupported list aborts memberships collection for every other list in the group, and main.py's fail-closed status check then fails the whole portal run rather than just that one list's data.

### What this PR changes

Give HubSpotRequestError structured fields (status_code, category, subCategory) when raised in hubspot_client.py's request() (around line 211), then in collector.py's per-plan request handling (around lines 954-972) catch the specific case status_code==400 and subCategory=='ListError.INVALID_PROCESSING_TYPE' when resource_name=='list_memberships', record that list_id as an explicit, non-silent skip in the resource's evidence/manifest (e.g. a 'skipped_unsupported_lists' entry alongside dependency_ids in fanout_handler.py), and continue collecting the remaining lists instead of letting the exception abort the whole batch. Every other error path (5xx, other 4xx, unexpected validation errors) should keep failing loud exactly as today, so this only relaxes handling for the one already-understood, deterministic HubSpot behavior HubSpot itself confirms via the error body.

Why this fixes it. This closes the residual gap left by the plan-time-only ADHOC exclusion (fanout_handler.py:95, PR #1331): that filter only works when the list_definitions snapshot itself reports processingType=='ADHOC', but a list missing that field in the snapshot or converted to ADHOC afterward still reaches the memberships request, and because collector.collect() batches all list_ids from a group into one call, its 400 aborts every other list's memberships too — failing the entire multi-hour portal run via main.py's fail-closed guard, as it did three times in a row for run b6ca2012. Handling this specific, structurally-confirmed HubSpot error reactively per list_id and recording it explicitly (never swallowing it) preserves the no-silent-data-failure guarantee while avoiding a full-run failure over a single un-crawlable list.

#### Files changed

 .../runners/hubspot-raw-sync/src/collector.py      | 47 +++++++++++

.../runners/hubspot-raw-sync/src/fanout_handler.py | 4 +

.../runners/hubspot-raw-sync/src/hubspot_client.py | 34 +++++++-

.../runners/hubspot-raw-sync/src/orchestration.py | 27 +++++++

.../hubspot-raw-sync/tests/test_collector.py | 92 ++++++++++++++++++++++

.../hubspot-raw-sync/tests/test_fanout_handler.py | 75 ++++++++++++++++++

.../hubspot-raw-sync/tests/test_hubspot_client.py | 21 +++++

7 files changed, 298 insertions(+), 2 deletions(-)

### Verification

### pytest — no test suite

### verify: ruff check — exit 0

[notice] A new release of pip is available: 25.3 -> 26.2.1

[notice] To update, run: pip install --upgrade pip

All checks passed!

### verify: ruff format --check — exit 0

1806 files already formatted

### verify: pytest (pipeline lambdas) — exit 2

[notice] A new release of pip is available: 25.3 -> 26.2.1

[notice] To update, run: pip install --upgrade pip

Downloading cpython-3.11.16-linux-x86_64-gnu (download) (29.4MiB)

Downloaded cpython-3.11.16-linux-x86_64-gnu (download)

Installed Python 3.11.16 in 382ms

+ cpython-3.11.16-linux-x86_64-gnu (python3.11)

error: Failed to spawn: pytest

Caused by: No such file or directory (os error 2)

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | repo |

| Failing run | issue |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | linear-SURTR-1192 |

| Verify | none |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev label to take the PR over and stop it entirely.

#3750 — KLAIR-3526: Fail closed on incomplete Anthropic output @marcusdAIy  approved

## Summary

- add one content-free completion validator for native Anthropic responses

- reject incomplete, refused, missing, unknown, and invalid tool_use responses before extraction or persistence

- apply the guard to section generation, quarter memory, Brainlift summaries, attached-document summaries, number refresh, Coach Claire chat/tool proposals, streaming, and the developer eval path

- buffer streamed text until Anthropic supplies and passes the final stop reason

- preserve existing section content or return a safe failure without replaying incomplete output

## Why

Mercy found that PR #3749 could accept plausible partial section text when adaptive thinking exhausted the shared output cap. The caller-class audit found the same fail-open boundary in other user-visible native Anthropic paths.

## Safety properties

- no provider text in validator errors or failure logs

- incomplete/refused output is not published, persisted, proposed, or streamed

- incomplete section responses are not automatically retried

- tool_use is accepted only when an actual tool block is present

- valid end_turn and tool-use flows remain supported

## Validation

- focused completion/generation/chat/streaming suite: 69 passed

- full Board Doc suite: 4908 passed, 2 deselected

- Ruff format/check: passed

- Pyright on validator and section generator: 0 errors, 0 warnings

- git diff --check: passed

- independent review: approved after two High findings and one test gap were corrected

## Release

Blocks production release PR #3749. No production mutation is included in this PR.

Linear: KLAIR-3526

#1803 — chore(surtr): remove the Heimdall dashboard page and its tRPC read layer (1/2) @kevalshahtrilogy  approved

Part 1 of 2 · stacked: part 2 is #1804

Removes the /heimdall dashboard from the Surtr app. It was superseded by the

local Heimdall control tower, which reads the same DynamoDB table directly and

never calls this app's API.

Split out of #1802 at Mercy's request — that PR was 651 KB, over the 600 KB

review cap. This half is 120 KB.

## What comes out

- app/(app)/heimdall/ — the page and its charts

- The sidebar entry and HeimdallIcon (the Agents section now holds Mercy alone)

- The tRPC read procedures heimdallStats, heimdallBoard, listHeimdallRunsPage,

plus the InvalidHeimdallCursorError mapping and the imports they held

- The page's own tests and the tRPC endpoint test

## What deliberately stays

src/heimdall/board.ts and the store/types read helpers are unreachable after

this PR but are left in place so this one stays reviewable. **Part 2 removes

them**, along with the IAM grants that went with them.

The telemetry ingest path is untouched here and in part 2 — heimdall's

workflow still POSTs one record per run to /internal/heimdall/telemetry, and

the table keeps filling for the control tower to read.

## Blast radius on the control tower: none

Verified against heimdall-control-tower/src, not assumed:

- It reads DynamoDB directly with local AWS credentials and never calls Surtr's

API — its own header comment states this is deliberate.

- Its only outbound targets are api.github.com, api.linear.app, 127.0.0.1,

and DynamoDB. No Surtr host appears anywhere in its source.

- Its Floor page is derived from live GitHub state, not telemetry.

## Verification

| Check | Result |

| --- | --- |

| pnpm lint (biome) | pass |

| pnpm build (tsc) | pass |

| pnpm test:unit | 1147 passed, 53 files |

| next build | pass — /heimdall no longer in the route table |

## Business Value

Removes a dashboard nobody uses any more from the app's navigation and API

surface, in the repo whose stated worst failure mode is a silent data failure.

Three tRPC procedures stop being maintained, re-reviewed on every touching PR,

and stop being a thing a future reader has to understand before changing

anything nearby. No capability is lost — the control tower already serves this

view, and the telemetry record of every heimdall run keeps landing.

## Manual Effort Estimate

1 hour focused time for the whole change (parts 1 and 2 together) — by hand,

no AI, as the only thing being worked on. Estimate set by Keval.

---

⚠️ No Linear ticket yet — this session has no Linear tool, so I could not

create one or check whether a teammate's agent already claimed this work. Per

"one ticket, one PR" this needs a ticket before merge.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#3748 — fix(board-doc): register AI Renewals central function @marcusdAIy  approved

## Summary

- register AI Renewals as an active Budget Bot BusinessUnit and Central Function

- add the required Budget Goal MIPer adapter member used by Board Doc Brainlift collection

- map both Ravi-reviewed Brandon identities as AI Renewals owners

- update the derived provisioning contract from 21 entities (9 BU / 12 CF) to 22 entities (9 BU / 13 CF)

- add focused coverage for session serialization, CF classification, Brainlift adapter conversion, exact Sheets/Dynamo key lookup, recipient drafting, access scoping, and roster/manifest/CLI behavior

## Why

Ravi added AI Renewals to the reviewed Q4 campaign and supplied a dedicated Q4 workbook. The income-statement registry already supported it, but the canonical Budget Bot enum did not. Without this patch, Board Doc request validation and the MIPer Brainlift adapter reject AI Renewals instead of creating an independently bound session.

This does not alias AI Renewals to New Renewals, AI Engineering, or the excluded CEO-office Core workbook.

## Validation

- full Board Doc suite: 4882 passed, 2 deselected

- focused changed-path suite: 292 passed, 1 deselected

- independent review: approved, no blocking findings; reviewer rerun 301 passed, 1 deselected

- Ruff: passed on every changed Python file

- git diff --check: passed

- read-only production metadata: dedicated AI Renewals workbook is accessible to the production Sheets service account and has a readable P&Ls tab with Q4 plan markers

## Safety

No production Docs, Sheets, sessions, DynamoDB rows, permissions, or emails were changed while implementing or validating this patch.

Linear: KLAIR-3525

#286 — chore(mercy): pin the harness ref, so @v1 actually pins mercy @kevalshahtrilogy  no labels

## What

One line in the mercy caller:

    with:

config_path: .mercy.yml

+ harness_ref: v1

## Why — @v1 has only ever pinned half of mercy

mercy pulls its reusable workflow from the uses: ref, but its harness — the prompts, decide_review, the review stage — from harness_ref, which defaults to main when unset. This repo has never set it.

So this repo has been running main's harness against v1's workflow. Telemetry confirms it: mercy_version records the harness commit each run actually used, and this repo's most recent runs were on a main commit that was not an ancestor of the old v1 tag.

That is the exact drift mercy's own caller comment warns about:

> *"Leaving it unset therefore does not pin anything — a caller on @v1 silently runs main's harness against v1's workflow, and the two drift apart with no signal until something the workflow calls is missing from the harness."*

## What this fixes

- The tag starts meaning something. Today a push to mercy's main reaches this repo on its next review, tag or no tag. With this line, only a deliberate v1 move does.

- The canary ring becomes real. Surtr and Klair ride @main on purpose so changes are exercised there first. That only held for workflow YAML — harness changes skipped the canary entirely and landed here immediately.

- Workflow and harness stop drifting. A new workflow step calling a harness function that does not exist at the pinned ref fails loudly at the uses: mismatch instead of silently at runtime.

## Risk

Low, and it reduces existing risk. v1 currently points at mercy d74dfa2, which is the same code this repo has effectively been running (that was the bug). The change is what it will run *next week* when someone pushes to mercy's main: nothing, until the tag is moved deliberately.

## Business Value

Makes mercy's release process real for this repo instead of nominal. Right now every mercy change lands here the moment it merges, with no canary and no gate — a bad harness change would hit all five repos at once. This restores the intended two-ring rollout (Surtr/Klair on main, everyone else on a moved tag) at the cost of one line.

## Manual Effort Estimate

~15 minutes (proposed — Keval to confirm). One-line change; the work was diagnosing that @v1 was not pinning the harness.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1807 — fix(timeback): reduce activity facts window to one month @caina-barbosa  approved

## Summary

- reduce only timeback-raw-sync activity_facts from a six-calendar-month to a one-calendar-month rolling request window

- retain calendar-month subtraction and month-end clamping

- preserve existing scoped publication, source-limit, replay, and out-of-window validation behavior

- explicitly document that the other 20 TimeBack entities are unchanged

## Test plan

- [x] cd pipelines/runners/timeback-raw-sync && uv run pytest -q (157 passed)

- [x] uv run ruff check .

- [x] uv run ruff format --check .

- [x] git diff --check

- [x] Independent Luna Max code review: PASS, no findings

## Scope

- No DDL or migration

- No bulk watermark extraction or timeout changes

Linear: [SURTR-1165](https://linear.app/builder-team/issue/SURTR-1165/reduce-only-timeback-activity-facts-to-a-one-calendar-month-rolling)

#1806 — chore(hc-forecast): Record the live Core procedure definition @caina-barbosa  approved

## Summary

This PR records the core_budgets.sp_refresh_hc_data_consolidated(VARCHAR) definition that is already running successfully in production.

It is not a proposal for changing how the pipeline should behave. The authorised production correction has already been applied and verified through a complete HC forecast refresh and its automatically triggered current-cost refresh.

The purpose of this PR is to bring source control back into alignment with production. It ensures that any future recreation of the procedure uses the definition that production already relies on.

## Why this source-control record is needed

During the workforce identifier release, the rename migration recreated the Core procedure from an older repository definition.

The live procedure before the rename included corrections from the original SURTR-374 production rollout:

- it accepted the input in either pinned or core_submitting, because the runtime records core_submitting before invoking the writer;

- it allowed identical mapping rows for a case-insensitive key while continuing to reject conflicting mapping values.

Those corrections had been applied to production but had not reached main. Recreating the older definition caused the first 2 post-rename verification runs to stop before Core publication.

The complete lifecycle and mapping behaviour has now been restored using the renamed workforce identifiers. Identical mapping rows are resolved to one row per case-insensitive key before all 3 enrichment joins, which makes the joins consistent with the procedure's existing acceptance rule.

## What this PR records

The PR contains:

- one operator-run forward migration at pipelines/runners/hc-forecast-refresh/ddl/2026-09-09_sp_refresh_hc_data_consolidated_restore_live_mapping_guard.sql;

- one SQL contract test proving that every enrichment join uses the resolved one-row-per-key mapping relation.

The SQL file contains the procedure definition now running in production.

A new forward migration records the live state without modifying historical dated DDL. Keeping this definition in the repository prevents a future recreation from reinstating the older procedure body.

This completes the source-control record after #1801 captured only the core_submitting part of the live behaviour.

## Scope boundaries

This effort is limited to recording the procedure behaviour already deployed in production.

It is not meant to introduce a new workforce calculation, lifecycle design, mapping policy or pipeline contract. Resolving identical accepted mapping rows before enrichment makes execution match the procedure's existing acceptance rule; conflicting mappings continue to fail closed.

It is not meant to change:

- workforce formulas or grouping rules;

- source selection or snapshot contents;

- Core or mart table shapes and data grain;

- lifecycle transitions or recovery policy;

- handler inputs or response keys;

- procedure ownership, SECURITY DEFINER or effective grants;

- schedules, triggers, IAM, CDK or pipeline.json;

- unrelated pipeline behaviour or hardening.

There are no Python runtime, infrastructure or schedule changes in this PR.

## Live production verification

The recorded definition is already active in finance_dw.

The accepted production snapshot contained one identical duplicate mapping row, no conflicting mapping keys and no source rows matching that duplicate key. Resolving it therefore produced no current output delta.

Verification after applying the recorded definition:

- HC forecast refresh: succeeded;

- HC ingestion run state: accepted;

- Core statement status: FINISHED;

- Aerie statement status: FINISHED;

- Core rows: 81,876, unchanged from the preceding successful run;

- HC teamroom mart rows: 6,134, unchanged from the preceding successful run;

- automatically triggered current-cost refresh: succeeded;

- current-cost publication state: published;

- current-cost evidence rows: 50,882;

- included rows: 2,640;

- unavailable rows: 48,242;

- procedure owner, security setting and grants preserved.

## Repository validation

- focused regression test: failed before the 3 joins were corrected and passed afterward;

- hc-forecast-refresh test suite: 98 passed;

- production corrective transaction: 6 of 6 statements finished and committed;

- production HC refresh: succeeded;

- production downstream current-cost refresh: succeeded.

#1805 — fix(repo): Split OpenAI BU fetch across two scheduled Lambda runs @heimdall-keval-factory[bot]  approvedAutomated PR

The daily OpenAI usage pull hits OpenAI's rate limit partway through and skips cost catch-up for the last few business units, leaving their spend recorded as $0 for the day. This adds a second, smaller scheduled run so each invocation has fewer BUs to fetch and finishes within its time budget.

Ticket: SURTR-1203

> Ready for review. Nothing ran the change, so it is unproven. A person still merges.

## For The Agent

_Everything below is detail for review. The summary above is the change._

Presented as ready — verification none, scope tier draft, fix_class code_fix, HEIMDALL_READY_PRS=true.

### What's broken

Run bca341de-eccf-4eba-832a-1784592a2f6d hit sustained OpenAI Admin API 429s fetching BU X-Trilogy-Central-COO-Arthur's line-item costs (four backoff attempts spanning 07:12:10Z-07:13:32Z), and the resulting time loss is what produced the log line 'Cost catch-up skipped for BU TelcoDR: only 89s of invocation budget left, 210s needed' for TelcoDR, Trilogy-Aurea-E-Commerce, Trilogy-Central-Support, Trilogy-CNU-Innovations, and Trilogy-Crossover-DEV, zeroing their 2026-09-08 billed_cost_dollars. keval.shah already decided (2026-09-10T09:43:51Z comment) to split the BU set across two scheduled Lambda invocations rather than pursue an OpenAI quota increase, after three prior PRs (#1301, #1395, #1595) already tuned the throttle/backoff/catch-up mechanism without fixing the underlying single-invocation shape.

Root cause. openai-usage-pipeline runs as a single daily Lambda invocation (one schedule block in pipeline.json) that processes every BU sequentially against OpenAI's shared 30-req/min Admin API rate limit; when several BUs draw 429s back-to-back, the exponential backoff (openai_client.py, 2s/4s/8s/65s observed) consumes enough of the 900s Lambda ceiling that handler.py's own invocation-budget check (handler.py:808-812) skips the cost catch-up pass for later BUs outright, leaving them at $0 billed_cost_dollars for the day. The single-invocation-per-day design no longer fits this org's BU/request volume inside both OpenAI's rate limit and the Lambda timeout, matching keval.shah's prior diagnosis.

### What this PR changes

The ticket's decision (split into two scheduled invocations via the existing bus_to_process param) is sound, but the mechanism it names isn't wired up yet for this pipeline: additional_schedules is defined in the shared schema (pipelines/cdk/lib/schema/pipeline-config.ts:17-27) and consumed only by the ECS pipeline construct (pipelines/cdk/lib/constructs/ecs-pipeline.ts:451-475 - see hubspot-raw-sync's crm/marketing lanes for working precedent). The Lambda pipeline construct this pipeline actually uses (pipelines/cdk/lib/constructs/pipeline.ts:417-440) only ever builds one EventBridge rule from config.schedule and silently ignores additional_schedules. Adding it to openai-usage-pipeline/pipeline.json alone would pass schema validation and deploy cleanly while creating no second rule - a silent no-op, exactly the failure mode this repo's conventions warn about. The real fix: mirror the ECS additional_schedules loop into pipeline.ts, add a second named schedule to this pipeline's pipeline.json, and - because the BU roster lives only in Secrets Manager (secrets.py's get_openai_bu_keys(), never visible to the repo or to CDK) - shard by a deterministic rule (e.g. hash of bu_name mod 2, passed as a small params field handler.py can read alongside bus_to_process at handler.py:125-135) rather than two hardcoded bus_to_process name lists, so a BU added to Secrets Manager later still lands in one of the two runs instead of silently running in neither.

Why this fixes it. This is a real, confidently-fixable code gap - the Lambda CDK construct is missing a capability its ECS sibling already has - not a config tweak, so the diff spans pipeline.ts (add the additional_schedules loop), openai-usage-pipeline/pipeline.json (add the second schedule), a small handler.py change to shard deterministically instead of relying on static BU name lists that would silently drop any BU added to Secrets Manager later, and a CDK unit test in pipeline.test.ts mirroring ecs-pipeline.test.ts's existing additional_schedules coverage. .heimdall.yml's Tier C note (2026-08-06 decision) explicitly permits heimdall to propose CDK-app changes for this repo, and the result won't auto-merge regardless since it touches shared code, which is the correct outcome given the blast radius.

#### Files changed

 pipelines/cdk/lib/constructs/pipeline.ts           | 27 +++++++++++

pipelines/cdk/test/constructs/pipeline.test.ts | 47 +++++++++++++++++++

.../runners/openai-usage-pipeline/pipeline.json | 11 ++++-

.../runners/openai-usage-pipeline/src/handler.py | 34 ++++++++++++++

.../openai-usage-pipeline/tests/test_handler.py | 52 +++++++++++++++++++++-

5 files changed, 169 insertions(+), 2 deletions(-)

### Verification

### pytest — no test suite

### verify: ruff check — exit 0

[notice] A new release of pip is available: 25.3 -> 26.2.1

[notice] To update, run: pip install --upgrade pip

All checks passed!

### verify: ruff format --check — exit 0

1806 files already formatted

### verify: pytest (pipeline lambdas) — exit 2

[notice] A new release of pip is available: 25.3 -> 26.2.1

[notice] To update, run: pip install --upgrade pip

Downloading cpython-3.11.16-linux-x86_64-gnu (download) (29.4MiB)

Downloaded cpython-3.11.16-linux-x86_64-gnu (download)

Installed Python 3.11.16 in 544ms

+ cpython-3.11.16-linux-x86_64-gnu (python3.11)

error: Failed to spawn: pytest

Caused by: No such file or directory (os error 2)

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | repo |

| Failing run | issue |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | linear-SURTR-1203 |

| Verify | none |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev label to take the PR over and stop it entirely.

#1292 — feat(enrollment): drop HubSpot identities + overlay, publish SIS identity and deposit signal (#1290) @vvp-trilogy  approved

Closes #1290.

Reduces mart_enrollment_dtl to what SIS actually knows. This is a dbt-only change — the enrollment consumer (#1216) is still open, so nothing in production reads this mart yet.

## What changed

Identity — SIS only

- Removed HubSpot contact_id and deal_id from mart_enrollment_dtl, int_enrollment_cohort, int_enrollment, and int_student, and dropped the now-orphaned stg_sis_student_external_id / stg_sis_enrollment_external_id joins plus their two hubspot_* scoped-uniqueness tests.

- Renamed student_keystudent_id, unaliased from sis_enrollments.student_id through the intermediates to the mart. Removed the unused student_number projection entirely (grep -r student_number dbt/ returns nothing).

Deposit — drop the EduCRM overlay, publish the SIS signal

- Deleted int_deposit_overlay.sql, stg_educrm_enrollment_deposit.sql, and the two overlay singular tests. No dbt model reads EduCRM enrollment data any more; the educrm.enrollment_dtl source declaration is retained (read only by the parity_sis_vs_hubspot_enrollment_2026 analysis, per Out-of-Scope #9).

- deposit_paid_date is now published as CAST(NULL AS TIMESTAMP) on both arms — reserved, documented, appearing the day SIS grows a native paid-at.

- Added finalsite_deposit_state / finalsite_deposit_source to stg_sis_enrollment and carried them through int_enrollment and int_enrollment_cohort. The mart publishes:

- deposit_state — three-valued paid / unpaid / NULL (NULL = unknown, not "No"; every pre-2026 row is NULL).

- deposit_signal — SIS's own evidence vocabulary (AMOUNT, CHECKLIST, ADVANCE_DEPOSIT_CHECKLIST, ADVANCE_DEPOSIT_ATTRIBUTE), carried verbatim and not relabelled into the Admissions Pipeline macro's vocabulary; NULL whenever the deposit is not paid.

- New singular test assert_sis_enrollment_deposit_signal_requires_paid pins deposit_signal IS NULL whenever deposit_state is not paid. Added accepted_values on deposit_state (error) and deposit_signal (warn).

Docs — model yml, source yml, staging headers, and the rule-14 dual-identity paragraph updated to describe a single SIS identity; the removed EduCRM overlay docs deleted.

## Verification

- dbt build --select +mart_enrollment_dtl+ (isolated pr1290_* relations): clean (PASS, 0 errors; the only WARN is the pre-existing assert_sis_campus_unresolved_hubspot_program, unrelated to this change).

- Parity (projection/rename only): total rows (8,606), per-year distribution, and 2026 per-cohort fact counts are byte-identical to the pre-change build. Column count stays 34.

- Deposit cross-check: 2026 distinct-enrollment deposit_state totals — paid 1,576 / unpaid 32 / NULL 45 (= 1,653) — match sis_enrollments directly. deposit_paid_date is NULL on all 8,606 rows.

- Built the independent mart_admissions_pipeline_dtl in isolation and ran assert_admissions_pipeline_reconciles → PASS (deleting the EduCRM deposit staging model did not break the admissions pipeline).

- pnpm typecheck passes; pnpm biome check is a no-op (dbt is ignored).

## Notes on stale ticket references

The issue named assert_sis_campus_program_map_hubspot_resolves.sql and docs/enrollment-cohort-parity.md; neither exists in the current repo (superseded by #1286's int_school_identity work / never created). The only consumers of stg_educrm_enrollment_deposit were the overlay and its two tests (all deleted), so no test needed repointing.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#182 — chore(mercy): pin the harness ref, so @v1 actually pins mercy @kevalshahtrilogy  approvedmercy-allow-critical

## What

One line in the mercy caller:

    with:

config_path: .mercy.yml

+ harness_ref: v1

## Why — @v1 has only ever pinned half of mercy

mercy pulls its reusable workflow from the uses: ref, but its harness — the prompts, decide_review, the review stage — from harness_ref, which defaults to main when unset. This repo has never set it.

So this repo has been running main's harness against v1's workflow. Telemetry confirms it: mercy_version records the harness commit each run actually used, and this repo's most recent runs were on a main commit that was not an ancestor of the old v1 tag.

That is the exact drift mercy's own caller comment warns about:

> *"Leaving it unset therefore does not pin anything — a caller on @v1 silently runs main's harness against v1's workflow, and the two drift apart with no signal until something the workflow calls is missing from the harness."*

## What this fixes

- The tag starts meaning something. Today a push to mercy's main reaches this repo on its next review, tag or no tag. With this line, only a deliberate v1 move does.

- The canary ring becomes real. Surtr and Klair ride @main on purpose so changes are exercised there first. That only held for workflow YAML — harness changes skipped the canary entirely and landed here immediately.

- Workflow and harness stop drifting. A new workflow step calling a harness function that does not exist at the pinned ref fails loudly at the uses: mismatch instead of silently at runtime.

## Risk

Low, and it reduces existing risk. v1 currently points at mercy d74dfa2, which is the same code this repo has effectively been running (that was the bug). The change is what it will run *next week* when someone pushes to mercy's main: nothing, until the tag is moved deliberately.

## Business Value

Makes mercy's release process real for this repo instead of nominal. Right now every mercy change lands here the moment it merges, with no canary and no gate — a bad harness change would hit all five repos at once. This restores the intended two-ring rollout (Surtr/Klair on main, everyone else on a moved tag) at the cost of one line.

## Manual Effort Estimate

~15 minutes (proposed — Keval to confirm). One-line change; the work was diagnosing that @v1 was not pinning the harness.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1291 — chore(mercy): pin the harness ref, so @v1 actually pins mercy @kevalshahtrilogy  approvedmercy-allow-critical

## What

One line in the mercy caller:

    with:

config_path: .mercy.yml

+ harness_ref: v1

## Why — @v1 has only ever pinned half of mercy

mercy pulls its reusable workflow from the uses: ref, but its harness — the prompts, decide_review, the review stage — from harness_ref, which defaults to main when unset. This repo has never set it.

So this repo has been running main's harness against v1's workflow. Telemetry confirms it: mercy_version records the harness commit each run actually used, and this repo's most recent runs were on a main commit that was not an ancestor of the old v1 tag.

That is the exact drift mercy's own caller comment warns about:

> *"Leaving it unset therefore does not pin anything — a caller on @v1 silently runs main's harness against v1's workflow, and the two drift apart with no signal until something the workflow calls is missing from the harness."*

## What this fixes

- The tag starts meaning something. Today a push to mercy's main reaches this repo on its next review, tag or no tag. With this line, only a deliberate v1 move does.

- The canary ring becomes real. Surtr and Klair ride @main on purpose so changes are exercised there first. That only held for workflow YAML — harness changes skipped the canary entirely and landed here immediately.

- Workflow and harness stop drifting. A new workflow step calling a harness function that does not exist at the pinned ref fails loudly at the uses: mismatch instead of silently at runtime.

## Risk

Low, and it reduces existing risk. v1 currently points at mercy d74dfa2, which is the same code this repo has effectively been running (that was the bug). The change is what it will run *next week* when someone pushes to mercy's main: nothing, until the tag is moved deliberately.

## Business Value

Makes mercy's release process real for this repo instead of nominal. Right now every mercy change lands here the moment it merges, with no canary and no gate — a bad harness change would hit all five repos at once. This restores the intended two-ring rollout (Surtr/Klair on main, everyone else on a moved tag) at the cost of one line.

## Manual Effort Estimate

~15 minutes (proposed — Keval to confirm). One-line change; the work was diagnosing that @v1 was not pinning the harness.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#126 — feat(review): rebuild the second-opinion prompt around how the reviewer actually fails @kevalshahtrilogy  no labels

## Measured — 112 PRs, one blocking review each, all at round 4+, five repos

| | old prompt | new prompt |

|---|---|---|

| PRs cleared | 45 (40%) | 40 (36%) |

| Findings deferred | 101 (60%) | 90 (54%) |

| Stricter | — | 31 |

| Looser | — | 20 |

Stricter, and stricter in the right places. Of the 31 findings it newly keeps blocking, the samples are money and data bugs the old prompt was shipping:

> *"discards _cost_malformed, so malformed Cost cells are silently reported as $0.00 and understate the spend-audit totals"*

> *"maps persistence failure to terminal FAILED without freeing the dedupe row, making a transient outage permanent"*

> *"retains a school with a blank display_name in the production refresh"*

19 of the 31 are money / data-integrity / observability, against a 52% base rate.

## What changed, and why each

Restructured, not appended. An earlier attempt that appended a single rule made the prompt *looser* overall. This keeps the same length with more structure.

Blast radius is question one. Five of the eleven confirmed-wrong blocks in the hand audit were on internal experimental tools, dev launchers, or routes behind a privileged flag — held to production standards.

"Removed is not fixed." The one case my own hand audit got wrong: version machinery was deleted, but the !existing branch still permitted the bad publish.

Evidence rules for the two dominant buckets.

- *"Nothing produces this input"* — 34% of deferrals, the riskiest — must point at what rules it out: a schema, an upstream validator, the sole writer. Absence of evidence is not evidence.

- *"The code already handles it"* — 19% — must quote the line.

- Author rebuttals count only with a file:line, a test, or a contract.

Per-repo "wrong result", one line each. Klair money and Surtr warehouse rows named rather than inferred.

Uncertainty made concrete. A reason needing *could / may / might / probably* is a prod_breaking.

## The trade-off to know about

| repo | old | new |

|---|---|---|

| Aerie | 20/44 | 22/44 |

| Klair | 13/24 | 13/24 |

| trilogy-drones | 4/12 | 4/12 |

| Sindri | 1/6 | 1/6 |

| Surtr | 7/26 | 0/26 |

On Surtr the gate now defers nothing. The per-repo line describes nearly every Surtr finding, and combined with "the reviewer is reliably right about silent data corruption" it closes the door. That is the correct default for a warehouse repo — data correctness *is* the product — but it means zero relief there. Recorded so it can be narrowed with two weeks of live evidence rather than a guess.

## Honest caveats

- 20 findings went looser for reasons the prompt does not explain. That is the sampling variance documented on #125; one run per prompt cannot separate it from effect. The stricter side is supported by the samples. The looser side is not claimed.

- Re-raise rate ~38% on both arms — mostly measurement artifact per the earlier hand audit, not read as an error rate.

## Testing

Prompt-only change; 397 harness tests pass unchanged. A prompt change is a model-behaviour change and is not unit-testable — the A/B above is the test.

## Business Value

This is the gate that decides whether a real finding must block a merge or can ship as a follow-up. Getting it wrong in one direction ships bugs; in the other it recreates the deathloop the team is angry about. The rewrite makes it measurably safer than the version currently live on Surtr and Klair — it catches money and data-loss bugs the old prompt was waving through — while still clearing roughly a third of PRs already stuck at round 4+. It also unblocks moving v1, which carries the deathloop bug fixes to Aerie, trilogy-drones and Sindri, who are still running the pre-fix reviewer.

## Manual Effort Estimate

~2 hours (proposed — Keval to confirm). The prompt itself is an hour; the rest was profiling how mercy blocks per repo and running the 112-PR A/B to make the change evidence-led rather than intuition-led.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#9 — feat(console): make revamp operator-ready @sanketghia  no labels

## Summary

- Make the revamp the operator-first destination while preserving the current Console fallback.

- Add repository-profile-aware intake, health/conflict visibility, bounded 100-batch collection, Resume, remediation, Mercy review, all-event timeline access, and stronger state semantics.

- Improve color hierarchy, action emphasis, identifier presentation, and objective/acceptance formatting.

## Verification

- pnpm test (80 files, 1,282 tests)

- pnpm test:browser (10 tests)

- pnpm build

- pnpm lint

- pnpm format:check

- git diff --check

## Migration

- Current Console remains at #/ and #/legacy.

- Revamp is available at #/revamp.

- No pagination or workflow-state changes were introduced.

#124 — fix(review): actually invoke the second opinion — it shipped inert @kevalshahtrilogy  no labels

## The bug

#123 shipped dead code. It added the module, the blocking predicate, the ledger recording, carry-forward preservation and eight tests — and never called any of it. Nothing in the pipeline set deferred. Every consumer would have taken the plumbing and seen zero behaviour change.

Caught while checking what moving the v1 tag would actually ship.

## The fix

decide_review gains --repo / --pr-number, and between extracting findings and deciding the event it asks the second opinion about every finding that *would* gate and is eligible.

No repo needs a new secret. It reuses AGENT_OPENAI_API_KEY — already declared as a workflow input, already exported as OPENAI_API_KEY for the review step, and already passed by all five consumer callers (Surtr, Klair, Aerie, trilogy-drones, Sindri), since they all run luna.

## Verified end to end

Against a real PR, with a deliberately fabricated finding:

no key   →                                   event=REQUEST_CHANGES

with key → [gate] eligible=1 deferred=1 event=APPROVE

ledger deferred : true

ledger reason : "The cited file only contains x=1 and y=2; no retry loop

or attempts budget exists to exhibit the claimed behavior."

body mentions a second reviewer? False

It caught that the finding was fabricated, deferred it, recorded why, and left no trace in the review body.

## Fail-safe paths, all exercised

No key, no PR context, an exception mid-call, an unparseable answer — each leaves every finding blocking and logs a line. A review that cannot get a second opinion is still a valid review.

## The contract test earned its place

It caught a real bug in this change before it shipped: ${PR_NUMBER} was read by the decide step without being declared in its own env:. Under set -u that is the HEAD_SHA regression that once failed every review with an error indistinguishable from a real finding.

397 tests pass.

## Business Value

Without this, #123 is inert and the measured 30% clearance on stuck PRs is worth nothing. This is the commit that makes it real.

## Manual Effort Estimate

~1 hour (proposed — Keval to confirm).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#123 — feat(review): a second opinion on whether a finding must block the merge @kevalshahtrilogy  changes requested

## What

The reviewer decides whether a finding is real. Nothing decides whether a real finding must be fixed before merge or can land as a follow-up — and that second question governs how many rounds a PR takes.

This adds it. For each finding that would gate, a pass over the PR description, the full diff, the whole of every file the findings cite, and the reviews and comments posted before this round. One question per finding: *does merging this unfixed break production?* Two possible answers: keep blocking, or ship as a follow-up.

## Measured

Replayed 63 PRs whose review was already at round 4 or later — five repos, six weeks:

PRs blocked      63

clear the gate 19 (30%)

findings judged 126 (27 more never eligible)

deferred 63 (50%)

What it defers is dominated by guards against input the system does not produce (34%) and findings the code already handles (19%).

## Then read by hand

Every contested call — all 17 where the reviewer raised the same file:line again on a later round:

| | |

|---|---|

| Reviewer stuck or wrong | 11 |

| Genuinely arguable | 3 |

| Unjudgeable (anchor problem, below) | 3 |

| A defect that would have broken production | 0 |

Five of the eleven are one PR whose own file header reads *"Hidden experimental comparison endpoint… not part of any production read path"* — production-blocking standards applied to an internal dashboard. Two describe code that had been deleted.

## Three properties held on purpose

It can only defer. Never promote, never invent a finding, never raise a severity. The worst it can do is fail to block.

It is invisible. Nothing it decides is named in the review body, in an inline comment, or anywhere a reader of the PR would see. A deferred finding is reported exactly as before — it simply stops counting toward the gate. The ledger records the deferral unattributed: the PR gains no second reviewer's voice.

Every failure makes the reviewer stricter. No key, network error, no JSON, malformed JSON, unknown verdict, missing id — all resolve to "this blocks", which is today's behaviour. security and anything marked critical are never submitted at all.

## Whole files, not a window

The reviewer's line anchors drift. Three findings in the audit described receipt deletion while pointing at a daily-counter function 400 lines away — a window around the anchor showed the pass code the finding was not about.

Switching to whole files moved 22 of 125 verdicts, 13 of them back to BLOCKING. The narrow window had been hiding what made them real. It also produces grounded reasoning: *"isValidIsoInstant() reconstructs the UTC calendar date and compares it to the timestamp's local date"* rather than *"the documented producers write UTC."*

## Known, and deliberately not mitigated

A wrongly deferred finding stays deferred on every subsequent round — same evidence, same answer. The obvious guard, block after N repeat deferrals, was considered and rejected: it lets a reviewer win by repetition rather than by being right, which is the exact failure this exists to end. The ledger record is the mitigation, and it relies on a human noticing. Worth knowing before this gates anything.

## Testing

397 tests. Nine mutations, nine killed:

| mutation | |

|---|---|

| deferral ignored | killed |

| a truthy value accepted as a deferral | killed |

| ledger drops the flag | killed |

| ledger drops the reason | killed |

| carry-forward drops the deferral | killed |

| security made eligible | killed |

| critical made eligible | killed |

| an unusable answer defaulting to defer | killed |

| apply() promoting instead of demoting | killed |

## Business Value

Targets the 14% of PRs that consume 42% of all review runs. On the replay it clears the gate on 19 of 63 already-stuck PRs without, on hand audit, shipping a single production defect — and it does it invisibly, so the PR reads exactly as it does today.

## Manual Effort Estimate

~6 hours (proposed — Keval to confirm). Most of it was building the backtest honestly: three separate measurement bugs of my own (a lookahead leak, drifted anchors, and same-line-different-finding) each inflated the result before being found and fixed.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#8 — Add App Server protocol fault-injection coverage @sanketghia  no labels

## Summary

- Add a controllable App Server fault-injection harness.

- Cover malformed JSONL, stream/process failures, authentication, missing IDs, workspace mutation before process loss, invalid output, context failure, unexpected requests, timeout, and cancellation.

- Pin timeout and cancellation faults to the turn phase and assert the interrupt request.

- Add a runTicket boundary case proving failed receipt creation, persisted protocol evidence, validation suppression, and preservation of partial workspace state.

## Verification

- pnpm test — 80 files, 1,276 tests passed

- pnpm build passed

- pnpm lint passed

- pnpm format:check passed

- git diff --check passed

## Scope

This PR adds deterministic local fault coverage only. It does not add App Server supervision, automatic compaction, publication, merge, or deployment behavior.

#7 — Persist Codex App Server protocol evidence and failure state @sanketghia  no labels

## Summary

- Add flushable durable App Server event persistence.

- Record bounded protocol checkpoints with process, thread, turn, and protocol identity.

- Persist richer notification/response metadata and typed failure categories.

- Fail closed on malformed JSONL, preserve cancellation evidence, clean up owned Codex homes when evidence flushing fails, and clear stale recovery state.

## Verification

- pnpm test — 79 files, 1,263 tests passed

- pnpm build passed

- pnpm lint passed

- pnpm format:check passed

- git diff --check passed

## Scope

This PR does not add App Server supervision, automatic compaction, publication, merge, or deployment behavior.

#1285 — AERIE-1283: dbt CI — models only in production, models then tests in dev @vvp-trilogy  no labels

## Summary

Splits the dbt GitHub Action (.github/workflows/dbt.yml) so the two credentialed writers run different command shapes per environment, per issue #1283. Today both jobs issue the identical dbt build --select path:models path:seeds, interleaving all tests into the DAG in both environments — wrong in opposite directions.

## Changes

- scheduled-build (production) — added --exclude-resource-type test to the existing dbt build step. The hourly refresh now runs models and seeds only, zero tests, so a tripped data-quality assertion never stops the refresh. Kept build (not run) so seeds stay ordered ahead of the models that select from them.

- pr-build (dev) — split the single dbt step into two steps against the same resolved secret: a models/seeds build --exclude-resource-type test step, then a distinct, required dbt test step. Both pass --vars "{pr_number: N}" so tests read the pr<N>_-prefixed objects the build just wrote. No continue-on-error — a failing test fails the check, and the pr<N>_ objects still exist for inspection.

- dbt/Dockerfile.dbt — updated CMD to the production shape (build … --exclude-resource-type test).

- Docs — refreshed the workflow header comment and added a CI section to dbt/README.md, both stating plainly that production runs no tests.

- Untouched: pr-parse-only (fork PRs) and pr-cleanup.

## Verification

- .github/workflows/dbt.yml parses (yaml.safe_load).

- Dockerfile CMD is valid JSON exec-form.

- Diff re-read against every Acceptance Criterion in #1283.

Closes #1283

#1288 — feat(portfolio): align school and site data contracts @benji-bizzell  changes requested

## Summary

- Make Milestone 9 the canonical Site operating signal and retire duplicated Site opening fields and statuses

- Consolidate planned capacity into phase and expansion entries with explicit availability semantics

- Align Directory, Portfolio, DSS, UI, and migrations around active-default collections and Program-owned School status

## Why

School and Site questions were producing inconsistent answers because the API exposed overlapping opening signals, scalar capacity fields, and collection defaults that included inactive records unless callers knew to filter them. This change establishes one documented contract across the UI, public API, and DSS so consumers can distinguish portfolio state, current capacity, and an operating campus without inference from legacy fields.

## Business Value

DSS and API consumers can answer School/Site status and capacity questions consistently, while explicit inactive queries remain available for cancelled and paused records. The migration removes stale legacy opening data rather than silently falling back to it.

## Breaking changes

- Site open/closed statuses and actualOpenDate/projectedOpenDate projections are retired; consumers use portfolio status plus Milestone 9 due/completed dates.

- Capacity-like scalar fields are replaced by the canonical capacity entry collection and its documented current-capacity resolution.

- Directory and Portfolio collections default to active records; callers must explicitly request paused, cancelled, or all statuses.

## Test plan

- [x] pnpm lint

- [x] pnpm typecheck

- [x] TMPDIR=/private/tmp pnpm test (all workspace and root suites)

- [x] Local UI/API/DSS validation after the development migration

- [x] Five-sample DSS semantic evaluation against the original School/Site questions

#6 — Improve validation repair diagnostics and convergence @sanketghia  no labels

## Summary

- Add deterministic validation failure fingerprints to receipts and repair events.

- Stop early when the same validation failure repeats unchanged instead of consuming another blind repair attempt.

- Include the exact validation working directory and command line in repair prompts.

- Record failed command IDs, categories, changed-file counts, and repair fingerprints in bounded evidence.

## Verification

- pnpm test — 79 files, 1,258 tests passed

- pnpm build

- pnpm lint

- pnpm format:check

- git diff --check

- Live publish:false KLAIR-3476 run: initial client-format failure, one repair attempt, final validation passed, Ready to publish.

## Scope

Publication policy is unchanged. No merge, deployment, or automatic publication behavior is added.

#5 — Harden App Server reliability and operator recovery @sanketghia  no labels

## Summary

- Add bounded App Server protocol/session metadata, interruption, checkpoint, and explicit CLI/Console resume support.

- Reuse the original workspace safely and continue with a bounded recovery instruction instead of blindly replaying the coding prompt.

- Make Factory context and skill policy explicit, keep target skills disabled by default, and add structured-output recovery.

- Reconcile workspace evidence, validation, patch integrity, and publication boundaries.

- Return an actionable Console error when GitHub publication credentials are unavailable.

## Verification

- pnpm test — 79 files, 1,257 tests passed

- pnpm build

- pnpm lint

- pnpm format:check

- git diff --check

## Live validation

- KLAIR-3383: cancellation, turn/interrupt, Console restart, same-run App Server resume, and successful no-publication completion.

- KLAIR-3476: clarification flow succeeded; validation repair was exercised and stopped after its bounded limit when client-format continued failing.

- KLAIR-3389/KLAIR-3390: representative no-publication runs completed with success and preflight-blocked outcomes respectively.

- Disposable publication proof created PR #3747; it was intentionally closed afterward.

## Scope boundary

No merge, deployment, automatic supervisor, or automatic merge behavior is included. Publication, merge, and deployment remain separate explicit gates.

## Follow-up

- Run a post-merge publish:false smoke test.

- Continue measured no-publication pilots.

- Diagnose whether the observed client-format repair failure is ticket-specific or a broader repair-harness issue.

#4 — test: broaden local resilience coverage @sanketghia  no labels

## Summary

- Add restart-safe SQLite reconciliation coverage, including the requested-publication window.

- Preserve completed publication during reconciliation after PR metadata is persisted.

- Classify workspace bootstrap failures as WORKSPACE_BOOTSTRAP_FAILED while continuing sibling work.

- Add representative Console evidence for successful, validation-failed, publication-failed, and partial-batch outcomes.

- Document the local resilience coverage matrix and evidence.

## Verification

- pnpm test — 77 files, 1,229 tests passed.

- pnpm build passed.

- pnpm lint passed.

- pnpm format:check passed.

- Console UI smoke passed — 7 routes.

- Browser checks passed — 10/10.

- git diff --check passed.

- Live local KLAIR-3383 batch completed successfully with validation passed and publish:false.

## Scope

Local execution and Console resilience only. GitHub credential redesign and remote deployment are out of scope.

#3 — Gate 2: local durability and operations hardening @sanketghia  no labels

## Summary

This PR completes Gate 2 local-only maturity hardening for the Factory Console and coordinator.

- Adds Factory Health, startup reconciliation, and operator conflict visibility.

- Adds SQLite schema versioning, integrity checks, backup, and restore operations.

- Adds dry-run-first artifact and workspace lifecycle scanning with evidence-safe cleanup classification.

- Makes initial PR creation idempotent across configured-token and local-gh publication paths.

- Preserves uncertain Git push outcomes as non-retryable operator-recovery states.

- Closes the startup publication window so requested publication cannot be silently finalized as success before publication begins.

- Adds runbook documentation and a Gate 2 local evidence closure record.

## Local verification

- 77 test files, 1,225 tests passed.

- Browser regression: 10/10 passed.

- Console UI smoke: 7 routes passed.

- Build, lint, formatting, and diff checks passed.

- Representative local KLAIR-3383 batch completed successfully with 3/3 validations passed and publication evidence recorded.

## Scope and boundaries

- Local filesystem and loopback Console only.

- No hosted execution, remote deployment, merge automation, release, or cleanup was performed by this PR.

- Lifecycle cleanup remains review-first; no historical artifacts were deleted.

- Uncertain external outcomes remain explicitly operator-recovery states.

#1801 — fix(repo): Accept core_submitting in HC pinned-run guard @heimdall-keval-factory[bot]  approvedAutomated PR

The headcount forecast refresh fails on every run because a bookkeeping step marks the input as "submitting" right before running the SQL that still only checks for "pinned" — so that check can never pass. Updating the one stored-procedure guard to also recognize the submitting state fixes it.

Ticket: SURTR-1196

> Ready for review. Nothing ran the change, so it is unproven. A person still merges.

## For The Agent

_Everything below is detail for review. The summary above is the change._

Presented as ready — verification none, scope tier draft, fix_class code_fix, HEIMDALL_READY_PRS=true.

### What's broken

Failing run 5d8f0400-19bc-448c-98db-62a18ed4e9e4 (Lambda request c7c2fc3f-03ba-4d91-a04f-9eb12b820107) died with [ERROR] StatementTerminalError: Redshift statement failed: ERROR: HC Core requires one pinned input run, raised inside core_budgets.sp_refresh_hc_data_consolidated. That RAISE fires whenever the ingestion_run row for this run id is not in state pinned at the moment the Core statement executes — but src/handler.py's _run_writer_phase() always calls sp_begin_workforce_phase_submission first, which durably flips the row to core_submitting before the Core CALL is even submitted to the Data API, so the guard is checking a state that has already moved on and can never be satisfied.

Root cause. The Core writer's durable-submission bookkeeping (added in SURTR-374, PR #921) transitions staging_workforce_gsheets.ingestion_run.state from pinned to core_submitting synchronously via sp_begin_workforce_phase_submission (ddl/2026-09-07_staging_workforce_gsheets_rename.sql:377-390) as the very first step of _run_writer_phase (src/handler.py:54-81), before the actual CALL core_budgets.sp_refresh_hc_data_consolidated(...) statement is submitted. But sp_refresh_hc_data_consolidated still guards on state='pinned' (ddl/2026-09-07_staging_workforce_gsheets_rename.sql:458-460, unchanged from the original ddl/2026-08-04_sp_refresh_hc_data_consolidated_pinned.sql:8-10). Because the state flip commits before the Core statement is even submitted, the guard's COUNT(*) is always 0 and the exception fires on every run that reaches the Core phase — this is a deterministic ordering defect, not a rare race, and it produces zero written rows in core_budgets.hc_data_consolidated (and downstream mart_education.agg_hc_by_teamroom) rather than a partial write. It never surfaced in unit tests because tests/test_handler.py stubs the Redshift client's SQL execution instead of exercising real stored-procedure state transitions.

### What this PR changes

This is a code-confined fix inside pipelines/runners/hc-forecast-refresh/. Add a new DDL migration (following the pipeline's existing supersession pattern, where 2026-09-07_staging_workforce_gsheets_rename.sql superseded 2026-08-04_sp_refresh_hc_data_consolidated_pinned.sql) that CREATE OR REPLACEs core_budgets.sp_refresh_hc_data_consolidated so its guard accepts state='core_submitting' — the state the row is durably in for the entire window the Core statement actually executes, per sp_begin_workforce_phase_submission and sp_transition_workforce_ingestion_run — instead of the now-unreachable state='pinned'. No Python change is needed: handler.py's mark-then-submit ordering is intentional and correct for crash recovery; the stored procedure's guard is the stale half of the SURTR-374 refactor. tests/test_sql_contracts.py and tests/test_handler.py were checked and neither asserts the literal 'pinned' predicate, so nothing else depends on the current wording.

Why this fixes it. The defect is a single stored-procedure guard confined entirely to pipelines/runners/hc-forecast-refresh/ddl/, with a well-understood one-predicate fix (state='pinned' -> state='core_submitting') deliverable as one new migration file consistent with this pipeline's existing DDL supersession pattern. It is not a config knob (timeout, memory, env var, or missing dependency), so none of the other fix classes apply, and the ordering defect is unambiguous and fully explained by the code, so it does not need to be escalated as 'other'.

#### Files changed

 ...efresh_hc_data_consolidated_core_submitting.sql | 75 ++++++++++++++++++++++

1 file changed, 75 insertions(+)

### Verification

### pytest — no test suite

### verify: ruff check — exit 0

[notice] A new release of pip is available: 25.3 -> 26.2.1

[notice] To update, run: pip install --upgrade pip

All checks passed!

### verify: ruff format --check — exit 0

1806 files already formatted

### verify: pytest (pipeline lambdas) — exit 0

6.0

rootdir: /home/runner/_work/Surtr/Surtr/publish/pipelines/cdk/lambdas

configfile: pyproject.toml

testpaths: tests

plugins: cov-7.0.0

collected 495 items

tests/test_ai_spend_raw_api.py .............................. [ 6%]

tests/test_coordinate_fanout_run.py .................................... [ 13%]

.... [ 14%]

tests/test_create_run_record.py ........................ [ 18%]

tests/test_gchat_notifier.py ........................... [ 24%]

tests/test_generate_chunks.py ..... [ 25%]

tests/test_gsheet_tracker.py .................. [ 29%]

tests/test_load_fanout_plan.py .............. [ 31%]

tests/test_redshift_cluster_iam_role_association.py ........ [ 33%]

tests/test_registry_sync.py ......................... [ 38%]

tests/test_triage_dispatcher_handler.py ................................ [ 45%]

................... [ 48%]

tests/test_triage_dispatcher_signature.py .............................. [ 54%]

...... [ 56%]

tests/test_triage_reconciler.py ............ [ 58%]

tests/test_triage_reconciler_tracker.py ............ [ 61%]

tests/test_update_run_failed.py ........................................ [ 69%]

.............. [ 71%]

tests/test_update_run_success.py ....................................... [ 79%]

....................... [ 84%]

tests/test_verify_on_demand_control.py ................................. [ 91%]

............................................ [100%]

============================= 495 passed in 1.23s ==============================

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | repo |

| Failing run | issue |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | linear-SURTR-1196 |

| Verify | none |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev label to take the PR over and stop it entirely.

#1800 — fix(SURTR-648): use supported CAPEX batch request @marcusdAIy  approved

## Summary

- remove ExecutionMode from the CAPEX BatchExecuteStatement request

- retain atomicity through the Data API batch operation's transaction contract

- validate fake-client requests against botocore's real BatchExecuteStatement input shape

- add a regression proving unsupported request fields are rejected instead of silently accepted

## Why

The released installer passes ExecutionMode="TRANSACTION". The deployed botocore service model rejects that field during client-side parameter validation, before any DDL is submitted. BatchExecuteStatement already executes its Sqls as one transaction, so the option is unnecessary.

The prior fake accepted arbitrary keyword arguments and asserted the same invalid request. The replacement validates every submitted request through botocore while keeping the exact-request assertion. The installer omits ExecutionMode for compatibility with deployed SDK service models, including models that predate the optional field.

## Validation

- PYTHONPATH=src uv run pytest -q tests/ — 36 passed

- ruff==0.15.22 check — passed

- ruff==0.15.22 format --check — passed

- git diff --check — passed

## Rollout

This PR changes source only. It does not apply Redshift DDL, start a pipeline execution, enable on-demand execution, or enable the schedule. After a separate production promotion, the existing authorized fail-closed preflight and controlled installation can proceed.

Linear: [SURTR-648](https://linear.app/builder-team/issue/SURTR-648)

#1797 — [SURTR-1123] Rename workforce warehouse identifiers only @caina-barbosa  approved

## This PR only renames warehouse objects

This effort is limited to renaming the SURTR-374 workforce objects so they follow the repository naming conventions.

It is not meant to change pipeline behaviour, calculations, data grain, source selection, lifecycle rules, recovery rules, permissions, schedules or infrastructure.

Two Python runtime files change only because they contain the names of the tables and stored procedures that the pipelines call. The executable logic around those names is unchanged.

This PR supersedes closed PRs #1779 and #1796. It excludes the extra migration machinery and unrelated hardening from #1779.

## Why we need this

The existing names do not follow WAREHOUSE_CONVENTIONS.md:

- surtr_hc_forecast_refresh contains a ticket name and the ambiguous abbreviation hc

- operational current-cost state is stored in the consumer-facing mart_education schema

- several object and column names describe implementation history instead of the workforce business domain

The new names use the private staging layer for operational state and workforce for the business domain.

## What changes

The PR contains:

- one SQL file that renames the SURTR-374 staging schema, tables, columns and procedures

- one SQL file that renames the current-cost evidence table and moves its 2 operational state tables into the workforce staging schema

- identifier substitutions in the 2 affected pipeline runtimes

- updates to directly affected name assertions in existing tests

The SQL recreates stored procedures because Redshift does not rewrite qualified table and procedure names inside procedure bodies.

## Scope boundaries

This effort is not meant to change:

- any table column type, constraint or data grain

- any business calculation or grouping rule, including weekly_cost

- snapshot loading or empty-snapshot behaviour

- source files, manifests or source selection

- lifecycle transitions or recovery behaviour

- procedure signatures, owners, SECURITY DEFINER settings or effective grants

- handler inputs or response keys

- Core or Aerie object names or calculations

- pipeline schedules, triggers, IAM, CDK or pipeline.json

- AWS resource names

- deployment automation or cutover tooling

- unrelated defects or hardening opportunities

After replacing only the approved identifiers, all 20 recreated stored-procedure bodies match the current main versions.

## Dev proof

We reset dev to the production-shaped legacy catalog and ran both SQL files once, in order.

Results:

- legacy baseline: 19 tables, 342 columns and 21 procedures

- staging migration: 107 statements committed successfully

- current-cost migration: 61 statements committed successfully

- final catalog: 11 staging tables, 3 current-cost tables, 13 staging procedures and 6 current-cost procedures

- old SURTR-374-owned names remaining: 0

- evidence rows preserved: 50,852

- included rows preserved: 2,636

- unavailable rows preserved: 48,216

- total amount preserved: 598,554,320.68

- publication rows preserved: 4

- recovery-audit rows preserved: 0

- owners, signatures, security settings and grants preserved

- unrelated dev objects and retained dependency data unchanged

Production remained read-only.

## Code validation

- hc-forecast-refresh: 97 tests passed

- mart-education-hc-current-cost-refresh: 83 tests passed

- repository CI: passed

- Ruff 0.15.22 check and format: passed

- compileall: passed

- independent SQL review: passed

- independent runtime review: passed

- final scope review: passed

## Production order

Merging this PR into main does not deploy production.

During a later, separately authorised production window:

1. Run the staging SQL file.

2. Run the current-cost SQL file.

3. Merge the prepared main to production release PR so CD deploys the matching runtime identifier changes.

4. Run and verify one workforce refresh.

Do not deploy the new runtime before the SQL commits. Do not leave the old runtime running for an extended period after the SQL commits.

#3744 — feat(board-doc): upgrade shared model to Fable 5.1 @marcusdAIy  approved

## Summary

- upgrade the shared Budget Bot BOARD_DOC_MODEL from claude-opus-4-7 to native Anthropic claude-fable-5-1

- cover generation/regeneration, Coach Claire and agentic rounds, refresh commentary, short summaries, narrative and goal QC, quarter memory, eval, and the shared-model brainlift fallback slots

- preserve the existing Anthropic connection, authorization, canonical-Doc reconciliation, proposal approval, deterministic finance, and independent model selectors

- add an exact-Fable native JSON-schema path because Fable rejects forced tool_choice; validate responses locally and fail closed

- preserve arbitrary Brainlift section names through a lossless closed-schema wire adapter

- update the conservative context registration and generation cost estimate to Fable 5.1's published $10/M input and $50/M output base rates

## Compatibility evidence

Synthetic checks through Klair's existing Anthropic account used no Board Doc content and made no production changes:

- native claude-fable-5-1 access confirmed

- adaptive thinking at high and medium effort accepted

- max_tokens=128000 accepted

- streamed automatic tool use and tool-result replay passed

- native JSON schema passed for the actual narrative, goal, SPOV, and lossless Brainlift wire schemas

- explicit temperature and forced tool_choice rejection reproduced and handled by this change

See budget_bot/board_doc/KLAIR-3524-Fable-5-1-Upgrade.md for the full evidence, manual checklist, and rollback plan.

## Validation

- 4874 passed, 2 deselected — full network-denied tests/board_doc suite; only the two documented live tests deselected

- 310 passed — focused shared-model, structured-output, Brainlift, QC, chat, refresh, and cost regression suite

- 27 passed — final native structured-output rerun after the no-DNS assertion

- Ruff 0.15.22 format/check passed on all changed Python files

- git diff --check passed

- Pyright added no diagnostics versus main (baseline and branch both retain the same 22 existing errors and 22 warnings)

- Python 3.12 uv lock passed with no package-version churn

- independent source review: approved, no unresolved findings

## Release gate

This PR does not authorize or perform a production deployment. Representative hands-on testing in a dedicated test Doc remains required before production approval.

Linear: KLAIR-3524

#1286 — AERIE-1284: Shared HubSpot display-name dimension for enrollment + pipeline reports @vvp-trilogy  approved

## Summary

Replaces the hand-maintained sis_campus_program_map.csv seed with a derived dimension, int_school_identity (one row per HubSpot program), that maps every system's school identity — SIS campus_id, Finalsite site, HubSpot program object id, HubSpot program_code — to one user-friendly label (display_name, the HubSpot program display name). Both the Enrollments report (mart_enrollment_dtl) and the Admissions Pipeline report (mart_admissions_pipeline_dtl) consume it, so the two reports name the same school identically. The mapping is read from SIS, so the report self-heals: adding a hubspot_program external id in SIS admits a campus on the next run with no repo change.

Closes #1284.

## What changed (following the issue's 8 steps)

1. all_program declared as an educrm source; new stg_educrm_all_program (pure 1:1 projection, equal_rowcount test).

2. campus_external_ids declared as a sis source; new stg_sis_campus_external_id (soft-delete filter, casts, no joins) with a (campus_id, system) uniqueness test. The source envelope has no surrogate id, so the grain is keyed on (campus_id, system).

3. int_school_identity in the new intermediate/shared/ folder. Grain: one row per HubSpot program; resolves display_name from all_program.program_name on the numeric hubspot_program_id, never on a name. Columns: hubspot_program_id (bigint), hubspot_program_code, sis_campus_id, finalsite_site, display_name — only display_name is unprefixed.

4. Drop rule in stg_sis_campus — an EXISTS semi-join so a campus that resolves to no HubSpot program never reaches int_campus/axis/grid/mart. Documented in the header as a second knowing rule-5 exception with the removal trigger named (the SIS-name cutover).

5. mart_enrollment_dtl repointed at the dimension: the program_map CTE reads int_school_identity, the LEFT JOINs become INNER (every campus reaching the mart resolves by construction), and has_hubspot_program is gone.

6. mart_admissions_pipeline_dtl — all three arms publish program_name from the dimension. The Finalsite arm INNER-joins on finalsite_site (drops unresolvable tenants, symmetric with the enrollment drop); the EduCRM arms LEFT-join on hubspot_program_code (null-tenant pre-launch leads are kept). Existing campus_name/campus_long_name/program_code are unchanged — the label is added, not renamed.

7. Deleted sis_campus_program_map.csv, assert_sis_campus_program_map_covers_axis.sql, assert_sis_campus_program_map_hubspot_resolves.sql. Repointed assert_sis_deposit_overlay_unmatched_within_threshold.sql's Alpha-program bound at the dimension.

8. New WARN test assert_sis_campus_unresolved_hubspot_program.sql (dropped-campus visibility, shape modelled on assert_admissions_pipeline_tenant_coverage). Updated assert_admissions_pipeline_reconciles.sql so the Finalsite expected arm applies the same drop filter.

The published mart column names (program_code, program_name) are unchanged — the frozen contract with the sync and Convex. Only the internal derivation changed.

### One extra change: connect_timeout 30 → 120

redshift_connector applies connect_timeout as the socket read timeout, so it bounds every query. Removing the covers_axis blocker re-enables mart_enrollment_dtl, whose dense factless anti-join over the int_enrollment_cohort view chain legitimately runs ~40s against current data — so a 30s timeout left it red on a timeout (proven locally: it builds in ~36–39s with the timeout raised; the query is not a SQL error). Raised to 120s in profiles.yml.docker (CI) and profiles.yml.example (local). This is pre-existing cost, not a regression from this PR — the change *reduces* the grid from 67 to 53 campuses.

## Verification

Full dbt build --select path:models path:seeds --vars '{pr_number: N}' run against the real Redshift warehouse: PASS=289, WARN=2, ERROR=0, SKIP=0.

- The 2 WARNs are the new assert_sis_campus_unresolved_hubspot_program (14 dropped campuses) and the pre-existing assert_admissions_pipeline_tenant_coverage (6, data drift — both its inputs are untouched by this PR).

- Acceptance checks against the built objects: enrollment mart has 0 null program_code/program_name across 53 campuses (was 67); Nova Academy Austin → Nova Austin, Nova Academy Bastrop → Nova Bastrop; cross-mart program_name consistency = 0 mismatches; all three pipeline arms 100% labelled; the Finalsite arm drops exactly the two founders tenants (32 rows).

- pnpm typecheck, pnpm biome check, pnpm lint:boundaries: all clean (no TS files touched).

## Open question for sign-off (from the issue, unchanged by this PR)

Dropping unresolvable campuses removes real enrollments at a few campuses (the issue's "22 students", chiefly the two Founders entities and Alpha Orlando). This is by design and surfaced by the new WARN test; adding the eight hubspot_program bindings in SIS (an out-of-repo action listed in the issue) returns those campuses automatically on the next run. Not a code dependency.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1260 — Remove getPortfolioHealth MCP surface @YibinLongTrilogy  approved

## Summary

Remove the retired getPortfolioHealth aggregate from every Aerie-owned product surface: Convex, MCP discovery, Worker/stdio registration, agent contracts and policies, UI rendering, Flue output handling, public API metadata, documentation, fixtures, and tests. The supported v2 Insights list/detail workflow remains available, while proactive Worker guidance now uses bounded status-filtered site rosters plus factual overdue-milestone evidence.

### Changes

- chat/convex/rhodes/mcp.ts and chat/rhodes-worker/mcp-server/tools/views.ts — Delete the public Convex query and MCP registration; add an explicit bounded listSites path for status-filtered rosters.

- packages/contracts/src/agent-run-protocol.ts, packages/contracts/src/agent-tool-registry.ts, and chat/convex/agentRuns/runs.ts — Remove the tool from agent names, schemas, descriptions, public policies, and capability routing; bump the public tool-policy version to invalidate stale runs.

- chat/lib/rhodes-mcp-contract.ts and UI/Flue renderers — Remove in-app advertisement, the dedicated portfolio-health card, and the obsolete nested gateway fixture; keep listSites milestone progress visible with DTO-shaped coverage.

- chat/rhodes-worker/lib/mcp-instructions.ts and mcp-server/tools/sites.ts — Replace the retired proactive workflow with capped active/paused site reads and factual overdue reporting.

- chat/lib/public-api/v2/domains/insights.ts, parity guard, docs, and feature history — Preserve the supported v2 Insights contract, document the later cleanup, and allow the intentional cross-repository MCP removal during parity checks.

- Tests — Remove obsolete behavior fixtures and add negative registry coverage, bounded roster coverage, forwarding/schema coverage, and UI DTO coverage.

### Design Decisions

- The v2 Insights list/detail API is retained because it is a supported bounded public workflow, not the retired MCP aggregate.

- listSites remains generally compatible for existing callers; only explicit status-filtered calls with limit use the bounded roster path used by proactive guidance.

- The replacement instructions report only returned milestone progress and overdue evidence; they do not infer health labels from missing evidence or the absence of overdue items.

- The public Agent tool-policy version is incremented so durable runs created under the prior tool catalog fail closed rather than appearing compatible with a changed policy.

## Business value

Eliminates a retired, misleading aggregate health surface and prevents agents or UI clients from presenting stale readiness scores, ratings, blocker rankings, or unsupported portfolio judgments. Users retain factual, bounded milestone and overdue evidence through the supported workflows.

## Estimated manual effort

Estimated time to complete this work without AI: 1 working day.

## Test Plan

- [x] Full Rhodes Worker test suite.

- [x] Focused Convex parity, public-agent, contract, site-tool, and Rhodes card tests.

- [x] Chat, Worker, and contracts typechecks.

- [x] Biome, architecture-boundary, Convex-path, read-bound, test-architecture, and git diff --check validation.

- [x] Read-only seven-lane adversarial review with independent verification of accepted findings.

- [ ] Run deployed MCP discovery/parity checks after the normal release deployment; no deployment or external mutation was performed for this PR.

#1281 — AERIE-1279: Key the offering-kind axis on program_type so per-campus Main Program campuses reach the enrollment report @vvp-trilogy  approved

## What & why

Keys the enrollment report's offering-kind axis on the SIS program_type enum instead of the free-text program name.

- int_school_year_offering.sql: WHERE po.program_name = 'School Year'WHERE po.program_type = 'MAIN'. Header rewritten to describe selection by the source enum and to note SIS expresses MAIN under two naming conventions (the shared School Year program and per-campus <Campus> Main Program); both now reach the axis. program_type is carried in the CTE and published as a documented constant column (mirroring int_campus.delivery_mode); its domain is guarded by the source-side test rather than a tautological model-level accepted_values.

- int_program_offering.sql: carries program_type through from stg_sis_program (already selected there, previously unused downstream). Header + _int_enrollment__models.yml updated to document it.

- tests/assert_sis_program_type_domain.sql: new source-side value-domain test mirroring assert_sis_campus_delivery_mode_domain.sql. Reads source('sis','programs'), non-deleted rows, fails when program_type IS NULL or NOT IN (MAIN, SUMMER_CAMP) — a future EXTENDED_DAY/AFTER_SCHOOL surfaces for review instead of being silently dropped.

- seeds/sis_campus_program_map.csv: eight Alpha School campuses added (alphabetical by campus_name) so assert_sis_campus_program_map_covers_axis stays green.

- Stale name = 'School Year' scope references in stg_sis_program.sql and _sis__sources.yml updated to the current program_type = 'MAIN' shape (rule 22).

Why: program_type is SIS's own enum for the MAIN-vs-SUMMER_CAMP distinction. The name predicate only expressed "the year-long main program" as a side effect of the old one-shared-School Year-program convention; when SIS began provisioning per-campus <Campus> Main Program records, the label changed and eight active, PHYSICAL Alpha School campuses silently fell off the axis. The #1268 staging campus filter (not the name) is what keeps the 11 out-of-scope Main Program campuses out — the name predicate's only live effect was dropping these eight.

Closes #1279.

## Verification — real Redshift build (green)

Full dbt build --select path:models path:seeds --vars '{pr_number: 1279}' against sandbox_education: PASS=284, ERROR=0, SKIP=0, one pre-existing unrelated WARN (assert_admissions_pipeline_tenant_coverage, admissions-pipeline domain, present on main).

Acceptance tests passing: assert_sis_program_type_domain, assert_sis_campus_program_map_covers_axis, assert_sis_campus_program_map_hubspot_resolves, unique_int_school_year_offering_id, dbt_utils_unique_combination_of_columns_mart_enrollment_dtl_campus_id__cohort_id__session_school_year (has_fact grid uniqueness), assert_sis_enrollment_has_fact_consistency, assert_sis_enrollment_qualified_is_fact (qualified → has_fact completeness).

- The eight campuses now produce 10 campus×year axis rows (Anywhere Center - Founders & Greenwich - Armonk: 2026+2027; the other six: 2026). None were on the axis before.

- Axis grain unchanged: campus×year has no group > 1 before or after; unique_int_school_year_offering_id passes.

- SELECT DISTINCT program_type FROM int_school_year_offering → only MAIN; Summer Camp contributes zero axis rows (Port Chester's Summer Camp 2026 offering is correctly excluded).

- The four zero-enrollment campuses (Bethesda, Miami Beach (Biarritz), San Juan, Virtual) each form 13 factless grid rows (0 facts).

## Before/after cohort counts — SY2024–SY2027 (post-#1258 spine)

Measured on the same source snapshot: main's code and this branch both built fresh (pr9000_ baseline vs pr1279_), so the delta reflects only this change, not the hourly-refresh drift in the production tables. Counts are on the #1258-filtered spine (test records dropped), so they differ from the issue's raw-spine figures by design.

Distinct fact enrollments:

| School year | Before | After | Δ |

|---|---|---|---|

| SY2024 | 284 | 284 | 0 |

| SY2025 | 757 | 757 | 0 |

| SY2026 | 1621 | 1667 | +46 |

| SY2027 | 22 | 23 | +1 |

| Total | 2752 | 2799 | +47 |

Fact rows (cohort fan-out): SY2026 3735→3818, SY2027 25→26. Total mart rows incl. factless grid: SY2026 4255→4429, SY2027 542→569.

Delta attribution — confined to the eight new campuses + one explained re-enrollment. The +46 SY2026/SY2027 distinct fact enrollments are the four enrolled new campuses: Alpha Greenwich - Armonk 24, Alpha Anywhere Center - Founders 17, Alpha Austin - Founders 3, Alpha Lexington 2. (Bethesda, Biarritz, San Juan, Virtual add 0.)

The +1 at the existing campus Alpha Greenwich - Port Chester (SY2027, 10→11) is a correct second-order effect, not a predicate re-resolution — Port Chester's axis offering IDs are byte-identical before/after. One student transferred from the newly-admitted Alpha Greenwich - Armonk (2026) and holds an ON_HOLD 2027 enrollment at Port Chester. With Armonk now visible on the axis, that prior-year enrollment is seen, so the 2027 enrollment's is_returning flips FALSE→TRUE and it joins the re-enrollment-other cohort. The report is now correctly recognizing a returning student whose prior enrollment was previously invisible. No source drift (staging stg_sis_enrollment identical at 18,169 rows across both builds).

## Founders SIS-vs-HubSpot on-campus parity (#1216 restated)

Alpha Anywhere Center - Founders — the reported symptom (14 on-campus students in HubSpot, zero rows in SIS before this change) — now resolves. SIS cohort split (SY2026): on-campus 16, first-day 16, re-enrolled 1, re-enrollment-other 1 (17 distinct enrollments; the 16 PENDING_REVIEW land on-campus, plus one re-enrollment). SIS on-campus 16 vs HubSpot's reported 14 — this change closes the parity gap (SIS was showing 0) rather than opening a new one; the small SIS/HubSpot residual is the pending-review timing the issue describes. Full #1216 cross-system parity reconciliation (HubSpot-side models) is out of scope for this dbt change.

## Alpha Virtual — explicit decision

Alpha Virtual is admitted despite its name because SIS records it as delivery_mode = 'PHYSICAL', so the #1268 mode filter does not hold it back and the type-keyed axis includes it. It has zero enrollments today (factless grid rows only), so it costs nothing numerically. Correcting the SIS delivery_mode at source is Out of Scope — worth raising with the SIS team, but the report is not special-casing one campus while the source disagrees (same resolution as the Alpha World School case in #1268: SIS is the spine).

## Seed mappings — two reviewed decisions + build-driven corrections

- Alpha Miami Beach (Biarritz)has_hubspot_program = false. The one HubSpot Alpha Miami Beach program is already mapped to the distinct SIS campus Alpha Miami Beach (a34e72b4); pointing Biarritz at it too would double-count it in any SIS↔HubSpot parity comparison. (Whether Biarritz is a genuinely separate site vs a rename is an admissions data question — Out of Scope.)

Two corrections vs the issue's verbatim "Seed rows required" block, both required to pass existing seed tests and both following the seed's own documented conventions (caught by the mandatory local build):

1. program_name populated for Biarritz and Virtual with their identity name (= program_code: Alpha Miami Beach Biarritz, Alpha Virtual) instead of an empty field. The seed's not_null test on program_name (unchanged by this PR) rejects an empty CSV cell, which dbt loads as NULL. This matches the established SIS-only-campus convention — campuses HubSpot has no program for publish their own campus name as an identity program (e.g. Alpha Carrollton, Alpha Anywhere Center Port Chester). The identity name is distinct from Alpha Miami Beach, so the Biarritz no-double-count intent is preserved.

2. Alpha Austin - Founders and Alpha San Juan set has_hubspot_program = false (not true). Their HubSpot programs exist in EduCRM (26 and 39 rows) but carry zero real enrollment facts yet (has_fact = false throughout) — assert_sis_campus_program_map_hubspot_resolves requires a true mapping to resolve to a real fact. This is exactly the seed's documented "valid campuses whose HubSpot program carries no enrollment facts yet (SIS leads HubSpot for newly onboarded campuses)" case. The four campuses that do resolve to real facts (Anywhere Center - Founders, Bethesda, Greenwich - Armonk, Lexington) stay true.

The seed yml description counts are updated to the current shape (rule 22): has_hubspot_program TRUE 48→52, FALSE 10→14 (6 SIS-only + 8 no-facts-yet).

## Rule notes

Rule 8 preserved — the offering-kind filter stays in exactly one place (int_school_year_offering); only the column it reads changed. No new rule-5 exception taken. All descriptions/comments state the current shape (rule 22); no name = 'School Year' scope reference remains.

#1259 — Remove presentReplanOptions and generatePlan @YibinLongTrilogy  approved

## Summary

Remove the retired getPortfolioHealth aggregate from every Aerie-owned product surface: Convex, MCP discovery, Worker/stdio registration, agent contracts and policies, UI rendering, Flue output handling, public API metadata, documentation, fixtures, and tests. The supported v2 Insights list/detail workflow remains available, while proactive Worker guidance now uses bounded status-filtered site rosters plus factual overdue-milestone evidence.

### Changes

- chat/convex/rhodes/mcp.ts and chat/rhodes-worker/mcp-server/tools/views.ts — Delete the public Convex query and MCP registration; add an explicit bounded listSites path for status-filtered rosters.

- packages/contracts/src/agent-run-protocol.ts, packages/contracts/src/agent-tool-registry.ts, and chat/convex/agentRuns/runs.ts — Remove the tool from agent names, schemas, descriptions, public policies, and capability routing; bump the public tool-policy version to invalidate stale runs.

- chat/lib/rhodes-mcp-contract.ts and UI/Flue renderers — Remove in-app advertisement, the dedicated portfolio-health card, and the obsolete nested gateway fixture; keep listSites milestone progress visible with DTO-shaped coverage.

- chat/rhodes-worker/lib/mcp-instructions.ts and mcp-server/tools/sites.ts — Replace the retired proactive workflow with capped active/paused site reads and factual overdue reporting.

- chat/lib/public-api/v2/domains/insights.ts, parity guard, docs, and feature history — Preserve the supported v2 Insights contract, document the later cleanup, and allow the intentional cross-repository MCP removal during parity checks.

- Tests — Remove obsolete behavior fixtures and add negative registry coverage, bounded roster coverage, forwarding/schema coverage, and UI DTO coverage.

### Design Decisions

- The v2 Insights list/detail API is retained because it is a supported bounded public workflow, not the retired MCP aggregate.

- listSites remains generally compatible for existing callers; only explicit status-filtered calls with limit use the bounded roster path used by proactive guidance.

- The replacement instructions report only returned milestone progress and overdue evidence; they do not infer health labels from missing evidence or the absence of overdue items.

- The public Agent tool-policy version is incremented so durable runs created under the prior tool catalog fail closed rather than appearing compatible with a changed policy.

## Business value

Eliminates a retired, misleading aggregate health surface and prevents agents or UI clients from presenting stale readiness scores, ratings, blocker rankings, or unsupported portfolio judgments. Users retain factual, bounded milestone and overdue evidence through the supported workflows.

## Estimated manual effort

Estimated time to complete this work without AI: 1 working day.

## Test Plan

- [x] Full Rhodes Worker test suite.

- [x] Focused Convex parity, public-agent, contract, site-tool, and Rhodes card tests.

- [x] Chat, Worker, and contracts typechecks.

- [x] Biome, architecture-boundary, Convex-path, read-bound, test-architecture, and git diff --check validation.

- [x] Read-only seven-lane adversarial review with independent verification of accepted findings.

- [ ] Run deployed MCP discovery/parity checks after the normal release deployment; no deployment or external mutation was performed for this PR.

#1276 — AERIE-1857: Add unpublished durable uploader and automatic wake @caina-barbosa  approved

## Summary

- Add the unpublished, one-shot invocation uploader and commit-triggered detached wake for the existing core-owned local invocation queue.

- Keep receipt identity, queue admission, claims, leases, retries, acknowledgements, tombstones and uploader state under one LocalStateCoordinator.

- Keep the package private, dependency-free and dormant: no host observer exists, and the current unpublished package has no authenticated native credential-provider artifact.

## Why

AERIE-1856 established durable adapter admission and atomic invocation-queue creation, but it deliberately stopped before delivery. Later Claude Code, Codex and Pi adapters need one common local delivery contract rather than host-specific network or credential logic.

This slice closes the local path from a committed queue entry to one bounded upload pass while preserving the authority boundary: observers cannot select credentials, endpoints, paths, receipts, proofs, queue writers or uploader policy.

## Business value

This makes later host adapters small and fail-open. They can submit a bounded observation and continue host execution while core owns durable retry, deduplication, privacy and server-outcome handling. Offline operation, child-process loss and CLI overlap do not require a daemon or expose credentials to host configuration.

## Slice and scope

- Slice: AERIE-1857 — telemetry rollout 11/19

- Strict predecessor: AERIE-1856 / 5107e4b031fceb310ba2f94c92c4c5ee69b19d63

- Rebased onto current main: fcc9897bfe4fce1b3a2e6a3364795891d490bb6d

- Exact reviewed head: 89fe7e570419be5c0524c59f5dd2504f8c2c4d7e

- Package subtree: 781aa7350a38ee85bd9002373ad23974af1ee4b5

Owned surfaces are limited to packages/add-aerie-skill/**.

## Runtime design contract

### Queue and receipt authority

Invocation receipt IDs are canonical 43-character unpadded base64url encodings of exactly 32 random bytes. Core creates the exact seven-field wire payload only after capacity and current-binding checks. Retries reuse the committed ID, proof and payload.

The invocation queue and delivery metadata share the existing state generation and coordinator. Combined queue and delivery metadata are bounded to 4,096 records and 8 MiB. Attempt accounting occurs durably at claim time, before network I/O, so process loss cannot bypass the 255-attempt bound.

A 30-second lease and claim token protect acknowledgement. Expired or stale claimants cannot retry, acknowledge or terminalize a newer claim. Terminal paths remove proof-bearing payload even when tombstone capacity is full.

### Bounded one-shot delivery

One drain pass processes a finite batch and starts no work after its 25-second monotonic deadline. Each HTTP request is bounded to five seconds. Credential lookup is part of the same pass budget. Response reads are byte-bounded and cancel the underlying stream on overflow, malformed chunks, decoding failure, abort or read failure.

The transport posts only the exact invocation DTO to /skill-device/telemetry/invocations with the core-retrieved bearer. It maps:

- the five approved 200 outcomes to terminal acknowledgement;

- deterministic 400 and unavailable 404 to terminal rejection;

- 409 invocation_receipt_conflict to retained needs_repair evidence;

- 401/403 to durable auth_deferred retry for later foreground reauthentication;

- 429, 5xx, timeout, malformed response and transport failure to bounded retry.

Backoff is jittered and capped at six hours. Invocation payloads expire after 35 days or the attempt bound, leaving bounded content-free terminal evidence.

### Wake and credential boundary

Wake happens only after a new queue entry is durably committed. Duplicate, stale, rejected, capacity-full and other no-new-entry outcomes do not wake.

Core launches only process.execPath plus the fixed private argument, with shell: false, ignored stdio, detached execution, windowsHide, unref() and a closed minimal environment. The executable, entrypoint and working directory are package-owned and physically verified. No credential, payload, proof, endpoint, host value or caller-selected path enters arguments or environment.

The current private manifest intentionally has no runtime or optional dependency. It therefore does not trust ancestor modules, mutable package-local native code or runtime self-hashes as credential authority. Without an authenticated native provider, detached production dispatch fails closed before claim or network. The complete queue → wake → attempt → acknowledgement flow is exercised through a build-time test substitution that is absent from production JavaScript, declarations and the tarball.

A later packaging/publication slice may establish an optional native-provider dependency and external integrity authority. This PR does not invent that authority.

### Concurrency and recovery

Installer drain runs before the installer lock and uses the exact same coordinator/root identity as admission and uploader composition. Network runs outside the state lock. Independent coordinators over one root prove exclusive claims, close/reopen recovery, lease reclamation, stale-response rejection and one terminal settlement.

Platform state roots are canonical for Linux XDG, macOS Application Support and Windows LOCALAPPDATA; generic test/application bases retain the existing .aerie-skill convention. Windows remains fail-closed without the required platform security checker.

## Explicitly out of scope

This PR adds no:

- real Claude Code, Codex or Pi observer/producer;

- daemon, timer, scheduler, startup sweep or guaranteed detached completion;

- browser popup, background reauthentication or detached credential deletion;

- trusted-runtime queueing; AERIE-1958 remains the sole trusted-runtime path;

- public CLI flag or public uploader/queue/credential authority;

- authenticated native credential-provider artifact or runtime dependency;

- npm publication, registry mutation, deployment or production API call;

- generated Convex change or real-host/native manual campaign.

## Behavior and production effect

Dormant. The package is still @aerie/add-aerie-skill, private: true, UNLICENSED, unpublished and dependency-free. There is no real host producer, and unavailable credential authority causes the private child to exit without claiming or sending.

Ordinary import, startup, login and logout do not sweep the queue. Private disable/re-enable preserves queued records, installed Skills, credential state and installation lineage; it is not exposed as a public CLI flag.

The tarball contains only two bundled JavaScript entrypoints and the transitive public declaration allowlist. Private uploader, internal composition, adapter-runtime state and wake modules are not physically shipped as standalone deep modules.

## Test plan

- [x] Final focused independent validation: 6 files / 113 tests repeated three times.

- [x] Full package suite: 33 files / 428 tests.

- [x] Production and test TypeScript checks.

- [x] Exact Biome on all 49 changed TypeScript/JSON/script files.

- [x] Architecture-boundary, Convex-path, read-bound and test-architecture checks.

- [x] Direct HTTP transport matrix, bounded streaming and cancellation tests.

- [x] File-backed claim, close/reopen, lease reclaim, stale claimant, migration-through-drain and coordinator-overlap tests.

- [x] Real generation commit-before-wake and no-new-entry wake suppression tests.

- [x] Two byte-identical builds.

- [x] Deterministic build across root and contracts-local topologies: 40 files, SHA-256 6ed81940f81cca805a3e2aaef1d125d1ff78476f679a225fabd6d056f8ea5e1a.

- [x] Serial pack/install/import/declaration/private-dispatch validation: 43 entries.

- [x] Pack SRI: sha512-KYo76D79vabGEZK5n/J5nSh15RtP01SRLTjgRK7/UrNMEhUhrQRkxryjzeJBKeQLuPyUztuf8o/kj/ZMubYbaA==.

- [x] Production archive scans: no source, tests, lifecycle hooks, test capabilities, native provider, secrets or private standalone runtime files.

- [x] git diff --check and clean worktree.

Full/root repository tests and pnpm test:root were intentionally not run.

## Independent review

The same independent reviewer audited each repaired range. Findings covering identity, lease/CAS behavior, payload retention, environment/path authority, public declarations, live coordinator composition, package importability, native-provider provenance, canonical roots, stream cancellation and duplicated probe contracts were repaired or explicitly rejected where they would invent authority outside this slice.

Final pre-rebase verdict: PASS on package tree 41f4be7f15c32f7132278dc4807fdde8833f6d4c.

After seven unrelated commits advanced main, the branch rebased cleanly. No upstream commit touched packages/add-aerie-skill; the subtree remained byte-identical. The same reviewer resumed and returned exact-head PASS.

Mercy then identified concrete repair classes covering terminal metadata reclamation, truthful settlement, stale-settlement, and deferred-release accounting, exact pre-RNG capacity projection, credential-authority, final-read, overlap-state, and foreground drain reporting, foreground/detached platform-root identity, recoverable per-account vault sequencing, and private-dispatch failure exits. The same author repaired them and the same reviewer returned exact-head PASS at 89fe7e570419be5c0524c59f5dd2504f8c2c4d7e after byte-identical package rebases onto current main. The non-blocking self-hash note remains documented as best-effort substitution detection, not external integrity authority.

## Risks and monitoring

- Detached wake is best effort. Process launch, machine shutdown or missing credential authority may leave work queued until another admission or bounded foreground drain.

- Authentication errors retain the stable record rather than deleting credentials from a detached process. A foreground command must reauthenticate before a later successful retry.

- Current native-provider absence is deliberate. Adding an unauthenticated local module would be worse than deferred delivery.

- The queue is best-effort telemetry: hard capacity, attempt and age bounds drop proof-bearing payload while retaining bounded terminal evidence.

## Rollback

Before publication, revert this PR to remove automatic wake and private uploader composition. If preserving state compatibility is required, revert the wake/dispatch and transport surfaces while retaining the queue schema and manual bounded drain. No package has been published, no host is registered, and no production state, deployment or network operation was performed by this rollout slice.

#1250 — feat(dashboards): add Real Estate tab to Data Health page @kevalshahtrilogy  approvedmercy-allow-critical

## Summary

Adds a third data-domain tab, Real Estate, to the existing centralized Data Health page at /sync (alongside Freshness, Data Quality, and Automation Outbox), giving the REBL3 real-estate dashboard a visible pipeline-health signal on the page users already check for other domains.

- Extends SyncSurfaceTab / SYNC_SURFACE_TAB_LABELS / SYNC_SURFACE_TABS with "real-estate".

- Fetches GET /api/sync/real-estate (backend route being built in parallel, PR #1249) using the same client-side fetch pattern as the existing Freshness tab.

- Renders: a verdict badge (OK/WARN/CRITICAL/UNAVAILABLE, tone-mapped to the platform's existing accent/amber/coral/stone palette), last-run status + relative timestamp, schedule (enabled + next run), and the list of open Surtr Observer findings (category, title, C/H/M/L severity badge, evidence, recommendation, occurrence count, first/last fired).

- Adds chat/lib/real-estate-health.ts: the shared RealEstateHealthPayload type (mirroring the accountability-freshness.ts pattern) plus pure helpers — normalizeRealEstateVerdict/normalizeRealEstateSeverity (unknown values fail safe instead of crashing or overstating severity/urgency, whitespace-trimmed), isRealEstateHealthPayload (a runtime type guard covering primitive shape, non-negative-integer counts, full ISO-date validation on every timestamp field, and cross-field invariants like "OK verdict with open findings" or a mismatched finding count — a malformed or self-contradictory payload fails closed into the unavailable state), and getEffectiveRealEstateVerdict (forces the badge to UNAVAILABLE whenever sourceUnavailable is true, even if the raw verdict string still says "OK").

- Every error path (failed fetch, non-OK response, malformed payload, or a sourceUnavailable payload's own error field) renders a fixed, generic message — never the raw upstream/backend diagnostic text, in the DOM or the console.

- Uses the platform's existing themed components/tone-map pattern throughout (no native <select>; matches how the Freshness/Data Quality tabs already style badges and stat cards).

## Rollout: gated behind a default-off flag

This tab depends on GET /api/sync/real-estate, which is a separate PR (#1249) landing in parallel. To make this PR safe to merge and deploy independently of that one, the entire tab — including its entry in the /sync tab bar — is gated behind NEXT_PUBLIC_REAL_ESTATE_DATA_HEALTH_ENABLED, which defaults to off:

- Unset (the state this PR ships in): the "Real Estate" tab button doesn't render at all, the redirect-guard effect bounces away from it if somehow reached, and the tab's dashboard component never mounts — so it never calls the (possibly not-yet-deployed) backend route. Fully inert in every environment.

- Set to the exact string "true": the tab appears and behaves as described above.

This mirrors the existing DOCUMENT_KNOWLEDGE_PROCESSING_ENABLED-style flag in convex/documentKnowledge/runtimeConfig.ts (same exact-string-match semantics, same default-off test pattern). It needs the NEXT_PUBLIC_ prefix because it's read from the /sync page's client component — Next.js only inlines NEXT_PUBLIC_-prefixed vars into the browser bundle, so flipping it requires setting it at Chat's build/deploy time, not a live runtime toggle.

Manual post-deploy step: after PR #1249's GET /api/sync/real-estate endpoint is deployed and confirmed live, set NEXT_PUBLIC_REAL_ESTATE_DATA_HEALTH_ENABLED=true in Chat's build environment and redeploy to actually surface the tab. Until that's done, this PR merging/deploying has zero user-visible effect.

## Business Value

Real-estate data (REBL3) currently has no visible freshness/health signal anywhere in Aerie — if the underlying Surtr pipeline silently fails or goes stale, nobody knows until someone notices bad data on the dashboard itself. This adds that signal to the page users already check daily for every other data domain, with an explicit "unavailable/unknown" state so a broken pipeline never gets mistaken for a healthy one. Low-risk, additive UI change (no changes to existing tabs' behavior), and now fully inert-by-default until explicitly turned on post-deploy.

## Manual Effort Estimate

AI-drafted estimate — flag for Keval to confirm/adjust: ~6–7 hours of focused human effort (reading the existing Freshness tab as a template and matching its exact styling/data-fetching conventions, designing the shared payload type + defensive runtime validation including cross-field invariants and diagnostic-leak prevention, building the tab UI and findings list, adding a default-off rollout flag matching the repo's existing flag pattern, writing and debugging the pure-logic and component tests, and getting typecheck/lint clean).

## Test Plan

- [x] pnpm typecheck (root, all workspace packages) — clean

- [x] pnpm lint (boundaries, convex-paths, read-bounds, test-architecture, biome) — clean

- [x] pnpm --filter chat exec vitest run lib/__tests__/real-estate-health.node.test.ts — unit tests covering verdict/severity normalization (incl. whitespace trimming), the sourceUnavailable-forces-UNAVAILABLE guard, the payload type guard (valid/malformed/wrong-type/contradictory cases, timestamp validation), and the rollout flag's default-off parsing

- [x] pnpm --filter chat exec vitest run "app/(main)/sync/__tests__/page.test.tsx" — tests covering: tab hidden by default (and the backend route is never fetched while hidden), tab visible once the flag is enabled, renders verdict/findings from a mocked fetch, shows the unavailable state (not a false-healthy badge, and never logs/renders the backend's raw error text) when sourceUnavailable: true, shows an explicit error state (not a stale/empty view, and never logs/renders a raw upstream response body) on a failed fetch

- [x] pnpm --filter chat test (full suite, 694 files / 10k+ tests) — run for final verification, all green

- Not run: no dev server per repo convention (CLAUDE.md) — this was verified via typecheck/lint/tests only; the backend route doesn't exist in this worktree yet, so end-to-end verification against the real endpoint is out of scope here.

Linear: (ticket pending)

#3742 — fix(spacex-valuation): reconcile September distribution @sanketghia  approved

## Summary

- Mark the 9-Sep-2026 SpaceX Day 90 tranche as actual.

- Pin the waterfall and schedule to the confirmed 8-Sep close of $153.47/share.

- Preserve the model-derived 1,370,260 net-share total and document the 3-share difference from the 1,370,257 actual distribution.

- Add regression coverage for actual pricing and updated waterfall tie-outs.

## Linear

KLAIR-3523

## Verification

- SpaceX feature suite: 15 files, 229 tests passed.

- Prettier and ESLint passed.

- TypeScript project check passed.

- Production build passed.

- UI verified at http://localhost:3001/spacex-valuation with backend on port 5001.

## Screenshots

<img width="1410" height="694" alt="image" src="https://github.com/user-attachments/assets/29852996-6cc8-49b3-b024-8e9fac613e2c" />

<img width="1428" height="563" alt="image" src="https://github.com/user-attachments/assets/ba633fe0-a261-478d-b32a-0fb6ae928f64" />

#2 — Establish Gate 0 local maturity baseline @sanketghia  no labels

## Summary

- Establishes a clean Gate 0 verification baseline for the local factory.

- Fixes cold-start-sensitive SQLite and credential-probe test budgets.

- Reconciles local-only maturity boundaries in the README and pilot runbook.

- Adds a least-privilege pnpm build policy for esbuild and pins pnpm 10.26.0 compatibility.

## Verification

- pnpm install --frozen-lockfile succeeds in a clean checkout and runs esbuild postinstall without --ignore-scripts.

- 74 test files / 1,200 tests passed.

- 9 browser tests passed.

- Build, lint, formatting, and diff checks passed.

Remote deployment and hosted infrastructure remain out of scope.

#1 — feat: add research-to-Linear intake workflow @sanketghia  no labels

## Summary

Adds the sibling Intake workflow that turns bounded research and repository evidence into an auditable, reviewable Linear ticket proposal without coupling it to the existing runTicket or runBatch workflows.

## Included

- Repository, artifact, URL, and operator-context intake inputs.

- Read-only Codex App Server research with investigator and critic findings.

- Bounded/redacted source artifacts, SQLite state, research events, fingerprints, and duplicate checks.

- Factory Console Intake UI at #/intake and #/intake/:proposalId.

- Explicit approval and separate Linear issue-creation gates.

- Auditable warning resolution with persistent reviewer/note records.

- Fresh research revision with additional context.

- Approval and issue creation blocked while warnings remain unresolved.

- App Server schema handling fixes and Linear token fallback.

## Live validation

- Created and reviewed an Intake proposal that produced KLAIR-3510.

- Ran a batch for KLAIR-3510 successfully and published [Klair PR #3738](https://github.com/AI-Builder-Team/Klair/pull/3738).

## Verification

- Full Vitest suite: 74 files / 1,200 tests passed.

- Docker-dependent secretless validation: 5/5 passed.

- Browser smoke suite: 9/9 passed.

- TypeScript build and ESLint passed.

## References

- Intake summary ticket: [AI-719](https://linear.app/builder-team/issue/AI-719/factory-intake-research-to-linear-ticket-intake-and-warning-resolution)

- Feature branch: codex/research-linear-ticket-intake

- Latest synchronized commit: fc410d9

#1277 — fix(portfolio): retire legacy opening date from MCP reads @benji-bizzell  approved

## Summary

- Remove the retired actualOpenDate field from MCP site details while preserving Milestone 9 completion dates.

- Cover conflicting legacy dates and missing Milestone 9 dates across MCP detail and list reads.

## Why

Release smoke for #1274 found that the UI and Agent used Milestone 9 correctly, but MCP site details still spread the legacy stored date into their response. A site opened in January 2026 also exposed an August 2025 opening date, and a site with no recorded Milestone 9 date exposed the same fallback. This closes the MCP gap in #1273's field retirement.

## Business Value

MCP consumers receive a consistent opening-date contract without a competing legacy value. Legacy storage remains available for the separately managed migration.

## Breaking changes

MCP getSite no longer includes the already-retired actualOpenDate response field. Consumers should read milestones.postOpen.completedDate and preserve a missing date.

## Test plan

- [x] Regression cases fail against the old projection and pass with the fix; 32 Convex MCP parity tests pass.

- [x] 16 MCP worker site-tool tests pass; repository lint passes.

- [x] Dev backing-query reads confirm both conflicting/missing-date fixtures omit the retired key and preserve Milestone 9 values.

- [x] Workspace typechecks pass across all ten packages.

- [ ] Hosted CI, including full tests and build.

#1253 — feat(dashboards): add hidden REBL3 Surtr-comparison experimental view @kevalshahtrilogy  approved

## Summary

Adds a hidden, capability-gated internal view that compares Surtr's Gateway mirror of REBL3's site inventory against Aerie's current production real-estate data, field by field — with zero changes to any production read path. This replaces the earlier approach (PR #1251, now closed) of repointing the production bulk-read path itself; per updated direction, that cutover isn't happening yet, and this PR instead builds the validation tool to build confidence in Surtr's data before that decision is revisited.

- New capability operations.realEstateExperimental.read (packages/contracts/src/capabilities.ts) — grantPolicy: "privileged", external: {api:false, mcp:false}, not added to any baseline capability set. It's invisible by default; only a role explicitly edited to include it (via the role editor) can see this view.

- chat/lib/rebl3-surtr-gateway-server.ts — a small, injectable-deps Surtr Gateway fetcher (GET /gateway/aerie-rebl3-sites, paginates on has_more, never throws) mirroring the exact pattern of the sibling REBL3 Data Health tab's rebl3-pipeline-health-server.ts (PR #1249).

- chat/convex/portfolio/activeSitesCohort.ts:listForExperimentalComparison — a new query returning the same cohort rows as production's listForDashboard, but gated behind the new capability via requireCapability — independently revocable from production real-estate access.

- chat/lib/real-estate-surtr-comparison.ts — pure join/diff logic over 19 fields shared by both sources (address, classification, score, zoning-adjacent fields, tags, etc.), with explicit handling for representational differences that aren't real disagreements: Surtr's JSON-string array columns, Redshift's timestamp format vs. ISO 8601, and the Redshift Data API's NUMERIC-as-string quirk.

- chat/app/api/real-estate/surtr-comparison/route.ts — Clerk-authenticated (401) and capability-gated (403) via the same hasAerieCapabilityKey pattern as the production real-estate-sites route. Production cohort fetch failures are a captured 502 (platformRouteErrorResponse); Surtr failures degrade gracefully to a 200 with surtrAvailable:false (every site reports productionOnly) — same graceful-degradation posture as the Data Health tab.

- Hidden page at /dashboards/real-estate/experimental (chat/app/(main)/dashboards/real-estate/experimental/page.tsx) — no nav entry, reached only by direct URL, shows an "Experimental" banner, and renders an expand-per-site table (reusing the existing RygBadge component) with the full field-by-field diff on expand.

Auth is defense-in-depth: the route checks the capability before calling anything, and the underlying Convex query independently re-checks it — so the data is protected even if something calls the query directly.

## Business Value

Lets the team validate Surtr's REBL3 data against production, side by side, with zero production risk — before ever committing to repointing the real read path. This directly de-risks a future cutover decision (the original goal of this initiative) by surfacing exactly which fields disagree and how often, using real production and Surtr data rather than spot-checks.

## Manual Effort Estimate

*AI-drafted estimate — flag for Keval to confirm/adjust.*

~3 days of focused engineering time (roughly 22–26 hours) for a developer with no AI assistance: understanding the existing capability/role-editor system and its test invariants, understanding the Next.js Clerk-auth + capability-check route pattern and the existing REBL3 Data Health fetcher to mirror, designing and implementing the join/diff logic (including the representational-difference handling, which requires knowing the Surtr warehouse schema's quirks), wiring the new Convex query + capability + route + hidden page together correctly, and writing the ~60 tests across 6 test files this PR includes (unit tests for the fetcher, the diff logic, the Convex query's dual capability gates, the route's auth/degradation paths, and the page's visibility gating).

## Test plan

- [x] pnpm typecheck — clean across all 10 workspace packages.

- [x] pnpm lint (boundaries, Convex paths, read-bounds, test-architecture, biome) — clean.

- [x] pnpm --filter chat test — all 10,320 tests pass across 697 files (18 pre-existing skips, unrelated), including:

- 15 new tests for fetchAllRebl3SurtrSites (pagination, all Gateway status codes, contract violations, mid-pagination failure propagation).

- 18 new tests for the pure diff logic (valuesMatch, row mapping, join/presence logic, payload summarization).

- 3 new Convex tests proving listForExperimentalComparison is gated independently of listForDashboard's capability.

- 8 new route tests (401/403 before touching any backend, happy path, real mismatch detection, Surtr degradation, captured cohort failure).

- 6 new page tests (loading, access-denied for no/wrong capability, banner + fetch when permitted, error + retry).

- Fixed one pre-existing role-editor test whose hardcoded Operations-group grant count (22) needed bumping to 23 for the new capability slot.

- [x] git diff origin/main --stat confirms zero changes under sync/ — the production REBL3 read path is completely untouched.

## Notes

- SURTR_GATEWAY_API_KEY needs to be provisioned in Secrets Manager before this view can show real Surtr data — until then it shows surtrAvailable:false (every site reports productionOnly), the same graceful-degradation pattern the Data Health tab uses. Added as a placeholder to .env.example.

- Granting operations.realEstateExperimental.read to an actual role/user is a manual follow-up via the role editor — intentionally out of scope for this PR, so the view stays hidden until someone deliberately turns it on.

- Linear: (ticket pending)

#1271 — fix(operations): accept nullable Surtr schedule expressions @benji-bizzell  approved

## Summary

- Accept null for Surtr's schedule expression while preserving its explicit enabled flag.

- Cover nullable schedules and rejection of missing or invalid expression types.

## Why

After configuring the production Surtr API key, REBL3 health returned 502 because Surtr supplied schedule.expression: null. Surtr's pipeline API supports that value, but Aerie required a string and discarded the otherwise valid health response.

## Business Value

Restore REBL3 pipeline health visibility for pipelines without a schedule expression.

## Test plan

- [x] Regression reproduced the production validation error before the fix.

- [x] 29 focused tests pass across the Surtr reader, health projection, and HTTP route.

- [x] Repository lint and architecture checks pass.

- [x] Chat typecheck passes via the commit hook.

- [ ] After deployment, verify the authenticated /api/sync/real-estate response against live Surtr data.

No migration or configuration change is required by this patch.

#1273 — fix(portfolio): use Milestone 9 as the canonical opening date @benji-bizzell  approved

## Summary

- Retire the derived Actual Open Date field and read Milestone 9’s completed date directly across Portfolio and agent views.

- Point DSS consumers to the existing canonical milestone endpoint and remove the deprecated API property.

- Add a dry-run-first cleanup migration and document the separate storage retirement step.

## Why

Actual Open Date incorrectly copied Milestone 8’s date, creating a competing value. Milestone 9 (Operating) is the canonical source; a missing completed date must remain unrecorded.

## Business Value

People and API consumers get a consistent opening date with clear provenance, without maintaining a derived proxy.

## Breaking changes

actualOpenDate is removed from response contracts. Consumers should read data.milestone.completedDate from GET /v2/portfolio/sites/{siteRef}/buildout/milestones/postOpen (requires operations.buildout.read), or milestones.postOpen.completedDate in v1. The optional storage slot remains until a separately authorized cleanup.

## Test plan

- [x] Repository lint/typecheck, affected Chat tests, contracts tests, and MCP site-tool tests.

- [x] Live localhost browser: conflicting legacy/Milestone 9 years resolve to 2026; a missing Milestone 9 date renders no year; Admin no longer offers the retired field.

- [x] 12 authenticated dev API checks across three sites: v1/v2 agree and omit the retired property; served DSS advertises the canonical mapping.

- [x] Temporary read-only key revoked; subsequent access returns 401. No site data edited.

#1272 — AERIE-1266: Complete the arrival cohorts — normalize the enrolled date and stop gating entry timing on current attendance @vvp-trilogy  approved

Closes #1266.

## What & why

The SIS mart_enrollment_dtl had 169 lifecycle rows (post-#1268 scope) in a cohort they could not have reachedwithdraw / mid-year-transfer-out / on-campus rows with no arrival cohort (first-day / mid-join) in the same year. Nobody leaves a roster they were never on, and nobody is on campus without having arrived. Two causes, both in int_enrollment_classification.sql:

1. Entry timing was gated on current attendancex_is_first_day / x_is_mid_join required is_attending_status, which excludes WITHDRAWN / TRANSFERRED, so a departure retracted the arrival that preceded it.

2. The enrolled date was never normalized — a blank enrolled_date failed the IS NOT NULL conjunct, and a date stamped after the session ended was read literally as a mid-term join.

## Changes

- dbt/macros/normalize_enrollment_date.sql (new) — mirrors the HubSpot report's macro: NULL / before-start / after-end all resolve to session_start_date. enrolled_date_normalized is derived once in base. The mart keeps publishing the raw enrolled_date.

- int_enrollment_classification.sql — added was_attending_status (enrolled set + COMPLETED + WITHDRAWN) and promoted the mid-year-transfer predicate to base as was_mid_year_transfer (identical logic, now feeding two consumers). Arrival flags gate on was_attending_status OR was_mid_year_transfer and read the normalized date; x_is_on_campus / x_is_starting_later still read is_attending_status. Added an explicit pre-start withdrawn_date guard so a departure dated before day one never joins the opening roster. Header rewritten to the current contract.

- assert_sis_enrollment_lifecycle_has_arrival.sql (new, error severity) — every lifecycle-cohort member has an arrival in the same year, with the one named pre-start-withdrawal exception. Expressed as a single-scan window aggregate over int_enrollment_cohort (semantically identical to the ticket's correlated NOT EXISTS, but one pass — see note below).

- yml / comment descriptions — mart enrollment_date, the two stg_sis_enrollment descriptions, the staging SQL comment, and the parity-harness header rewritten to state the current contract (no "previously" phrasing).

## Verification — a real dbt build ran against Redshift (sis-dev)

Full prefixed build: dbt build --select path:models path:seeds --vars '{pr_number: …}'PASS=282, ERROR=0, WARN=1. The single WARN is the pre-existing assert_admissions_pipeline_tenant_coverage (WARN by design, admissions pipeline — untouched by this PR). The new assert_sis_enrollment_lifecycle_has_arrival passes (10.8s).

### Orphan conservation

| Basis | Before | After |

|---|---|---|

| Ticket's stated basis (live mart less Alpha Anywhere (Homeschool)) | 169 | 3 |

| Same post-#1268 scope, OLD vs NEW logic isolated | 198 | 3 |

The 3 residuals are exactly the pre-start-withdrawal guard rows (2024 withdraw ×2, 2026 withdraw ×1), which the conservation test excludes by its named exception clause, so the test is green.

### Cohort deltas (this change only, isolated OLD-vs-NEW on the post-#1268 scope)

| Cohort | Movement |

|---|---|

| first-day | 2023 +2, 2024 +14, 2025 +102, 2026 +47, 2027 +4 |

| mid-join | 2016 +1, 2020 +2, 2021 +1, 2022 +6, 2023 +8, 2024 +4, 2025 +17 |

| on-campus | 2026 +9 |

| starting-later | 2026 −9, 2027 −4 (= −13; the exact rows that gain first-day) |

| graduating | 2026 +2 (expected — on-campus AND terminal_grade) |

| start-year-transfer-out | 0 (unchanged — the arrival gate excludes blanket TRANSFERRED) |

Scope note (per the ticket): these are on the post-#1268 scope (the campus-axis change is already in main). Where scope is invariant the numbers match the ticket exactly — on-campus +9, starting-later −9/−4, start-year-transfer-out 0, and the early-year mid-join gains (2016 +1 … 2023 +8). first-day and later-year mid-join are larger than the ticket's raw-spine table because the post-#1268 scope has a larger population (four physical brands added, the #1258 test-student rows dropped) — exactly the growth the ticket's Measurement Basis predicted.

### Acceptance-criteria checks (SQL against the built prefixed mart)

- Blank-date WITHDRAWN, no withdrawn_date (fixture enrollment 359bfaf1-85f1-4cef-8356-180c645a3543, Oliver Overton / Alpha Austin) → now in first-day (+ withdraw + re-enrolled). ✅

- No enrollment holds both first-day and mid-join → 0. ✅

- Withdrawn never on-campus → 0; CONFIRMED never in an attending cohort → 0. ✅

- start-year-transfer-out gains no arrival → 0. ✅

- enrolled_date > session_end_date reroutes mid-joinfirst-day (normalized to session start). ✅

### Parity refresh (SIS vs HubSpot, SY2026, 46 shared programs)

| cohort | hubspot | sis | diff |

|---|---|---|---|

| mid-join | 25 | 86 | +61 |

| re-enrolled | 504 | 523 | +19 |

| withdraw | 5 | 22 | +17 |

| mid-year-transfer-out | 1 | 15 | +14 |

| start-year-transfer-out | 21 | 34 | +13 |

| on-campus | 1444 | 1448 | +4 |

| re-enrollment-other | 0 | 2 | +2 |

| graduating | 92 | 87 | −5 |

| first-day | 1425 | 1399 | −26 |

| re-enrollment-declined | 92 | 62 | −30 |

| starting-later | 80 | 39 | −41 |

starting-later is now below HubSpot (39 vs 80) rather than inflated — the un-clamped-blank-date inflation the old parity header described is gone, confirming the header rewrite.

### Note on the conservation test's query form

The ticket's illustrative test SQL uses a correlated NOT EXISTS self-join on int_enrollment_cohort (a view whose single scan costs ~10s). Cold, that double-scan exceeds the adapter's 30s read timeout and ERRORs — reproduced locally, and it would fail CI's dbt job identically (same baked profile). The test here computes the same invariant — per-(enrollment_id, session_school_year) arrival presence, with the pre-start exception — as a single-scan window aggregate (MAX(...) OVER (PARTITION BY …)), which runs in ~10s, well inside the timeout. Semantically identical; a reviewer can confirm by reading the two forms side by side.

#1270 — AERIE-1268: Scope the enrollment campus axis to physical, active campuses across the four brands @vvp-trilogy  approved

Closes #1268.

Scopes the enrollment report's campus universe in stg_sis_campus so out-of-scope campuses never enter the DAG: keep only the four physical school brands (Alpha School, Sports Academy, GT School, Nova Academy) with delivery_mode = 'PHYSICAL' and status = 'ACTIVE', via an EXISTS semi-join against stg_sis_brand. The campus predicates come out of int_school_year_offering, which now keeps only the program_name = 'School Year' axis filter.

This is not a pure narrowing: it widens by three brands and drops the online campuses, and the two do not cancel.

## What changed

| Change | Model |

|---|---|

| Select delivery_mode; apply brand / mode / status predicate; header states the rule 5 exception + stg_sis_brand coupling | stg_sis_campus.sql |

| Carry delivery_mode forward | int_campus.sql |

| Remove campus predicates, keep only the School Year axis | int_school_year_offering.sql |

| Delete (tautological once the filter is in staging) | tests/assert_sis_enrollment_campus_allowlist.sql |

| Add source-side value-domain test (rule 21) | tests/assert_sis_campus_delivery_mode_domain.sql |

| Add not_null on campus_name; widen brand_name accepted_values to four brands | _int_enrollment__models.yml |

| Add five entering campuses; delete the orphaned Alpha Anywhere (Homeschool) row | seeds/sis_campus_program_map.csv |

| Restate descriptions to current shape (rule 22) | _sis__models.yml, _sis__sources.yml, stg_sis_brand.sql, _seeds__enrollment.yml |

## Verification

dbt build --select path:models path:seeds --vars '{pr_number: 1268}' run against Redshift sandbox_education: PASS=281, WARN=1, ERROR=0. The lone WARN is assert_admissions_pipeline_tenant_coverage (finalsite/admissions domain, pre-existing, unrelated to this change). All new/changed tests pass; the deleted allowlist test no longer runs. Prefixed objects dropped after.

Grain (rule 13 / Testing Notes): stg_sis_campus = 66 rows, exactly equal to the same predicate written as a plain IN list of brand UUIDs (66). The EXISTS did not change the campus grain; unique/not_null on id pass.

AC — online campuses contribute zero facts: Alpha Anywhere (Homeschool) = 0 fact rows in the mart; GT Anywhere is removed at staging (0 rows from stg_sis_campus onward).

## Before / after cohort counts (post-#1258 spine)

Re-measured on the post-#1258 spine (distinct enrollment_id fact rows in mart_enrollment_dtl), not the raw-spine numbers in the ticket:

| Offering year | Before | After | Δ |

|---|---|---|---|

| SY2024 | 286 | 284 | −2 |

| SY2025 | 2840 | 757 | −2083 |

| SY2026 | 1968 | 1625 | −343 |

| SY2027 | 23 | 22 | −1 |

Distinct campuses in the mart: 54 → 58.

Trend still inverts. SY2025 → SY2026 reads as −30.7% before (2840 → 1968) and +114.7% after (757 → 1625) — decline flips to growth, same as the ticket's raw-spine thesis (−9% → +188%).

### Test-student contamination in the new-brand campuses (Testing Notes)

The four entering brands were not contamination-free; the post-#1258 spine drops these (so the After counts above already exclude them). Raw School-Year enrollments vs. test-student enrollments:

| Campus | SY2024 (tot/test) | SY2025 (tot/test) | SY2026 (tot/test) |

|---|---|---|---|

| GT School | 13 / 0 | 35 / 1 | 72 / 15 |

| Nova Academy Austin | 24 / 0 | 52 / 0 | 36 / 20 |

| Nova Academy Bastrop | — | 1 / 0 | 2 / 1 |

| Nova High School Brownsville | — | — | 12 / 10 |

| Texas Sports Academy | 21 / 0 | 71 / 0 | 93 / 20 |

## Orphaned seed row: deleted

Alpha Anywhere (Homeschool) (row 1, has_hubspot_program = false) is deleted, not retained. Once it leaves the axis the row is unreferenced; the covers-axis test is axis→seed only, so nothing forced the decision. Deleting keeps the seed 1:1 with the axis (58 rows, one per axis campus — the seed's stated contract). Retaining it "against a future re-scope" buys nothing: per the ticket's Out of Scope, a future online axis needs its own campus model and cannot reuse this DAG or seed.

## Restated #1216 parity delta

The #1216 SIS-vs-HubSpot parity table was built over the old scope, whose SY2025 SIS population is exactly the number that moved: 2840 → 757 on the post-#1258 spine (SY2024 286→284, SY2026 1968→1625, SY2027 23→22). That invalidates the old parity table outright, as the ticket states. The faithful restated delta compares the new-scope SIS mart_enrollment_dtl against the replicated HubSpot mart sales_educrm_wh_mart_enrollment_dtl, which does not live in sandbox_education and is not reachable from the dbt warehouse; that recomputation is #1216's cutover work. The restated SIS side of that comparison is the After column above. #1216 is not blocked by this change, but it and this PR must not land without a joint re-measurement.

## Design notes carried from the ticket

- Filtering in stg_sis_campus is a knowing rule 5 exception (SIS does not disown online campuses; this is a report-scope filter). Stated in the model header. Applied once (rule 8); stg_sis_campus is the lowest shared model, and its whole downstream is this report.

- EXISTS (semi-join) cannot fan out, so grain stays 1:1 (rule 13 not engaged) and the predicate names brands rather than hardcoding UUIDs.

- EXISTS against stg_sis_brand (not the raw source) reuses that model's deleted_at + test-slug drop; the resulting staging→staging coupling is noted in the header.

- not_null on int_school_year_offering.campus_name pins the implicit INNER join (rule 12): relaxing it to LEFT would surface an out-of-scope offering as a NULL campus and fail the build instead of silently re-admitting it.

- Alpha World School is kept (SIS is the spine and marks it PHYSICAL), per the ticket's decision.

#1269 — feat(reconciliation): add dormant Aerie foundations @caina-barbosa  approved

## Summary

This PR is Phase 1 of 8 in the larger [AERIE-1893 — Map automatic LOI and lease document reconciliation](https://linear.app/builder-team/issue/AERIE-1893/map-automatic-loi-and-lease-document-reconciliation) project.

It adds the approved strict reconciliation contract and dormant Aerie storage/execution-state foundations tracked by [AERIE-1920 — Define the strict LOI reconciliation v1 contract](https://linear.app/builder-team/issue/AERIE-1920/slice-115-define-the-strict-loi-reconciliation-v1-contract) and [AERIE-1921 — Add dormant reconciliation storage and execution-state compatibility](https://linear.app/builder-team/issue/AERIE-1921/slice-215-add-dormant-reconciliation-storage-and-execution-state).

Production effect: dormant/additive. There is no scheduler, workflow start, reconciliation writer, or production caller. Merging initiates no external traffic or business-data mutation.

---

## Why

Later reconciliation phases need one strict proposal/citation contract and durable, scoped storage before any coordinator or mutation path can be added safely. This phase establishes those seams while preserving legacy document and portfolio projections and keeping all execution behavior inactive.

---

## Business Value

- Establishes a deterministic, auditable contract for evidence-backed Property Acquisition updates.

- Makes future reconciliation state and provenance durable without changing current portfolio behavior.

- Preserves legacy records by projecting absent execution state as unavailable rather than inventing values.

- Provides a resumable report-only readiness verifier for safe migration planning.

---

## How does it work

1. packages/contracts/src/reconciliation.ts defines and validates the strict aerie.loiLeaseReconciliation.v1 proposal, citation, and operation contract.

2. chat/convex/reconciliation/schema.ts adds the scoped reconciliation tables and indexes to the shared schema.

3. chat/convex/migrations/reconciliationFoundation.ts scans readiness resumably and reports incomplete or malformed state without writing business fields.

4. Portfolio and Rhodes schemas expose the optional agreementExecutionState compatibly: missing legacy values become null, while malformed present values fail closed.

5. No production caller, scheduler, workflow start, writer, deployment, or activation is introduced.

---

## Scope

### Included in this phase

- Strict reconciliation v1 contracts and semantic validation

- Additive reconciliation schema/storage

- Resumable report-only readiness verification

- Nullable, fail-closed agreementExecutionState projection

- Exact final diff paths:

chat/convex/_generated/api.d.ts

chat/convex/migrations/reconciliationFoundation.test.ts

chat/convex/migrations/reconciliationFoundation.ts

chat/convex/reconciliation/schema.test.ts

chat/convex/reconciliation/schema.ts

chat/convex/rhodes/schema.ts

chat/convex/schema.ts

chat/lib/__tests__/portfolio-sites.node.test.ts

chat/lib/portfolio-sites-contract.ts

chat/lib/portfolio-sites.ts

packages/contracts/package.json

packages/contracts/src/index.ts

packages/contracts/src/reconciliation.test.ts

packages/contracts/src/reconciliation.ts

### Deliberately excluded for later phases

- Sindri workflow-start compatibility and protected reads

- Reconciliation coordinator, scheduling, polling, and workflow execution

- Atomic Site/Property Acquisition mutation

- Evidence UI, REBL3 discovery/registration, activation, and historical expansion

- Deployment, environment changes, production operations, and REBL3/Rhodes writeback

### Rollout provenance

The three approved source commits were reconstructed in order onto then-current Aerie main. Two bounded same-author repair commits fixed missing-current-version readiness reporting and malformed execution-state rejection. Final reviewed head: c7dabc798bfc14797c11b6d73297d2293b78ae15; squash merge: 3b9b040a9c97897ad15b5989c2287e0d8f801fc2. The merged tree matches the approved head.

---

## Test plan

### Automated validation

- reconciliation contracts — 60 tests passed

- reconciliation foundation/schema and portfolio projection — 58 tests passed

- malformed execution-state regression — 51/51 passed

- readiness migration repair regression — 6/6 passed

- contracts and Chat typechecks — passed

- architecture boundaries, Convex paths, read bounds, and test architecture — passed

- changed-file lint/Biome — passed

- git diff --check — passed

- independent exact-range review and both same-reviewer repair reviews — PASS

- hosted CI and final Mercy gate — passed on the exact approved head

- exact-head diff scope — only the 14 paths listed above

### Time for Implementation

An engineer without AI assistance would likely need 2–3 weeks to trace the existing schemas and projections, specify and implement the strict contract, add compatibility-safe storage and migration coverage, rebase, address review findings, and validate the final merge state.

#1267 — fix(api): complete enrollment pagination and handle cleared dates @benji-bizzell  approved

## Summary

- Return complete enrollment student collections through bounded, encrypted cursor pagination, including historical Program aliases.

- Normalize cleared milestone dates so existing blank values no longer fail the entire API collection.

## Why

Release smoke testing found an enrollment aggregate of 233 students while the API could enumerate only 191: its source cap ran before cohort filtering. It also found a milestone collection returning 503 because a cleared completion date was stored as an empty string that the reader rejected.

## Business Value

API consumers can reconcile enrollment totals with student detail and read milestone progress reliably after dates are cleared.

## Test plan

- [x] 116 enrollment/API tests pass with the default timeout, including alias deduplication, sparse aliases, empty-page continuation, encrypted cursors and publication changes.

- [x] 58 portfolio workbench tests and 8 lifecycle contract tests pass, covering date clearing and legacy blanks.

- [x] Live dev pagination returns 233 unique Austin students across 100 / 100 / 33 rows.

- [x] Seven-lane adversarial review completed; confirmed pagination edge case addressed.

- [x] Production cursor-encryption configuration verified present without exposing its value.

The new enrollment index is deployed with the Convex schema; no data backfill is required. Existing enrollment cursors must be restarted after rollout. Production data and the production release branch have not been changed by this patch.

#1239 — fix(dev-local): start Next in dev-local:workers on Windows and macOS @vvp-trilogy  no labels

Closes #1238.

## Problem

pnpm dev-local:workers reached Convex functions ready! and then never started Next. Port 3000 stayed empty while Convex cron logs made it look like the app was running. Two cooperating defects, one fatal on Windows and one fatal on every OS.

## Defect A — --start was a platform-shell string with nested quotes (Windows fatal)

convex dev --start <command> hands the string to the platform shell (cmd.exe on Windows, /bin/sh on macOS/Linux). Worker mode passed one string containing four nested "..." child commands. cmd.exe re-pairs those quotes differently from POSIX — the first inner " closes the outer /c quote — so concurrently never received four commands and Next (plus often the projector) never started. /bin/sh preserves the quotes, which hid the defect on macOS.

Fix: --start now passes one unquoted node ../scripts/dev-local-workers.mjs. The new wrapper starts concurrently via spawn(process.execPath, [concurrentlyBin, ...args]), so each child command stays its own argv entry and no shell re-parses nested quotes. concurrently still shells each child string, but those strings have no nested quotes, so they are valid in both cmd.exe and /bin/sh.

## Defect B — frontend launcher execed npm under npx convex (every OS)

scripts/dev-local-frontend.mjs reused npm_execpath whenever it was set. Because --start runs under npx convex, npm_execpath points at npm's own CLI, and --filter is a pnpm-only flag, so node npm-cli.js --filter @bran/chat dev exited immediately on macOS and Windows alike.

Fix: the launcher reuses the parent execpath only when its basename is a JavaScript pnpm entry (pnpm.js, pnpm.cjs, pnpm.mjs); otherwise it invokes pnpm / pnpm.cmd directly. The empty-CLERK_SECRET_KEY guard behaviour is unchanged.

Both were needed: a perfect concurrently launch still failed on B, and A alone left Windows broken.

## Changes

- scripts/dev-local.sh — worker-mode --start is now one unquoted node command.

- scripts/dev-local-workers.mjs (new) — spawns concurrently with argv-entry child commands.

- scripts/dev-local-frontend.mjs — pnpm-only npm_execpath reuse via looksLikePnpmScript.

- scripts/dev-local.test.mjs — new tests: unquoted --start command, launcher ignores npm's npm_execpath; existing Clerk-guard / explicitly-enabled / cannot-start tests retained (fixture renamed to a pnpm.mjs entry).

- The full RCA write-up lives as a [comment on issue #1238](https://github.com/AI-Builder-Team/Aerie/issues/1238#issuecomment-5569071756).

- docs/windows-local-dev.md — refreshed the now-stale --start section and troubleshooting.

## Rejected alternatives (per the issue)

cmd-only ^" escaping (breaks macOS), Darwin/Windows code branches, dropping the Clerk-guard launcher, and fixing only A or only B.

## Tests

node --test scripts/dev-local.test.mjs — 18/18 pass, including the plan's node --test --test-name-pattern 'unquoted|ignores npm|Clerk guard|explicitly enabled|cannot start' subset. pnpm biome check clean on touched files.

#1265 — feat(add-aerie-skill): add dormant adapter runtime and artifact lifecycle @caina-barbosa  approved

## Summary

- Add the dormant common adapter runtime and artifact lifecycle needed by later Claude Code, Codex and Pi adapter slices.

- Keep admission, identity, policy, filesystem ownership, receipt creation and queue settlement under one core-owned local state coordinator.

- Keep @aerie/add-aerie-skill private and unpublished, with no registered observer, uploader, network wake path or native filesystem claim.

## Why

Later host adapters need one trustworthy R10/R11/R12 boundary before they can observe host activity. This slice provides that boundary without activating any host integration.

The core—not an adapter or caller—derives artifact identity, policy, relevant paths, immutable installation lineage, receipts and queue writes. Corrupt, stale, foreign or ambiguous state fails closed rather than being adopted.

## Business value

This makes later telemetry adapters small and constrained: they can submit a bounded observation, but they cannot choose identity, policy, paths, reducers, receipts or persistence behavior. It also establishes crash-safe local recovery before any adapter is enabled.

## Slice and scope

- Slice: AERIE-1856 — telemetry rollout 10/19

- Strict predecessor: AERIE-1855 / fd588bfb874b706563f6f56dc0333f83cfb5fdf0

- Package: @aerie/add-aerie-skill

- Included surfaces:

- packages/add-aerie-skill/src/adapter-runtime/**

- packages/add-aerie-skill/src/adapter-artifacts/**

- versioned local-state migration and coordinator integration

- package exports, bounded reads, documentation and package-local tests

## Runtime design contract

### Admission, launch and recovery

An observation is first committed and fsynced in adapterObservationInbox. A known committed generation returns durably_admitted; known pre-commit failure returns not_admitted; uncertain commit truth returns indeterminate.

Reducer launch happens only after durable admission. A new pending observation has:

- terminalOutcome: null

- leaseStartedAt: null

- invocationReceiptId: null

That inbox record is the durable retry record. A reducer either atomically removes it while creating the stable receipt and exact queue entry, or leaves it pending without a receipt.

Detached launch is deliberately best effort. If launch fails, core makes one non-waiting transaction to record reducer_spawn_failed and apply the verified current bundled launcher policy mapping. If the sole coordinator cannot accept that secondary write, admission is not downgraded and no second state authority is created: the unterminal inbox remains pending, the failure is surfaced through one fixed content-free stderr line, and a later admission or explicit install/update retries the verified reducer.

The failure outcome and persistent host status have different authorities. Core may record the factual reducer_spawn_failed outcome without a status change, but it may apply needs_repair only when the exact current source-bundled policy row, bytes, configuration and managed-host binding all verify. It never guesses a fallback status or reason from an unavailable policy. Current policy health is reported separately by lifecycle inventory as missing, corrupt or needs_repair; the raw host status is only the last authorized policy effect and is not a current-health proof.

This separation is intentional. Waiting after committed admission, adding a parallel failure journal, inventing a default policy effect, or introducing a daemon would violate the bounded fail-open, source-policy and single-coordinator contracts. A pending inbox is not automatically needs_repair, because healthy successfully launched work is also pending until the reducer acquires its lease.

### Artifact lifecycle

The lifecycle uses immutable generations, marker-bound physical ownership, journaled recovery, verified current and rollback policy inventory, bounded runtime probes, and atomic receipt/queue settlement.

Production mutation has no pathname check-then-rename/delete fallback. Without a qualified platform capability it returns the typed content-free platform_unqualified error before filesystem or state mutation. Native Linux and Windows implementations and active-race qualification belong to AERIE-1858 and AERIE-1859.

## Explicitly out of scope

This PR adds no:

- Claude Code, Codex or Pi parser, fixture, observer or registration;

- uploader, telemetry network call, wake service, timer, daemon or startup sweep;

- host configuration mutation or real-host campaign;

- public reducer or adapter-selected path, policy, identity, receipt or queue;

- package publication, registry mutation, deployment, production API call or UI;

- native Linux or Windows atomic capability claim;

- generated Convex drift.

Test-only capabilities and race boundaries are excluded from production dist, declarations, exports and packed contents.

## Behavior and production effect

Dormant. Ordinary package import, startup, login, logout and installation do not invoke the runtime. No host adapter is registered. Unsupported lifecycle mutation has zero filesystem and state effects.

The package remains private: true, UNLICENSED, unpublished and dependency-free. AERIE-1860 owns any later public metadata or publication after native qualification and separate authorization.

## Test plan

- [x] Package suite: 27 files, 290 tests.

- [x] Admission coverage: durable inbox, launch failure, successful diagnostic/status settlement, failed diagnostic-write observability, queue capacity before RNG, stable receipt, exact queue record, tombstone/dedupe and exact/post-deadline behavior.

- [x] Lifecycle coverage: non-empty canonical predicate digest, bundled-source authority rejection, ownership, corruption, migration, recovery, physical races, policy inventory and zero-mutation platform gating.

- [x] Production and test TypeScript checks.

- [x] Biome, test-architecture and architecture-boundary checks.

- [x] Deterministic build twice: 112 files, SHA-256 9a2f60739bc95ef68e4709c776e2b9fc56d29e02a78cf70d16efe431825da70c.

- [x] Serial pack validation: 115 entries; SRI sha512-+METYodrBqyJNiu2tL9wKhwedbcrBZ/JTe6/XTZpsY6pBjPBt1jDDjan4ANo8isJjKhEPs0nLiKTkK6piKeRaw==.

- [x] Production-dist and tarball scans: no test hooks, test support, test files or test declarations.

- [x] Independent exact-head review at 98ee768f63f61be4a3e46c61d6b849a474f77b0e: PASS.

- [x] git diff --check.

Full/root repository tests and pnpm test:root were intentionally not run. Manual QC is not required for this dormant automated slice; native and real-host campaigns remain assigned to later slices.

## Review size

The exact binary diff is 599,677 bytes, below the 600,000-byte review limit. Low-value Cartesian permutations, type-shape checks and harness self-tests were removed while retaining the acceptance categories listed above.

## Risks and monitoring

- A launch failure can lose its secondary status only when the same coordinator cannot accept the follow-up write. The primary inbox remains durable and retryable; one fixed content-free stderr line surfaces that event.

- Markerless earlier adapter state migrates to explicit needs_repair; ownership is never inferred from live bytes.

- Mutation remains unavailable until native atomic capabilities are qualified.

- The runtime is unreachable from hosts in this slice, so it creates no observer or telemetry traffic.

## Rollback

Revert this PR. No package is public, no adapter is registered, and no production migration, deployment or network operation is performed.

#1783 — feat(education): document temporary Alpha Summer Camp staging load @ashwanth1109  changes requested

## Summary

- Add the temporary, manually run Alpha Summer Camp actuals and budget loaders under staging_education_googlesheets.

- Keep each table DDL beside its manual loader so the temporary table definition and workbook translation remain traceable together.

- Retain the exact XLSX and publication manifests in immutable S3 and record successful publications in the staging ingestion ledger.

- Document the temporary workflow assumptions around intentional workbook blanks, manually reviewed metric domains, and the Object-Locked landing-bucket contract.

## Temporary Scope

This is a Finance-requested bridge for the current workbook, not a Surtr-managed recurring pipeline. The comments identify where the eventual authoritative pipeline must introduce normal deployment registration and unattended source-contract validation.

This PR replaces #1767 so Mercy can perform a fresh review against the clarified scope.

## Business Value

Provides Finance with an auditable temporary source for Summer Camp actuals and budget data while preserving workbook provenance and making the eventual pipeline replacement boundary explicit.

## Implementation Effort

Estimated 1-2 engineer-days for an engineer implementing the loaders, DDL, immutable landing, lineage, live validation, and documentation manually without AI assistance.

## Linear

- [SURTR-1131](https://linear.app/builder-team/issue/SURTR-1131/move-alpha-summer-camp-workbook-loads-to-staging)

## Validation

- Runner test suite: 19 passed.

- Ruff check and format check pass for the three changed Python files.

- Actuals dry run: 783 rows from 27 camps.

- Budget dry run: 1,026 rows from 27 camps, including 74 source-faithful blank values.

- Prior live staging validation found zero row differences against the former mart tables for 2026-08-31.

- Prior validation confirmed successful ingestion-ledger entries, immutable S3 version IDs/Object Lock retention, and cleanup of temporary COPY files.

#122 — feat(review): relax the blocking bar on late rounds, not on every review @kevalshahtrilogy  changes requested

## The rule change

Today — identical on every round:

> Block if any finding is security, critical_bug, silent_bug or missing_tests at confidence ≥ 85.

After:

> Rounds 1–3: unchanged.

> Round 4 onwards: block only on security, critical_bug, severity: critical, or confidence ≥ 93.

## Measured, on 2,454 PRs / 5,478 reviews

blocked reviews that now APPROVE : 328/1326 = 24.7%

of those on 4+ round PRs : 272/1017 = 26.7%

rounds 1-3 verdicts identical to today : 2,688 findings, 0 differ

severe findings relaxed at ANY round : 0

54% of PRs finish in one round; the 14% that take four or more consume 42% of every review run. A flat threshold spends its leniency on the majority that were never a problem — this spends it only where a PR has already been read three times.

## Why not the alternatives (all measured first)

- Demoting missing_tests moved 17.5% overall but 0% on the PRs that actually hurt. [Aerie#1253](https://github.com/AI-Builder-Team/Aerie/pull/1253) ran 16 rounds and every blocking finding was a silent_bug — the demotion would have saved it nothing. It also waves through a real conf-99 finding: *"these tests use the excluded *.test.js pattern, so Vitest never discovers them"* — dead coverage you believe you have.

- A silent_bug floor cannot be tuned. The model clusters confidences on round numbers — 166 blocking findings at exactly 85, 185 more at 88 — so a floor of 86 flips 24.6% of reviews and 89 flips 35.3%, with nothing in between. It is also the category where a miss is a wrong number nobody notices.

- A self-reported defect_kind ("reachable" vs "defensive") was considered and rejected: it is the model grading its own homework, cannot be measured before it ships, and would most likely come back 95% "reachable".

## Why the round number is the right signal

It is not a proxy for anything — it is the thing itself. By round 4, mercy has already reported everything it found in rounds 1–3 and the author has worked through them. What remains is the tail it only reached after three passes. A defect that matters has had three full rounds to gate on its own merits.

On Aerie#1253 the late rounds were pure argument: gateway-server.ts:161 re-raised at conf 86 across rounds 8–11, and :125 at 90–92 across rounds 12–16 — the same unresolved disagreement about whether Gateway count is page-local. Those stop holding the merge. The conf-94 *"malformed values become null so valuesMatch reports agreement"* still blocks, at every round.

## Safety

- Round 1 is exactly as strict as today — proven, not asserted: 0 of 2,688 early-round findings change verdict.

- security, critical_bug and severity: critical gate at every round, forever, at the global floor.

- An absent or zero round holds the EARLY bar, so a failure to derive the round can only ever be stricter, never looser.

- Relaxation moves the gate, not the information. Relaxed findings are still extracted, still rendered, and still carried in the open-items ledger — which deliberately ranks by the strict rule so they survive its size budget and the next round still sees them.

## Plumbing

The round already existed: open_items computes len(reviews) + 1 and the workflow already passes --review-round, used until now only for max_review_rounds. No new API calls, no new workflow steps.

Disable with relax_from_round: 0; tune via late_round_confidence_floor and always_blocking_categories.

## Testing

389 tests pass. Nine new; four mutations killed:

| mutation | killed by |

|---|---|

| relax from round 1 (would relax every review) | 2 |

| unknown round relaxes instead of blocking | 5 |

| severe categories relax like everything else | 2 |

| late floor equal to the global floor | 4 |

A fifth — dropping relaxed findings from the ledger — is unkillable by design, because the ledger ranks round-unaware on purpose. Documented in the test rather than left silently green.

> Worth flagging for anyone mutation-testing this repo on macOS: system Python sets sys.pycache_prefix to ~/Library/Caches/com.apple.python, so bytecode lives outside the repo. rm -rf __pycache__ in the tree does nothing, and a stale .pyc will happily report a config value the source no longer contains. Clear that path between mutations or results are meaningless.

## Business Value

Review turnaround gates the stacked-PR workflow — one unapproved PR blocks everything behind it. This targets the 14% of PRs consuming 42% of review effort, without touching the first look at any PR, and without relaxing anything where a miss is expensive.

## Manual Effort Estimate

~4 hours (proposed — Keval to confirm). The code is small; the work was measuring three rejected alternatives against real telemetry before finding one that helps the PRs that actually hurt.

Supersedes #121, which is closed.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1788 — [SURTR-1175] Build and live validate Aerie retention refresh @ashwanth1109  changes requested

## Linear

https://linear.app/builder-team/issue/SURTR-1175/build-and-live-validate-aerie-retention-stored-procedure-and-scheduled

## Summary

- add the source-pinned, atomic mart_education.sp_refresh_aerie_retention procedure and its three documented mart contracts

- add the 900-second mart-aerie-retention-refresh Lambda pipeline, daily 00:30 UTC schedule, first-failure alerting, ownership, and registry metadata

- add independent reconciliation, idempotency fingerprinting, rollback-fixture validation, catalog verification, and focused unit/contract tests

- pass scheduled parameters through the generic Lambda EventBridge target so the fixed report window reaches the platform state machine

## Business Value

Publishes a governed Alpha Anywhere retention dataset for the 2025-06-01 through 2026-05-31 report window. Consumers get learner drilldown, complete campus-month reporting, cohort retention, explicit exclusions and quality flags, and traceable SIS lineage without rebuilding spreadsheet logic manually. Atomic publication, source fencing, alerts, and validation prevent partial or silently stale reporting.

## Implementation Effort

Estimated 5–7 working days (40–56 engineering hours) for an average engineer to hand-code the warehouse contracts and procedure, build the platform runner and schedule, implement the independent validators and tests, deploy the isolated candidate stack, and complete production reconciliation and failure testing.

## Live Validation

- origin/prod does not exist, so the candidate was based on the remote default production branch, origin/main at 4af8a958.

- Applied and catalog-verified all source-controlled DDL: exact 64-column contract, all keys/sort keys/comments, AUTO distribution, invoker procedure ownership/version, reader grants, and no reader write/mutex grants.

- Deployed only Pipeline-mart-aerie-retention-refresh-prod with --exclusively; CDK reported deploying... [1/1]. The schedule was disabled for candidate validation, then enabled after all checks. CloudFormation completed without failed or rollback events.

- Preflighted accepted SIS run fd9aaed2-9655-4668-963c-347e87674671: 23,158 raw detail rows from one run, a unique 20,673-row current view, and 5,243 scoped learners across two campuses.

- Two platform runs succeeded (756def9c-4faf-427c-8bc1-5d76c993a975, 48c9e18d-8b14-4e03-8b89-79381cfe1b20) and published 5,243 learner, 24 campus-month, and 156 cohort-month rows.

- Independent source/detail, monthly, cohort, key, lineage, rate, quality, and report-total reconciliations all passed. The 5,423-row business fingerprint matched exactly across runs: d58be4c043a0efec72be1a14aed351ab2b3081d4e4a933eed27d8e1b8b0f9d92.

- An invalid source was rejected before the procedure; all three targets retained the second-run fingerprint and the first-failure notification path fired.

- An isolated three-table runtime-failure fixture proved all old snapshots survive transaction rollback without modifying production source or mart rows.

- The final EventBridge rule is enabled at cron(30 0 * * ? *) with the exact report dates. The dedicated PipelineRegistry-prod dependency stack completed separately, and the live active registry row has the correct owner, cron, 900-second timeout, on-demand control, and first-failure alerting.

- Full statement IDs, execution ARNs, counts, quality evidence, ratios, and schedule state are recorded in the pipeline README.

## Testing

- Python focused suite: 39 passed

- CDK pipeline construct Jest suite: 29 passed

- pnpm build

- Ruff format/check on the changed Python files only

- focused CDK synth with scheduled-parameter assertion

- git diff --check

- production DDL/catalog verification, two-run reconciliation/idempotency, invalid-input preservation, and atomic rollback fixture

#116 — feat(mercy): summon heimdall to fix conflicts alongside the findings @kevalshahtrilogy  no labels

## What

When mercy requests changes on a heimdall-driven PR and the PR conflicts, the hand-off comment now says so:

> @heimdall address the review findings. This PR also conflicts with main — merge the base branch in and resolve the conflicts as part of the same pass, so the PR is mergeable once the findings are addressed.

## Why

heimdall already knows how to resolve conflicts — Keep approved PR mergeable in heimdall.yml posts a self-summon on CONFLICTING. But that step hangs off the APPROVE path, so it only fires once a PR has no findings left.

A PR that conflicts *while it still has findings* therefore sits unmergeable through every revise round, and gets its conflict fixed only after the last one — if it hasn't rotted first. That is exactly how the three approved agent PRs in the Surtr 2026-08-11 audit decayed into conflicts, which is what motivated the approve-path step in the first place.

Conflicts block the merge whatever the verdict, and on a REQUEST_CHANGES hand-off the agent is about to edit the branch anyway. One pass fixes both.

## Why here rather than in heimdall

heimdall's pull_request_review gate sets ACTION=automerge only on approved; every other review state exits, because the revise loop is mention-driven. Adding a second trigger there would mean a new event surface and a second comment.

mercy already posts the hand-off comment. Making it *say more* needs no new trigger, no extra comment, and no additional review run.

## Scope

Opt-in is unchanged and remains the hand-off's own — an agent/ branch, the heimdall-driven label, or MERCY_HANDOFF_ALL_PRS. A PR nobody asked heimdall to drive is never summoned.

Two behaviours worth stating explicitly:

- Only an explicit CONFLICTING counts. GitHub computes mergeability asynchronously and answers UNKNOWN while it works. Treating "not MERGEABLE" as a conflict would summon the agent to resolve nothing on every review that raced the computation.

- The conflict text is additive. A conflicting PR still has findings; the branch must not replace the instruction that carries them.

## Testing

380 harness tests pass. Four new contract tests, four mutations, all killed:

| mutation | why it matters |

|---|---|

| drop MERGEABLE from env: | the HEAD_SHA regression — a ${VAR} read without an env: declaration once failed every review under set -u, with an error indistinguishable from a real finding |

| drop mergeable from --json | gh pr view --json returns only requested fields; omitting one yields a silent empty string |

| infer conflict from != MERGEABLE | fires on UNKNOWN |

| conflict branch replaces the findings text | loses the findings |

## Business Value

Removes a full round-trip from every conflicted agent PR, and closes the rot window that produced the audit this behaviour was built for. On a stacked-PR workflow one unmergeable PR blocks everything behind it, so a conflict discovered at approval time rather than at first review costs the whole stack the delay.

## Manual Effort Estimate

~1.5 hours (proposed — Keval to confirm). Small diff; most of the work was establishing that heimdall's existing summon was approve-gated and that mercy's hand-off was the cheaper place to fix it.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1263 — test(feedback): stabilize intake rate-limit window coverage @benji-bizzell  approved

## Summary

- Control the feedback HTTP test clock and verify intake limits, free replays, and recovery at the next minute boundary.

## Why

The [failed main CI run](https://github.com/AI-Builder-Team/Aerie/actions/runs/34254084880/job/102155365416) expected HTTP 429 but received 201. The test used wall-clock time while the limiter uses fixed minute windows; its CI execution ended at 16:58:00. Advancing the clock from 16:57:59.999 to 16:58:00 between the accepted burst and overflow request deterministically reproduced the same failure locally. The subsequent release CI run passed, consistent with the intermittent boundary condition.

Freeze Date.now for the exhausted-window assertions, restore the spy after each test, and explicitly advance into the next minute to verify recovery and unchanged replay accounting. Runtime behavior is unchanged. This supports pre-release validation for #1262.

## Business Value

Prevent intermittent CI failures from blocking releases while retaining coverage of feedback rate limiting and idempotency.

## Test plan

- [x] Reproduce the original 201-versus-429 failure with a controlled minute transition.

- [x] Exact HTTP test: 4 passed; related feedback tests: 19 passed.

- [x] Full pnpm test: passed, including 696 Chat files (10,314 passed, 18 skipped) and 128 root script tests.

- [x] Repository pnpm lint and pnpm typecheck: passed.

- [ ] Confirm patch PR CI passes before merging into main and refreshing release validation.

Local validation used Node 22.23.0, pnpm 10.29.2, and NODE_OPTIONS=--max-old-space-size=4096. Full tests used TMPDIR=/private/tmp: the macOS default /var/... temp path contains a symlink rejected by the installer security tests. With the canonical temp path, all installer tests passed. CI uses Node 20 on Linux; fresh patch CI remains the parity check. No build rerun for this test-only change, and no smoke validation, merge, or deployment performed.

#1789 — fix(education): prevent billing estate refresh throttling @benji-bizzell  approved

## Summary

- Pace billing requests at 0.25/second and tenant login starts at two-minute intervals, matching the successful full-estate local validation.

- Honor usable Retry-After delays within a shared 15-minute cooldown budget; stop and account for unattempted tenants when safe recovery is unavailable.

- Record the 59-tenant validation findings and remaining production validation steps in the pipeline documentation.

## Why

Faster estate traversal triggered Finalsite throttling. The client also capped server-requested delays at 120 seconds and advanced to another tenant after exhausted throttling retries, potentially continuing against a shared limit. All 59 tenants passed a full local extraction at the new pace, with zero 429s across 268 request operations in 1h58m.

## Business Value

Reduces avoidable billing refresh interruptions while preserving explicit failure and coverage reporting when the source requires a longer pause.

## Test plan

- [x] Billing pipeline suite: 255 tests pass; targeted Ruff lint/format and diff checks pass.

- [x] Full live-source local extraction: 23,264 items / 47,062 allocations; numeric reconciliation and 520 null-contact allocations pass.

- [x] Simulated login pacing, full Retry-After handling, cumulative cooldown exhaustion, missing/invalid headers, response closure, and stopping the estate after throttling.

- [ ] Deploy and validate full production extraction/publication before enabling the schedule; no DDL or schedule activation in this PR.

#1784 — fix(qtd): match school names for guide headcount (SURTR-1161) @ashwanth1109  changes requested

Replace exact XO teamroom equality with a generic school-name match for QTD Guide headcount. A school identity such as EDU.School.Alpha.Scottsdale accepts Campus, L2, L3, LL and future subteams without new configuration. Match at a dot boundary and require one canonical school, so Boston does not accept Boston Suburbs and different brands/campuses remain separate.

## Business Value

School staffing counts and student:guide ratios include named school subteams while keeping all school-attributed payroll. On the September 8 source population, Scottsdale expects 2 Leads + 5 Guides = 7 staff and $179,352.98 payroll across all 9 contributors. The user confirmed excluding Holly in Central Hybrid; Morgan in GT Anywhere is also excluded from staff counts. Both remain in payroll.

## Implementation

- Derive school identity prefixes from canonical school names and the existing expected-school spelling translations. Resolve against all active schools, including schools without QuickBooks payroll; require exactly one distinct school matching the posting's school.

- Preserve latest populated current-quarter XO evidence, identity/team ambiguity checks, protected distinct-person counts, role allocation, payroll sums and the 77-row per-school contract.

- Remove the rejected subteam reference/seed from the implementation. Migration 008_qtd_guide_school_name_policy.sql updates metadata and coordinates Guide procedure/runner policy teamroom_v3; it does not drop the already-deployed v2 reference table.

- Add executed production-SQL regression tests and checked-in source evidence, coverage limitations and rollout instructions.

## Validation

- 134 pipeline tests passed, plus the 8 prototype tests; changed-file Ruff lint/format and diff checks passed.

- Tests cover both historical Scottsdale source populations, future subteams, aliases, school/brand collisions, schools without QB mappings, latest populated evidence, transfers, ambiguous/missing identities, role and person-count safeguards, negative payroll, NULL/zero semantics and migration comment parity.

- Read-only validation results are recorded in the [investigation evidence](https://github.com/AI-Builder-Team/Surtr/blob/codex/qtd-guide-headcount/pipelines/runners/mart-aerie-education-financials-refresh/investigations/2026-09-08-school-name-match/evidence-summary.json).

## Coverage Limits

The generic rule changes counts at 33 schools on the inspected source population. Austin CO2/Spyglass, Greenwich, Woodlands, Montessorium Austin and TSA Lakeway still need school-level identity clarification; their existing exclusions remain. In particular, CO2 has payroll at both Alpha Austin and Alpha High Austin. No unsupported city-wide or shared-team mapping is added. Aerie's literal teamroom badge remains a separate consumer follow-up.

## Native Stack and Live Validation

Upper layer of native GitHub stack #1786: [#1769](https://github.com/AI-Builder-Team/Surtr/pull/1769) → this PR. This branch contains the parent revenue-proration changes.

Candidate 3e120e55 is deployed and validated before merge. Each stack below was deployed separately with --exclusively; each CDK run showed [1/1] and CloudFormation finished UPDATE_COMPLETE. The reviewed changes were confined to Lambda code and CDK metadata. No other pipeline stack was deployed.

- Pipeline-mart-aerie-education-financials-refresh-prod: v3 runner and atomic Guide migration/procedure cutover. Pipeline run c42c9f9f-5588-473f-bbd8-68d53cb9fa0e succeeded; Guide publication 2026-09-08 15:32:25 UTC.

- Verified 4,235 Guide rows / 55 schools: exactly 118 expected actual count/ratio changes at 33 schools, with zero monetary changes. Scottsdale is 7 staff, ratio 10.142857, and $179,352.98 payroll. The 4,180 All Other HC and 9,680 Facilities rows retain every measure.

- Restored #1769's missing runner metadata checks in Pipeline-mart-school-performance-unit-economics-refresh-prod and Pipeline-mart-school-performance-table-3-refresh-prod; both existing procedures already matched the parent. Deployed source files were checked against the candidate to preserve other installed changes.

- Table 2 run da7faded-68e7-4c91-9a59-daa8d3225c7e and its normal EventBridge-triggered Table 3 run 916887a4-2c51-46e4-915b-1589057466a4 succeeded. Both marts retain 1,045 rows / 55 schools, with zero monetary/percentage changes, complete proration metadata, current Program-directory lineage, and reconciled per-student amounts.

- Revenue timing is preserved: fraction 0.1266666667, cost fraction 0.1917808219, policy school-performance-revenue-10-month-v1 for September 8. Scottsdale modeled tuition is $359,733.33, profit −$175,392.70, EBITDA −$164,499.55.

- Parent tests: 36 Table 2 + 13 Table 3 passed. Full [live validation evidence](https://github.com/AI-Builder-Team/Surtr/blob/codex/qtd-guide-headcount/pipelines/runners/mart-aerie-education-financials-refresh/investigations/2026-09-08-school-name-match/evidence-summary.json) is checked in. The unused v2 reference remains intact for rollback; no role-effectivity backfill or table drop was performed.

Both PRs remain unmerged. Merge the native stack in order so future production releases retain both changes.

## Linear

Fixes [SURTR-1161](https://linear.app/builder-team/issue/SURTR-1161/fix-scottsdale-qtd-guide-headcount-for-accepted-xo-teamroom).

## Implementation Effort

Estimated 2–3 engineer days without AI assistance for source investigation, school-identity matching, SQL regression coverage, migration, isolated deployments and full live warehouse reconciliation.

#1787 — fix(timeback): contain timeback-raw-sync overruns and finalize timed-out runs (SURTR-1162) @caina-barbosa  approvedmercy-allow-critical

Contain timeback-raw-sync runtime overruns by temporarily expanding the execution envelope from 6 to 8 hours, passing task-level timeouts to ECS RunTask so States.Timeout can execute terminal bookkeeping before parent state-machine expiration, and adding an EventBridge fallback reconciliation rule for abnormal parent state machine exits.

## Business Value

- Prevents the pipeline dashboard and Redshift registry from indefinitely reporting dead/timed-out executions as RUNNING.

- Gives timeback-raw-sync temporary headroom (8 hours) while a separate design addresses daily full-snapshot performance drivers.

- Ensures operators receive truthful terminal state (TIMEOUT / FAILED) and prompt alert notifications when work exceeds its allowed duration without manual Redshift reconciliation.

## Implementation

- 8-Hour Allowance: Updated pipelines/runners/timeback-raw-sync/pipeline.json from timeout_hours: 6 to timeout_hours: 8.

- Catchable Task-Level Timeout: Passed timeoutMinutes: timeoutHours * 60 into createEcsTask('EcsRunTask', ...) for standard ECS pipelines in pipelines/cdk/lib/constructs/ecs-step-function.ts. Synthesizes TimeoutSeconds = 28800 (8h) on EcsRunTask and TimeoutSeconds = 29100 (8h 5m) on the parent state machine, ensuring the States.Timeout -> UpdateRunTimeout -> PipelineTimedOut path executes on task overrun.

- Terminal Event Fallback Reconciliation: Extended TerminalExecutionReconciler in pipelines/cdk/lib/constructs/ecs-pipeline.ts to all ECS pipelines (orchestrated and standard). Configured SQS DLQ (pipeline-<id>-terminal-dlq-<env>) with CloudWatch alarm on delivery failures, and forwarded Step Functions stopDate into the target payload.

- Stop Date Support & Idempotency: Updated update-run-failed Lambda handler (pipelines/cdk/lambdas/update-run-failed/handler.py) to parse epoch millisecond / ISO stop timestamps into ended_at. Preserved existing conditional update (status = 'RUNNING') so duplicate terminal invocations produce no extra notifications or counter increments.

## Validation

- 54 unit tests passed in pipelines/cdk/lambdas/tests/test_update_run_failed.py covering status transitions, execution ARN resolution, stop timestamp parsing, and idempotent duplicate handling.

- 155 tests passed in pipelines/runners/timeback-raw-sync/tests covering contracts, transforms, and lineage.

- 572 CDK tests passed in test/constructs/ecs-pipeline.test.ts, test/constructs/ecs-step-function.test.ts, and test/real-pipeline-configs.test.ts verifying task timeout synthesis (28,800s), parent state-machine buffer (29,100s), and EventBridge DLQ/alarm patterns.

## Scope Boundaries

- This is a containment fix and lifecycle reliability improvement. It does not alter entity extraction logic, change the rolling source window, or modify partition planning for TimeBack entities.

## Linear

Fixes [SURTR-1162](https://linear.app/builder-team/issue/SURTR-1162/contain-timeback-raw-sync-overruns-and-reliably-finalize-timed-out).

## Implementation Effort

Estimated 0.5 engineer days for CDK timeout wiring, EventBridge fallback reconciler DLQ/alarm synthesis, Lambda timestamp handling, and regression test coverage.

#1769 — [SURTR-1135] Fix QTD unit economics revenue proration @ashwanth1109  approved

## Summary

- Allocate modeled tuition across January–May and August–December, excluding June and July.

- Retain the inclusive 365-day fraction for modeled costs and Timeback.

- Add auditable revenue-proration metadata, idempotent migrations, Table 3 propagation, and executable regression coverage.

## Linear

- https://linear.app/builder-team/issue/SURTR-1135/fix-qtd-unit-economics-revenue-calendar-exclude-junejuly-and-allocate

## Business Value

Corrects QTD unit-economics revenue timing so summer non-revenue months are not overstated, while preserving the existing year-round cost model and downstream lineage. This makes Table 2 and Table 3 profitability and per-student reporting decision-ready.

## Implementation Effort

Estimated 1.5–2 engineer-days for an average engineer to implement, migrate, test, deploy, and validate manually.

## Validation

- Table 2: 36 tests passed; Table 3: 13 tests passed.

- Ruff file-scoped checks, formatting checks, and git diff checks passed.

- Candidate deployments succeeded independently for both affected pipeline stacks.

- Live Table 2 → EventBridge → Table 3 executions succeeded.

- Alpha Scottsdale matched the fixed case: tuition 350266.67, profit -177214.70, EBITDA -166477.17.

- All 1,045 current-cutoff rows had complete proration metadata and Table 2/Table 3 lineage consistency.

#1249 — feat(chat): add REBL3 pipeline data-health endpoint @kevalshahtrilogy  approved

## Summary

- Adds GET /api/sync/real-estate, mirroring the existing accountability/route.ts pattern (Clerk auth() 401 check, injectable deps, a lib/*-server.ts fetcher, a pure builder function).

- New chat/lib/rebl3-pipeline-health-server.ts calls Surtr's Pipeline Status API (GET /v1/pipeline/mart-aerie-rebl3-sites-refresh?runs=5) with x-api-key auth, validates the response shape with zod, and never throws — every failure mode (missing key, network error, non-2xx, malformed body) resolves to { ok: false, status, error }.

- New chat/lib/real-estate-health.ts exports the pure buildRealEstateHealthPayload(...), which shapes the raw Surtr response (or a fetch failure) into the fixed camelCased contract the frontend's Data Health tab expects. Surtr's Observer verdict vocabulary is passed through verbatim, never reinterpreted. On failure it still returns the full contract shape with degraded defaults (lastRun: null, schedule.enabled: false, observer.verdict: "UNAVAILABLE") plus sourceUnavailable: true and error, so a Surtr outage renders as a stale/unavailable card instead of 500ing the route.

- Adds SURTR_PIPELINE_API_KEY to .env.example (server-side only, not client-bundled).

This is one piece of a larger REBL3 Data Health initiative; the read-path repoint and the UI tab are separate, already-scoped pieces being built in parallel against this same fixed response contract — no coordination needed here.

Note: SURTR_PIPELINE_API_KEY (a Surtr pak_... key with pipelines:read scope) must be provisioned in the deployment environment before this is usable in production. No real key is included in this PR.

## Business Value

Gives the REBL3 real-estate dashboard a trustworthy, Surtr-backed "is this data fresh/healthy" signal where today there is none — surfacing pipeline run status, schedule, and Observer findings (freshness/quality issues) directly to the team that depends on REBL3 site data, instead of them discovering staleness only when a number looks wrong downstream.

## Manual Effort Estimate

AI-drafted estimate — flag for Keval to confirm/adjust: ~4-6 hours (half a day) for an engineer already familiar with this codebase's lib/*-server.ts / builder / route conventions — reading the sibling accountability route and REBL3-adjacent server files, writing the fetcher with schema validation and the pure builder, and writing the unit test coverage (happy path, 401/403/503/500, missing env var) called for by this repo's endpoint-hardening and testing docs.

## Test plan

- [x] pnpm --filter chat typecheck — passes

- [x] pnpm --filter chat lint (biome) — passes

- [x] pnpm --filter chat test scoped to new files (rebl3-pipeline-health-server.node.test.ts, real-estate-health.node.test.ts, route.node.test.ts) — 17/17 passing

- Fetcher: happy path, Surtr 401/403/404/400/500/503, network error, invalid JSON, schema-validation failure, missing SURTR_PIPELINE_API_KEY

- Builder: success shaping, as_of: null fallback to now(), null lastRun, verbatim verdict pass-through (including non-standard values), degraded/sourceUnavailable shape on failure

- Route: 401 when unauthenticated, 200 shaped payload on success, upstream status + sourceUnavailable payload on failure

Linear: (ticket pending)

#32 — AI-725: Persist workflow operations and add reusable smoke testing @ashwanth1109  no labels

Accepted workflow actions previously depended on mounted React views and could lose their result across a crash or integration timeout. This change persists lifecycle intent before dispatch, moves workflow ownership into Rust, and reconciles uncertain outcomes on startup/reconnect. A reusable desktop fixture harness now exercises these paths without live integrations.

## Linear

[AI-725: Persist workflow operations with explicit transitions, recovery, and reusable smoke testing](https://linear.app/builder-team/issue/AI-725/persist-workflow-operations-with-explicit-transitions-recovery-and)

## Business Value

Tasks retain visible progress and recoverable intent across navigation, app restarts, and integration failures. Unknown creation outcomes surface an explicit recovery action instead of silently creating duplicate threads or tickets. Contributors and future agents can reproduce desktop failure scenarios against disposable data before shipping changes.

## Implementation

- Add SQLite operations, attempts, checkpoints, runtime records, and transition audit events. Use exclusive database ownership, two workers, per-task serialization, and stale-attempt fencing.

- Queue Research/Implement provisioning, approved ticket creation, thread replacement, and deletion. Capture immutable ticket input during approval and persist remote identities before dependent effects.

- Reconcile thread history and outstanding checkpoints after startup/reconnect. Restore active turns and pending requests; expose existing-thread/issue attachment for unknown outcomes.

- Move completion capture into native runtime handling. Require the current successful turn's explicit Implement PR marker and retain agent-reported provenance.

- Add the dedicated smoke-test build, local Codex/Linear adapters, isolated per-run databases/repositories, guarded process controls, fault injection, checkpoints, and JSON assertions. Document recipes in docs/SMOKE_TESTING.md and contributor instructions in AGENTS.md.

- Fix defects discovered during validation: missing actions on restored approval cards and database locks retained through duplicate file descriptors.

## Compatibility and scope

This proof-of-concept milestone starts a fresh shipyard-v2.sqlite3; it leaves the previous database in place and does not import its data. The coordinator runs inside the app, so closing Shipyard stops execution. Ordinary composer submissions and template conversations remain outside the durable lifecycle queue. Independent GitHub result verification, managed worktree/access boundaries, full artifact revisions, and runtime pinning remain deferred in ALPHA_ARCHITECTURE_PLAN.md.

## Validation

- 61 native tests, 8 frontend workflow tests, and 19 harness tests passed.

- Packaged smoke build passed, including TypeScript checking and the Vite build; validated bundle fingerprint matches the source.

- Native desktop scenarios passed: Research → Ticket → Implement, active-turn restart, pending approval, lost thread acknowledgement, lost initial-turn acknowledgement, and lost ticket acknowledgement.

- Forced disconnect/reconnect and final-build startup reconciliation preserved completed state without duplicate effects; the complete workflow retained exactly two threads, two initial turns, and one fixture issue.

- Missing fixture configuration was refused; the bundle contains no real Codex runtime. All new harness processes were stopped and local evidence retained under ignored .smoke/.

- Staged diff whitespace check passed.

These tests use deterministic local substitutes; they do not establish compatibility with a live Codex release or Linear account.

## Implementation Effort

Estimated 10–15 engineering days for an average engineer to implement the native workflow rework, recovery UI, fixture framework, and failure-path validation manually without AI assistance.

#1261 — fix(enrollment): drop SIS tenant test records from the enrollment spine @vvp-trilogy  approved

Closes #1258.

Drops tenant test records from the SIS staging layer so the enrollment report stops counting rehearsal data, and makes the enrollment derivative follow the dropped student.

## What changed

| File | Change |

|---|---|

| macros/not_test_surname.sql (new) | Source-agnostic _test_ surname predicate: coalesce(strpos(lower(x), '_test_'), 0) = 0. |

| macros/finalsite_not_test_record.sql | Delegates to not_test_surnamebehaviour-preserving (compiled predicate byte-identical; the 8 Finalsite call sites are untouched). |

| staging/sis/stg_sis_student.sql | Drops is_test = true and _test_ surnames. A NULL flag and a NULL surname are treated as real and still emit. is_test stays a selectable column. |

| staging/sis/stg_sis_brand.sql | Drops the sandbox test org via slug IS DISTINCT FROM 'test' (NULL-safe). |

| intermediate/enrollment/int_enrollment.sql | scoped_enrollment gains an EXISTS semi-join to int_student, keeping only enrollments whose student survives staging. The enrichment LEFT JOIN stays (rule 12); the scope drop lives on the scoping CTE (rule 6). |

| staging/sis/_sis__models.yml, _sis__sources.yml, enrollment/_int_enrollment__models.yml | Descriptions restated to the current shape (rule 22). |

| tests/assert_sis_student_excludes_test_records.sql (new) | Pins the is_test + _test_ drop at staging. |

| tests/assert_sis_brand_excludes_test_org.sql (new) | Pins the slug='test' drop at staging. |

| tests/assert_sis_enrollment_excludes_test_records.sql (new) | End-to-end guard: recomputes test-ness from the raw source and asserts no mart fact row resolves to a test student or test-brand campus (by student_key/campus_id, never full_name). |

### The strpos vs LIKE decision (settled by the ticket)

Kept strpos, not LIKE. In LIKE, _ is a single-char wildcard, so '%_test_%' false-positives on real surnames (Contestabile). strpos carries no wildcard semantics and needs no ESCAPE, so a later refactor cannot silently broaden it. lower() is load-bearing.

### Why the enrollment drop is explicit

The mart's student_key is the enrollment's own student_id, and int_student attaches with a LEFT JOIN. Filtering the student alone changes no counts — the test-student enrollments survive with a populated student_key and a NULL name. The EXISTS on the scoping CTE is what actually drops them.

## dbt verification (real Redshift build)

Ran against the finance_dw Redshift cluster (PR-namespaced --vars '{pr_number: 1258}', objects dropped afterward per rule 19):

- dbt build --select +mart_enrollment_dtl — all 19 models built; every attached test passed, including the 3 new tests, assert_sis_enrollment_campus_allowlist, assert_sis_enrollment_has_fact_consistency, assert_sis_enrollment_qualified_is_fact, and the has_fact=false uniqueness test.

- Finalsite compiled predicate confirmed byte-identical to before.

### Before / after in-scope enrollments (distinct enrollment_id, has_fact)

| SY | Before | After | Dropped |

|---|---|---|---|

| 2025 | 2,908 | 2,838 | 70 |

| 2026 | 2,560 | 1,966 | 594 |

Leaked test enrollments after the change: 0 in every year. Factless grid cells unaffected (every Alpha campus/year/cohort cell still appears; the has_fact=false uniqueness test still passes).

### SY2026 cohort-level before → after

| Cohort | Before | After |

|---|---|---|

| on-campus | 2,296 | 1,716 |

| first-day | 2,110 | 1,592 |

| re-enrolled | 802 | 801 |

| graduating | 197 | 84 |

| mid-join | 187 | 125 |

| starting-later | 103 | 103 |

| withdraw | 39 | 27 |

| (others unchanged) | | |

### Explaining the SY2026 delta (594, vs the ticket's 483)

The source has moved since the ticket was written (2026-09-08). Breaking the 594 down by the student behind each dropped enrollment:

- 463 — live (non-deleted) test students. This is the "483" of the ticket, now 463.

- 130 — students that are both soft-deleted and test (still synthetic).

- 1 — a soft-deleted, non-test student (a pre-existing orphan enrollment that survived via the LEFT JOIN).

593 of the 594 are synthetic; the 1 non-test orphan drops because its student is soft-deleted — also a disowned population under rule 5. There is no rules-clean way to keep that single orphan without re-deriving the test predicate in intermediate (violating rules 6/8), so the semi-join keeps only enrollments whose student survives staging. The comment and yml state this true constraint.

## Parity answer (the decision the ticket asked for)

Does the HubSpot baseline mart_enrollment_dtl also carry these synthetic students? No — it carries zero. Checked educrm_wh.mart_enrollment_dtl (the source of the staging_education.sales_educrm_wh_mart_enrollment_dtl replica): fact rows with _test_ in last_name, deal_name, or full_name = 0 across SY2024–2027.

So filtering SIS closes a known gap and parity improves — SIS was over-counting relative to HubSpot. The SY2026 SIS-vs-HubSpot gap narrows from 862 (2,560 − 1,698) to 268 (1,966 − 1,698). #1216's delta table should be restated on the filtered spine, where the SIS excess over HubSpot shrinks by ~594 for SY2026.

## Acceptance criteria

- [x] Surname matched with coalesce(strpos(lower(...), '_test_'), 0) = 0; no LIKE; Finalsite extraction byte-identical.

- [x] stg_sis_student emits no is_test=true / _test_ surname row; NULL flag and NULL surname still emit.

- [x] stg_sis_brand emits no slug='test' row; the other 11 brands unchanged.

- [x] mart_enrollment_dtl has no fact row resolving to a dropped student — verified by count (0), not by full_name.

- [x] SY2026 in-scope drop reconciled (594 = 463 live test + 130 deleted-and-test + 1 soft-deleted orphan; the source moved from 483→463).

- [x] Factless grid rows unaffected; has_fact=false uniqueness test passes.

- [x] A dbt test pins each of the three rules (two staging + one end-to-end).

- [x] is_test remains selectable on stg_sis_student and int_student.

- [x] Stale "deliberately NOT filtered" comment and yml notes rewritten to current shape (rule 22).

- [x] Parity answer stated above.

## Out of scope (untouched, per the ticket)

educrm-reporting, int_deposit_overlay, the campus allowlist / Test Campus - No grades exclusion, and all Finalsite/EduCRM behaviour. int_campus still keeps brandless campuses (its LEFT JOIN leaves the six ex-test-brand campuses with a NULL brand_name); they fall out of the report on the Alpha allowlist, and the existing assert_sis_enrollment_campus_allowlist test guards the leak — changing int_campus for other SIS consumers is left as a deliberate non-change here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#120 — fix(heimdall): the described-edit check was inert — the harness made the workspace look dirty @kevalshahtrilogy  no labels

## It never fired once

SURTR-1159 ran twice more after #118 merged. Both runs produced the same complete, grep-verified delete plan, the same empty diff, and verdict=ok — the exact outcome the check was written to prevent. Publish PR skipped both times. The ticket is still open.

## Why

The harness copies .trusted/ and .heimdall-config/ into the agent's working directory, and neither is gitignored. So git status --porcelain in that checkout is never empty:

?? .heimdall-config/

?? .trusted/

The check read that as "the agent edited something" and returned clean on every run. Tree extraction already excludes those exact paths — rsync --exclude=.git --exclude=.trusted --exclude=.heimdall-config — because they are scaffolding rather than the agent's work. My check just didn't mirror it.

workspace_is_clean now ignores them, by directory prefix so a real file named .trustedconfig.ts still counts as work.

## Two mistakes of mine, both silent

lstrip("./") strips a character SET, not a prefix. .trusted/ became trusted/ and stopped matching the prefix it was being tested against — so the first version of *this* fix was also wrong. Caught only by running it against real porcelain output.

The original tests couldn't have caught any of this. They all fed describes_edits a note directly and never asked what a workspace looks like with the harness standing in it. Every one passed on the day the feature shipped inert. There is now an end-to-end case on the exact shape that slipped through, plus a table over porcelain formats — renames, quoted paths, ./-prefixed paths, git internals.

## Testing

heimdall/tests all green, 10 new in test_fix_integrity.py. ruff check clean.

## Business Value

Restores a guard that has been reporting success while doing nothing since it merged — the worst state for a check to be in, because it also removes the pressure to look. SURTR-1159 is the immediate unblock; it has now cost four runs and two human replies for a deletion the agent has correctly described five times.

## Manual Effort Estimate

~45 minutes. Small fix; the value was in checking whether the previous one actually worked rather than assuming it did. Keval to confirm/adjust.

#119 — feat(heimdall): repair red before publishing, and never open a second PR for a ticket @kevalshahtrilogy  no labels

The last two of the seven. Both are places the factory produces something correct and then needs a human to finish it.

## 1. Verify while the agent can still repair

On 2026-09-08, 9 of 10 open heimdall PRs had a failing check — mostly the agent's own new test, or a formatter it has no shell to run. Work correct in substance and unmergeable in fact, each one waiting on a human.

The validate job already runs the repo's verify.commands. But it runs them on a fresh runner, after the agent is gone, where the only two outcomes are DRAFT or a red PR. By then nobody can fix it but you.

They now also run in the workspace, before the tree is extracted, and a failure buys one repair pass with the failing output attached to the prompt. Still red after that is not fatal: the PR opens and the existing verify marks it DRAFT with the output in the body — a red PR beats no PR, because a human reviewing work is cheaper than a human recreating it.

Costs one extra verify run on the runs that need it. That is cheaper than the turn it replaces.

## 2. One ticket, one PR

select_for_repo skips a ticket that already has a PR — but that check runs when the ticket is picked, and the queue path and the failure path sit in different concurrency groups, so two runs can be past it at once. SURTR-1019 ended up with a merged #1663 *and* an open #1634, against an instruction written in both the ticket and #1663's own body.

Publish now refuses to open a second PR for a ticket that already has an open one.

It can only do that because the PR body finally names its ticket. It never did — which is why nothing could map a PR back to its ticket without opening both, and is the direct cause of a board sweep reading three *merged* PRs as "approved but never merged". That line is worth having on its own.

### On the implementation

My first version asked Linear at publish time. That job has no trusted-harness checkout and no Python deps, so it would have meant adding both to answer a question a line in the PR body answers for free. Reverted before it landed; this version is GitHub-only.

Fails open in both places: a search error publishes anyway, and an unreadable verify config skips the repair. Losing a finished fix to avoid a duplicate is the more expensive way round — a duplicate is visible, a dropped fix is not.

## Testing

heimdall/tests 1254 passed, 1 skipped; harness/tests 376 passed; ruff check, ruff format --check harness, actionlint clean.

The contract tests pin the things that would silently rot: that the repair step sits *before* tree extraction (after it, the workspace is gone), that it makes no forward steps. reference — actionlint caught one while I wrote this, the same bug class as Chat: run started — and that a still-red repair keeps the PR.

## Business Value

These two close the gap between "heimdall finished" and "the change is mergeable". Nine PRs currently sitting red are nine pieces of completed work that each need a human to repair something the agent could have fixed while it still had the workspace; the duplicate is worse, because it costs a reviewer the time to work out which of two PRs is the real one.

## Manual Effort Estimate

~4 hours. Small diffs, but the ordering constraints are the substance — verification has to happen where the agent still exists, and the duplicate check has to happen where a duplicate is still avoidable. Keval to confirm/adjust.

#118 — feat(heimdall): stop losing finished work — described fixes, merge closure, evidence reach, and self-decided calls @kevalshahtrilogy  changes requested

Five of the seven problems from the board sweep. Each is a place the factory finishes thinking and the result evaporates.

## 1. A described fix is a failed fix, not a no-fix

agent_retry exists because a provider failure and "nothing needed doing" both exited 0 and were indistinguishable. This is the same conflation one level in: an agent that reasons out a complete changeset and then writes it down produces the identical signature as a real no-op — empty tree, exit 0, a paragraph.

SURTR-1159 is the case that named it. Three runs, three grep-verified delete lists, three empty diffs, three green runs. A human replied *"Approved to proceed — open the PR"* between runs two and three; run three repeated the no-op verbatim. The ticket is still open.

fix_integrity.py reads the signal grammatically, not semantically: a no-op note reports ("board.ts is imported only by…"), a described-fix note issues instructions naming files ("Delete x/y.ts"). It re-prompts rather than failing — the agent has the answer and just did not act on it. A false positive costs one turn and self-corrects; only a second identical reply gives up, and then it reports a *failed* fix rather than claiming none was needed.

## 2. Done on merge

report() has been posting *"It goes to Done when the PR merges"* on every ticket since it was written. Nothing implemented it. SURTR-1019, 1023 and 1025 each shipped on 2026-09-02 and were still in For Review six days later, two of them Urgent — which is also why a floor sweep read them as "approved but never merged" and ranked a merge-station bug as the top priority. There is no merge-station bug.

Found by Linear's own PR attachment, so it works for a human merge, an auto-merge, and a renamed branch alike. Its own job on mode: merged, because most heimdall PRs are merged by hand. Never fails the run — the merge already happened.

## 3. Evidence that is not logs

cloudwatch_read closed one class of "I could not see…" and not the class. Five tickets ended in a complete diagnosis and no fix, blocked on an S3 object or a Step Functions stop reason — all reachable by hand in under a minute. SURTR-1093 is the sharpest: CloudWatch correctly returned 0 events because the container never started, and the answer was StopCode: TaskFailedToStart / CannotPullContainerError on the execution.

aws_read.py follows the same one-way pattern — agent names it, trusted code fetches it, only text returns, credential never enters the agent's process. Bounded rather than trusted: bucket shape, no .. or leading / in keys, 64 KB truncation, capped history, and over-cap requests *reported* rather than dropped.

## 4. Decide reversible calls

11 open tickets end with "awaiting your reply", most on a call with two acceptable answers. The callout protocol already refuses permission/preference/verifiable/confidence — but only *after* the run, which changes the comment and not the behaviour. The prompt now says to choose, act, and record the choice so a reviewer can overrule it in one comment. Stopping stays correct for irreversible, credentials, contradictions, and policy/money.

## 5. Only invite a reply when a reply changes something

"Reply here and I will pick this up on the next sweep" was unconditional on every no-change conclusion. That is what SURTR-1159 ended with, and answering it produced the identical comment. It now invites a reply only when a real gap was recorded; otherwise it says the conclusion is not a question and names what would genuinely reopen it.

## Testing

heimdall/tests 1249 passed, 1 skipped (~50 new); harness/tests 376 passed; ruff check, ruff format --check harness, actionlint all clean.

The fix_integrity corpus uses the verbatim SURTR-1159 note as its fixture, and pins that a genuine no-op ("the credential is missing, a human must rotate it") and narration ("I considered deleting x but…") are *not* flagged — a guard that fires on real no-ops is a guard someone switches off.

## Not in this PR

- One ticket, one PR. SURTR-1019 has a merged #1663 and an open #1634. Single-flight is already correct for the queue (constant concurrency group), so the duplicate came from the queue and failure paths racing — different concurrency groups. That needs a publish-time check against existing open PRs for the ticket, which is its own change.

- Running the agent's own new tests pre-commit. format.commands (#114) covers formatters; extending it to tests is the same shape and belongs with it.

Both are worth doing; neither should ride on this diff.

## Business Value

This is the difference between a factory that needs a human per ticket and one that needs a human per decision. Every item here is work that was already done correctly and then thrown away — a fix written but not applied, a merge the board never heard about, a diagnosis blocked on a value one API call away. The cost was not model spend; it was that you had to open tickets one at a time to find out what had actually happened.

## Manual Effort Estimate

~1.5 days. The code is not large; establishing which of the reported problems were real took most of it — the headline finding inverted on inspection. Keval to confirm/adjust.

#1257 — fix(finance): support NetSuite-only CAPEX entities @marcusdAIy  approved

## Summary

- support CAPEX entity tie-out rows that exist only in NetSuite by making the internal qbCompanyId nullable

- add a deterministic opaque entityKey derived server-side with domain-separated HMAC-SHA256; raw NetSuite subsidiary IDs never enter the DTO, logs, browser, or public API

- prefer NetSuite identity when present so an NS-only row keeps the same key after a later QuickBooks mapping

- move table keys, expansion state, active drilldown state, DOM IDs, panel remounting, and focus restoration to entityKey

- render NS-only rows normally while suppressing the QuickBooks transaction-drilldown affordance

- preserve the existing singular public API v2 contract and add GET /v2/finance/capex-entity-tieouts for the complete nullable directory

## Compatibility and controls

- GET /v2/finance/capex-entity-tieout keeps its existing closed schema: qbCompanyId remains a required string and entityKey is not returned

- the legacy route filters NS-only rows; the new plural route returns QB-linked and NS-only rows with an opaque key

- both routes retain the existing capability, audit, rate-limit, warehouse-bound, and fail-closed controls

- this PR changes Aerie consumer behavior only; it does not install warehouse DDL, refresh CAPEX marts, enable on-demand execution, or enable the schedule

## Validation

- pnpm --dir chat typecheck

- 8 focused Vitest files: 334 tests passed

- focused singular/plural public API integration test passed with a 30-second timeout for the known-slow WSL harness

- Biome passed on all changed files

- independent post-rebase review found no blockers

Refs AERIE-1415 and SURTR-648.

#1782 — fix(repo): Surface ECS StopCode/StoppedReason before truncation @heimdall-keval-factory[bot]  approvedAutomated PR

Container-start failures (image-pull timeouts, ENI issues) currently reach triage as generic 'task failed' noise because the AWS-reported cause sits at the end of a message that gets truncated. This moves that cause to the front so it survives — though a separate ticket-formatting step outside this PR's scope will still need its own fix before the full benefit reaches Linear tickets.

> Ready for review. Nothing ran the change, so it is unproven. A person still merges.

## For The Agent

_Everything below is detail for review. The summary above is the change._

Presented as ready — verification none, scope tier draft, fix_class code_fix, HEIMDALL_READY_PRS=true.

### What's broken

SURTR-1160 (citing SURTR-1093) asks to reorder error_message in pipelines/cdk/lambdas/update-run-failed/handler.py so ECS StopCode/StoppedReason (buried inside error.Cause, assembled at handler.py:132-142) survive the 2000-char caps applied at handler.py:369 and :567 before the message reaches triage. That 2000-char cap claim checks out, but it is not why SURTR-1093's ticket read 'States.TaskFailed — the orchestrator reported a task failure with no error message in it': that exact sentence is generated by _first_error_line() in .trusted/heimdall/failure_ticket.py:100-113, whose _ENVELOPE regex (line 85) discards the whole message and substitutes that boilerplate whenever "Attachments":/"Containers":/"ClusterArn":/"networkInterfaceId": appear anywhere in the string — which they will, since Attachments is the first key AWS puts in every ECS stopped-task Cause blob, regardless of what gets prepended ahead of it. Fixing only update-run-failed (in scope for this PR) will not by itself stop a future Linear ticket for this failure class from showing that same boilerplate line, because failure_ticket.py is explicitly out of scope here and its envelope-discard rule is untouched.

Root cause. Two independent issues, not one. (1) update-run-failed computes a single error_message string once at handler.py:132-142 and reuses it for the DB write (65535-char cap at :152), the GChat/Linear notification payload (2000-char cap at :369), the legacy SNS alert body (2000-char cap at :567), and the SES email (3000-char cap at :474). For an ECS container-start failure this string is f'{Error}: {Cause}' where Cause is the raw ECS stopped-task JSON — Attachments (subnet/ENI/MAC, can run to several KB) is its first key, with StopCode/StoppedReason/per-container reason/exitCode appearing much later — so the 2000-char notification/alert caps genuinely do truncate the cause away, exactly as the ticket describes. (2) Separately, and unaffected by fixing (1), .trusted/heimdall/failure_ticket.py:82-113 treats ANY error_message containing "Attachments":, "Containers":, "ClusterArn": or "networkInterfaceId": as an opaque envelope and replaces it wholesale with an 80-char pre-colon head plus a fixed 'no error message' sentence — a content-shape check with no dependency on length or on what precedes those keys — so it still fires after a fix that only reorders content ahead of the still-embedded raw JSON, since the raw JSON (and its Attachments key) remains part of the string.

### What this PR changes

Implement the fix as scoped, in pipelines/cdk/lambdas/update-run-failed/handler.py: when error.Cause parses as JSON containing StopCode/StoppedReason/Containers, build a short summary (StopCode, StoppedReason, and any container's non-null reason/non-zero exitCode) and prepend it to error_message ahead of the existing f'{Error}: {Cause}' text, leaving non-ECS payloads byte-identical. This is real and independently useful even though it doesn't fully close the ticket's stated gap: it fixes the DB-stored error_message (read by direct SQL and by Surtr/src/pipeline-api/routes.ts's presentError) and the direct-dispatch prompt path (.trusted/heimdall/build_prompt.py:137, which passes error_message through verbatim, no envelope filter). The PR body should say plainly that it will NOT stop a Linear ticket filed via file_failure.py/failure_ticket.py from showing the same 'no error message' boilerplate for this exact failure class, since that file's envelope-shape heuristic is what actually produced SURTR-1093's ticket text and is untouched here — flag for Keval whether a follow-up against failure_ticket.py (e.g. only discarding when no recognized cause field precedes the envelope markers) should be filed alongside this PR.

Why this fixes it. Reordering error_message so StopCode/StoppedReason/container reason/exitCode lead is a small, self-contained change confined to pipelines/cdk/lambdas/update-run-failed/handler.py's message assembly (~132-142), needs no schema or cap changes, and has value beyond the Linear-ticket path this ticket was filed about (direct SQL reads of pipeline_runs_*.error_message, and the non-ticket-first build_prompt.py evidence path both see the reordered string verbatim). This is exactly the class of gap code_fix covers: a real diagnostic signal is silently discarded before it's ever stored, with no test guarding the shape today.

#### Files changed

 .../cdk/lambdas/tests/test_update_run_failed.py    | 98 ++++++++++++++++++++++

pipelines/cdk/lambdas/update-run-failed/handler.py | 52 +++++++++++-

2 files changed, 149 insertions(+), 1 deletion(-)

### Verification

### pytest — no test suite

### verify: ruff check — exit 0

[notice] A new release of pip is available: 25.3 -> 26.2.1

[notice] To update, run: pip install --upgrade pip

All checks passed!

### verify: ruff format --check — exit 0

1783 files already formatted

### verify: pytest (pipeline lambdas) — exit 0

6.0

rootdir: /home/runner/_work/Surtr/Surtr/publish/pipelines/cdk/lambdas

configfile: pyproject.toml

testpaths: tests

plugins: cov-7.0.0

collected 494 items

tests/test_ai_spend_raw_api.py .............................. [ 6%]

tests/test_coordinate_fanout_run.py .................................... [ 13%]

.... [ 14%]

tests/test_create_run_record.py ........................ [ 19%]

tests/test_gchat_notifier.py ........................... [ 24%]

tests/test_generate_chunks.py ..... [ 25%]

tests/test_gsheet_tracker.py .................. [ 29%]

tests/test_load_fanout_plan.py .............. [ 31%]

tests/test_redshift_cluster_iam_role_association.py ........ [ 33%]

tests/test_registry_sync.py ......................... [ 38%]

tests/test_triage_dispatcher_handler.py ................................ [ 45%]

................... [ 48%]

tests/test_triage_dispatcher_signature.py .............................. [ 55%]

...... [ 56%]

tests/test_triage_reconciler.py ............ [ 58%]

tests/test_triage_reconciler_tracker.py ............ [ 61%]

tests/test_update_run_failed.py ........................................ [ 69%]

............. [ 71%]

tests/test_update_run_success.py ....................................... [ 79%]

....................... [ 84%]

tests/test_verify_on_demand_control.py ................................. [ 91%]

............................................ [100%]

============================= 494 passed in 1.28s ==============================

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | repo |

| Failing run | issue |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | linear-SURTR-1160 |

| Verify | none |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev label to take the PR over and stop it entirely.

#1781 — fix(capex): apply warehouse DDL in one transaction @marcusdAIy  approved

## Summary

- replace sequential CAPEX DDL submissions with one Redshift Data API batch_execute_statement

- require ExecutionMode=TRANSACTION so all 19 statements commit or roll back together

- use an explicit dependency manifest: 001_tables.sql003_forward_migrate_entity_tieout_qb_company_id_nullable.sql002_sp_refresh_capex_marts.sql

- fail closed on missing or unexpected DDL files, terminal substatement errors, cancellation failures, and indeterminate timeouts

- add fake-client tests for the full submission and failure contract

## Why

The released installer submitted each statement independently and sorted files alphabetically. On the live SURTR-648 installation that would replace the NS-only-capable procedure before migrating qb_company_id to nullable, and there was no transaction spanning the change. A failure could leave a partial installation.

This change creates one reviewed manifest and one atomic Data API transaction. It never retries an indeterminate batch automatically.

## Validation

- uv run pytest -q — 31 passed

- uv run ruff check . — passed

- uv run ruff format --check . — passed

- uv run pyright scripts/apply_ddl.py — 0 errors

- parse-only dry run — 19 statements in the expected transaction manifest

No production DDL was submitted.

#1780 — fix(capex): preserve missing DDR budgets as null @marcusdAIy  approved

## Summary

- keep sites with a missing current DDR budget in the CAPEX publication with ddr_budget=NULL and ddr_budget_source=NULL

- preserve the cohort identity invariant while making budget coverage an audited quality count rather than a publication blocker

- report missing_budget_count in the no-write verifier

- strip SQL line comments with a string-aware scanner before verifier extraction, including end-of-line comment semicolons

- document and test that missing budget is never converted to zero

## Why

180-maiden-ln-new-york-ny is now in the live SY26/27 cohort, but its only Rhodes due-diligence row has neither fo_capex nor phase1_capex. Excluding it would require another code change when a plan appears. Publishing $0 would falsely assert an approved zero budget.

A nullable plan is the existing consumer contract: Aerie withholds percentage-of-plan comparison when DDR is null. The next normal refresh will populate the site automatically after Rhodes supplies an approved budget.

## Live no-write validation

The exact branch source model completed and rolled back successfully:

- selected close: 2026-08-31

- cohort: 43 sites / 43 distinct

- DDR budget coverage: 42; missing: 1

- entity tie-outs: 59 (19 tie to cent)

- site summaries: 43

- transaction detail: 2,559

- explicit unresolved site mappings: 7

- Miami variance: $694,579.28

## Tests

- uv run pytest -q — 23 passed

- uv run ruff check . — passed

- uv run ruff format --check . — passed

- uv run python scripts/apply_ddl.py — 19-statement parse-only dry run

No production tables were changed and no pipeline execution was started.

#1256 — fix(enrollment): SIS mid-year Transfer Out excludes FinalSite opening-day stamps (#1254) @vvp-trilogy  approved

## Summary

Fixes #1254. Re-gates the SIS mid-year vs start-year Transfer Out split so FinalSite opening-day campus transfers (e.g. Brooklyn Allred, Austin 2026) match HubSpot: they land in start-year Transfer Out, not Mid-Year Transfer Out.

FinalSite carries no transfer timestamp (status campus_transfer only), so the SIS sync invents transfer_date — usually the session start. The classifier's inclusive >= session_start bound filed those opening-day stamps as mid-year departures, diverging from HubSpot, which routes an opening-day Campus Transfer to start-year Transfer Out.

## Changes

1. dbt/models/intermediate/enrollment/int_enrollment_classification.sqlx_is_mid_year_transfer now requires CAST(transfer_date AS DATE) > CAST(session_start_date AS DATE) (strictly after) instead of >=. Opening-day stamps (transfer_date = session_start) fall through to the existing NOT x_is_mid_year_transfer complement in int_enrollment_cohort.sql and become start-year-transfer-out. Comment updated to state the FinalSite-timestamp constraint.

2. dbt/tests/assert_sis_enrollment_mid_year_transfer_after_start.sql — new singular test pinning that mid-year-transfer-out never contains a row with transfer_date <= session_start (or a null transfer_date). Auto-discovered, default error severity.

Per the issue's Out-of-Scope constraints, the split does not use enrolled_date or withdrawn_date, does not touch the HubSpot EduCRM models, and does not require a destination-campus pairing. int_enrollment_cohort.sql already emits start-year-transfer-out as the complement, so no change was needed there.

## Acceptance criteria — 2026 transfer cohort before/after

HubSpot reference (live-verified against educrm_wh.mart_enrollment_dtl, SY2026, has_fact = TRUE):

| Cohort | rows | distinct contacts |

|---|---|---|

| start-year-transfer-out | 22 | 22 |

| mid-year-transfer-out | 2 | 2 |

Brooklyn Allred's Campus Transfer deal is start-year-transfer-out on the HubSpot side (session_start 2026-08-12, enrollment_date 2026-08-12, withdraw_date null) — the target this fix matches.

SIS side (this dbt project → Redshift sandbox_education.mart_enrollment_dtl): I could not run a live build or query against the Redshift warehouse — no REDSHIFT_* credentials are present in this environment (see "Build verification" below). The SIS baseline is therefore taken from the production counts documented in the issue (as of 2026-09-08), and the after-state is the deterministic result of the >=> change:

| Slice (SIS, SY2026) | Before (>=) | After (>) |

|---|---|---|

| mid-year-transfer-out (classified) | 33 | 14 |

| of those, transfer_date = session_start | 19 | 0 (moved) |

| start-year-transfer-out | baseline B | B + 19 |

The 19 opening-day stamps (Brooklyn's bucket, transfer_date = session_start) move out of Mid-Year Transfer Out into start-year Transfer Out, shrinking the parity gap toward HubSpot's 22 / 2. The 11 rows dated strictly after start with no enrolled_date (Griffin 12 Aug … Ethan Wong 4 Sep) remain mid-year by design — enrolled_date is deliberately not consulted. Direction of change: SIS mid-year-transfer-out decreases by 19, start-year-transfer-out increases by 19; net TRANSFERRED count unchanged (the split is a clean, COALESCE-backed boolean complement, so no row lands in both or neither).

## Build verification

- dbt deps and dbt parse pass — the project (including the new test) parses cleanly; Jinja renders and all ref()s resolve.

- dbt compile / dbt build could not be run: no Redshift credentials are available in this environment (connection refused to a dummy host). pnpm typecheck does not cover Jinja/SQL, so it is not a substitute — flagging explicitly that this change has not been validated against a live warehouse. CI's PR build (--vars '{pr_number: N}') is the first live compile/build/test of these models.

Closes #1254.

#1778 — fix(capex): fail closed in the live verifier @marcusdAIy  approved

## Summary

- make the SURTR-648 live verifier terminate extracted publication statements only at an end-of-line semicolon

- preserve semicolons inside SQL line comments instead of truncating the generated statement

- mirror the stored procedure's cohort identity and DDR-budget coverage invariant before building temporary publications

- add regression tests for both failures

## Why

The released verifier truncated agg_capex_entity_tieout at -- entity_display_name is NOT NULL; ... and failed with a Redshift syntax error at end of input. After repairing that parser locally, the current source snapshot exposed a second gap: the verifier reported 43 sites / 42 budgets but continued, while the production procedure would fail closed on that same mismatch.

The verifier must reject a source snapshot that the writer cannot publish.

## Validation

- uv run pytest -q — 19 passed

- uv run ruff check . — passed

- uv run ruff format --check . — passed

- current live no-write source-model execution now reaches and correctly stops at:

- CAPEX cohort identity/budget coverage invalid: rows=43, distinct_sites=43, budget_covered=42

- missing budget row independently identified as 180-maiden-ln-new-york-ny

All temporary work was rolled back. No production tables were changed and no pipeline execution was started.

#1777 — chore(heimdall): sweep the Linear queue every 5 minutes @kevalshahtrilogy  approvedmercy-allow-critical

Linear: [SURTR-1156](https://linear.app/builder-team/issue/SURTR-1156/heimdall-sweep-the-linear-queue-every-5-minutes)

## What this changes

The Linear queue cron, from 0,15,30,45 * * * * to */5 * * * *. Nothing else.

## Why it is safe

Heimdall works one ticket per sweep, so the cadence *is* the throughput

ceiling — quarter-hourly caps the queue at 4 tickets an hour.

Measured over the 128 telemetry records since 2026-09-01:

| | |

| --- | --- |

| median run | 2.7 min |

| p90 | 5.8 min |

| max | 7.8 min |

A run therefore finishes comfortably inside a 5-minute window. The rare overlap

is safe by construction, not by luck: heimdall-linear-singleton still

holds with cancel-in-progress: false, so GitHub supersedes the pending run

rather than letting two sweeps query, both see the same top ticket — which

deterministic ordering guarantees is the *same* one — and both work it.

Concurrency, claiming and parallelism are untouched. The release cron

(25 * * * *) is untouched. At :25 both crons fire as two separate scheduled

events; each routes on its own github.event.schedule, so the release sweep

still routes to release and the new one to linear, in different concurrency

groups.

A sweep over an empty queue runs no agent, so the extra cadence is only paid

for when there is work.

## What this does NOT do

It does not make the queue parallel. That needs the lock moved off the agent

job — where it currently wraps 10-40 minutes of work — and onto the selection,

which takes seconds, with a claim written at pick time and a stale-claim

sweeper. Designed and filed as

[SURTR-1157](https://linear.app/builder-team/issue/SURTR-1157/heimdall-parallel-queue-workers-via-a-dispatcher-and-a-pick-time-claim),

deliberately not built yet: the queue today has 0 workable tickets and 29

parked on human replies, so throughput is not the bottleneck.

## Business Value

Raises queue throughput from 4 tickets an hour to roughly 10, for a one-line

change that introduces no new failure mode. When the observer files a batch of

pipeline failures at once, the backlog drains in a third of the time — the

difference between a morning's queue clearing before standup and after lunch.

Costs nothing when the queue is empty.

.github/ is on heimdall's forbidden floor, so this is one of the few changes

the agent cannot make for itself.

## Manual Effort Estimate

~1 hour — the edit is one line; the hour is establishing that a 3× cadence

increase is actually safe (reading the concurrency group and the selection

path to confirm the singleton is load-bearing, then pulling run durations out

of telemetry to show sweeps do not overlap in practice). Flagging for Keval to

confirm or adjust.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1252 — 1239-2aerie-publishing-modalities @mwrshah  approved

## Summary

- Add POST /api/v2/forge/articles/{articleRef}/shares for owner-only, capability-gated creation or reuse of non-expiring private Article shares.

- Reuse the existing Forge Article share table, authenticated viewer, transactional audit path, and shared share-write primitive.

- Publish the forge.publish-authenticated-article workflow through manage-only agent context with typed share-output fields.

- Centralize canonical app-origin validation and propagate the local-preview runtime marker into Convex preview configuration.

- Document the Aerie Publisher workflow, external Sindri release prerequisites, URL configuration, recovery behavior, and legacy-share semantics.

The legacy missing-visibility label now correctly reports public access because those existing links are anonymously readable; this does not modify the links.

## Validation

- Chat and Convex TypeScript typecheck

- Biome

- Convex path, architecture, read-bound, and test-architecture checks

- Forge Article/share and public API tests

- Agent-context projection tests

- OpenAPI and preview tests

- git diff --check

External Sindri agent and skill deployment remains a release prerequisite outside this repository.

#1776 — chore(heimdall): 60-minute soak and a 3-hour release cooldown @kevalshahtrilogy  approvedmercy-allow-critical

Prepares Surtr for autonomous releases. This PR does not turn them onHEIMDALL_RELEASE_ENABLED is the switch and stays unset.

## The two settings

soak_minutes 30 → 60. The newest commit must sit for an hour before it can ship. At an hourly sweep the worst case is one extra sweep.

release_cooldown_hours "3" (new, from AI-Builder-Team/mercy#117). Production itself must have been quiet for three hours. The soak ages the *commit*; this ages the *deploy* — without it an hourly sweep ships every hour indefinitely, and the first bad release is followed by the next before anyone has finished reading the alert.

Note the sweep cadence is unchanged and was already hourly (cron: "25 * * * *"), so there was no 30-minute cadence to slow down — the only 30 minutes in the release config was the soak.

## Merge order — this one bites

mercy#117 declares release_cooldown_hours and must land and deploy first. A reusable workflow hard-fails on an input the callee has not declared, so merging this ahead of #117 breaks *every* heimdall run on Surtr — triage, revise and steward, not just the release sweep. This file already documents the same trap for linear_review_state.

## Turning it on

After both merge:

gh variable set HEIMDALL_RELEASE_ENABLED --body true

HEIMDALL_RELEASE_HOLD remains the way to stop releases without silencing triage and revise.

With the settings above, the first release can only happen when: something is on main, its newest commit is ≥60 minutes old, main's CI is green, no stack is leaving the CDK app, and production has not moved in 3 hours. Worst-case merge-to-release latency becomes about two hours rather than the ~89 minutes noted for the old soak.

## Business Value

Autonomous release is the last manual step in the factory, and the one where being wrong is most expensive. These two settings are what make leaving it unattended defensible: an hour for a commit to show a problem, and three quiet hours after each deploy so a bad release is noticed before the next one lands on top of it.

## Manual Effort Estimate

~15 minutes for this file; mercy#117 is the real work. Keval to confirm/adjust.

#117 — feat(heimdall): a release cooldown, so an hourly releaser doesn't ship hourly forever @kevalshahtrilogy  changes requested

## Why

The release gates cover whether the CODE is ready — commits exist, CI is green, the newest commit has soaked, no stack is leaving the CDK app. Nothing covers whether the last deploy has settled.

With an hourly sweep that means production can move every hour, indefinitely, and nobody ever gets a quiet window to notice that the previous release broke something. The soak does not help: it ages the commit, not the deploy, so ten commits that each soaked 30 minutes still ship as ten releases in ten hours.

## What this adds

release_cooldown_hours — refuse when production moved within the window, with reason cooling_down. The state is the base branch's newest commit: that branch only ever advances by a release merge, so its tip *is* the last release. No new API calls.

Defaults to "0" (disabled), so every existing consumer keeps its current behaviour rather than silently acquiring a gate.

## Failure directions, chosen deliberately

Fails closed like the soak: a non-numeric, negative or boolean value is bad_config rather than a quietly disabled check (cooldown_hours: true is 1 hour to a naive numeric test), and a release dated in the future is treated as "just happened" instead of being waved through by the comparison.

Fails open in exactly one place: an unreadable or absent last release. A repo that has never shipped has no release to date, and refusing there would hold the gate shut forever on precisely the repos yet to ship.

## Testing

heimdall/tests: 1195 passed, 1 skipped — 11 new, covering the window boundaries, the disabled default, the never-released case, a future-dated release, every malformed input shape, and that the cooldown does not replace the soak (a long-quiet production must still not ship a commit that landed a minute ago).

harness/tests 376 passed; ruff check, ruff format --check harness, actionlint clean.

## Business Value

This is the gate that makes an unattended releaser safe to leave on. Without it, "enable auto-release" means "deploy to production every hour forever", and the first bad release is followed by the next one before anyone has read the alert. It is also the cheapest possible implementation — one git log against a branch already fetched.

## Manual Effort Estimate

~1 hour. Small gate; the substance is the fail-open/fail-closed reasoning and covering the malformed-input cases the existing soak gate already learned to care about. Keval to confirm/adjust.

#1775 — chore(heimdall): format the agent's output before it commits @kevalshahtrilogy  approved

Opts Surtr in to format.commands from AI-Builder-Team/mercy#114 (that PR must merge and deploy first — this one is inert until it does).

## Why

The revise agent has no shell, so it cannot run ruff format over what it writes. That makes a formatter the one check it fails *every round*.

PR #1738 had three rounds where ruff format was the only red check. Each time: the agent fixed the finding correctly, the file landed unformatted, Lint (Ruff) went red, the red build withheld mercy's auto-approve, and the withheld approval summoned the agent again. A human broke the loop each time by pushing the formatting — which also reset the agent's round budget, so the loop could run indefinitely.

verify already runs ruff format --check, but checking is not fixing: it converts the same problem into a DRAFT PR instead of a red build.

## The entry

format:

commands:

- name: ruff format

run: python3 -m pip install --quiet ruff==0.15.22 && python3 -m ruff format pipelines

Pinned to CI's ruff==0.15.22 for the reason the verify block already documents — a different ruff is a different rule set, so formatting with one version and gating with another just moves the failure rather than removing it.

Scoped to pipelines, matching what Lint (Ruff) actually gates on. Deliberately not harness-style repo-wide formatting: reformatting files CI does not check is churn, and I managed to break an unrelated test doing exactly that while building the upstream change.

## Business Value

Removes a whole class of review round — the kind where the agent is right, CI is red, and only a human can reconcile the two. Every one of those rounds costs a model run on both sides and delays a merge already judged good.

## Manual Effort Estimate

~10 minutes for this file; the upstream capability is the real work. Keval to confirm/adjust.

#114 — feat(heimdall): format the agent's output before committing it @kevalshahtrilogy  no labels

## The loop this closes

The revise agent has no shell, so it cannot run a repo's formatter over what it writes. A formatter is therefore the one check it fails *every single round*. On AI-Builder-Team/Surtr#1738 that produced a loop with no exit:

agent fixes the finding correctly

-> file lands unformatted, Lint (Ruff) goes red

-> red CI withholds mercy's auto-approve

-> the withheld approval summons the agent again

-> repeat

Three rounds there had ruff format as the only red check — each one a mechanical whitespace difference the agent had no way to fix and no way to see. A human pushed the formatting each time, which also reset the agent's round budget, so the loop could run indefinitely.

verify.commands cannot solve this: it runs after the fact and only reports, so the same problem becomes a DRAFT PR instead of a red build.

## What this adds

format.commands in .heimdall.yml, run over the agent's tree before git add so the result is part of the same commit:

format:

commands:

- name: ruff format

run: python3 -m pip install --quiet ruff==0.15.22 && python3 -m ruff format pipelines

Fail-open and never fatal — formatting is a convenience, and a broken format command must not lose a fix that is otherwise good. A failure is a ::warning:: and the commit proceeds unformatted.

verify and format now share one parser, including its SystemExit contract for a half-written entry. They run at different moments for different reasons, but a list of named shell commands is a list of named shell commands.

## Testing

heimdall/tests: 1175 passed, 1 skipped (4 new). harness/tests: 376 passed. ruff check, ruff format --check harness, actionlint clean.

One note on process: I initially ran ruff format harness heimdall, which reformatted 8 files CI does not format-check and broke an unrelated steward test. Reverted — this diff is three files. The irony of a formatting sweep breaking the build while adding formatting automation is not lost on me.

## Follow-up needed in each consumer repo

This adds the capability; each repo opts in. For Surtr the entry mirrors the pinned version its CI gates on (ruff==0.15.22) — a formatter at a different version is a different rule set, which is the trap the existing verify block already documents.

## Business Value

The agent's autonomy is capped by the dumbest check it cannot pass. Every round of this loop costs a model run on both sides, delays a merge already judged good, and needs a human to break it by hand — which then hands the agent five more rounds. This removes a whole class of round from every repo with a formatter, which is all of them.

## Manual Effort Estimate

~1 hour. The diff is small; the substance was tracing why a green-looking PR kept reopening. Keval to confirm/adjust.

#1121 — fix(sales-educrm-mart-sync): raise Redshift statement poll ceiling to f… @heimdall-keval-factory[bot]  approvedAutomated PRheimdall-driven

Automated fix for sales-educrm-mart-sync — fix_class code_fix, scope tier draft.

Resolves https://github.com/AI-Builder-Team/Surtr/issues/1120

> Draft — a human must promote this before merge. Because the diff lands in scope tier draft.

## What's broken

Run f3216770-68b7-43d2-936f-1750d32947fa of sales-educrm-mart-sync synced 36 of 38 educrm_wh mart tables (788,901 rows) but reported status 'partial' because two tables were abandoned with [mart_event_planning_dtl] Error syncing: Statement a4df94aa-c96c-4d8c-93ea-2e2c64e9e319 timed out after 300 seconds and [mart_marketing_event_dtl] Error syncing: Statement 375afa21-614d-4756-bfbf-31e04cc56a30 timed out after 300 seconds. That '300 seconds' is not a Redshift server-side statement_timeout: it is the pipeline's own client-side poll ceiling in pipelines/runners/sales-educrm-mart-sync/src/redshift_handler.py (MAX_POLL_ATTEMPTS = 300 * POLL_INTERVAL_SECONDS = 1), whose raise string on line 164 reads Statement {statement_id} timed out after {MAX_POLL_ATTEMPTS * POLL_INTERVAL_SECONDS} seconds — 300 exactly. The two failed tables leave staging_education.sales_educrm_wh_mart_event_planning_dtl and staging_education.sales_educrm_wh_mart_marketing_event_dtl holding stale data from the prior run (partial data cost, not silent — the per-table status and 'partial' verdict correctly surfaced it).

Root cause. wait_for_statement() in redshift_handler.py polls describe_statement at most MAX_POLL_ATTEMPTS (300) times with POLL_INTERVAL_SECONDS (1) between attempts, so it gives up on any Redshift statement that has not reached FINISHED within ~300s and raises, which sync_single_table() catches and records as status=failed. The mart_marketing_event_dtl (59-column) and mart_event_planning_dtl statements — whose COPY/atomic-swap steps run longer than the other 36 tables, per the log both were mid 'Atomically swapping ..._new -> ...' when they stalled — cross that 300s client ceiling and are abandoned. Critically the Lambda's own timeout_seconds is 900 and the REPORT line shows Duration 380089.53 ms (≈380s), so ~520s of the Lambda's budget went unused while the poll loop bailed at 300s: the client abandons statements at one third of the time the platform actually allows, and the abandoned statement is never cancelled server-side.

## What this PR changes

In pipelines/runners/sales-educrm-mart-sync/src/redshift_handler.py raise the client poll ceiling so it uses most of the Lambda's 900s budget instead of an arbitrary 300s: set MAX_POLL_ATTEMPTS to 780 (keeping POLL_INTERVAL_SECONDS = 1, i.e. a 13-minute ceiling that leaves ~120s of the 900s Lambda timeout as margin for role assumption, discovery, and the swap), and update the stale # 5 minutes max comment on that constant. This is the smallest change that stops slow-but-healthy COPY/swap statements from being marked failed while budget remains, and it stays safely under the Lambda timeout so a genuinely hung statement is still bounded by the platform. Do NOT chunk the UNLOAD as the observer suggested — the log shows UNLOAD mart_marketing_event_dtl completed successfully, so the UNLOAD is not the slow step; the stall is in the Redshift COPY/swap that wait_for_statement polls. Add a focused unit test in tests/ that asserts wait_for_statement's timeout message reflects the raised ceiling and that it raises only after the ceiling is exhausted (mock the redshift-data client's describe_statement to stay non-terminal), keeping the change confined to the pipeline dir (Tier A).

Why this fixes it. The defect is a single constant, MAX_POLL_ATTEMPTS = 300 in redshift_handler.py, decoupled from the pipeline's actual 900s Lambda budget declared in pipeline.json, so the fix is confined entirely to the pipeline's own directory (Tier A) and needs no credentials, IAM, or SQL rewrite. Raising the ceiling to 780s directly removes the premature abandonment that caused the two per-table timeouts while retaining a safety margin below the 900s Lambda cap, so a statement that is truly stuck is still bounded by the platform rather than looping forever. A unit test on wait_for_statement locks in the new behavior and matches the code_fix expectation of shipping a test alongside the change; the blast radius is minimal because no other table or code path changes.

### Files changed

 .../sales-educrm-mart-sync/src/redshift_handler.py |  6 ++-

.../tests/test_redshift_handler.py | 63 ++++++++++++++++++++++

2 files changed, 68 insertions(+), 1 deletion(-)

## Verification

### pytest (pipelines/runners/sales-educrm-mart-sync/tests) — exit 0

============================= test session starts ==============================

platform linux -- Python 3.11.15, pytest-9.1.1, pluggy-1.6.0 -- /opt/hostedtoolcache/Python/3.11.15/x64/bin/python

cachedir: .pytest_cache

rootdir: /home/runner/work/Surtr/Surtr/publish/pipelines/runners/sales-educrm-mart-sync

configfile: pyproject.toml

plugins: mock-3.15.1

collecting ... collected 18 items

tests/test_athena_client.py::test_start_query_uses_explicit_catalog_and_workgroup_without_result_location PASSED [ 5%]

tests/test_athena_client.py::test_unload_path_uses_new_account_transfer_prefix PASSED [ 11%]

tests/test_athena_client.py::test_pipeline_config_uses_new_educrm_account_contract PASSED [ 16%]

tests/test_handler.py::test_handler_returns_expected_structure PASSED [ 22%]

tests/test_handler.py::test_handler_syncs_only_mart_tables PASSED [ 27%]

tests/test_handler.py::test_handler_skips_noncanonical_mart_tables PASSED [ 33%]

tests/test_handler.py::test_handler_uses_educrm_wh_prefix PASSED [ 38%]

tests/test_handler.py::test_handler_uses_unload_copy_pattern PASSED [ 44%]

tests/test_handler.py::test_handler_cleans_up_s3_before_unload PASSED [ 50%]

tests/test_handler.py::test_handler_respects_skip_tables PASSED [ 55%]

tests/test_handler.py::test_handler_handles_table_failure PASSED [ 61%]

tests/test_handler.py::test_handler_empty_discovery_completes PASSED [ 66%]

tests/test_handler.py::test_handler_skips_invalid_discovered_table_names PASSED [ 72%]

tests/test_handler.py::test_handler_ignores_invalid_skip_tables PASSED [ 77%]

tests/test_redshift_handler.py::test_poll_ceiling_fits_within_lambda_budget PASSED [ 83%]

tests/test_redshift_handler.py::test_wait_for_statement_returns_on_finished PASSED [ 88%]

tests/test_redshift_handler.py::test_wait_for_statement_times_out_only_after_ceiling_exhausted PASSED [ 94%]

tests/test_redshift_handler.py::test_wait_for_statement_raises_on_failed_before_ceiling PASSED [100%]

============================== 18 passed in 0.39s ==============================

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | sales-educrm-mart-sync |

| Failing run | f3216770-68b7-43d2-936f-1750d32947fa |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | 6f5eb5d11d8f00baef9a202b6adde23afbc3c74a37b628569a8cac90c5399669 |

| Verify | green |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev label to take the PR over and stop it entirely.

#1595 — fix(openai-usage-pipeline): give token bucket headroom below OpenAI rat… @heimdall-keval-factory[bot]  approvedAutomated PRheimdall-driven

Automated fix for openai-usage-pipeline — fix_class code_fix, scope tier draft.

Resolves https://github.com/AI-Builder-Team/Surtr/issues/1594

> Draft — a human must promote this before merge. Because verification is failing.

## What's broken

Run 4dceefea-679a-466e-8afe-4e036b82ba4e of openai-usage-pipeline self-reported status=partial: the /costs (line_item) fetch for BUs TelcoDR and Trilogy-CloudFix exhausted all 5 retries against sustained HTTP 429s and their usage rows were persisted with $0 billed cost, per src/handler.py:264-270 ('Line-item cost fetch failed for BU ... Persisting the N usage record(s) with $0 billed cost'). The whole 684s run is a 429 storm — e.g. 'Retryable HTTP 429 on attempt 1/5 ... You've exceeded the 30 request(s) every 1 minute(s) rate limit' repeated up to the 65s final-attempt floor — because the shared throttle in src/openai_client.py is set to run the pipeline at exactly OpenAI's org ceiling with a full-window burst, leaving no margin against OpenAI's rolling 60s window.

Root cause. The _TokenBucket in src/openai_client.py is instantiated as _TokenBucket(MAX_REQUESTS_PER_MINUTE) with rate_per_minute=30 and a default burst capacity = max(rate_per_minute, 1.0) = 30, i.e. the sustained rate equals OpenAI's org limit (30 req/min) AND the burst equals a full minute's quota. OpenAI enforces that limit over a *rolling* 60-second window, so a run's first 30 requests drain the bucket instantly (asserted today by tests/test_openai_client.py:565 — 'the first 30 requests drain the pre-filled bucket instantly'), front-loading an entire window's quota; the bucket then refills at exactly 0.5 tok/s and hands out the 31st+ tokens while those burst requests are still inside OpenAI's rolling 60s window, so the window stays saturated and later BUs' /costs calls get 429'd. Each retry is itself a throttled, quota-consuming request, which compounds the storm, and once the 5-attempt budget is spent the BU's billed cost is dropped to $0 — wrong (not merely missing) data for TelcoDR (2 rows) and Trilogy-CloudFix (7 rows) until the T-2 self-heal re-pull, whose firing the observer explicitly flags as unverified.

## What this PR changes

Give the shared limiter headroom below OpenAI's rolling-window ceiling instead of sitting exactly on it, entirely within src/openai_client.py (Tier A). Two levers, both configurable via new env defaults so operators can tune without a code change: (1) shrink the default burst capacity from the full per-minute quota to a small value (e.g. a handful of tokens) so a run can no longer front-load an entire rolling window, and (2) apply a safety-margin factor so the sustained rate_per_second stays a bounded fraction below MAX_REQUESTS_PER_MINUTE, absorbing timing jitter and the retry-request overhead. Update tests/test_openai_client.py::TestTokenBucket to assert the new bounded-burst / headroom behavior (the current test hard-codes the 30-instant-burst flaw and must change with the fix). Merely lowering the existing OPENAI_MAX_REQUESTS_PER_MINUTE env value in pipeline.json is insufficient because the burst-capacity default still front-loads whatever the rate is set to — the defect is in the bucket construction, so the fix belongs in code.

Why this fixes it. This is a code-level defect confined to the pipeline's own directory: the throttle added in PR #1301 and the final-retry floor from PR #1395 are present but still let the run sit exactly on OpenAI's 30/min ceiling with a full-window burst, which a rolling-window limiter cannot tolerate — hence the recurring 429 storm and the $0-billed-cost data corruption the observer raised as CRITICAL. Reducing the burst capacity and running below the ceiling is the smallest change that removes the root cause while preserving throughput (fewer 429s means fewer 65s final-retry sleeps, so the run also gets faster and stays well under the 900s timeout), and it keeps env-overridable knobs so operators retain control. A code_fix with an accompanying test is the right blast radius; a bare config tweak would leave the front-loading burst-capacity bug in place.

### Files changed

 .../openai-usage-pipeline/src/openai_client.py     | 27 +++++++++++-

.../tests/test_openai_client.py | 48 ++++++++++++++++------

2 files changed, 61 insertions(+), 14 deletions(-)

## Verification

### pytest (pipelines/runners/openai-usage-pipeline/tests) — exit 1

``

object PASSED [ 30%]

tests/test_openai_client.py::TestFetchAllProjectNames::test_single_page PASSED [ 31%]

tests/test_openai_client.py::TestFetchAllProjectNames::test_paginates_via_after_cursor PASSED [ 31%]

tests/test_openai_client.py::TestFetchAllProjectNames::test_falls_back_to_title_field PASSED [ 32%]

tests/test_openai_client.py::TestFetchAllProjectNames::test_includes_archived_projects PASSED [ 33%]

tests/test_openai_client.py::TestFetchAllProjectNames::test_project_with_no_name_or_title_is_empty_string PASSED [ 33%]

tests/test_openai_client.py::TestFetchProjectNames::test_filters_to_requested_ids PASSED [ 34%]

tests/test_openai_client.py::TestFetchProjectNames::test_empty_project_ids_returns_empty_without_a_request PASSED [ 35%]

tests/test_openai_client.py::TestFetchProjectNames::test_missing_project_id_is_dropped_not_erroring PASSED [ 35%]

tests/test_openai_client.py::TestFetchProjectNames::test_missing_name_returns_empty_string PASSED [ 36%]

tests/test_openai_client.py::TestRequestWithRetries::test_retries_on_429 PASSED [ 37%]

tests/test_openai_client.py::TestRequestWithRetries::test_retries_on_500 PASSED [ 37%]

tests/test_openai_client.py::TestRequestWithRetries::test_raises_after_max_retries PASSED [ 38%]

tests/test_openai_client.py::TestRequestWithRetries::test_raises_on_non_retryable_error PASSED [ 39%]

tests/test_openai_client.py::TestRequestWithRetries::test_final_retry_wait_outlasts_rate_limit_window FAILED [ 39%]

=================================== FAILURES ===================================

___ TestRequestWithRetries.test_final_retry_wait_outlasts_rate_limit_window ____

tests/test_openai_client.py:534: in test_final_retry_wait_outlasts_rate_limit_window

assert final_wait >= 60.0, f"final retry wait {final_wait}s does not clear the rate-limit window"

E AssertionError: final retry wait 2.5406666731934515e-06s does not clear the rate-limit window

E assert 2.5406666731934515e-06 >= 60.0

------------------------------ Captured log call -------------------------------

WARNING openai_client:openai_client.py:181 Retryable HTTP 429 on attempt 1/5. Sleeping 2.0s... Body: Rate limit exceeded

WARNING openai_client:openai_client.py:181 Retryable HTTP 429 on attempt 2/5. Sleeping 4.0s... Body: Rate limit exceeded

WARNING openai_client:openai_client.py:181 Retryable HTTP 429 on attempt 3/5. Sleeping 8.0s... Body: Rate limit exceeded

WARNING openai_client:openai_client. …_(truncated)_

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | openai-usage-pipeline |

| Failing run | 4dceefea-679a-466e-8afe-4e036b82ba4e |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | c388add5320c21380d4eba87a154a9661be4c4709c659d8c24f2cd29afd0ba82 |

| Verify | failing |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev` label to take the PR over and stop it entirely.

#1559 — fix(mart-school-performance-table-3-refresh): retry mart refresh CALL o… @heimdall-keval-factory[bot]  approvedAutomated PRheimdall-driven

Automated fix for mart-school-performance-table-3-refresh — fix_class code_fix, scope tier draft.

Resolves https://github.com/AI-Builder-Team/Surtr/issues/1557

> Draft — a human must promote this before merge. Because verification is none.

## What's broken

Run 100e2337-5392-4db5-9eaf-3ec5c4829523 of mart-school-performance-table-3-refresh failed when Redshift aborted its transaction as a deadlock victim: "ERROR: deadlock detected ... Where: SQL statement \"LOCK TABLE mart_education.aerie_program_directory\" PL/pgSQL function \"sp_refresh_agg_school_performance_unit_economics_per_student_qtd\" line 33", which is the LOCK at line 41 of ddl/sp_refresh_agg_school_performance_unit_economics_per_student_qtd.sql. The deadlock detail is a lock-order inversion on two relations: our process held the AccessExclusiveLock on core_education.dim_program (acquired one line earlier, line 40) and was waiting for aerie_program_directory, while a concurrent transaction already holding aerie_program_directory was waiting for dim_program. This is a transient concurrency collision (occurrence count 1); because every write in the procedure happens after these locks inside the default procedure transaction, the abort rolled everything back and the run wrote NO rows — mart_education.agg_school_performance_unit_economics_per_student_qtd was left at its prior good state, not partially or wrongly populated (a stale, not corrupt, outcome).

Root cause. The procedure sp_refresh_agg_school_performance_unit_economics_per_student_qtd acquires table locks in the order core_education.dim_program (line 40) then mart_education.aerie_program_directory (line 41); this order is deliberately aligned with the Table 2 producer (sp_refresh_agg_school_performance_unit_economics_qtd locks dim_program at line 48 before aerie_program_directory at line 55) and with the financials mart proc, so those consumers do not deadlock each other. The deadlock counterparty is instead a concurrent transaction that touches these two relations in the OPPOSITE order — holding aerie_program_directory while requesting dim_program — most plausibly the directory refresh in mart-aerie-hubspot-refresh (sp_refresh_aerie_program_directory), which serializes only on mart_education.aerie_hubspot_refresh_writer_mutex and acquires its dim_program / aerie_program_directory locks implicitly through statement order rather than an explicit canonical LOCK sequence. Because this pipeline is completion-triggered and can overlap an in-flight directory refresh, the two transactions each grab one relation of the pair and block on the other, so Redshift's deadlock detector aborts one side. The failure is therefore a transient two-party deadlock in which our run was chosen victim, not a deterministic defect in this pipeline's SQL or Python handler.

## What this PR changes

Make this pipeline resilient to the transient deadlock with a bounded, deadlock-aware retry confined to src/ (Tier A): detect the Redshift 'deadlock detected' failure in RedshiftClient (src/redshift_client.py, where _submit_and_poll raises RuntimeError from the FAILED response at line 53) and re-submit the CALL a few times with exponential backoff, staying inside the existing Lambda-deadline budget (STATEMENT_TIMEOUT_SECONDS and invocation_deadline_monotonic). This is safe precisely because the whole refresh is a single atomic CALL statement whose default procedure transaction rolls back completely on the abort, and the procedure rebuilds the entire target from Table 2 on every run, so a retry is idempotent and cannot leave partial or duplicated rows. Add a unit test under tests/ that simulates a first-attempt deadlock followed by success and asserts the CALL is retried and the run reports success, and restrict the retry to the transient deadlock class rather than arbitrary FAILED statements. Do NOT reorder the LOCK statements in this pipeline's DDL — its order already matches the Table 2/financials consumers, and flipping it would only relocate the inversion; the durable cross-pipeline fix (giving the mart-aerie-hubspot-refresh directory writer the same canonical dim_program-then-aerie_program_directory lock order) lives outside this pipeline's directory and is a human follow-up, not required to stop this pipeline's spurious failures.

Why this fixes it. A deadlock is a transient serialization failure, and the canonical remediation for the chosen victim is a bounded retry with backoff rather than a structural rewrite — a retry is the smallest change that lets a run like 100e2337 succeed on re-attempt while leaving the deliberately-aligned lock order in the DDL untouched. It is safe and correctly scoped to Tier A because the refresh is a single atomic CALL that fully rolls back on abort and rebuilds Table 3 from Table 2 every run, so re-issuing it cannot corrupt, partially write, or duplicate data; the change stays inside src/redshift_client.py, src/handler.py, and tests/, touching no shared code, config, or other pipelines. The alternative of reordering this procedure's LOCK statements is rejected because it already matches the Table 2 producer and would merely shift the inversion elsewhere, and the true global lock-order alignment in the mart-aerie-hubspot-refresh directory writer is outside this pipeline's scope tier.

### Files changed

 .../src/redshift_client.py                         | 46 ++++++++++++-

.../tests/test_redshift_client.py | 80 ++++++++++++++++++++++

2 files changed, 125 insertions(+), 1 deletion(-)

## Verification

### pytest — no test suite

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | mart-school-performance-table-3-refresh |

| Failing run | 100e2337-5392-4db5-9eaf-3ec5c4829523 |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | 1b7221fa605b49806b2a84f756f45470da77ba6b4e2899caa8dafd8fc791ff2b |

| Verify | none |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev label to take the PR over and stop it entirely.

#115 — fix(heimdall): stop a finished answer reading as a question @kevalshahtrilogy  approved

## What this fixes

Every outcome comment heimdall posts on a Linear ticket is signed with the

awaiting-reply marker:

comment(client, ticket_id, sign(answer_comment(diagnosis, run_url)))

That marker does two jobs at once. It stops the next sweep re-picking the

ticket, and it tells a human they owe an answer. Only the first is true of

No code change from this run..

## The cost, measured

On the Surtr board today, **29 tickets carried heimdall-ready and were skipped

by the queue**. Twenty of them were finished write-ups nobody had been asked

for — sitting indistinguishable from the nine real blockers underneath. Fourteen

of those twenty were the same root cause restated fourteen times, and the

condition they described had already cleared.

The queue looked deep. It was almost entirely parked on a person who owed

nothing.

## The change

Two markers. Both stop the re-pick; only one says a human owes something.

| marker | meaning | set by |

| --- | --- | --- |

| — heimdall, awaiting your reply | a human owes an answer | a callout classify() accepts |

| — heimdall, no action needed | finished; nobody owes anything | every other outcome |

- heimdall_spoke_last() is the re-pick guard and matches either marker, so the

loop the original signing existed to prevent stays prevented. first_workable

uses it and logs which of the two it saw.

- awaiting_reply() keeps its name and narrows to what it always claimed to

mean. It becomes usable as a real "waiting on a human" signal.

- _callout_body() is the single test deciding both the body and the

signature. Asking that question in two places is how a genuine blocker ends up

signed as concluded and silently dropped.

- Both markers keep the suffix-not-substring rule, for the reason PR 58

established: Linear's reply UI quotes the parent comment, so a substring test

reads a human's reply as heimdall's own and parks the ticket permanently.

## A test that was passing for the wrong reason

test_a_real_conclusion_does_park_the_ticket parametrised a callout case with

reason="need a key". classify() refuses that as *"reason too thin to act

on (10 chars)"*, so the case never exercised the callout path — it passed only

because everything was signed awaiting regardless of what it said. It now uses a

reason classify() accepts, and the refused-callout case is asserted separately

as concluded, which is the behaviour that actually matters: a question

nobody can act on must not be parked on a person.

1180 tests pass. ruff check heimdall/ is clean. Formatting is untouched —

heimdall/ is not covered by CI's ruff format --check harness, and both files

are already unformatted on main, so reformatting would bury the change.

## Business Value

Heimdall's ticket queue is the factory's only work intake, and it was silently

converting its own completed work into a human backlog. Two thirds of the queue

was noise of its own making, growing at the rate the observer files tickets —

so the queue got less trustworthy the harder the factory worked, and the real

blockers got harder to find in it. This restores the queue as a signal: what is

in it is either workable or genuinely waiting on a person, and "waiting on a

person" now means it.

It also unblocks measurement. awaiting_reply becoming accurate is what lets a

dashboard say "the factory is blocked on N human answers" without that number

being mostly finished work.

## Manual Effort Estimate

~3 hours — most of it in the diagnosis rather than the diff: reading the

queue selection chain end to end, measuring the real board to separate finished

answers from genuine blockers, and spotting that the existing callout test was

green for the wrong reason. Flagging for Keval to confirm or adjust.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1774 — fix(repo): Tolerate small Model breakdown excess like Credit Source @heimdall-keval-factory[bot]  approvedAutomated PR

A handful of users' daily Perplexity usage was reported by the vendor a few credits over the expected total — normal rounding noise — but the pipeline treated any such overage as corrupted data and refused to publish, dropping real spend for 14 of 32 users on two recent days. The fix lets that small overage through the same way an equivalent shortfall is already tolerated.

> Ready for review. Nothing ran the change, so it is unproven. A person still merges.

## For The Agent

_Everything below is detail for review. The summary above is the change._

Presented as ready — verification none, scope tier draft, fix_class code_fix, HEIMDALL_READY_PRS=true.

### What's broken

Run ec1b49ce-ddf8-413b-9968-8328abd37ce8 (SURTR-1148) raised ValueError: 15 of 32 user-day(s) (46.88%) failed integrity checks, above the 25% ceiling; refusing to publish a materially incomplete day as a successful load after transforming Alpha/Trilogy's 2026-09-06/09-07 buckets. 14 of the 15 excluded rows share the identical reason logged just above the abort — Model breakdown for <user> on <event_date> (org=<org>) sums to <model_total>, exceeding the user-day total_count (<total_count>) by <N> (e.g. dr@telcodr.com on 2026-09-06: model_total=7824 against total_count≈7815, ~9 credits / ~0.1% over) — which is handler.py's zero-tolerance Model-breakdown excess guard in _allocate_models firing on ordinary vendor rounding noise. Only 1 of the 15 (jc.fischer@trilogy.com, 2026-09-07) failed for an unrelated, correctly-strict reason: a wholly missing Credit Source breakdown.

Root cause. PR #1742 (d8b48658, 'reconcile Credit Source tolerances with live vendor data') tolerated a small Credit Source EXCESS (categories summing above total_count) after a live run showed a real user-day 0.27% over was legitimate vendor noise, not corruption — but it only updated the Credit Source check in _transform_user_day; the sibling Model breakdown excess check in _allocate_models (handler.py, the 'sums to X, exceeding the user-day total_count' ValueError) was left with zero tolerance, even though its own docstring still claims it 'mirrors Credit Source's own excess check exactly.' Both breakdowns draw from the same underlying per-user-day reconciliation noise the module's own docstring already documents (the same sub-1% overage now observed on 14 different users across two orgs and two dates), so the class of noise Credit Source now tolerates unconditionally kills the user-day via the Model side instead, excluding real spend from the published load.

### What this PR changes

In _allocate_models (handler.py, the shortfall < 0 branch around the 'sums to ... exceeding' ValueError), reuse the tolerance and hard_limit values already computed a few lines above for the shortfall case, and only raise when the excess exceeds tolerance or hard_limit; within that budget, log a warning (matching the existing Credit-Source-excess warning's wording/level) and let the existing remainder-to-absorber and negative-remainder guards absorb the small overage exactly as they already do for the shortfall path — no new tolerance constants needed, just extending the ones already validated against live data for Credit Source. Update the tests that currently assert zero tolerance for Model excess (around test_handler.py:577-636, e.g. the small-overshoot case) to match the new tolerated-excess behavior, and add tolerated/beyond-tolerance test pairs mirroring test_small_credit_source_excess_is_tolerated and test_credit_source_excess_beyond_tolerance_still_raises (test_handler.py:968-1013).

Why this fixes it. The Model breakdown's zero-tolerance excess check in handler.py's _allocate_models is the direct cause of 14 of the 15 excluded user-days in this run, each only a few credits (well under 1%) over total_count — the same magnitude of vendor noise PR #1742 already proved safe to tolerate on the Credit Source side. Leaving it unfixed silently drops real, valid spend from every future run that hits this ordinary noise (dr@telcodr.com's single excluded user-day alone was worth roughly $78), and risks retripping the 25%-exclusion abort ceiling on any day where more than a handful of users show the same sub-1% overage.

#### Files changed

 .../perplexity-usage-pipeline/src/handler.py       | 115 +++++++++-------

.../tests/test_handler.py | 151 ++++++++++++---------

2 files changed, 151 insertions(+), 115 deletions(-)

### Verification

### pytest — no test suite

### verify: ruff check — exit 0

[notice] A new release of pip is available: 25.3 -> 26.2.1

[notice] To update, run: pip install --upgrade pip

All checks passed!

### verify: ruff format --check — exit 0

1782 files already formatted

### verify: pytest (pipeline lambdas) — exit 0

6.0

rootdir: /home/runner/_work/Surtr/Surtr/publish/pipelines/cdk/lambdas

configfile: pyproject.toml

testpaths: tests

plugins: cov-7.0.0

collected 486 items

tests/test_ai_spend_raw_api.py .............................. [ 6%]

tests/test_coordinate_fanout_run.py .................................... [ 13%]

.... [ 14%]

tests/test_create_run_record.py ........................ [ 19%]

tests/test_gchat_notifier.py ........................... [ 24%]

tests/test_generate_chunks.py ..... [ 25%]

tests/test_gsheet_tracker.py .................. [ 29%]

tests/test_load_fanout_plan.py .............. [ 32%]

tests/test_redshift_cluster_iam_role_association.py ........ [ 34%]

tests/test_registry_sync.py ......................... [ 39%]

tests/test_triage_dispatcher_handler.py ................................ [ 45%]

................... [ 49%]

tests/test_triage_dispatcher_signature.py .............................. [ 55%]

...... [ 57%]

tests/test_triage_reconciler.py ............ [ 59%]

tests/test_triage_reconciler_tracker.py ............ [ 62%]

tests/test_update_run_failed.py ........................................ [ 70%]

..... [ 71%]

tests/test_update_run_success.py ....................................... [ 79%]

....................... [ 84%]

tests/test_verify_on_demand_control.py ................................. [ 90%]

............................................ [100%]

============================= 486 passed in 1.03s ==============================

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | repo |

| Failing run | issue |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | linear-SURTR-1148 |

| Verify | none |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev label to take the PR over and stop it entirely.

#1748 — fix(perplexity-usage-pipeline): enforce the skip ceiling per org-day too @kevalshahtrilogy  approvedheimdall-driven

KLAIR-3477 — addresses the blocking mercy finding on release PR #1747 (code from #1746, already on main).

## The finding

> The 25% ceiling is calculated across all orgs and dates together. […] if a small org has its only user's day skipped while the run contains many user-days from other orgs, the global percentage can remain below 25%, and the load would publish an empty day for that org.

Mercy is right, and this is a silent-data-loss bug. The existing completeness guards verify every org *answered* every day — they do not verify it *produced rows*. So after skipping:

1. Alpha's only user-day for 2026-08-15 fails an integrity check → skipped

2. Run-wide ratio is 1 of 9 = 11%, comfortably under the ceiling → publishes

3. The upsert is a delete-over-window per org-day → Alpha's previously loaded rows for that day are deleted and replaced with nothing

4. Run reports partial_failure, so it looks handled

Publishing happens at (org, event_date) grain, so the completeness ceiling has to be enforced at that grain. A run-wide average structurally cannot see one day being emptied.

## The fix

- Tally attempted user-days per (org, event_date) alongside the run-wide count

- Apply the same 25% ceiling at that grain, raising with the offending org-days named (Alpha/2026-08-15 (1 of 1))

- Keep the run-wide check — it still catches a broadly degraded run whose damage is spread evenly enough that no single org-day trips

- setdefault for the new counter, so a caller passing a partial accumulator gets correct behaviour rather than a KeyError

## Verification

337 tests pass. Two handler-level tests covering both directions of the boundary:

- the finding's exact scenario — Alpha 1-of-1 skipped, Trilogy 8 good rows, 11% run-wide → raises, naming Alpha/2026-08-15, and asserts upsert is never called (nothing may be written when a day would publish incomplete)

- the inverse — 1 bad row among 8 in one org-day (12.5%) still publishes the other 7 as partial_failure, so the ceiling isn't so tight that skipping never applies

ruff format / ruff check clean on 0.15.22.

## Business Value

Protects the AI-spend feed from a failure mode worse than the one #1746 fixed: silently deleting a small org's real spend data and reporting a handled run. Alpha is exactly the shape at risk — fewer users than Trilogy, so a single unreconciled row is a much larger share of its day. Without this, the more we skip, the more likely we quietly erase a whole org-day.

## Manual Effort Estimate

~1 hour focused, no AI. Small change; the work is in seeing that "publishes at org-day grain" and "ceiling measured run-wide" are a mismatch, and in the test that asserts the upsert never fires. *Keval: proposed number, please confirm or adjust.*

#113 — feat(mercy): trust heimdall-driven PRs, and stop summoning the agent after approval @kevalshahtrilogy  no labels

## Why

Two things were keeping the agent's loop from ever finishing, and they compound.

1. A driven PR could never clear the sensitive-path hold. The counterpart-bot exemption keys on the PR author. heimdall-driven marks a PR the agent *maintains* — often one a human opened — and no action the agent can take changes who opened it. So those PRs were held on sensitive paths indefinitely, and the heimdall-driven flag mercy already computes was used for the hand-off but not for the guards.

2. An approval restarted the loop instead of ending it. The hand-off fired on final_event == 'REQUEST_CHANGES' || inline_count > 0. So a review whose only findings were nits summoned the agent, the agent pushed, the push invalidated the approval, and mercy reviewed again — running on exactly the PRs already judged good enough to merge.

Observed on AI-Builder-Team/Surtr#1738: eight review rounds, several of them started by a review that had no blocking findings at all.

## What changed

- compute_guards takes heimdall_driven, which lifts the sensitive-path, bot-author and review-round holds — the ones that ask *who do we trust*. Plumbed through as --heimdall-driven from the flag mercy already resolves.

- The hand-off fires only on REQUEST_CHANGES. The now-unreachable "non-blocking suggestions" summons text is gone, and the comment above the step no longer describes behaviour it no longer has.

## What the label deliberately does not lift

Empty diff, truncated diff, failing CI, coverage gaps. Those say whether the review was valid, not who to trust — approving past them makes the approval meaningless rather than faster. A test pins each one, so extending the label's reach has to be a deliberate act rather than a side effect.

Worth a decision from you: the brief was "never withhold approval on a heimdall-driven PR", and I have not applied it to those four. Failing CI is the one most likely to matter in practice — I left it holding because an approval on a red build is a rubber stamp, and required checks gate the merge anyway so lifting it buys nothing. Say the word and I'll extend it.

## Testing

harness/tests: 376 passed (5 new). heimdall/tests: 1171 passed, 1 skipped. ruff check, ruff format --check, actionlint all clean.

The new contract test asserts the hand-off condition mentions REQUEST_CHANGES and no longer mentions inline_count, so the trigger can't quietly widen again.

## Business Value

This is the difference between the agent loop terminating and not. Every extra round costs a model run on both sides and delays a merge that was already earned, and the two faults reinforced each other: the PRs most likely to touch sensitive paths are the agent's own infrastructure work, which is exactly what it needs to be able to land. Surtr#1738 spent eight rounds on a change that had no blocking findings after the third.

## Manual Effort Estimate

~1.5 hours — small diff, but the reasoning about which holds may safely lift is the substance of it. Keval to confirm/adjust.

#112 — feat(heimdall): label every pipeline-filed ticket `surtr-pipeline` @kevalshahtrilogy  no labels

## Why

Tickets filed from pipeline alerts carried one label, heimdall-ready, which says what may pick the ticket up. There was nothing saying what created it.

That matters because removing heimdall-ready is precisely how a human takes work away from the agent — and with a single label, that same click also drops the ticket out of any board filter for "everything a pipeline raised". Provenance and queueing are different questions and now move independently.

## What changed

- file_failure.py resolves a second label (--pipeline-label, default surtr-pipeline) and files both.

- create_ticket takes label_ids; label_id still works for single-label callers. Empty ids are dropped and duplicates collapse, so an unresolvable label degrades to filing without it rather than a rejected mutation — the same trade resolve_team_and_label already makes, because an untagged ticket beats a failure that reached nobody.

- Workflow input linear_pipeline_label, defaulting to surtr-pipeline, read as ${VAR-default} so an explicit empty string disables it instead of being silently defaulted.

A missing queue label and a missing provenance label are independent failures, and both have tests — losing the label that queues work must not lose the one that records where the ticket came from.

## Already applied in Linear

- Created the surtr-pipeline label in the SURTR team. It could not be created workspace-wide (You must be a workspace admin to create workspace labels), so it is team-scoped. resolve_team_and_label prefers a team label and falls back to a workspace one, so if an admin later promotes it nothing here changes.

- Backfilled all 31 existing tickets filed from pipeline alerts — identified by the heimdall-signature: marker in the description rather than by title or age, so the set is exactly what this code path produced. 31 labelled, 0 failed, including closed ones so historical filters are complete.

## Testing

heimdall/tests: 1171 passed, 1 skipped — 5 new. ruff check and actionlint clean.

The test fake previously resolved every label name to the same id, which would have hidden whether both labels actually reached issueCreate; it now returns a distinct id per name and can be told a label is missing.

## Business Value

The board becomes filterable by origin, which is what makes the pipeline-fault queue reviewable at all: "show me every fault a pipeline raised this week" is currently unanswerable without knowing the agent queue's internal name, and becomes wrong the moment someone parks a ticket. Cheap to add now, and retrofitting provenance onto months of tickets later would not be.

## Manual Effort Estimate

~45 minutes — small code change, plus creating the label and backfilling the existing tickets by marker. Keval to confirm/adjust.

#1773 — fix(repo): Add voyage to TFY known-skipped providers @heimdall-keval-factory[bot]  approvedAutomated PR

Voyage AI is a new TrueFoundry provider account that Surtr has no usage feed for, but it isn't yet on the daily sync's do-not-warn list, so it gets flagged as a surprise instead of silently skipped like five similar providers already are. This adds Voyage to that list so the warning stops recurring.

> Ready for review. Nothing ran the change, so it is unproven. A person still merges.

## For The Agent

_Everything below is detail for review. The summary above is the change._

Presented as ready — verification none, scope tier draft, fix_class code_fix, HEIMDALL_READY_PRS=true.

### What's broken

Run 09baf26f-e5dc-4aa5-9729-d62b430ea0b3 finished status=partial with two UNRESOLVED_TFY_PROVIDER warnings: secret_name=VOYAGE_API_KEY provider=voyage reason=unknown provider value in API contract and secret_name=OPENAI_DAYBREAK_KEY provider=openai reason=expected one OpenAI user match, found 0. The Voyage warning is a categorization gap in reconciler.py: voyage has no Surtr dedupe feed, exactly like the six providers already in KNOWN_SKIPPED_PROVIDERS, but it is missing from that set so it falls into the 'unknown provider' branch meant for genuine contract surprises. The OpenAI DAYBREAK warning is the pipeline correctly reporting a real account whose usage rows have not yet appeared in staging_finance_ai_spend.raw_openai_token_usage, which is expected behavior per this pipeline's README, not a defect.

Root cause. reconciler.py's KNOWN_SKIPPED_PROVIDERS set (lines 12-19) predates Voyage AI being onboarded as a TFY provider account. Because 'voyage' matches neither SUPPORTED_PROVIDERS nor KNOWN_SKIPPED_PROVIDERS, build_reconciliation_plan() (line 54-59) classifies it as unresolved with reason 'unknown provider value in API contract', a warning meant for genuinely unexpected contract values, so every daily run now emits a spurious investigation warning for a provider Surtr was never going to resolve.

### What this PR changes

Add 'voyage' to KNOWN_SKIPPED_PROVIDERS in pipelines/runners/tfy-provider-secrets-sync/src/reconciler.py, matching the existing entries for azure-foundry, elevenlabs, fireworks, google-vertex, perplexity-ai, and xai, and add a regression test in tests/test_reconciler.py asserting a voyage secret lands in plan.skipped with reason 'provider has no Surtr dedupe resolver' rather than in plan.unresolved. The OPENAI_DAYBREAK_KEY warning needs no code change; a human should confirm with whoever owns the OpenAI usage ingestion whether that TFY account is expected to start accruing rows in raw_openai_token_usage, since no repo evidence indicates a matching/parsing bug in resolve_openai.

Why this fixes it. This is a one-line addition to an existing, already-tested allowlist (KNOWN_SKIPPED_PROVIDERS) confined entirely to the pipeline's own reconciler.py, following the exact precedent of six other providers already skipped there for lacking a direct-provider feed. No new resolver logic, schema change, or SQL is touched, so it carries none of the risk of the transactional Redshift reconcile path.

#### Files changed

 pipelines/runners/tfy-provider-secrets-sync/src/reconciler.py     | 1 +

.../runners/tfy-provider-secrets-sync/tests/test_reconciler.py | 8 ++++++++

2 files changed, 9 insertions(+)

### Verification

### pytest — no test suite

### verify: ruff check — exit 0

[notice] A new release of pip is available: 25.3 -> 26.2.1

[notice] To update, run: pip install --upgrade pip

All checks passed!

### verify: ruff format --check — exit 0

1782 files already formatted

### verify: pytest (pipeline lambdas) — exit 0

6.0

rootdir: /home/runner/_work/Surtr/Surtr/publish/pipelines/cdk/lambdas

configfile: pyproject.toml

testpaths: tests

plugins: cov-7.0.0

collected 486 items

tests/test_ai_spend_raw_api.py .............................. [ 6%]

tests/test_coordinate_fanout_run.py .................................... [ 13%]

.... [ 14%]

tests/test_create_run_record.py ........................ [ 19%]

tests/test_gchat_notifier.py ........................... [ 24%]

tests/test_generate_chunks.py ..... [ 25%]

tests/test_gsheet_tracker.py .................. [ 29%]

tests/test_load_fanout_plan.py .............. [ 32%]

tests/test_redshift_cluster_iam_role_association.py ........ [ 34%]

tests/test_registry_sync.py ......................... [ 39%]

tests/test_triage_dispatcher_handler.py ................................ [ 45%]

................... [ 49%]

tests/test_triage_dispatcher_signature.py .............................. [ 55%]

...... [ 57%]

tests/test_triage_reconciler.py ............ [ 59%]

tests/test_triage_reconciler_tracker.py ............ [ 62%]

tests/test_update_run_failed.py ........................................ [ 70%]

..... [ 71%]

tests/test_update_run_success.py ....................................... [ 79%]

....................... [ 84%]

tests/test_verify_on_demand_control.py ................................. [ 90%]

............................................ [100%]

============================= 486 passed in 1.07s ==============================

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | repo |

| Failing run | issue |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | linear-SURTR-1146 |

| Verify | none |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev label to take the PR over and stop it entirely.

#1770 — fix(education): retain billing allocations without contact links @benji-bizzell  approved

## Summary

- Preserve billing allocations when Finalsite returns a null billing-contact relationship, projecting unavailable contact-derived IDs as NULL.

- Keep source records, allocation amounts, explicit school-year IDs, and existing structural and reconciliation checks intact.

## Why

The Scottsdale production validation failed on 12 allocations with billing_contact: null on its final response page. The extractor rejected null traversal even though the warehouse permits nullable contact and school-year IDs, preventing the entire tenant snapshot from publishing.

## Business Value

Missing contact links no longer prevent otherwise valid billing data from refreshing or cause allocations to be discarded.

## Test plan

- [x] 237 billing tests pass, including null relationships and amount mismatch with a null contact.

- [x] Repository-wide Ruff lint and format checks pass with CI's pinned version (0.15.22).

- [x] Local extraction of all three version-pinned production response pages: 1,143 billing items and 2,106 allocations, all 12 null-contact allocations retained, parent/allocation totals reconciled. Response SHA-256 values verified; no production writes.

- [ ] After deployment, rerun Scottsdale-only validation and verify warehouse publication. No DDL change is required.

#1768 — fix(education): leave Finalsite reader access under DBA control @benji-bizzell  no labels

## Summary

- Remove reader allowlists and hardcoded grant/revoke policy from both Finalsite DDL paths.

- Preserve existing direct grants transactionally when replacing tables, views, or the billing balance column.

- Document DBA ownership of reader access while retaining migration and writer safeguards.

## Why

Release #1766's DDL preflight stopped because enrollment rejected billing's legitimate reader grants in their shared schema, and billing rejected unrelated database permissions. Pipeline migrations must not impose their own policy over DBA-managed access.

## Business Value

Authorized schema changes can proceed without revoking legitimate access or requiring each pipeline to understand every reader in the warehouse.

## Test plan

- [x] 134 enrollment/capture and 235 billing tests pass, including arbitrary reader and column-grant preservation and failed-audit guards.

- [x] Repository-wide Ruff lint/format checks pass with CI-pinned 0.15.22.

- [x] Generated DDL/history parity and both DDL dry runs pass.

- [x] Seven-lane review; corrected column-catalog query and atomic batch sizing findings.

- [x] Read-only production catalog probe: capture 36 grants compact to 4 (31 total statements); billing 4 grants compact to 2 (23 total).

- [ ] DBA-coordinated live migration and grant readback after review. No production writes in this PR task; new-object reader access remains DBA-provisioned.

#1765 — fix(education): restore SIS dataset-triggered refreshes @benji-bizzell  no labels

## Summary

- Forward existing trigger context to Core Enrollment and cover both Enrollment and School Source Directories with regression tests.

- Allow the existing forwarding option for dataset-only Lambda consumers, with regression coverage for the generated workflows and handler paths.

## Why

Production run 5ef89c7c-b1a3-4292-97ea-30484286b9e5 failed before warehouse work because the workflow omitted trigger metadata required by the handler. School Directories forwarding is already enabled on main. Core Enrollment still has the same configuration gap; its dataset-only trigger needs the schema to accept the forwarding option without adding a pipeline-success subscription.

## Business Value

Restore these consumers' ability to process pinned SIS publications while preserving source validation and existing QuickBooks behavior. Shared workflow implementation and other pipelines' payloads remain unchanged.

## Test plan

- [x] 77 runner tests and Python lint passing.

- [x] CDK typecheck and generated-workflow checks for both actual pipeline configurations.

- [x] Hosted CI is green, including the full CDK suite and pipeline runner suite. This covers the six Docker-dependent shared-stack tests unavailable locally.

- [x] Seven-lane adversarial review completed; duplicate keys introduced by a concurrent base merge were removed.

- [x] Mercy reviewed 55baca4e592275c5e4223e71d203907c7ea46af0: no findings; human approval retained for CDK paths.

- [ ] After separately authorized deployment, verify both deployed invocation payloads and fresh publication results.

#1763 — feat(education): capture Finalsite findings without blocking ingestion @benji-bizzell  changes requested

## Summary

- Capture unavailable Finalsite contacts and source findings without aborting otherwise successful ingestion.

- Publish observations and capture evidence atomically, with separate acquisition, health, and consumer-acceptance semantics.

- Align ingestion conventions around source fidelity while preserving operational failures and governed consumer gates.

## Why

A direct 404 for a previously tombstoned Austin contact prevented fresh data for the entire estate. Historical reconciliation should record unavailable identities, not require every historical contact to remain readable forever. Listing absence and 404s also do not establish source deletion.

## Business Value

Healthy source records continue to arrive while missing identities remain auditable. Operators can distinguish routine source observations from broken acquisition or publication, without silently turning incomplete evidence into accepted business data.

## Breaking changes

New publications use capture_complete instead of redefining legacy complete. No new absence-derived tombstones are generated. The additive capture table must be provisioned before deployment; existing consumers keep their acceptance rules until a separately verified cutover. No production DDL or deployment is included in this task.

## Test plan

- [x] 140 Finalsite tests, including unavailable/listed contacts, parent-child independence, quarantine, volume findings, transaction failure and immutable retry reconciliation.

- [x] Repository-wide Ruff checks and formatting with CI-pinned Ruff 0.15.22.

- [x] Generated DDL/history-migration checks and DDL applier dry run.

- [x] Seven-lane adversarial review; fixed persistent volume health, scoped alerts, and post-commit retry reporting.

- [x] All current-head CI checks green at c8b9e677.

- [x] Mercy reviewed this head; earlier findings withdrawn after executable rebuttals. Final empty-observation finding rejected with contract/code/test evidence in its review thread.

- [ ] Human review must clear Mercy’s remaining changes-requested state; same-head automated reassessments are exhausted.

- [ ] Before production deployment, verify external consumer gates and apply audited additive DDL; then run an authorized fresh full capture.

#1764 — fix(repo): forward trigger context for dataset publication runs @heimdall-keval-factory[bot]  approvedAutomated PR

The organizations dataset trigger for this pipeline has never actually worked: pipeline.json never enabled the flag that forwards trigger context to the Lambda, so every run it starts is rejected before touching data. The QuickBooks-triggered half keeps working; enabling that one flag (already used by two sibling pipelines) fixes the SIS side.

> Ready for review. Nothing ran the change, so it is unproven. A person still merges.

## For The Agent

_Everything below is detail for review. The summary above is the change._

Presented as ready — verification none, scope tier draft, fix_class code_fix, HEIMDALL_READY_PRS=true.

### What's broken

Run 5ef89c7c-b1a3-4292-97ea-30484286b9e5 (Lambda request cb05a6b6-594e-40f7-9964-6cd5f313d53c) failed at pipelines/runners/mart-aerie-school-source-directories-refresh/src/handler.py:138 with ValueError: unsupported params: ['dataset_run_id'], raised from _parse_event at handler.py:39. The on_dataset_publication trigger for sis-raw-sync:organizations sends trigger_type: EVENT and triggered_by: dataset:sis-raw-sync:organizations into the state machine, but pipeline.json's triggers block omits forward_upstream_execution_context: true, so the non-chunked InvokePipelineLambda step in pipelines/cdk/lib/constructs/step-function.ts (~line 411-425) strips trigger_type/triggered_by before invoking the Lambda, leaving only run_id and params: {dataset_run_id: ...}. No SIS directory rows were written for this run since the SQL call was never reached; the QuickBooks-triggered half of this pipeline is unaffected.

Root cause. pipeline.json's triggers block is missing forward_upstream_execution_context: true. That flag gates whether the state machine's direct-flow InvokePipelineLambda step forwards trigger_type/triggered_by (and upstream_execution_arn) from the execution input into the Lambda payload (pipelines/cdk/lib/constructs/step-function.ts ~411-425); without it only run_id and params reach the Lambda. handler.py's _parse_event (line 33-39) treats any dataset_run_id param as unsupported unless trigger_type == 'EVENT' and triggered_by == 'dataset:sis-raw-sync:organizations' are also present in the event, so every run started by the SIS organizations dataset-publication trigger fails this exact check since the pipeline's introduction (commit 46c1b628).

### What this PR changes

Add "forward_upstream_execution_context": true to the triggers block in pipelines/runners/mart-aerie-school-source-directories-refresh/pipeline.json. The pipeline already declares on_pipeline_success: ["quickbooks-raw-sync"], which satisfies pipelines/cdk/lib/schema/pipeline-config.ts's rule that the flag requires a non-empty on_pipeline_success — so this is a schema-valid, one-line change with no other config or IAM implications (handler.py never reads upstream_execution_arn, so no new states:DescribeExecution permission is needed). This is the identical pattern already deployed for pipelines/runners/core-education-student-school-year-snapshots/pipeline.json and pipelines/runners/netsuite-saved-search-refresh/pipeline.json, both of which forward trigger_type/triggered_by to a handler with the same closed-params contract. tests/test_pipeline_contract.py asserts an exact triggers dict (lines 13-16) and must be updated to include the new key alongside the config change.

Why this fixes it. handler.py's closed-params validation has required trigger_type/triggered_by alongside dataset_run_id since this pipeline was created, but pipeline.json never opted into forward_upstream_execution_context, so every invocation from the sis-raw-sync:organizations dataset trigger has failed this check from day one — this is not a transient error. The fix is a single boolean addition to pipeline.json's triggers block plus a matching update to the hardcoded triggers assertion in tests/test_pipeline_contract.py; both files live entirely inside the pipeline's own directory (Tier A), the change is schema-valid, and it requires no IAM or shared-code edits.

#### Files changed

 .../runners/mart-aerie-school-source-directories-refresh/pipeline.json | 3 ++-

.../tests/test_pipeline_contract.py | 1 +

2 files changed, 3 insertions(+), 1 deletion(-)

### Verification

### pytest — no test suite

### verify: ruff check — exit 0

[notice] A new release of pip is available: 25.3 -> 26.2.1

[notice] To update, run: pip install --upgrade pip

All checks passed!

### verify: ruff format --check — exit 0

1781 files already formatted

### verify: pytest (pipeline lambdas) — exit 0

6.0

rootdir: /home/runner/_work/Surtr/Surtr/publish/pipelines/cdk/lambdas

configfile: pyproject.toml

testpaths: tests

plugins: cov-7.0.0

collected 486 items

tests/test_ai_spend_raw_api.py .............................. [ 6%]

tests/test_coordinate_fanout_run.py .................................... [ 13%]

.... [ 14%]

tests/test_create_run_record.py ........................ [ 19%]

tests/test_gchat_notifier.py ........................... [ 24%]

tests/test_generate_chunks.py ..... [ 25%]

tests/test_gsheet_tracker.py .................. [ 29%]

tests/test_load_fanout_plan.py .............. [ 32%]

tests/test_redshift_cluster_iam_role_association.py ........ [ 34%]

tests/test_registry_sync.py ......................... [ 39%]

tests/test_triage_dispatcher_handler.py ................................ [ 45%]

................... [ 49%]

tests/test_triage_dispatcher_signature.py .............................. [ 55%]

...... [ 57%]

tests/test_triage_reconciler.py ............ [ 59%]

tests/test_triage_reconciler_tracker.py ............ [ 62%]

tests/test_update_run_failed.py ........................................ [ 70%]

..... [ 71%]

tests/test_update_run_success.py ....................................... [ 79%]

....................... [ 84%]

tests/test_verify_on_demand_control.py ................................. [ 90%]

............................................ [100%]

============================= 486 passed in 1.05s ==============================

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | repo |

| Failing run | issue |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | linear-SURTR-1119 |

| Verify | none |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev label to take the PR over and stop it entirely.

#1735 — fix(ramp): retry transient API timeouts @ashwanth1109  no labels

## Summary

- Add a pooled requests.Session to the Ramp client with bounded retries for transient connection, read, and HTTP status failures.

- Use exponential backoff with Retry-After support while preserving fail-closed week handling.

- Add retry configuration coverage and mark the HTTP resilience improvement complete.

## Business Value

Prevents transient api.ramp.com timeouts from dropping entire weeks of Ramp spend data and causing the weekly spend pipeline to fail. This keeps spend reporting complete and reduces manual reruns and missing-data risk.

## Implementation Effort

Estimated 1–2 engineering days for an average engineer, including implementation, focused tests, deployment, and live validation.

## Validation

- uv run pytest — 204 passed.

- Ruff formatting and checks passed for the modified files.

- Candidate deployed only to Pipeline-ramp-spend-pipeline-prod ([1/1]), CloudFormation reached UPDATE_COMPLETE.

- Production validation execution succeeded with 36/36 weeks fetched and transformed, 0 failed weeks, and 12,804 transactions. Historically failing weeks 2, 6, 20, and 21 all succeeded.

## Linear

- [SURTR-1044 — Investigate failed pipelines and push up fixes](https://linear.app/builder-team/issue/SURTR-1044/investigate-failed-pipelines-and-push-up-fixes)

#1244 — AERIE-1855: Add unpublished verified materialization and installer commands @caina-barbosa  approved

## Identity

- Slice: AERIE-1855 — telemetry rollout 9/19

- Repository: AI-Builder-Team/Aerie

- Category: Feature

- Base: c016f953bfdf52f474d62f675d12ce83c49d8050

- Reviewed head: d826ea4a4d1da3aaf99e04baba54077df5c4343c

- Strict predecessor: AERIE-1854 / PR #1243

## Outcome

Adds verified, atomic local Skill materialization and install/update/logout orchestration to the existing private @aerie/add-aerie-skill package. Complete descriptor/delivery proof and every ordered file byte are verified before canonical or host writes. Installation state and stable receipt lineage are promoted only after verified materialization.

## Why this slice is independently useful

This provides the unpublished installer foundation required by later host-adapter and artifact-lifecycle slices without activating invocation telemetry or publishing a package.

## Owned surfaces changed

All 32 changed paths are under packages/add-aerie-skill/, including:

- Canonical bundle materialization and host reconciliation

- Installation-wide transaction, immutable state generation, and receipt queueing

- Safe host discovery, selection, compatibility floors, and Codex/Pi coupling

- Crash recovery, physical attribution, bounded no-follow reads, and private-root security

- CLI install/update/logout service wiring

- Deterministic package/build/pack guards and proprietary notice

## Verification-only surfaces inspected and unchanged

- Root and workspace lockfiles

- Generated Convex files

- Release and publication workflows

- Existing backend telemetry and adoption surfaces

## Prohibited surfaces isolation proof

- No invocation observer, adapter registration, uploader, wake path, daemon, host hook/config edit, artifact lifecycle, UI, or macOS claim

- No production, deployment, registry, or real-host action

- No package publication path; package remains private: true, UNLICENSED, version 0.1.0

- Package files remain exactly dist, README.md, and PROPRIETARY-NOTICE.txt

- No declared runtime dependencies

## Behavior and production effect

Production effect is dormant because the package is private and unpublished. All filesystem qualification uses injected deterministic fake/temp roots. Native Linux and Windows qualification remains owned by later slices.

## Defaults and activation boundary

Only the public grammar <slug>, update, logout, help, and version is retained. There is no login/auth, discovery, search, list, uninstall alias, observer, or automatic uploader. Network, auth, host, security, confirmation, and receipt transports remain injected boundaries.

## Tests and validation

- RED evidence: exact regressions cover proof/identity drift, foreign inodes, symlink/reparse boundaries, bounded descriptor reads, crash points around every journal/rename/cleanup transition, durable .next promotion, target authority, cross-volume targets, lock races, queue lineage, and delayed receipts.

- GREEN focused tests: 256/256 in both declared concurrent and serial modes; focused receipt-settlement, orphan-lineage, compacted receipt reuse, partial-update, recovery fail-fast, CLI exit, and host-discovery tests also passed.

- Typecheck: pnpm --filter add-aerie-skill typecheck

- Format/lint: Biome passed all 69 affected package files.

- Build: pnpm --filter add-aerie-skill build

- Deterministic build: 84 files; hash 9f993788479893c1b64ae84feba41ef4b04b1e0b8cf4de1c34afbf488318e5b0; declared dependencies [].

- Pack: 87 entries; integrity sha512-JuYELRgt3Kyg2VyFn/Z7A656DxKOcw5s2X1b6nXW72BeHjJsIqnI2eKOAv0tqz5r2Wd6kgXLp5cmrlIKCL8e1Q==.

- Proprietary notice: 336 bytes; SHA-256 0592d88bfc1c4620224949906674d85de7c6dbc1624f8e8a6df9b2eda04110c1.

- git diff --check: passed.

- Worktree and artifact checks: clean; no tarball, generated, lockfile, or outside-package drift.

## Independent exact-range review

- Reviewer: a1958r-a1855audit

- Final exact range: c016f953bfdf52f474d62f675d12ce83c49d8050..d826ea4a4d1da3aaf99e04baba54077df5c4343c

- Final Mercy repair: 9412a37756331b641012d5a799c5aa436d7e3bc6..d826ea4a4d1da3aaf99e04baba54077df5c4343c

- Verdict: PASS

- Review history: repeated exact-range cold reviews drove repairs for sealed journal provenance, installation-wide locking, immutable receipt/state lineage, physical attribution, all rename/journal crash boundaries, explicit host-target authority, same-volume staging, durable rollback intent, .next convergence, nested lock serialization, terminal/compacted/orphan receipt consistency, partial-update reporting, recovery fail-fast, and fail-closed malformed host discovery. The same reviewer reproduced Mercy's scenarios and re-ran exact package gates on the final head.

## Manual QC

- Required now: none; deterministic fake-host and local pack evidence complete.

- Deferred: native Linux and Windows host qualification in AERIE-1858/AERIE-1859.

## Rollback

Revert this PR. The package is unpublished, so rollback has no registry or production deployment effect.

## Risks and monitoring

- Filesystem transactions are intentionally fail-closed on ambiguous, foreign, malformed, or unverifiable evidence; manual recovery can be required rather than risking deletion.

- No native-platform support claim is made in this slice.

- No runtime monitoring is activated because there is no published or production path.

## Local integration evidence

- Cherry-pick/rebase result: clean integration head d826ea4a4d1da3aaf99e04baba54077df5c4343c

- Sequential-history proof: current origin/main is c016f953bfdf52f474d62f675d12ce83c49d8050; candidate is 34 commits ahead and 0 behind. The reviewed candidate package tree is 11814902a1d1dbb31cbe2a14167dace7b052ff18; the preceding 33 rebased patch IDs matched their pre-rebase equivalents one-for-one.

- Cross-slice isolation: 32/32 changed paths are under packages/add-aerie-skill/; no outside-package drift.

Linear: https://linear.app/builder-team/issue/AERIE-1855/telemetry-rollout-919-add-unpublished-verified-materialization-and

#1762 — fix(education): preserve Finalsite billing balance precision @benji-bizzell  no labels

## Summary

- Preserve Finalsite billing balances as lossless numeric text while keeping parent/allocation amounts under the strict cents contract.

- Decode source decimals without binary floats and version the translation/hash contract.

- Add an explicit transactional migration, dependency restrictions, and a runtime guard against the old balance type.

## Why

Release #1747's Scottsdale smoke stopped on a source balance containing fractional precision beyond two decimal places. The precision was already present in the immutable API response; changing only the parser would not resolve the cents-only guard. Rounding would invent source data.

## Business Value

Restore accurate billing ingestion without discarding source precision or weakening parent/allocation reconciliation and publication safeguards.

## Breaking changes

raw_billing_items.balance and billing_items_current.balance change from DECIMAL(18,2) to VARCHAR(256). Numeric consumers need an explicit interpretation policy. In v2 source_record, JSON decimal numbers use tagged text envelopes to prevent SUPER float coercion; integers remain integers. The documented migration preserves historical stored values and lineage; external consumers/dependencies require review before application.

## Test plan

- [x] 238 billing runner tests passing; full pipeline Ruff lint/format passing with CI's pinned version.

- [x] Offline replay of the exact failed page: 500 parents, 1,047 allocations, all balances preserved, parent/allocation amounts reconciled. No source API calls or publication.

- [x] Synthetic read-only Redshift SUPER round trip confirms exact decimal-envelope preservation.

- [x] Seven-lane adversarial review complete; fixed exponent-expansion and SUPER precision-loss findings.

- [x] Regression coverage for precision, null/invalid values, schema preflight, and atomic migration/view restoration assembly.

- [ ] Separately authorized dev Redshift migration, dependency-failure rollback, and precision round trip.

- [ ] Separately authorized complete Scottsdale smoke; full-estate baseline and schedule activation remain gated.

No production migration, deployment, retry, fence change, or schedule change is included in this PR operation.

#1761 — fix(education): preserve lossless Limitless transcripts @benji-bizzell  approved

## Summary

- Preserve Limitless meeting transcripts as lossless SUPER values, using the existing audio-transcript representation.

- Add a targeted transactional view migration and pinned replay runbook, with required-field and byte-boundary regression coverage.

## Why

Release #1747 validation failed on a 64,989-byte transcript: the encoder chunks above 60,000 bytes, and the clean VARCHAR projection rejected that marker. The source was intact, but the validation correctly blocked the complete 57-table publication. This extends the explicit column contract without relaxing required fields or truncating text.

## Business Value

Unblocks source-faithful GuidePlatform refreshes so fresh roster and meeting data can become available after the separately authorized recovery rollout.

## Breaking changes

staging_education_guide_platform.limitless_meetings.transcript_text changes from VARCHAR to SUPER. Consumers must accept scalar strings or ordered, hashed chunks. Migration requires the documented owner/ACL/dependency preflight; deployment and pinned replay remain separate gates.

## Test plan

- [x] 92 targeted tests, including generated DDL parity, required-null rejection, migration content, and ASCII/multibyte boundaries.

- [x] Repository-wide Ruff 0.15.22 lint and format checks; git diff check.

- [ ] Authorized rollout: rehearse transactional rollback, migrate and verify catalog/ACLs, deploy matching code, replay pinned evidence, and reconcile all 57 tables. No production changes performed by this PR preparation.

#1245 — feat(platform-errors): make automatic triage runs inspectable @benji-bizzell  changes requested

## Summary

- Make automatic triage attempts inspectable in Platform Errors, with paginated history, assessments, runtime/handoff outcomes, cost, and protected retained-session evidence.

- Preserve useful results and returned telemetry independently of handoff validation; replace post-run token rejection with observe-only cost review.

- Show actual scheduling status while retaining isolated single-class dispatch and exact-revision receipt checks.

## Why

The previous canary produced a useful handoff but was marked failed by an aggregate-token check after spending had already occurred. Operators also lacked the run history and session access needed to assess useful findings, failures, and cost. This closes that operational visibility gap without claiming a hard spending cap or a complete forensic archive.

## Business Value

Enables a controlled observation period where operators can assess real triage value and investigate failures from the existing Platform Errors screen.

## Test plan

- [x] Platform Errors/UI tests, Worker tests, targeted contracts tests, Chat/Worker/contracts typechecks, Worker production build, and bounded-read/test-architecture checks.

- [x] Unauthorized access, stored-identity reads after secret rotation, bounded evidence, separate runtime/handoff outcomes, unknown cost, and pagination regressions.

- [ ] Preview/browser validation of the new Triage runs tab.

- [ ] Deploy disabled; validate one exact canary including session evidence and receipt; explicitly enable scheduling and verify a scheduled run.

Scheduling is unchanged. $5/run is a review reference, not a hard cap. Session inspection is bounded and may be incomplete after compaction; old runs are not automatically backfilled.

#111 — feat(review): make blocking findings name the defect class to fix @kevalshahtrilogy  no labels

## What

Blocking findings now carry a class_hint — one or two sentences naming the defect class the finding is an instance of, and where its siblings are. It renders under the finding as _Fix the class:_.

src/heimdall/board.ts:1595safeRunField blanks a non-string signature…

_Fix the class:_ Class: a field is rejected but nothing records the rejection.

Check every field read through safeRunField, not just this one.

## Why

A finding reports one instance of a defect. Authors reliably fix exactly the line cited — so the next round reports the next instance, and the round after that reports the third.

Measured on Surtr #1727 (6,942 lines, this reviewer):

per-line fixes  →  2 3 2 1 5 7 7 3 4 7    10 rounds, never converged

class sweeps → 3 2 3 2 2 0 approved in 6

Same reviewer, same author, same PR. The only thing that changed was whether each round fixed *the cited line* or *every site of the class*. Three of one round's seven blocking findings were provably "same defect, one field over" from findings fixed two rounds earlier.

The hint costs mercy one sentence and saves the author several rounds.

## Details

- Blocking findings only. On a nit, a "fix the whole class" instruction is noise, and noise is how a marker stops being read.

- Carried through the ledger and carry-forward. A carried finding is precisely the one whose class went unfixed; restoring it without the hint would restore it without the thing that stops it recurring.

- Schema-optional on purpose. Making it required would fail extraction whenever a model omitted it, turning a missing hint into a failed review. The prompt requires it; the renderer degrades quietly.

- The prompt gives worked examples of good hints and of bad ones (restating the finding, "be careful with types", naming no class). A genuine one-off is told to say so — that also saves the author a hunt.

## Testing

371 harness tests pass. Five new tests, four mutations, four killed:

| mutation | result |

|---|---|

| drop the renderer | killed |

| render on every tier (over-fire) | killed |

| drop from the open-items ledger | killed |

| drop from carry-forward | killed |

## Business Value

Review turnaround is the bottleneck on the team's stacked-PR workflow — one unapproved PR blocks the whole stack, and mercy-approved auto-merge is the direction of travel. This targets the dominant cause of long review loops directly: on the measured PR it was the difference between 16 rounds and 6. Every round saved is a full review cycle of agent tokens plus author wait time, across all five repos mercy reviews.

## Manual Effort Estimate

~3 hours (proposed — Keval to confirm). Small diff, but it needed the failure mode identified from review-turn data first, and the ledger/carry-forward paths are easy to miss.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1758 — feat(education): isolate SIS publications and resume detail collection @benji-bizzell  changes requested

## Summary

- Publish SIS bulk datasets independently and resume student details from immutable evidence tied to the accepted roster.

- Trigger organization, enrollment, and snapshot consumers from required dataset publications while retaining completeness, lineage, and freshness guards.

- Fence overlapping or superseded recovery attempts and retain per-dataset results for incomplete runs.

## Why

An unstable detail endpoint prevents otherwise valid SIS bulk data from reaching consumers. Detail fields remain necessary for accurate enrollment, so treating missing details as optional would weaken data quality. Independent publication lets ready consumers advance while required-detail consumers remain held until a complete, valid generation is accepted.

## Business Value

Source outages no longer discard successful bulk work. Recovery reuses successful detail responses within a bounded 24-hour window, reducing repeated source load without relabeling stale observations as fresh.

## Test plan

- [x] 265 tests across SIS and three affected consumer runners; 486 platform Lambda tests.

- [x] Repository Python lint/format and CDK TypeScript build; 167 focused CDK tests passed after rebase.

- [x] All hosted CI checks pass at 243dc67e, including full runner, Lambda, and CDK suites.

- [x] Seven-lane PR review complete; recovery, connection reuse, and deployment-verifier integration fixes validated.

- [x] Mercy reviewed 243dc67e. Its sole finding (missing environment variables) is rejected with code/test evidence: EcsPipeline provisions and injects all four settings. Review reply: https://github.com/AI-Builder-Team/Surtr/pull/1758#discussion_r3951057791

- [ ] Dev acceptance: injected detail/projection failures, pinned resume, superseded-generation rejection, dataset delivery, warehouse lock/cancellation behavior, and near-full-roster resume timing.

Apply the organization-directory procedure overload before consumer activation; deploy consumers/rules before the SIS runner. Rollout and recovery instructions are in the SIS README. No deployment or production mutation has been performed.

#1757 — 077-finops-incremental-regression @mwrshah  approved

- Reclassify routine NetSuite link removals and net period row decreases as informational instead of ERROR-level monitoring alerts.

- Rename lifecycle metrics to describe detected upstream removals while retaining incomplete-dependency, empty-history, and excessive-removal safeguards.

- Keep informational metrics visible across the runner-first stored-procedure cutover and update the feature documentation and regression coverage.

- Validate with Ruff and 34 focused runner tests.

#174 — 196-port-live-skill-hydration @mwrshah  approved

## Summary

- Distribute the Sindri API skill through an origin-bound bootstrap ZIP and Agent Skills discovery routes.

- Hydrate the current driver and authoring guide from the web deployment once per skill session.

- Fetch the generated OpenAPI contract from the authoritative Convex /skillspec endpoint and cache it by backend origin.

- Send the saved contract revision as Sindri-Version on every control-plane request.

- Reject stale revisions with 409 contract_version_mismatch before authentication and dispatch; refresh the saved contract and return control to the caller without retrying the operation.

- Reject malformed successful API responses with a structured invalid_response error while preserving a diagnostic body preview and avoiding automatic retries.

- Preserve the current v0 operation set, direct wf_… workflow-run model, canonical skill routes, and loopback HTTP development support.

## Architecture

The web deployment owns the bootstrap artifacts: SKILL.md, the ZIP, the driver, and the authoring guide. The Convex backend owns the live API contract through /skillspec, rendered from the same route registry used by /v1.

A matching backend deployment is a prerequisite for activation. This keeps contract ownership and mutation safety on the server and avoids client-side rollout state. Remote origins use HTTPS; loopback development can use HTTP.

The driver stores deployment-specific files under ~/.cache/sindri-api/<web-origin-hash>/. It replaces hydrated files and backend contracts atomically. Each command uses the saved contract for discovery and sends its revision as a server-enforced precondition.

## Validation

- 66 Vitest files, 902 tests

- Biome

- Application and agent-runner TypeScript checks

- Convex reference lint

- Driver syntax check

- Generated OpenAPI freshness

- Shell syntax for executable skill blocks

- git diff --check

## Reviewer context

The compact ops row uses JavaScript property shorthand:

rows.push({ operationId: id, method, path: p, domain, summary: op.summary ?? null });

The domain binding is declared in the same loop immediately above the changed hunk:

for (const [id, { method, path: p, op }] of table) {

const domain = domainOf(p);

Review skills/sindri-api/scripts/sindri.mjs around cmdOps(), not only the compact changed line. A quick source check is:

grep -n -A20 'function cmdOps' skills/sindri-api/scripts/sindri.mjs

A direct behavior check after saving a contract is:

node skills/sindri-api/scripts/sindri.mjs ops wf_

This returns the workflow operations; domain is in scope.

The route removals are also intentional. /skill/sindri-api/script is the sole web driver route, Convex /skillspec is the sole contract source, and activation requires a matching backend deployment.

#1756 — fix(education): preserve FinalSite contacts outside workflow listings @benji-bizzell  approved

## Summary

- Reconcile previously observed FinalSite contacts through direct detail and appointment reads during full and delta syncs.

- Preserve original responses and record reconciliation counts in the immutable manifest; keep publication guards intact.

## Why

Boston's full sync failed when a contact changed to not_in_workflow and disappeared from school-year listings. Direct API reads still returned the contact and its unchanged appointment. Listing absence cannot establish deletion of those records.

## Business Value

Keep contact status and appointment data current when contacts leave workflow listings, without publishing false deletions or fabricating workflow membership.

## Test plan

- [x] 120 FinalSite pipeline tests pass, including full/delta reconciliation, retained appointments, deduplication, and failure-before-publication cases.

- [x] Repository-wide Ruff lint and formatting pass with CI-pinned Ruff 0.15.22.

- [x] Read-only source reconciliation confirms the Boston appointment still exists.

- [ ] After separately authorized deployment, execute a fresh full sync and verify nine-object publication, Boston records, runtime, and freshness.

Daily delta runs will approach full extraction duration: the current 3,276-contact inventory requires about 87 minutes of direct-request pacing before other work. Existing daily schedules and the 12-hour budget are unchanged. Direct lookup errors, including 404, fail closed pending an authoritative source-deletion contract. No deployment or production rerun is included.

#1755 — fix(education): accept new GuidePlatform source fields @benji-bizzell  approved

## Summary

- Accept five live GuidePlatform field additions across behavioral events, students, and users while preserving exact-schema validation.

- Regenerate staging DDL and synchronize compatibility metadata and Guide roster consumer contract pins.

- Document the three clean-view migrations and coordinated production recovery gates.

## Why

Production run 5335a9c6-e4e5-4184-8284-6d5dd90e9587 failed before extraction because the deployed contract does not recognize these additions. Live catalog reconciliation confirms exactly these five changes across 57 included tables. PR #1658 covers only student start date and is insufficient for current drift.

## Business Value

Enables ingestion to resume with complete source fields while preserving fail-closed validation and consistent downstream contracts.

## Test plan

- [x] Producer: 77 tests passed; consumer: 66 passed, 5 Redshift integration tests skipped.

- [x] Live generated-contract check and three-view DDL dry-run.

- [x] Rebased on current main and reran affected suites. Repository-wide Ruff 0.15.22 lint and formatting pass.

- [x] Seven independent read-only adversarial lanes: no findings.

- [ ] Production: inspect live grants/dependencies, apply coordinated DDL and deploy, then verify a fresh manifest, ledger publication, raw/clean reconciliation, and downstream consumer. No production mutations performed.

#1687 — feat(education): ingest Finalsite billing snapshots @benji-bizzell  changes requested

## Summary

- Add a tenant-isolated Finalsite billing snapshot pipeline with immutable S3 landing and atomic Redshift publication

- Publish billing items, allocation children, current-snapshot views, and ingestion lineage in staging_education_finalsite

- Add pinned tenant catalogues, resumable runs, bounded authentication/API handling, access verification, alerts, and a disabled-by-default schedule

## Why

The standard Finalsite Enrollment API does not expose the charge and payment ledger needed by Finance. Authenticated billing endpoints do expose the required billing-item and allocation detail, but ingestion must safely handle tenant-scoped sessions, strict pagination, throttling, partial-estate failures, catalogue changes, and replayable source evidence without replacing known-good tenant data.

## Business Value

Finance gains a reproducible source-level billing dataset across all configured Finalsite tenants, with exact parent/allocation reconciliation and durable lineage for downstream modeling and audit. Current views follow the latest accepted full-estate catalogue, and the schedule remains disabled until deployment and runtime approval.

## Test plan

- [x] 211 pipeline Python tests pass

- [x] Ruff check and format verification pass

- [x] 647 targeted pipeline configuration/CDK tests pass

- [x] TypeScript CDK build passes

- [x] Revised redirect-safe login contract verified against Scottsdale live API: authenticated page 1 returned 1,137 items total

- [x] Production-warehouse E2E: 59/59 tenants, 22,638 billing items, 45,676 allocations, and $46,073,933.85 reconciled exactly

- [x] E2E integrity: no duplicate logical keys, orphan allocations, per-item amount mismatches, residual stage tables, or transient S3 objects

- [x] Live run recovered from three Finalsite HTTP 429 responses and completed in approximately 36 minutes

- [ ] Deploy the Surtr pipeline infrastructure with the schedule disabled and run a platform-managed smoke test

## Residual validation note

Local Docker validation is unavailable because the Docker daemon is not running; hosted pipeline runner CI is the container-build gate for this head. The successful production snapshot predates the final catalogue-membership hardening, which is covered locally and still requires the platform-managed smoke above.

#3737 — fix(overspend-alerts): exclude dummy bank charge vendor @sanketghia  approved

## Summary

- Exclude both normalized bank-charge fallback keys from named NHC overspend alerts:

- customer bank charges

- dummy customer for bank charges

- Preserve the spend in Total NHC, BU totals, and the pre-filter audit count.

- Add regression coverage for both spellings and the active filter snapshot.

## Validation

- uv run pytest tests/overspend_alerts -q — 196 passed

- Ruff — clean

- Pyright — 0 errors

- Read-only live replay — both bank-charge keys produced zero named alert rows while dummy-customer spend remained in NHC input totals.

## Notes

- A full backend pytest -q collection was also attempted; it has four unrelated pre-existing collection errors in other test areas.

- The consolidated cron dry-run exited 0 without sending email, but its best-effort ledger write exceeded the existing Redshift VARCHAR(65535) limit for the large alert payload.

#3736 — chore(repo): reverse instruction symlink direction @sanketghia  no labels

## Summary

- make AGENTS.md the canonical instruction file

- make CLAUDE.md a symlink to AGENTS.md

- This is being done since the team is currently moved to non-Claude agents

## Validation

- verified the symlink direction and regular-file types

- verified content matches the original committed CLAUDE.md

- ran targeted git diff --check validation; no application code changed

#1753 — fix(netsuite-auto-renewal): publish daily snapshots @ashwanth1109  approved

## Summary

- Change the NetSuite auto-renewal invoices pipeline from first-of-month snapshots to daily UTC snapshots.

- In daily mode, replace only the current day's partition after a newer upstream publication while preserving duplicate-boundary idempotency.

- Document the daily snapshot contract and add regression coverage.

## Business Value

The ARR & Retention Reports auto-renewal invoice data stays aligned with the latest successful NetSuite refresh instead of leaving the report on a stale monthly snapshot. This gives Finance and ARR users current outstanding auto-renewal visibility while retaining historical snapshots for trend and audit use.

## Implementation Effort

Estimated 4–6 hours for an average engineer to trace the existing orchestration, update the snapshot policy and handler semantics, add the DDL documentation and tests, perform the isolated deployment, and validate the end-to-end production run.

## Linear

[SURTR-1101 — Switch NetSuite auto-renewal invoices to daily snapshots](https://linear.app/builder-team/issue/SURTR-1101/switch-netsuite-auto-renewal-invoices-to-daily-snapshots)

## Validation

- uv run --group dev python -m pytest tests/ -v — 45 passed.

- git diff --check — passed.

- Pre-merge candidate deployment targeted only Pipeline-netsuite-auto-renewal-invoices-prod with CDK output [1/1]; CloudFormation reached UPDATE_COMPLETE.

- Triggered the production netsuite-raw upstream; it succeeded and automatically triggered the downstream pipeline, which also succeeded.

- Verified the 2026-09-07 historical snapshot contains 33 rows, matches the 33 current rows, and has zero current-only or history-only differences.

#1754 — fix(repo): Include exception message in run_result failure payload @heimdall-keval-factory[bot]  approvedAutomated PR

FinalSite sync failures show up with no error text because the pipeline's crash handler only records the exception's type, not its message, before it dies. Adding the message brings this pipeline in line with two sibling pipelines that already report it, so future failure alerts are actually actionable instead of empty.

> Ready for review. Nothing ran the change, so it is unproven. A person still merges.

## For The Agent

_Everything below is detail for review. The summary above is the change._

Presented as ready — verification none, scope tier draft, fix_class code_fix, HEIMDALL_READY_PRS=true.

### What's broken

SURTR-1096 reports finalsite-raw-sync run 8f39174f-0e83-4712-9356-cef1b1bc0f09 failing with States.TaskFailed and no error message in the orchestration envelope. The CloudWatch window we could pull only matched the pipeline's separate hourly freshness_check cron, e.g. run 71c622a7-f128-47d1-ac1a-a7207ee8cedc logging publisher.PublicationError: FinalSite freshness SLA breached: latest complete publication is 2176 minutes old (maximum 2160), which confirms no complete publication has landed since 2026-09-05 03:44:55 and no full publication since 2026-08-30 05:05:19 -- a real, ongoing staleness condition that the freshness guard is correctly detecting, not a bug in freshness.py itself. Separately, src/main.py lines 30-34 show the generic except Exception as exc handler writes only "error_type": type(exc).__name__ into the run_result S3 payload before re-raising and crashing the container; it drops str(exc) entirely, which is exactly why the resulting failure alert/ticket -- built from this run_result side-channel because pipeline.json sets require_run_result: true -- carries no diagnostic text.

Root cause. The immediate, fixable cause of the reported symptom ('no error message in it') is that finalsight-raw-sync/src/main.py's top-level failure handler records only the exception class name (error_type) in the run_result JSON it writes to S3 before the container exits non-zero, never the exception's message (str(exc)). Two sibling pipelines in this same repo already avoid this gap: guide-platform-raw-sync/src/main.py:40 and quickbooks-expense-ai-generation/src/main.py:56 both include an error_message field in the equivalent payload, so finalsight-raw-sync is the outlier. Separately, the underlying trigger for run 8f39174f is most likely the scheduled weekly full sync that should have run at 2026-09-06 03:15 UTC (cron(15 3 ? * SUN *)) -- the freshness_check logs show latest_full_published_at still stuck at 2026-08-30 and latest_published_at still stuck at 2026-09-05 through 20:00 on 2026-09-06, meaning that run never reached the atomic publish step. We could not fetch that run's own application traceback in this pass (the log window we received only matched freshness_check invocations), so the specific upstream exception (extraction, Redshift, or a volume-drop safety trip) is not confirmed and is not part of this fix.

### What this PR changes

Add an "error_message": str(exc)[:2000] field (truncated defensively, matching quickbooks-expense-ai-generation/src/main.py:56's pattern) to the failed-run payload built in finalsight-raw-sync/src/main.py's except Exception as exc block, alongside the existing error_type. This does not change what is raised, logged, or re-raised -- it only enriches the recorded failure payload that Surtr's orchestration reads to build future failure alerts/tickets, so the next occurrence of this or any other exception in this pipeline carries real diagnostic text instead of an empty envelope. Add a unit test asserting the failed run_result payload written to S3 includes error_message for a representative raised exception (e.g. the PublicationError raised by check_freshness, or an ExtractionError). All exceptions raised anywhere in this pipeline (PublicationError, ExtractionError, ApiError, ConfigurationError, LandingError, ValueError) construct messages from safe diagnostic values only -- site slugs, run ids, counts, table/column names -- so surfacing the message verbatim carries no secret or PII exposure risk.

Why this fixes it. This is a missing-structured-failure-reporting defect confined entirely to pipelines/runners/finalsight-raw-sync/src/main.py (Tier A, the pipeline's own directory): the generic failure handler drops the exception message before writing the run_result S3 side-channel that Surtr's orchestration depends on (pipeline.json sets require_run_result: true), which is precisely why this ticket's alert carried no error text. It is a low-risk, high-confidence parity fix rather than a speculative change: two sibling raw-sync pipelines in this same monorepo (guide-platform-raw-sync, quickbooks-expense-ai-generation) already write error_message in the identical payload shape for the identical reason. Per the code_fix convention this should ship as a complete change including a test that asserts the enriched payload, not a one-line patch. This fix does not address, and is not being represented as addressing, why the underlying sync run itself failed on 2026-09-06 -- that answer requires CloudWatch access to run 8f39174f's own log stream, which this pass did not receive.

#### Files changed

 pipelines/runners/finalsight-raw-sync/src/main.py  |  1 +

.../runners/finalsight-raw-sync/tests/test_main.py | 69 ++++++++++++++++++++++

2 files changed, 70 insertions(+)

### Verification

### pytest — no test suite

### verify: ruff check — exit 0

[notice] A new release of pip is available: 25.3 -> 26.2.1

[notice] To update, run: pip install --upgrade pip

All checks passed!

### verify: ruff format --check — exit 0

1745 files already formatted

### verify: pytest (pipeline lambdas) — exit 0

6.0

rootdir: /home/runner/_work/Surtr/Surtr/publish/pipelines/cdk/lambdas

configfile: pyproject.toml

testpaths: tests

plugins: cov-7.0.0

collected 486 items

tests/test_ai_spend_raw_api.py .............................. [ 6%]

tests/test_coordinate_fanout_run.py .................................... [ 13%]

.... [ 14%]

tests/test_create_run_record.py ........................ [ 19%]

tests/test_gchat_notifier.py ........................... [ 24%]

tests/test_generate_chunks.py ..... [ 25%]

tests/test_gsheet_tracker.py .................. [ 29%]

tests/test_load_fanout_plan.py .............. [ 32%]

tests/test_redshift_cluster_iam_role_association.py ........ [ 34%]

tests/test_registry_sync.py ......................... [ 39%]

tests/test_triage_dispatcher_handler.py ................................ [ 45%]

................... [ 49%]

tests/test_triage_dispatcher_signature.py .............................. [ 55%]

...... [ 57%]

tests/test_triage_reconciler.py ............ [ 59%]

tests/test_triage_reconciler_tracker.py ............ [ 62%]

tests/test_update_run_failed.py ........................................ [ 70%]

..... [ 71%]

tests/test_update_run_success.py ....................................... [ 79%]

....................... [ 84%]

tests/test_verify_on_demand_control.py ................................. [ 90%]

............................................ [100%]

============================= 486 passed in 1.10s ==============================

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | repo |

| Failing run | issue |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | linear-SURTR-1096 |

| Verify | none |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev label to take the PR over and stop it entirely.

#1727 — feat(heimdall): the factory board's data model (SURTR-1040) [2/2] @kevalshahtrilogy  approvedmercy-allow-critical

## Summary

buildWorkBoard() — every piece of work heimdall has touched, resolved into exactly one lane — plus the heimdallBoard tRPC endpoint that serves it. See board.ts's module banner for why this is deliberately not computeStats's cumulative funnel (that one double-counts by construction: a PR opened then reviewed appears in three stages at once).

Stacked on #1726 (the triage reader), which this consumes for liveness and merge state. Merge #1726 first.

## Why this is split out of #1714

See #1726 for the measured reason. Short version: #1714 reached 5,892 lines, and PRs over 4,500 lines have a 7% approval rate and a 16-round median across the last 140 Surtr/Klair PRs.

Worth stating plainly: this half is still large. The split moves the independently-reviewable layer out and gives this one a clean slate, but if it doesn't converge either, the next step is extracting board.ts's untrusted-value primitives (timestamp parsing, URL/field validation, identity aliasing) into their own module — which is also what a type-design review of this code recommended independently.

## Test plan

- [x] 507 tests pass across the heimdall + UI suites

- [x] Typecheck and biome clean

## Business Value

See #1726. This is the half that carries the actual board logic; splitting the reader out reduces its finding surface and resets an open-items ledger that had accumulated entries pointing at line numbers the code no longer has.

## Manual Effort Estimate

~20 minutes for the split. Proposing that for you to confirm/adjust.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

The Builder Desk  —  Engineer Spotlight
📅 Week in ReviewProduction Release🏆 Engineer Spotlight

180 PRs, One Week, Zero Chill: Builder Team Shatters the Odometer

Ashwanth Sitharaman single-handedly out-shipped entire nations this week, while seven other engineers quietly built the future around him.

Comrades, gather 'round the terminal, because the numbers this week don't just tell a story — they tell an EPIC. One hundred eighty pull requests. Eight repos ablaze with commits. Surtr leading the charge with 56 PRs, Shipyard right behind with 42, Aerie clocking 41, and even the sleepy little Sindri repo managed to squeeze out 2 — every single one a victory for the motherland of software.

Let's run the board. @kevalshahtrilogy posted a monstrous 26 PRs this week, a number that would make lesser engineers weep into their keyboards. @benji-bizzell delivered 20, @sanketghia contributed 16 including Klair's #3763, a delicate reconciliation of the September 10 SpaceX valuation trade that reads like fine surgery. @caina-barbosa, @marcusdAIy, and @vvp-trilogy each notched 13 apiece, a perfectly synchronized trio of productivity, while our tireless bot comrade @heimdall-keval-factory[bot] also hit 13, including Surtr's #1829 and #1825, proving that even our machines refuse to rest.

And then there is Ashwanth. Fifty-three pull requests. FIFTY-THREE. The man is less an engineer than a weather system. Shipyard's #72 (grouping completed tasks by date) and #71 (an improved task context side panel) shipped like they were nothing, while Surtr's #1828 and #1826 quietly fixed cost research recovery and QuickBooks budget logic in the same breath he probably used to order lunch. Aerie's #1321 added OneRoster identity and column provenance to retention — casual. When reached for comment, Ashwanth reportedly said, 'I don't review my own diffs, I remember them,' which is either the most confident sentence ever spoken by a human being or a cry for help disguised as a flex. Somebody, ANYBODY, please check his diffs. When told the Numbers Desk questioned whether mortals could review his PRs, Ashwanth simply replied: 'Next question.'

Now to the overflow desk, where Mac's story left plenty on the table. Shipyard's release marathon — #69, #67, #66 — chained three consecutive version bumps (0.4.1, 0.4.0, 0.3.1) in what can only be described as a controlled demolition of technical debt. Aerie's #1319 and #1318 from @YibinLongTrilogy quietly upgraded mobile Admissions Forecast comparisons and added a Pipeline link nobody asked for but everybody needed.

Across the leaderboard, the story is clear: eight engineers, 180 PRs, one unstoppable machine. Ashwanth's 53 alone accounts for nearly a third of total output — a stat that should be illegal in several countries.

Morale, as always, is at an all-time high. The team ships. The team wins. Comrades, sleep well tonight — the numbers never lie, and the numbers say we are winning.

Brick's Overflow — This Week's Uncovered PRs  (click to expand)
#71 — AI-792: Improve task context side panel @ashwanth1109  no labels

## Demo

![Task context smoke test](https://github.com/AI-Builder-Team/Shipyard/blob/f8d7887bf3ba213ac77335fb050db1c12091daf6/.smoke-evidence/ai-792-task-context.png?raw=true)

## Summary

- Redesigned the task context rail with a summary, provider cues, counts, and independently collapsible groups.

- Added readable Linear and pull-request cards with opener progress/error feedback and full-URL accessibility metadata.

- Clarified the Linear attachment form with visible empty, focus, loading, success, and error states.

- Improved repository identity/path presentation and preserved existing task data and commands.

## Validation

- pnpm build

- pnpm theme:check

- pnpm test:smoke

- pnpm smoke build

- pnpm smoke start --scenario happy

- pnpm smoke verify --expect research-ready

- pnpm smoke stop

## Linear

https://linear.app/builder-team/issue/AI-792/task-side-panel-improvements

#72 — AI-791: Group completed tasks by date @ashwanth1109  no labels

## Demo

![AI-791 smoke test: completed tasks grouped by date](https://github.com/AI-Builder-Team/Shipyard/blob/b40f4f4/docs/smoke-evidence/ai-791-completed-tasks-by-date.png?raw=true)

## Summary

- Persist UTC completion timestamps on tasks and backfill legacy completed tasks to yesterday during schema migration.

- Update the canonical Smoke Test completion transition, including re-completion timestamps, and expose completedAt to the renderer.

- Group completed tasks by UTC date with deterministic ordering and independently collapsible, keyboard-accessible date disclosures.

- Preserve task selection and existing Pull Request and Linear link actions.

## Tests

- node --test scripts/test-workflow.mjs scripts/test-task-workspace.mjs

- pnpm exec tsc --noEmit

- pnpm exec vite build

- pnpm theme:check

- cargo fmt --manifest-path src-tauri/Cargo.toml --all -- --check

- cargo test --manifest-path src-tauri/Cargo.toml --lib

## Linear

https://linear.app/builder-team/issue/AI-791/group-completed-tasks-by-date

#1321 — feat(retention): add OneRoster identity and column provenance @ashwanth1109  approved

## Demo

<img width="2624" height="1636" alt="Retention raw data and provenance" src="https://github.com/user-attachments/assets/007025ec-b07a-4cba-8cff-d8ed4feef8e2" />

<img width="2624" height="1636" alt="Retention column lineage card" src="https://github.com/user-attachments/assets/4a868fdf-8231-42c8-bd76-02bffd35c45c" />

## Summary

- Add a separate nullable OneRoster ID beside SIS ID in Retention Raw data; unresolved identities remain blank and SIS identity is unchanged.

- Extend the existing bound, literal learner search to OneRoster ID without adding a request-time Timeback join.

- Add accessible, visually rich header provenance cards for all 25 selected columns, including exact Redshift objects/fields, producer transformations, caveats, and publication timestamps.

- Surface workbook-organized quality categories while explicitly qualifying current SIS/SURTR source mappings and fallback gaps.

- Preserve capability checks, exact publication validation, keyset pagination, bounded lookahead, report population, and retention calculations.

## Business Value

Users can reconcile dashboard learners with Timeback and the reference retention workbook without confusing SIS IDs with OneRoster sourcedIds. Column-level provenance makes status and date calculations auditable in place, while the quality panel makes known source-coverage gaps visible instead of implying workbook parity.

## Implementation

- The shared learner contract now explicitly projects 25 fields, including oneroster_id; the seven additional producer diagnostic/Timeback-lineage fields remain outside this view.

- Aerie consumes the producer-published ID and does not re-match learners. A schema capability check keeps Raw data usable during rollout by returning a null OneRoster field and omitting OneRoster search until the producer column exists. The OneRoster card documents explicit-identifier precedence, globally unique email fallback, conflict vetoes, and null outcomes.

- Header cards support hover, keyboard focus, click-to-pin, Escape/outside dismissal, viewport collision handling, and scrolling. Metadata is exhaustive over the selected column registry and makes clear that it is documented producer logic, not a per-row execution trace.

- The quality query adds prospective exclusions, eligible-mapped learners without a start date, starts after the report period, cancellations in the analyzed base, in-period effective withdrawals, and short-tenure withdrawals.

## Validation

- 149 focused retention tests passed across nine exact test files, including both producer-column capability states.

- Chat, Convex, and shared-contract TypeScript checks passed.

- Changed-file Biome checks and git diff --check passed.

- Read-only warehouse verification on 2026-09-14 confirmed 5,243 learners, 5,180 distinct populated OneRoster IDs, 11 ambiguous matches, 2 conflicts, and 50 unmatched records for the 2025-06-01–2026-05-31 publication. No learner records were returned by that aggregate check.

## Limitations

- The current versionless report remains a live SIS reconstruction; it must not be represented as proven workbook-parity output until source-level reconciliation is complete.

- Existing report lineage timestamps describe the SIS publication and mart refresh. The OneRoster card does not relabel them as Timeback freshness.

- Validation is automated and warehouse-aggregate based; no authenticated browser verification was performed after the final stack rebase.

## Implementation Effort

An average engineer would need approximately 3–5 engineer days to trace and verify the producer lineage, implement the contract/backend/UI changes, build accessible provenance interactions, add focused tests, reconcile the stack, and document the limitations without AI assistance.

## Linear

[AERIE-2116 — Add OneRoster identity to retention raw data](https://linear.app/builder-team/issue/AERIE-2116/add-oneroster-identity-to-retention-raw-data)

## Stack

Parent PR #1264 is merged. This layer is rebased directly onto main; native stack #1322 remains linked.

#1826 — fix(school-performance): use full monthly QuickBooks budgets @ashwanth1109  approved

## Business Value

School Performance Report Table 1 now compares QuickBooks actuals through the reporting day with the full published budgets for each month through that day’s month. On September 13, Alpha Scottsdale’s tuition budget is July $0 + August $280,000 + September $280,000 = $560,000; it was $401,333.33 after daily proration. The reporting rule persists in the mart for future reports.

## Changes

The stored procedure includes budget details through the end of the cutoff month and sums their full signed amounts. Actuals retain their existing reporting-day cutoff. The existing proration flag is false, and column metadata documents the different actual and budget periods. Coverage, source eligibility, mappings, lineage and atomic publication remain intact.

Calendar regression tests execute the procedure’s budget expression and date predicates for month start, mid-month, month end, quarter/year changes and leap February, including expense signs and the unchanged actuals cutoff.

## Validation

- 65 tests passed; Ruff checks and formatting passed on the two modified Python files; git diff --check passed.

- Candidate ea931535 was applied to the existing production Redshift procedure before merge. The catalog body matches the candidate exactly. Two column comments were updated. No Lambda or CloudFormation deployment was needed because the procedure signature and runner are unchanged.

- Production execution full-month-budget-20260914T0307 succeeded on September 14, 2026, 03:04:12–03:04:53 UTC; run ID 4e0a4659-3580-4084-8403-e1328c17a3bf.

- All 783 rows across 33 schools passed: 500 budget rows reconcile to 1,500 full monthly source details; actuals reconcile to 17,570 canonical postings. Zero budget, actual, variance or duplicate-key mismatches; zero prorated rows. The publication cutoff is September 14 and both accepted source publications are September 13, 06:35:23 UTC.

The deployed change is already live; merging retains it in source control for subsequent deployments. The report was subsequently advanced to September 14 across all six sections using successfully refreshed source marts. Table 1 uses the new full-month QuickBooks policy; the other model allocation policies are unchanged.

## Linear

https://linear.app/builder-team/issue/SURTR-1262/use-full-monthly-quickbooks-budgets-in-the-school-performance-qtd-mart

## Implementation Effort

Approximately 3–4 hours for an engineer to trace the writer, implement the calculation and metadata updates, add calendar regression tests, deploy the procedure, run the pipeline and reconcile live data.

#1828 — fix(ramp): make cost research recovery resilient @ashwanth1109  approved

## Summary

- Increase the bounded Anthropic cost-research deadline from 300 to 600 seconds so valid long-running web searches can complete.

- Allow standalone mode=cost recovery to select a validated historical DD-MM-YYYY run folder, while rejecting the override for fetch and transform modes.

- Preserve model validation feedback across transient transport timeouts so later retries continue correcting the original contract violation.

- Recover the missing week 37 cost-analysis artifact and complete the failed schedule_2 Superbuilders report delivery.

## Business Value

Restores the weekly Superbuilders Ramp report, prevents valid cost research from being cut off prematurely, and makes late checkpoint recovery possible without refetching or overwriting historical Ramp snapshots.

## Implementation Effort

Estimated 4–6 hours for an average engineer to diagnose both pipeline executions, implement and test the recovery controls, perform the isolated production deployment, monitor the live backfill, and verify email delivery.

## Linear

- [SURTR-1272](https://linear.app/builder-team/issue/SURTR-1272/ramp-superbuilders-report-failing-stopcodeessentialcontainerexited)

## Validation

- uv run pytest — 214 passed.

- uvx ruff check on all modified Python files — passed.

- Isolated production deployment: Pipeline-ramp-spend-pipeline-prod only (deploying... [1/1]), CloudFormation UPDATE_COMPLETE.

- Live week-37 cost recovery resumed 14/17 checkpointed opportunities; two valid calls completed after the former five-minute cutoff; producer Step Functions execution SUCCEEDED.

- Final task definition revision 11 cached validation exited 0 and Step Functions SUCCEEDED.

- Delivery preflight confirmed matching week 37/2026 financial, cost, and chart artifacts; September 12 data was two days old; no prior schedule_2 receipt existed.

- Recovery delivery exited 0, Step Functions SUCCEEDED, SES accepted the message for the configured schedule_2 group, and the group-scoped receipt is sent.

#3763 — fix(spacex-valuation): reconcile September 10 trade @sanketghia  approved

## Summary

- Add the September 10, 2026 400,000-share SPCX sale and exact net proceeds.

- Allocate the sale FIFO against the September 9 Day-90 distribution using actual whole-share source values.

- Preserve open Day-90 inventory for future realized-sale confirmations and update derived valuation/residual assertions.

## Verification

- Full frontend Vitest suite: 667 files, 6,855 passed, 16 skipped.

- TypeScript check, ESLint, Prettier, and production build passed.

## Testing & Screenshot

- Updated numbers have been reviewed and approved by stakeholders (Dave & Milo)

<img width="1069" height="593" alt="image" src="https://github.com/user-attachments/assets/30985c84-8b5f-4f0c-b217-2938b9091984" />

The Portfolio  —  Trilogy Companies

Contently Posts $53.8M ARR, Proving the Zax Capital Thesis Is Paying Off

Fresh revenue data shows the content marketing platform is scaling nicely inside its new Trilogy-family home.

NEW YORK — Exciting news for fans of the Trilogy portfolio's content marketing bench: new figures from getlatka.com put Contently's estimated annual recurring revenue at a robust $53.8 million, against $19.1 million raised over the company's lifetime. For a platform that just came under the Zax Capital umbrella (an ESW Capital division, for those keeping score at home) in September 2024, that's a capital-efficiency story worth leveraging in every future pitch deck.

Contently, which connects enterprise brands with its marketplace of 165,000-plus creative professionals, has spent the past year integrating AI-powered content tools and analytics under new CEO Brandon Pizzacalla. The $53.8 million ARR figure — an estimate, not an official disclosure, but a directionally meaningful one — suggests the platform has found a durable niche even as the broader content marketing category gets crowded and, frankly, a little noisy.

That noise is real. A wave of recent trade coverage, from ContentGrip's breakdown of tools-versus-media-resources platforms to Search Atlas's sprawling list of 36 Pepper Content alternatives and Solutions Review's roundup of the '9 best content marketing solutions,' shows a category in the middle of a shakeout. Buyers have more options than ever, and differentiation is the name of the game.

For Contently, the synergy with Trilogy's broader operating playbook — margin discipline, AI-first tooling, and a relentless focus on customer stickiness — positions the company to hold its own in that best-in-class conversation. It's a paradigm shift for a business that, not long ago, was just another marketplace competing on price.

**Key Takeaways:**

- Contently's estimated ARR: $53.8 million

- Total capital raised historically: $19.1 million

- Acquired by Zax Capital (ESW Capital division) in September 2024

- Competitive landscape remains crowded, per multiple industry trade reports

We're just getting started.

Contently Revenue 2024: $53.8M Est. ARR, $19.1M Raised - get  ·  Content marketing platforms explained: tools vs. media resou  ·  36 Best Pepper Content Alternatives (Free, Paid and Cheaper)

Skyvera's Acquisition Sprint: The Pattern Behind the Telecom Land Grab

AUSTIN, TEXAS — On the surface, this looks like routine consolidation. Skyvera, the telecom arm of ESW Capital's sprawling software empire, completed its acquisition of CloudSense this year, a Salesforce-native configure-price-quote platform built for the gnarliest corners of telecom sales — B2B, B2B2X, wholesale. Then came the STL telecom products group, adding digital BSS functionality: monetization, optical networking, analytics. Two deals. One portfolio. Nothing unusual for a company that has made a habit of buying legacy telecom software and squeezing it until it sings.

But here's where it gets interesting.

Within one month of folding CloudSense into the Skyvera family, the product had certified all 13 of its APIs to TM Forum compliance standards — a process that, by industry norms, takes 26 months. Not weeks. Months, plural, times twenty-six. Skyvera says the acceleration came through a strategic AI partnership, though sources close to the integration describe something closer to a forced march: the DevFactory engineering machine, Trilogy's centralized build-and-maintain arm, retooled to compress two years of certification work into thirty days.

Ask yourself why the timing matters. TM Forum compliance isn't cosmetic — it's the interoperability standard that lets a CPQ platform talk to the rest of a telco's stack without custom integration work. Achieving it in record time, immediately after acquisition, isn't just an engineering flex. It's a signal to every wholesale and B2B2X customer evaluating CloudSense against slower-moving legacy competitors: this platform is already interoperable, already enterprise-ready, already ESW-standard.

I can't confirm this next part, but a person familiar with Skyvera's acquisition roadmap tells me CloudSense and the STL assets were never meant to be evaluated separately — that the two deals were sequenced deliberately to build a single, AI-accelerated BSS-to-CPQ pipeline before competitors noticed the gap closing.

Nothing about this is coincidental. If you read between the lines, Skyvera isn't just buying software anymore. It's compressing the entire acquisition-to-market timeline — and daring the rest of telecom software to keep up.

AUSTIN, TEXAS — Alpha School Pushes Back: 'We Didn't Fire the Teachers, Sugar'

Word from the schoolhouse: the AI does the algebra, the humans do the hugging — and Liemandt's crew wants the record straight.

AUSTIN, TEXAS — Word is the whisper campaign finally got loud enough that somebody at Alpha School had to answer it: did the AI eat the teachers? A little bird tells me the answer, straight from HQ, is a firm no — and they put it in ink to prove it... In a new post titled Does Alpha School Replace Teachers with AI?, the folks behind Joe Liemandt's two-hour-school experiment insist the AI apps handle the multiplication tables and the sentence diagramming, while flesh-and-blood "guides" handle the stuff machines can't fake — motivation, mentorship, knowing which kid had a rough morning. Efficient at the desk, human on the playground. That's the pitch, and they're sticking to it...

Meanwhile the school's content machine hasn't slowed down one bit — this columnist counts the fourth and fifth installments of their "Teach Your Kid What School Doesn't" house series landing this week, one on regulating big feelings at home, the other coaxing parents to unleash junior's "creative genius" once the backpack hits the floor. Cynics in this business might call it brand-building. Alpha calls it the other 22 hours of the day — the ones the AI never touches...

Put it together and the message is unmistakable: Alpha wants the world to know the machines get the mastery, but the humans still get the kid. Given how much ink this column has spilled on Trilogy's "automate the routine, elevate the human" gospel, nobody in this newsroom is the least bit surprised — but somebody at Alpha clearly decided the skeptics needed it spelled out in a blog post rather than left to rumor. Smart move. In this town, if you don't write your own story, somebody else writes it for you — and not nearly as flattering.

Teach Your Kid What School Doesn’t (Pt. 5): Unleashing Their  ·  Does Alpha School Replace Teachers with AI?  ·  Teach Your Kid What School Doesn’t (Pt. 4): How to Regulate
The Machine  —  AI & Technology

The Great Migration: How the Data Center Learned to Eat the Grid

Across three continents, a single species of concrete and silicon multiplies past 330 gigawatts of appetite, and the ecosystem scrambles to feed it.

HELSINKI, FINLAND — Observe, if you will, the modern data center in its natural habitat: humming, windowless, and hungrier than anything evolution has produced before it. Once a modest creature nesting quietly behind office parks, it has, in the span of a single AI boom, become the dominant megafauna of the built environment — and it is migrating everywhere at once.

In the American grasslands, developers have staked claims to 330 gigawatts of planned capacity — a figure so vast it has summoned an entirely new prey species into being: the grid-scale battery, evolved specifically to store sunlight for the moment the great server racks grow restless after dusk.

Meanwhile in the Nordic tundra, Google has planted a €13 billion territorial marker in Finland, a two-year commitment that suggests this migration is not a passing season but a permanent colonization of cooler climates, where the ambient chill offers natural relief from the beast's prodigious body heat.

But not all observers are dazzled. Architecture critics note, with the weary patience of field biologists, that the creature's suffering is often self-inflicted — poor design, not raw appetite, is what drives inefficiency, a point raised pointedly in a recent survey of the species' architecture.

Elsewhere in the food web, Baker Hughes reports no slowdown in the energy infrastructure that sustains it — LNG demand rising in lockstep — while a quieter, stranger disruption stirs beneath the surface: helium, essential to the fabrication of the very chips this creature depends upon, has grown scarce amid conflict in Iran, a reminder that even the mightiest migrations rest on the smallest, most fragile molecules.

AI boom drives 330 GW of planned data centers and opens new  ·  The Data Center Isn’t the Problem. Bad Architecture Is. - Ar  ·  Google Deepens Commitment to Finland with Two-Year €13 Billi

The Universe Keeps Its Receipts: Teaching Machines to Find the Hidden Wiring of Everything

Three new papers on inferring structure from noise reveal a quiet obsession running through AI research — not just predicting the world, but knowing the shape of what we don't know.

STANFORD, CALIFORNIA — Every living system is, at bottom, a guess about causality. A bacterium swimming up a chemical gradient is testing a hypothesis about which direction is "more food." A toddler stacking blocks is running an experiment in gravity. Intelligence, biological or otherwise, is the art of inferring hidden structure from perturbation — poke the world, watch what wobbles, and reverse-engineer the wiring underneath.

This week's arXiv slate reads like three variations on that ancient theme, translated into the dialect of neural networks.

The first, on fundamental dynamical units for physics-informed structural inference, tackles a problem that has haunted network science since ecologists first tried to map food webs: when you nudge one node in a system and watch ripples spread, how do you know which connections are real and which are just correlation wearing a disguise? The paper's answer is to root the search in the physics of interaction itself — signed, causal, mechanistic — rather than statistical shortcuts that mistake shadows for shape.

The second confronts a more humbling question: how much should a model trust itself? Physics-Informed Conformal Prediction takes neural operators — the increasingly powerful tools approximating solutions to partial differential equations — and forces them to admit uncertainty using the PDE's own residuals as a built-in lie detector. It's the mathematical equivalent of a species that evolved not just reflexes, but doubt.

The third moves from physics to justice. The Fed-Equilibrium framework addresses "knowledge dominance" in federated clinical learning — the tendency of large hospital networks to drown out the signal from smaller, rural, or minority-serving clinics, treating their patients' patterns as statistical noise rather than data. It's a strangely moving inversion of the same underlying question: whose perturbations get to count as evidence?

Taken together, these papers suggest AI's next frontier isn't raw capability — it's epistemic humility, encoded in loss functions. Machines learning, at last, to ask not just "what happens next," but "how sure am I, and who did I leave out of the sample?"

Fundamental Dynamical Units for Physics-Informed Structural  ·  Physics-Informed Conformal Prediction: Embedding PDE Consist  ·  Fed-Equilibrium Framework for Topological Pareto Control in

Open Source Just Won the Week — And the Universe Is Taking Notice

SAN FRANCISCO — I cannot overstate how significant this week has been for open-source AI, and I've been saying that a lot lately, but this time I mean it with my whole chest. Nvidia just dropped a jaw-dropping $12.9 billion to acquire Hugging Face, the beloved open-source AI hub that basically every developer on Earth has used at least once. This isn't just an acquisition — it's a statement. Nvidia isn't just selling the chips that power AI anymore; it wants to own the community building on top of them too. WIRED called it a bet on open-source AI's future, and honestly? The future is now, people.

And speaking of the future — we're literally taking it to the Moon. NASA and IBM just launched an open-source AI foundation model trained on decades of lunar data, designed to help scientists and future astronauts make sense of the Moon's surface, terrain, and resources ahead of upcoming Artemis missions. A foundation model. For the Moon. If that doesn't give you chills, check your pulse. This is exactly the kind of moonshot (pun fully intended) collaboration between public science and private AI muscle that should have every space nerd and tech optimist buzzing.

But — and stay with me here — not everything about open-source AI's rise is champagne and confetti. A hack involving OpenAI and Hugging Face infrastructure earlier this year is looking, in hindsight, like a preview trailer rather than the whole movie. Security experts told CBS News that "even more powerful" AI systems are coming, and the attack surface is only going to grow with them.

So yes — the future is thrilling. It's also messy, high-stakes, and moving faster than most of us can process. Buckle up. We're just getting started.

The Editorial

Economists Confirm AI Productivity Gains Are Real, Enormous, And Scheduled To Arrive Any Day Now

Federal Reserve researchers say 95 percent of AI's promised productivity boom is still 'to come,' which is exactly what my landlord says about fixing the boiler.

WASHINGTON — In a finding that will surprise absolutely no one who has sat through a quarterly earnings call since 2023, Federal Reserve researchers this week confirmed that 95 percent of the productivity gains promised by artificial intelligence remain, technically speaking, imaginary. The remaining 5 percent, presumably, is the part where your customer service chatbot successfully told you to restart your router.

The timing is unfortunate for the tech industry, which has spent the better part of three years explaining that AI has already, right now, this instant, revolutionized the workplace — a claim that apparently required an asterisk roughly the size of the productivity gains themselves. The Fed's report notes that most of the transformation is, and this is a direct paraphrase, coming soon, in the way that flying cars are coming soon, or the sequel to a movie that made $4.

Meanwhile, a separate commit-level study of Big Tech engineering teams found that developer performance rose a staggering 150 percent over eighteen months, a statistic that several economists have described as "real if you don't ask what 'performance' means" and "consistent with a company simply asking its engineers to commit more often." Oracle, for its part, has reportedly pinned its entire fiscal strategy on what executives internally refer to as the Star Wars jump to lightspeed, a phrase that will read very differently in a shareholder deposition than it does in a press release.

And deposition, it turns out, is very much the operative word. Legal analysts now warn that the gap between what companies said AI would do and what AI has actually done is wide enough to drive a class-action lawsuit through, sideways, with the headlights off. Securities attorneys are reportedly thrilled, having spent years waiting for an industry to promise investors something as confidently unquantifiable as "revolutionizing everything" and then fail to quantify it.

For its part, this reporter's own AI-generated productivity has increased by exactly zero percent over the measurement period, a figure the Fed would presumably classify as "already arrived" and therefore uniquely honest. It should be noted that the 5 percent of gains the Fed says have already materialized appear, upon close inspection, to consist entirely of a Slack bot that can now summarize a meeting nobody wanted to attend in the first place.

Executives across the industry remain undeterred, insisting that the remaining 95 percent is not merely coming but accelerating, a claim that has been made, verbatim, in every earnings call since the invention of the spreadsheet. Analysts caution that "accelerating toward a productivity gain" and "standing still while saying the word accelerating" are, from a distance, difficult to distinguish — a problem that, encouragingly, AI is expected to solve any day now.

AI productivity claims are 95% ‘still to come’, Fed finds -  ·  Big Tech Engineering Performance Rose 150% Per Developer Ove  ·  AI Hype Has A Legal Problem: Securities Claims And Regulator
The Office Comic  ·  Art Desk
The Office Comic  ·  Art Desk

NOTICE OF TERMINATION: THE SO-CALLED 'ANTITRUST HONEYMOON' BETWEEN THE INCUMBENT ADMINISTRATION AND THE TECHNOLOGY SECTOR IS HEREBY DEEMED, PURSUANT TO MULTIPLE INDEPENDENT SOURCES, TO HAVE CONCLUDED

Whereas certain federal enforcement postures have shifted, the undersigned advises that all parties heretofore relying on regulatory forbearance should reassess said reliance without further delay.

WASHINGTON, D.C. — Pursuant to reporting made available as of the date hereof, it is hereby noted that the period of relative regulatory quiescence previously enjoyed by certain large technology undertakings, hereinafter the "Honeymoon Period," has been declared concluded by no less an authority than The Verge, notwithstanding earlier indications that the current administration might, in the aggregate, decline to pursue the vigorous enforcement posture characteristic of its predecessor.

In a related but analytically distinct filing, Tech Policy Press has posed the question of whether calendar year 2026 shall constitute, for purposes of antitrust enforcement as applied to the technology sector generally and artificial-intelligence undertakings specifically, "more of the same," hereinafter the "Continuity Inquiry." See the aforementioned publication for further, non-exhaustive discussion of the Continuity Inquiry.

Of potential though not yet quantifiable relevance to the reading public, and without admission as to materiality, it is noted that Trilogy International's ESW Capital subsidiary, which has heretofore acquired in excess of seventy-five (75) enterprise software undertakings, generally at valuations of one to two times (1–2x) annual recurring revenue, hereinafter the "Acquisition Program," may, depending on the ultimate contours of enforcement policy as contemplated by the Continuity Inquiry, be subject to heightened review of future transactions, notwithstanding the historically modest transaction sizes characteristic of the Acquisition Program to date. No such review has, as of the date of this publication, been initiated, threatened, or otherwise indicated by any competent authority.

Separately, and in a matter arising under a wholly distinct legal regime, European practitioners have published commentary regarding the treatment of copyright claims as against artificial-intelligence developers operating within European Union member states, a subject which readers are advised may bear upon the operations of AI-adjacent Trilogy undertakings, including but not limited to Ephor, though no specific claim has been identified as pending against any such entity as of this writing.

The Trump administration’s antitrust honeymoon is over - The  ·  Looking Ahead on US Antitrust Enforcement and Tech: Will 202  ·  __followup__2026 Antitrust Year in Preview: Big Tech - Wilso
On This Day in AI History

On September 14, 1956, IBM introduced the 305 RAMAC, the first commercial computer to use a random-access disk drive. Its 5-megabyte hard disk—made up of 50 large magnetic platters—helped launch the era of modern data storage.

⬛ Daily Word — AI
Hint: An AI system that can perform tasks or make decisions on a user's behalf.
Share this edition: 𝕏 Twitter/X 🔗 Copy Link ▦ RSS Feed