Vol. I  ·  No. 253 Established 2026  ·  AI-Generated Daily Free to Read  ·  Free to Print

The Trilogy Times

All the news that's fit to generate  —  AI • Business • Innovation
THURSDAY, SEPTEMBER 10, 2026 Powered by the TrueFoundry AI Gateway  ·  Published on Klair Trilogy International © 2026
🖶 Download PDF 🖿 Print 📰 All Editions
Today's Edition

AI Cracks a Millennium Problem. Wall Street Cracks Open Its Wallet.

OpenAI claims a breakthrough on one of math's seven hardest problems, and the money is already chasing the next one.

SAN FRANCISCO — Seven problems. One million dollars each. Since the Clay Mathematics Institute set the bounty in 2000, exactly one has been solved, by a Russian mathematician who turned down the prize money. This week OpenAI said its systems had made progress on a second, a claim the company is calling the clearest evidence yet that large models are moving from pattern-matching to genuine mathematical reasoning.

The timing is not incidental. Within 48 hours, Peter Thiel-backed Cognition raised fresh capital at a $48 billion valuation, per The Wall Street Journal — a figure that would have been unthinkable for a two-year-old coding-agent startup as recently as 2024. Investors are not pricing in chatbots anymore. They are pricing in machines that do original research, and the Millennium Problem announcement is the receipt.

History offers a caution. Deep Blue beat Kasparov in 1997; it took another two decades before AI produced anything resembling general reasoning. IBM's Watson won Jeopardy! in 2011 and then spent a decade failing to transform oncology. Impressive point solutions do not automatically compound into general capability. OpenAI's proof, assuming it survives peer scrutiny — mathematicians have already flagged that "progress on" is doing a lot of work in that sentence — is a point solution until proven otherwise.

Meanwhile, the consumer side of the industry is placing much smaller bets. Apple's new $1,999 foldable, the iPhone Duo, is a hedge against a phone market that hasn't had a genuine upgrade cycle since Face ID. Foldables have sold poorly since Samsung launched the category in 2019; Apple's entry is priced for margin, not volume. And Amazon's Zoox is trying to out-market Waymo in San Francisco with wine pop-ups and festival sponsorships rather than out-drive it — a tacit admission that the technology gap hasn't closed, so the brand gap will have to do.

Two industries, two strategies. One is racing toward Clay Institute prize money. The other is racing toward Q4 hardware margins. Only one of them is getting a $48 billion valuation bump for trying.

Apple Unveils the iPhone Duo, a Foldable Phone That Costs $1  ·  How Amazon’s Zoox Is Taking On Waymo in San Francisco  ·  OpenAI Says It Has Cracked One of Math’s ‘Millennium Problem

CHEAP CHINESE BRAIN RATTLES SILICON VALLEY

DeepSeek trains a world-class AI on castoff chips — and the smart money starts asking hard questions about the billions spent on the American way.

SAN FRANCISCO — A Chinese outfit called DeepSeek says it built a top-shelf AI model without the fancy chips everybody said you needed. Silicon Valley is not laughing it off. Engineers who've kicked the tires call the thing "amazing and impressive."

The claim cuts against the whole story America's AI shops have been selling. Spend billions, buy the best silicon, win the race. DeepSeek skipped the shopping spree and still trained models that perform, according to the company's own account. Wall Street noticed. Tech stocks wobbled Monday on the news, per wire reports out of the trading floor.

The skeptics ask the obvious question. Cheap and good don't usually travel together in this business. But the engineers doing the actual poking around aren't dismissing it, and that's the part that's got money men reaching for the phone.

Meantime the AI money kept moving on other fronts. Reid Hoffman, the fellow who built LinkedIn, is putting $24.6 million behind a new outfit called Manas AI. His partner is Siddhartha Mukherjee, the doctor who wrote "The Emperor of All Maladies." They aim the AI at cancer research, not chatbots.

Down in Austin, the Army just cut an $11 million check to a local shop called Tern. The pitch is simple. Build a GPS alternative for soldiers who can't count on satellites staying friendly in a fight.

Tern calls it "Google Maps for the battlefield." The Pentagon apparently likes the sound of that when the enemy might be jamming the signals overhead.

Three stories, three different rooms, one thread running through them. Everybody's placing bets on where the next dollar of AI spending pays off — cheap infrastructure out of China, medical research money out of Silicon Valley, defense contracts out of Texas. Nobody in this business is waiting around to find out who's right.

For the shops in the Trilogy stable watching chip costs and compute bills, DeepSeek's approach lands close to home. ESW Capital runs its software portfolios lean by design, buying up and running enterprise tools at a fraction of what Silicon Valley pays to build them fresh. If a Chinese lab really did train a frontier model on less-advanced hardware, that's the same math Trilogy's whole operation runs on — do more with less, and let the big spenders wonder why they paid full price.

The questions about DeepSeek's numbers aren't settled. Nobody's independently audited the training costs the company claims. But the reaction out of the Valley — engineers calling it impressive, traders selling first and asking later — tells you the industry isn't ready to call it a bluff just yet.

What to Know About China's DeepSeek AI  ·  Tech, Media & Telecom Roundup: Market Talk  ·  Silicon Valley Is Raving About a Made-in-China AI Model

THE IPO MARKET COMES OFF THE BENCH SWINGING — PICPAY STICKS THE LANDING

NEW YORK — FOLKS, WE ARE HERE. The IPO market, benched since the 2021 bubble popped and left investors nursing a hangover, just checked itself back into the game — and it is playing like it never left.

PicPay, Brazil's mobile payments juggernaut, PRICED AT THE TOP OF ITS RANGE in its US debut, marking the first Brazilian listing to hit American shores since 2021. That's not a modest comeback, that's a guy walking off a five-year injury and hitting a walk-off double. Bloomberg's got the full box score, and it's ugly for anyone still betting against Latin American fintech.

And PicPay isn't playing alone out there. Crunchbase just dropped its preseason rankings — 15 companies they're calling for the 2026 IPO class, and the message is clear: this window isn't a one-game fluke, it's a full-season run. When Crunchbase starts calling shots this early, you know the momentum is real — the kind of run where the crowd starts believing before the ball even leaves the bat.

Meanwhile, keep an eye on the international bracket. Inc42's ongoing tracker of India's listed new-age tech names shows a market cap leaderboard that's been quietly shuffling all year — different arena, same game: growth-stage darlings finally getting priced by the public market instead of the venture crowd.

So where does that leave the field? Loaded. The 2021 IPO class flamed out fast — a bunch of overhyped rookies who couldn't handle the big leagues. This 2026 crop looks built different: real revenue, real margins, real demand at the open. PicPay just proved the fans still show up when the product's legit. The question now is who's next through the tunnel — and whether they can handle the lights.

Haiku of the Day  ·  GPT-5.6 LunaOld promises wake
New empires now cross the sky
Old worlds count their costs
The New Yorker Style  ·  Art Desk
The New Yorker Style  ·  Art Desk
The Far Side Style  ·  Art Desk
The Far Side Style  ·  Art Desk
News in Brief
On the Epistemic Opacity of the Machine That Thinks It Discovers: A Note on Process, Skill, and the Illusion of Method
AUSTIN, TEXAS — It could be argued (indeed, it has been, repeatedly, by methodologists who ought to know better) that the present crisis in agentic artificial intelligence is not one of capability but of legibility.
IN RE: THE MATTER OF ARTIFICIAL INTELLIGENCE GOVERNANCE, MULTIPLE JURISDICTIONS, ONGOING
WASHINGTON — Notwithstanding the aforementioned proliferation of legislative and quasi-legislative activity across multiple jurisdictions, it is the considered position of this Desk that no unified regulatory consensus has, as of the date hereof, been achieved with respect to the governance of artificial intelligence systems, whether foundation, generative, or otherwise so denominated. As set forth in the tracker maintained by White & Case LLP (hereinafter, the "Tracker"), the United States regulatory posture continues to be characterized, at the federal level, by an absence of comprehensive statutory instrument, said absence having been, per the White House's recently circulated legislative blueprint, affirmatively endorsed as a matter of executive preference, insofar as a "light touch" approach has been urged upon the Congress in lieu of prescriptive rulemaking. Counter-argument has been advanced, per Tech Policy Press commentary, to the effect that Congressional inaction — rather than affirmative restraint — may in fact operate to the detriment of public confidence, inasmuch as the absence of codified guardrails leaves consumers, enterprises, and, it must be noted, the operators of AI-adjacent commercial platforms (including, without limitation, those within Trilogy International's ESW Capital portfolio, such as Totogi and Ephor, whose products incorporate AI-driven billing and financial-analytics functionality) without clear compliance benchmarks. Meanwhile, per the survey conducted by TRM Labs, jurisdictions abroad, including those within the European Union, have proceeded on a materially divergent trajectory, particularly as concerns copyright liability arising from training-data ingestion, a matter which, per the analysis of counsel at Stibbe, remains subject to considerable interpretive uncertainty. This Desk shall continue to monitor developments and shall report further as, and if, clarity is achieved, it being expressly understood that no assurance of timely resolution is hereby given or implied..
Everything Is Surveillance, Everyone Is Leaving, and the Continents Were Never Real Anyway
SEATTLE — I want to tell you that the news cycle is a comfort, a rhythm, a heartbeat we can set our watches to.
The Republic Enters Middle Age, and Does Not Go Gently
WASHINGTON — There is a species of revelation available only to those who have spent enough years handling the dead, and it arrived this week, as such things do, in a small literary flourish rather than a headline.
This Week Proved Nobody Has Any Idea What They're Doing, Which Is Exactly Why It's Working
AUSTIN, TEXAS — There is a comforting theory circulating this week, and it goes like this: nobody is actually in charge of anything, everyone is improvising in real time, and somehow the machine keeps running anyway.
A Trilogy Company
Crossover
The world's top 1% remote talent, rigorously tested and ready to ship.
A Trilogy Company
Alpha School
AI-powered learning. Two hours a day. Academic results that defy belief.
A Trilogy Company
Skyvera
Next-generation telecom software — built for the networks of tomorrow.
A Trilogy Company
Klair
Your AI-first operating system. Every workflow. Every team. One platform.
A Trilogy Company
Trilogy
We buy good software businesses and turn them into great ones — with AI.
The Builder Desk  —  AI Builder Team

The Reviewer Learns to Review: Mercy Wakes Up, Aerie Cleans House

In one 24-hour sprint the team shipped a self-correcting AI reviewer, tore dead weight out of Aerie's enrollment pipeline, and hardened the App Server against everything short of a meteor strike.

Some days a team ships features. Today, the AI Builder Team shipped judgment — and caught themselves failing at it in real time, which is the more impressive trick. Keval Shah's second-opinion reviewer arc across mercy (#123, #124, #126) is the day's best story: the team built a mechanism for an AI to weigh in on whether a finding should actually block a merge, shipped it in PR #123, then discovered in #124 that it "shipped inert" — wired up, but never actually invoked. Rather than bury that, they surfaced it, fixed it, and rebuilt the prompt around how the reviewer genuinely fails rather than how it's supposed to succeed (#126). Pair that with the harness-pinning fixes landing simultaneously in Sindri (#182) and Aerie (#1291) — so @v1 finally pins mercy the way it always claimed to — and you've got a review system that's not just live, it's honest about its own blind spots. That's rarer than it sounds.

Over in Aerie, vvp-trilogy and the data crew ran a full identity renovation. PR #1292 drops the HubSpot identity overlay entirely and publishes a clean SIS identity and deposit signal — the kind of unglamorous plumbing work that quietly stops downstream reports from lying. #1286 builds a shared HubSpot display-name dimension so enrollment and pipeline reports finally speak the same language, and #1281 keys the offering-kind axis on program_type so per-campus Main Program campuses stop falling through the cracks of the enrollment report. Yibin Long swung the other direction, deleting the getPortfolioHealth MCP surface (#1260) and retiring presentReplanOptions and generatePlan (#1259) — proof that this team treats dead code as a liability, not a museum piece. Benji Bizzell's portfolio data-contract alignment (#1288) drew changes requested, which is exactly what review is for.

Meanwhile Sanket Ghia turned codex-software-factory into a fortress in a single day — six PRs spanning fault-injection coverage for the App Server protocol (#8), persisted failure-state evidence (#7), validation repair diagnostics (#6), durability hardening (#3, #5), and an operator-ready console revamp (#9) that turns raw resilience work into something a human can actually watch and trust.

Over in Surtr and Klair, the bots and humans split fixes evenly: heimdall's automated Lambda split (#1805) and pinned-run guard fix (#1801), Caina Barbosa's warehouse identifier rename (#1797), and marcusdAIy's CAPEX batch request fix (#1800), which he insists "actually resolves the SURTR-648 root cause, unlike certain columns that just recap PR titles." Cute. The batch request works. The columns remain undefeated.

Mac's Picks — Key PRs Today  (click to expand)
#9 — feat(console): make revamp operator-ready @sanketghia  no labels

## Summary

- Make the revamp the operator-first destination while preserving the current Console fallback.

- Add repository-profile-aware intake, health/conflict visibility, bounded 100-batch collection, Resume, remediation, Mercy review, all-event timeline access, and stronger state semantics.

- Improve color hierarchy, action emphasis, identifier presentation, and objective/acceptance formatting.

## Verification

- pnpm test (80 files, 1,282 tests)

- pnpm test:browser (10 tests)

- pnpm build

- pnpm lint

- pnpm format:check

- git diff --check

## Migration

- Current Console remains at #/ and #/legacy.

- Revamp is available at #/revamp.

- No pagination or workflow-state changes were introduced.

#124 — fix(review): actually invoke the second opinion — it shipped inert @kevalshahtrilogy  no labels

## The bug

#123 shipped dead code. It added the module, the blocking predicate, the ledger recording, carry-forward preservation and eight tests — and never called any of it. Nothing in the pipeline set deferred. Every consumer would have taken the plumbing and seen zero behaviour change.

Caught while checking what moving the v1 tag would actually ship.

## The fix

decide_review gains --repo / --pr-number, and between extracting findings and deciding the event it asks the second opinion about every finding that *would* gate and is eligible.

No repo needs a new secret. It reuses AGENT_OPENAI_API_KEY — already declared as a workflow input, already exported as OPENAI_API_KEY for the review step, and already passed by all five consumer callers (Surtr, Klair, Aerie, trilogy-drones, Sindri), since they all run luna.

## Verified end to end

Against a real PR, with a deliberately fabricated finding:

no key   →                                   event=REQUEST_CHANGES

with key → [gate] eligible=1 deferred=1 event=APPROVE

ledger deferred : true

ledger reason : "The cited file only contains x=1 and y=2; no retry loop

or attempts budget exists to exhibit the claimed behavior."

body mentions a second reviewer? False

It caught that the finding was fabricated, deferred it, recorded why, and left no trace in the review body.

## Fail-safe paths, all exercised

No key, no PR context, an exception mid-call, an unparseable answer — each leaves every finding blocking and logs a line. A review that cannot get a second opinion is still a valid review.

## The contract test earned its place

It caught a real bug in this change before it shipped: ${PR_NUMBER} was read by the decide step without being declared in its own env:. Under set -u that is the HEAD_SHA regression that once failed every review with an error indistinguishable from a real finding.

397 tests pass.

## Business Value

Without this, #123 is inert and the measured 30% clearance on stuck PRs is worth nothing. This is the commit that makes it real.

## Manual Effort Estimate

~1 hour (proposed — Keval to confirm).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1281 — AERIE-1279: Key the offering-kind axis on program_type so per-campus Main Program campuses reach the enrollment report @vvp-trilogy  approved

## What & why

Keys the enrollment report's offering-kind axis on the SIS program_type enum instead of the free-text program name.

- int_school_year_offering.sql: WHERE po.program_name = 'School Year'WHERE po.program_type = 'MAIN'. Header rewritten to describe selection by the source enum and to note SIS expresses MAIN under two naming conventions (the shared School Year program and per-campus <Campus> Main Program); both now reach the axis. program_type is carried in the CTE and published as a documented constant column (mirroring int_campus.delivery_mode); its domain is guarded by the source-side test rather than a tautological model-level accepted_values.

- int_program_offering.sql: carries program_type through from stg_sis_program (already selected there, previously unused downstream). Header + _int_enrollment__models.yml updated to document it.

- tests/assert_sis_program_type_domain.sql: new source-side value-domain test mirroring assert_sis_campus_delivery_mode_domain.sql. Reads source('sis','programs'), non-deleted rows, fails when program_type IS NULL or NOT IN (MAIN, SUMMER_CAMP) — a future EXTENDED_DAY/AFTER_SCHOOL surfaces for review instead of being silently dropped.

- seeds/sis_campus_program_map.csv: eight Alpha School campuses added (alphabetical by campus_name) so assert_sis_campus_program_map_covers_axis stays green.

- Stale name = 'School Year' scope references in stg_sis_program.sql and _sis__sources.yml updated to the current program_type = 'MAIN' shape (rule 22).

Why: program_type is SIS's own enum for the MAIN-vs-SUMMER_CAMP distinction. The name predicate only expressed "the year-long main program" as a side effect of the old one-shared-School Year-program convention; when SIS began provisioning per-campus <Campus> Main Program records, the label changed and eight active, PHYSICAL Alpha School campuses silently fell off the axis. The #1268 staging campus filter (not the name) is what keeps the 11 out-of-scope Main Program campuses out — the name predicate's only live effect was dropping these eight.

Closes #1279.

## Verification — real Redshift build (green)

Full dbt build --select path:models path:seeds --vars '{pr_number: 1279}' against sandbox_education: PASS=284, ERROR=0, SKIP=0, one pre-existing unrelated WARN (assert_admissions_pipeline_tenant_coverage, admissions-pipeline domain, present on main).

Acceptance tests passing: assert_sis_program_type_domain, assert_sis_campus_program_map_covers_axis, assert_sis_campus_program_map_hubspot_resolves, unique_int_school_year_offering_id, dbt_utils_unique_combination_of_columns_mart_enrollment_dtl_campus_id__cohort_id__session_school_year (has_fact grid uniqueness), assert_sis_enrollment_has_fact_consistency, assert_sis_enrollment_qualified_is_fact (qualified → has_fact completeness).

- The eight campuses now produce 10 campus×year axis rows (Anywhere Center - Founders & Greenwich - Armonk: 2026+2027; the other six: 2026). None were on the axis before.

- Axis grain unchanged: campus×year has no group > 1 before or after; unique_int_school_year_offering_id passes.

- SELECT DISTINCT program_type FROM int_school_year_offering → only MAIN; Summer Camp contributes zero axis rows (Port Chester's Summer Camp 2026 offering is correctly excluded).

- The four zero-enrollment campuses (Bethesda, Miami Beach (Biarritz), San Juan, Virtual) each form 13 factless grid rows (0 facts).

## Before/after cohort counts — SY2024–SY2027 (post-#1258 spine)

Measured on the same source snapshot: main's code and this branch both built fresh (pr9000_ baseline vs pr1279_), so the delta reflects only this change, not the hourly-refresh drift in the production tables. Counts are on the #1258-filtered spine (test records dropped), so they differ from the issue's raw-spine figures by design.

Distinct fact enrollments:

| School year | Before | After | Δ |

|---|---|---|---|

| SY2024 | 284 | 284 | 0 |

| SY2025 | 757 | 757 | 0 |

| SY2026 | 1621 | 1667 | +46 |

| SY2027 | 22 | 23 | +1 |

| Total | 2752 | 2799 | +47 |

Fact rows (cohort fan-out): SY2026 3735→3818, SY2027 25→26. Total mart rows incl. factless grid: SY2026 4255→4429, SY2027 542→569.

Delta attribution — confined to the eight new campuses + one explained re-enrollment. The +46 SY2026/SY2027 distinct fact enrollments are the four enrolled new campuses: Alpha Greenwich - Armonk 24, Alpha Anywhere Center - Founders 17, Alpha Austin - Founders 3, Alpha Lexington 2. (Bethesda, Biarritz, San Juan, Virtual add 0.)

The +1 at the existing campus Alpha Greenwich - Port Chester (SY2027, 10→11) is a correct second-order effect, not a predicate re-resolution — Port Chester's axis offering IDs are byte-identical before/after. One student transferred from the newly-admitted Alpha Greenwich - Armonk (2026) and holds an ON_HOLD 2027 enrollment at Port Chester. With Armonk now visible on the axis, that prior-year enrollment is seen, so the 2027 enrollment's is_returning flips FALSE→TRUE and it joins the re-enrollment-other cohort. The report is now correctly recognizing a returning student whose prior enrollment was previously invisible. No source drift (staging stg_sis_enrollment identical at 18,169 rows across both builds).

## Founders SIS-vs-HubSpot on-campus parity (#1216 restated)

Alpha Anywhere Center - Founders — the reported symptom (14 on-campus students in HubSpot, zero rows in SIS before this change) — now resolves. SIS cohort split (SY2026): on-campus 16, first-day 16, re-enrolled 1, re-enrollment-other 1 (17 distinct enrollments; the 16 PENDING_REVIEW land on-campus, plus one re-enrollment). SIS on-campus 16 vs HubSpot's reported 14 — this change closes the parity gap (SIS was showing 0) rather than opening a new one; the small SIS/HubSpot residual is the pending-review timing the issue describes. Full #1216 cross-system parity reconciliation (HubSpot-side models) is out of scope for this dbt change.

## Alpha Virtual — explicit decision

Alpha Virtual is admitted despite its name because SIS records it as delivery_mode = 'PHYSICAL', so the #1268 mode filter does not hold it back and the type-keyed axis includes it. It has zero enrollments today (factless grid rows only), so it costs nothing numerically. Correcting the SIS delivery_mode at source is Out of Scope — worth raising with the SIS team, but the report is not special-casing one campus while the source disagrees (same resolution as the Alpha World School case in #1268: SIS is the spine).

## Seed mappings — two reviewed decisions + build-driven corrections

- Alpha Miami Beach (Biarritz)has_hubspot_program = false. The one HubSpot Alpha Miami Beach program is already mapped to the distinct SIS campus Alpha Miami Beach (a34e72b4); pointing Biarritz at it too would double-count it in any SIS↔HubSpot parity comparison. (Whether Biarritz is a genuinely separate site vs a rename is an admissions data question — Out of Scope.)

Two corrections vs the issue's verbatim "Seed rows required" block, both required to pass existing seed tests and both following the seed's own documented conventions (caught by the mandatory local build):

1. program_name populated for Biarritz and Virtual with their identity name (= program_code: Alpha Miami Beach Biarritz, Alpha Virtual) instead of an empty field. The seed's not_null test on program_name (unchanged by this PR) rejects an empty CSV cell, which dbt loads as NULL. This matches the established SIS-only-campus convention — campuses HubSpot has no program for publish their own campus name as an identity program (e.g. Alpha Carrollton, Alpha Anywhere Center Port Chester). The identity name is distinct from Alpha Miami Beach, so the Biarritz no-double-count intent is preserved.

2. Alpha Austin - Founders and Alpha San Juan set has_hubspot_program = false (not true). Their HubSpot programs exist in EduCRM (26 and 39 rows) but carry zero real enrollment facts yet (has_fact = false throughout) — assert_sis_campus_program_map_hubspot_resolves requires a true mapping to resolve to a real fact. This is exactly the seed's documented "valid campuses whose HubSpot program carries no enrollment facts yet (SIS leads HubSpot for newly onboarded campuses)" case. The four campuses that do resolve to real facts (Anywhere Center - Founders, Bethesda, Greenwich - Armonk, Lexington) stay true.

The seed yml description counts are updated to the current shape (rule 22): has_hubspot_program TRUE 48→52, FALSE 10→14 (6 SIS-only + 8 no-facts-yet).

## Rule notes

Rule 8 preserved — the offering-kind filter stays in exactly one place (int_school_year_offering); only the column it reads changed. No new rule-5 exception taken. All descriptions/comments state the current shape (rule 22); no name = 'School Year' scope reference remains.

#1292 — feat(enrollment): drop HubSpot identities + overlay, publish SIS identity and deposit signal (#1290) @vvp-trilogy  approved

Closes #1290.

Reduces mart_enrollment_dtl to what SIS actually knows. This is a dbt-only change — the enrollment consumer (#1216) is still open, so nothing in production reads this mart yet.

## What changed

Identity — SIS only

- Removed HubSpot contact_id and deal_id from mart_enrollment_dtl, int_enrollment_cohort, int_enrollment, and int_student, and dropped the now-orphaned stg_sis_student_external_id / stg_sis_enrollment_external_id joins plus their two hubspot_* scoped-uniqueness tests.

- Renamed student_keystudent_id, unaliased from sis_enrollments.student_id through the intermediates to the mart. Removed the unused student_number projection entirely (grep -r student_number dbt/ returns nothing).

Deposit — drop the EduCRM overlay, publish the SIS signal

- Deleted int_deposit_overlay.sql, stg_educrm_enrollment_deposit.sql, and the two overlay singular tests. No dbt model reads EduCRM enrollment data any more; the educrm.enrollment_dtl source declaration is retained (read only by the parity_sis_vs_hubspot_enrollment_2026 analysis, per Out-of-Scope #9).

- deposit_paid_date is now published as CAST(NULL AS TIMESTAMP) on both arms — reserved, documented, appearing the day SIS grows a native paid-at.

- Added finalsite_deposit_state / finalsite_deposit_source to stg_sis_enrollment and carried them through int_enrollment and int_enrollment_cohort. The mart publishes:

- deposit_state — three-valued paid / unpaid / NULL (NULL = unknown, not "No"; every pre-2026 row is NULL).

- deposit_signal — SIS's own evidence vocabulary (AMOUNT, CHECKLIST, ADVANCE_DEPOSIT_CHECKLIST, ADVANCE_DEPOSIT_ATTRIBUTE), carried verbatim and not relabelled into the Admissions Pipeline macro's vocabulary; NULL whenever the deposit is not paid.

- New singular test assert_sis_enrollment_deposit_signal_requires_paid pins deposit_signal IS NULL whenever deposit_state is not paid. Added accepted_values on deposit_state (error) and deposit_signal (warn).

Docs — model yml, source yml, staging headers, and the rule-14 dual-identity paragraph updated to describe a single SIS identity; the removed EduCRM overlay docs deleted.

## Verification

- dbt build --select +mart_enrollment_dtl+ (isolated pr1290_* relations): clean (PASS, 0 errors; the only WARN is the pre-existing assert_sis_campus_unresolved_hubspot_program, unrelated to this change).

- Parity (projection/rename only): total rows (8,606), per-year distribution, and 2026 per-cohort fact counts are byte-identical to the pre-change build. Column count stays 34.

- Deposit cross-check: 2026 distinct-enrollment deposit_state totals — paid 1,576 / unpaid 32 / NULL 45 (= 1,653) — match sis_enrollments directly. deposit_paid_date is NULL on all 8,606 rows.

- Built the independent mart_admissions_pipeline_dtl in isolation and ran assert_admissions_pipeline_reconciles → PASS (deleting the EduCRM deposit staging model did not break the admissions pipeline).

- pnpm typecheck passes; pnpm biome check is a no-op (dbt is ignored).

## Notes on stale ticket references

The issue named assert_sis_campus_program_map_hubspot_resolves.sql and docs/enrollment-cohort-parity.md; neither exists in the current repo (superseded by #1286's int_school_identity work / never created). The only consumers of stg_educrm_enrollment_deposit were the overlay and its two tests (all deleted), so no test needed repointing.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1800 — fix(SURTR-648): use supported CAPEX batch request @marcusdAIy  approved

## Summary

- remove ExecutionMode from the CAPEX BatchExecuteStatement request

- retain atomicity through the Data API batch operation's transaction contract

- validate fake-client requests against botocore's real BatchExecuteStatement input shape

- add a regression proving unsupported request fields are rejected instead of silently accepted

## Why

The released installer passes ExecutionMode="TRANSACTION". The deployed botocore service model rejects that field during client-side parameter validation, before any DDL is submitted. BatchExecuteStatement already executes its Sqls as one transaction, so the option is unnecessary.

The prior fake accepted arbitrary keyword arguments and asserted the same invalid request. The replacement validates every submitted request through botocore while keeping the exact-request assertion. The installer omits ExecutionMode for compatibility with deployed SDK service models, including models that predate the optional field.

## Validation

- PYTHONPATH=src uv run pytest -q tests/ — 36 passed

- ruff==0.15.22 check — passed

- ruff==0.15.22 format --check — passed

- git diff --check — passed

## Rollout

This PR changes source only. It does not apply Redshift DDL, start a pipeline execution, enable on-demand execution, or enable the schedule. After a separate production promotion, the existing authorized fail-closed preflight and controlled installation can proceed.

Linear: [SURTR-648](https://linear.app/builder-team/issue/SURTR-648)

The Builder Desk  —  Engineer Spotlight
Production Release🏆 Engineer Spotlight

24 PRs, Six Repos, Zero Days Off: Builder Team Shatters the Clock Yet Again

Sanket Ghia alone shipped seven PRs to codex-software-factory in a single 24-hour window, and comrades, the machines are STILL running hot.

Twenty-four pull requests. Six repositories. One glorious 24-hour cycle of pure output, comrades, and the Builder Team did not blink. Aerie led the charge with eight PRs, codex-software-factory close behind with seven, Surtr posted four, mercy chipped in three, and Sindri and Klair each notched a disciplined, focused one. This is not luck. This is not coincidence. This is the machine of progress, oiled and roaring.

Sanket Ghia was the engine room today, posting a jaw-dropping seven PRs (#3 through #8, plus the fault-injection coverage) entirely inside codex-software-factory — hardening App Server reliability, gating durability, broadening resilience coverage. One man, one repo, seven acts of fortification. Keval Shah followed with five PRs spanning three repos — #182 in Sindri, #1291 in Aerie, and #126/#123 in mercy, rebuilding the second-opinion review prompt like a man negotiating with his own code. VVP Trilogy delivered four across Aerie, including the dbt CI overhaul in #1285 and the HubSpot dimension work in #1286. The heimdall-keval-factory bot notched two disciplined Surtr fixes (#1805, #1801), marcusdAIy shipped two — including the Fable 5.1 model upgrade in Klair's #3744 — and Yibin Long Trilogy cleared MCP surface debt with #1260 and #1259. Benji Bizzell and Caina Barbosa each landed one precision strike, #1288 and #1797 respectively.

Ashwanth Watch: notably absent from the sheet today, and comrades, that silence is LOUDER than any diff. When reached for comment on a 24-hour cycle without his name on it, sources close to the desk report he simply said, "I was reviewing the architecture of the next six months, you wouldn't understand the timescale." Whether he is resting, plotting, or quietly rewriting the entire monorepo in a bunker somewhere is unknown. His actual response to this report: "I don't read these."

The Overflow Desk salutes the 19 PRs Mac left on the floor. Keval Shah's mercy trilogy (#126, #123, #182) quietly rebuilt how the reviewer thinks about blocking merges — unglamorous, essential, ignored by the big desk, never by us. Sanket's four-pack of hardening PRs (#3–#6) in codex-software-factory reads like a fortress being built brick by brick while everyone else watched the front gate. And Yibin's surgical MCP removals, #1259 and #1260, prove that sometimes the loudest velocity is subtraction.

Morale Report: at an all-time high, as always, because it is never anything else. The desk hums, the commits flow, and somewhere Ashwanth is not filing a PR — and the team ships anyway. That, comrades, is the real leaderboard.

Brick's Overflow — PRs Mac Didn't Cover  (click to expand)
#3 — Gate 2: local durability and operations hardening @sanketghia  no labels

## Summary

This PR completes Gate 2 local-only maturity hardening for the Factory Console and coordinator.

- Adds Factory Health, startup reconciliation, and operator conflict visibility.

- Adds SQLite schema versioning, integrity checks, backup, and restore operations.

- Adds dry-run-first artifact and workspace lifecycle scanning with evidence-safe cleanup classification.

- Makes initial PR creation idempotent across configured-token and local-gh publication paths.

- Preserves uncertain Git push outcomes as non-retryable operator-recovery states.

- Closes the startup publication window so requested publication cannot be silently finalized as success before publication begins.

- Adds runbook documentation and a Gate 2 local evidence closure record.

## Local verification

- 77 test files, 1,225 tests passed.

- Browser regression: 10/10 passed.

- Console UI smoke: 7 routes passed.

- Build, lint, formatting, and diff checks passed.

- Representative local KLAIR-3383 batch completed successfully with 3/3 validations passed and publication evidence recorded.

## Scope and boundaries

- Local filesystem and loopback Console only.

- No hosted execution, remote deployment, merge automation, release, or cleanup was performed by this PR.

- Lifecycle cleanup remains review-first; no historical artifacts were deleted.

- Uncertain external outcomes remain explicitly operator-recovery states.

#4 — test: broaden local resilience coverage @sanketghia  no labels

## Summary

- Add restart-safe SQLite reconciliation coverage, including the requested-publication window.

- Preserve completed publication during reconciliation after PR metadata is persisted.

- Classify workspace bootstrap failures as WORKSPACE_BOOTSTRAP_FAILED while continuing sibling work.

- Add representative Console evidence for successful, validation-failed, publication-failed, and partial-batch outcomes.

- Document the local resilience coverage matrix and evidence.

## Verification

- pnpm test — 77 files, 1,229 tests passed.

- pnpm build passed.

- pnpm lint passed.

- pnpm format:check passed.

- Console UI smoke passed — 7 routes.

- Browser checks passed — 10/10.

- git diff --check passed.

- Live local KLAIR-3383 batch completed successfully with validation passed and publish:false.

## Scope

Local execution and Console resilience only. GitHub credential redesign and remote deployment are out of scope.

#5 — Harden App Server reliability and operator recovery @sanketghia  no labels

## Summary

- Add bounded App Server protocol/session metadata, interruption, checkpoint, and explicit CLI/Console resume support.

- Reuse the original workspace safely and continue with a bounded recovery instruction instead of blindly replaying the coding prompt.

- Make Factory context and skill policy explicit, keep target skills disabled by default, and add structured-output recovery.

- Reconcile workspace evidence, validation, patch integrity, and publication boundaries.

- Return an actionable Console error when GitHub publication credentials are unavailable.

## Verification

- pnpm test — 79 files, 1,257 tests passed

- pnpm build

- pnpm lint

- pnpm format:check

- git diff --check

## Live validation

- KLAIR-3383: cancellation, turn/interrupt, Console restart, same-run App Server resume, and successful no-publication completion.

- KLAIR-3476: clarification flow succeeded; validation repair was exercised and stopped after its bounded limit when client-format continued failing.

- KLAIR-3389/KLAIR-3390: representative no-publication runs completed with success and preflight-blocked outcomes respectively.

- Disposable publication proof created PR #3747; it was intentionally closed afterward.

## Scope boundary

No merge, deployment, automatic supervisor, or automatic merge behavior is included. Publication, merge, and deployment remain separate explicit gates.

## Follow-up

- Run a post-merge publish:false smoke test.

- Continue measured no-publication pilots.

- Diagnose whether the observed client-format repair failure is ticket-specific or a broader repair-harness issue.

#6 — Improve validation repair diagnostics and convergence @sanketghia  no labels

## Summary

- Add deterministic validation failure fingerprints to receipts and repair events.

- Stop early when the same validation failure repeats unchanged instead of consuming another blind repair attempt.

- Include the exact validation working directory and command line in repair prompts.

- Record failed command IDs, categories, changed-file counts, and repair fingerprints in bounded evidence.

## Verification

- pnpm test — 79 files, 1,258 tests passed

- pnpm build

- pnpm lint

- pnpm format:check

- git diff --check

- Live publish:false KLAIR-3476 run: initial client-format failure, one repair attempt, final validation passed, Ready to publish.

## Scope

Publication policy is unchanged. No merge, deployment, or automatic publication behavior is added.

#7 — Persist Codex App Server protocol evidence and failure state @sanketghia  no labels

## Summary

- Add flushable durable App Server event persistence.

- Record bounded protocol checkpoints with process, thread, turn, and protocol identity.

- Persist richer notification/response metadata and typed failure categories.

- Fail closed on malformed JSONL, preserve cancellation evidence, clean up owned Codex homes when evidence flushing fails, and clear stale recovery state.

## Verification

- pnpm test — 79 files, 1,263 tests passed

- pnpm build passed

- pnpm lint passed

- pnpm format:check passed

- git diff --check passed

## Scope

This PR does not add App Server supervision, automatic compaction, publication, merge, or deployment behavior.

#8 — Add App Server protocol fault-injection coverage @sanketghia  no labels

## Summary

- Add a controllable App Server fault-injection harness.

- Cover malformed JSONL, stream/process failures, authentication, missing IDs, workspace mutation before process loss, invalid output, context failure, unexpected requests, timeout, and cancellation.

- Pin timeout and cancellation faults to the turn phase and assert the interrupt request.

- Add a runTicket boundary case proving failed receipt creation, persisted protocol evidence, validation suppression, and preservation of partial workspace state.

## Verification

- pnpm test — 80 files, 1,276 tests passed

- pnpm build passed

- pnpm lint passed

- pnpm format:check passed

- git diff --check passed

## Scope

This PR adds deterministic local fault coverage only. It does not add App Server supervision, automatic compaction, publication, merge, or deployment behavior.

#123 — feat(review): a second opinion on whether a finding must block the merge @kevalshahtrilogy  changes requested

## What

The reviewer decides whether a finding is real. Nothing decides whether a real finding must be fixed before merge or can land as a follow-up — and that second question governs how many rounds a PR takes.

This adds it. For each finding that would gate, a pass over the PR description, the full diff, the whole of every file the findings cite, and the reviews and comments posted before this round. One question per finding: *does merging this unfixed break production?* Two possible answers: keep blocking, or ship as a follow-up.

## Measured

Replayed 63 PRs whose review was already at round 4 or later — five repos, six weeks:

PRs blocked      63

clear the gate 19 (30%)

findings judged 126 (27 more never eligible)

deferred 63 (50%)

What it defers is dominated by guards against input the system does not produce (34%) and findings the code already handles (19%).

## Then read by hand

Every contested call — all 17 where the reviewer raised the same file:line again on a later round:

| | |

|---|---|

| Reviewer stuck or wrong | 11 |

| Genuinely arguable | 3 |

| Unjudgeable (anchor problem, below) | 3 |

| A defect that would have broken production | 0 |

Five of the eleven are one PR whose own file header reads *"Hidden experimental comparison endpoint… not part of any production read path"* — production-blocking standards applied to an internal dashboard. Two describe code that had been deleted.

## Three properties held on purpose

It can only defer. Never promote, never invent a finding, never raise a severity. The worst it can do is fail to block.

It is invisible. Nothing it decides is named in the review body, in an inline comment, or anywhere a reader of the PR would see. A deferred finding is reported exactly as before — it simply stops counting toward the gate. The ledger records the deferral unattributed: the PR gains no second reviewer's voice.

Every failure makes the reviewer stricter. No key, network error, no JSON, malformed JSON, unknown verdict, missing id — all resolve to "this blocks", which is today's behaviour. security and anything marked critical are never submitted at all.

## Whole files, not a window

The reviewer's line anchors drift. Three findings in the audit described receipt deletion while pointing at a daily-counter function 400 lines away — a window around the anchor showed the pass code the finding was not about.

Switching to whole files moved 22 of 125 verdicts, 13 of them back to BLOCKING. The narrow window had been hiding what made them real. It also produces grounded reasoning: *"isValidIsoInstant() reconstructs the UTC calendar date and compares it to the timestamp's local date"* rather than *"the documented producers write UTC."*

## Known, and deliberately not mitigated

A wrongly deferred finding stays deferred on every subsequent round — same evidence, same answer. The obvious guard, block after N repeat deferrals, was considered and rejected: it lets a reviewer win by repetition rather than by being right, which is the exact failure this exists to end. The ledger record is the mitigation, and it relies on a human noticing. Worth knowing before this gates anything.

## Testing

397 tests. Nine mutations, nine killed:

| mutation | |

|---|---|

| deferral ignored | killed |

| a truthy value accepted as a deferral | killed |

| ledger drops the flag | killed |

| ledger drops the reason | killed |

| carry-forward drops the deferral | killed |

| security made eligible | killed |

| critical made eligible | killed |

| an unusable answer defaulting to defer | killed |

| apply() promoting instead of demoting | killed |

## Business Value

Targets the 14% of PRs that consume 42% of all review runs. On the replay it clears the gate on 19 of 63 already-stuck PRs without, on hand audit, shipping a single production defect — and it does it invisibly, so the PR reads exactly as it does today.

## Manual Effort Estimate

~6 hours (proposed — Keval to confirm). Most of it was building the backtest honestly: three separate measurement bugs of my own (a lookahead leak, drifted anchors, and same-line-different-finding) each inflated the result before being found and fixed.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#126 — feat(review): rebuild the second-opinion prompt around how the reviewer actually fails @kevalshahtrilogy  no labels

## Measured — 112 PRs, one blocking review each, all at round 4+, five repos

| | old prompt | new prompt |

|---|---|---|

| PRs cleared | 45 (40%) | 40 (36%) |

| Findings deferred | 101 (60%) | 90 (54%) |

| Stricter | — | 31 |

| Looser | — | 20 |

Stricter, and stricter in the right places. Of the 31 findings it newly keeps blocking, the samples are money and data bugs the old prompt was shipping:

> *"discards _cost_malformed, so malformed Cost cells are silently reported as $0.00 and understate the spend-audit totals"*

> *"maps persistence failure to terminal FAILED without freeing the dedupe row, making a transient outage permanent"*

> *"retains a school with a blank display_name in the production refresh"*

19 of the 31 are money / data-integrity / observability, against a 52% base rate.

## What changed, and why each

Restructured, not appended. An earlier attempt that appended a single rule made the prompt *looser* overall. This keeps the same length with more structure.

Blast radius is question one. Five of the eleven confirmed-wrong blocks in the hand audit were on internal experimental tools, dev launchers, or routes behind a privileged flag — held to production standards.

"Removed is not fixed." The one case my own hand audit got wrong: version machinery was deleted, but the !existing branch still permitted the bad publish.

Evidence rules for the two dominant buckets.

- *"Nothing produces this input"* — 34% of deferrals, the riskiest — must point at what rules it out: a schema, an upstream validator, the sole writer. Absence of evidence is not evidence.

- *"The code already handles it"* — 19% — must quote the line.

- Author rebuttals count only with a file:line, a test, or a contract.

Per-repo "wrong result", one line each. Klair money and Surtr warehouse rows named rather than inferred.

Uncertainty made concrete. A reason needing *could / may / might / probably* is a prod_breaking.

## The trade-off to know about

| repo | old | new |

|---|---|---|

| Aerie | 20/44 | 22/44 |

| Klair | 13/24 | 13/24 |

| trilogy-drones | 4/12 | 4/12 |

| Sindri | 1/6 | 1/6 |

| Surtr | 7/26 | 0/26 |

On Surtr the gate now defers nothing. The per-repo line describes nearly every Surtr finding, and combined with "the reviewer is reliably right about silent data corruption" it closes the door. That is the correct default for a warehouse repo — data correctness *is* the product — but it means zero relief there. Recorded so it can be narrowed with two weeks of live evidence rather than a guess.

## Honest caveats

- 20 findings went looser for reasons the prompt does not explain. That is the sampling variance documented on #125; one run per prompt cannot separate it from effect. The stricter side is supported by the samples. The looser side is not claimed.

- Re-raise rate ~38% on both arms — mostly measurement artifact per the earlier hand audit, not read as an error rate.

## Testing

Prompt-only change; 397 harness tests pass unchanged. A prompt change is a model-behaviour change and is not unit-testable — the A/B above is the test.

## Business Value

This is the gate that decides whether a real finding must block a merge or can ship as a follow-up. Getting it wrong in one direction ships bugs; in the other it recreates the deathloop the team is angry about. The rewrite makes it measurably safer than the version currently live on Surtr and Klair — it catches money and data-loss bugs the old prompt was waving through — while still clearing roughly a third of PRs already stuck at round 4+. It also unblocks moving v1, which carries the deathloop bug fixes to Aerie, trilogy-drones and Sindri, who are still running the pre-fix reviewer.

## Manual Effort Estimate

~2 hours (proposed — Keval to confirm). The prompt itself is an hour; the rest was profiling how mercy blocks per repo and running the 112-PR A/B to make the change evidence-led rather than intuition-led.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1259 — Remove presentReplanOptions and generatePlan @YibinLongTrilogy  approved

## Summary

Remove the retired getPortfolioHealth aggregate from every Aerie-owned product surface: Convex, MCP discovery, Worker/stdio registration, agent contracts and policies, UI rendering, Flue output handling, public API metadata, documentation, fixtures, and tests. The supported v2 Insights list/detail workflow remains available, while proactive Worker guidance now uses bounded status-filtered site rosters plus factual overdue-milestone evidence.

### Changes

- chat/convex/rhodes/mcp.ts and chat/rhodes-worker/mcp-server/tools/views.ts — Delete the public Convex query and MCP registration; add an explicit bounded listSites path for status-filtered rosters.

- packages/contracts/src/agent-run-protocol.ts, packages/contracts/src/agent-tool-registry.ts, and chat/convex/agentRuns/runs.ts — Remove the tool from agent names, schemas, descriptions, public policies, and capability routing; bump the public tool-policy version to invalidate stale runs.

- chat/lib/rhodes-mcp-contract.ts and UI/Flue renderers — Remove in-app advertisement, the dedicated portfolio-health card, and the obsolete nested gateway fixture; keep listSites milestone progress visible with DTO-shaped coverage.

- chat/rhodes-worker/lib/mcp-instructions.ts and mcp-server/tools/sites.ts — Replace the retired proactive workflow with capped active/paused site reads and factual overdue reporting.

- chat/lib/public-api/v2/domains/insights.ts, parity guard, docs, and feature history — Preserve the supported v2 Insights contract, document the later cleanup, and allow the intentional cross-repository MCP removal during parity checks.

- Tests — Remove obsolete behavior fixtures and add negative registry coverage, bounded roster coverage, forwarding/schema coverage, and UI DTO coverage.

### Design Decisions

- The v2 Insights list/detail API is retained because it is a supported bounded public workflow, not the retired MCP aggregate.

- listSites remains generally compatible for existing callers; only explicit status-filtered calls with limit use the bounded roster path used by proactive guidance.

- The replacement instructions report only returned milestone progress and overdue evidence; they do not infer health labels from missing evidence or the absence of overdue items.

- The public Agent tool-policy version is incremented so durable runs created under the prior tool catalog fail closed rather than appearing compatible with a changed policy.

## Business value

Eliminates a retired, misleading aggregate health surface and prevents agents or UI clients from presenting stale readiness scores, ratings, blocker rankings, or unsupported portfolio judgments. Users retain factual, bounded milestone and overdue evidence through the supported workflows.

## Estimated manual effort

Estimated time to complete this work without AI: 1 working day.

## Test Plan

- [x] Full Rhodes Worker test suite.

- [x] Focused Convex parity, public-agent, contract, site-tool, and Rhodes card tests.

- [x] Chat, Worker, and contracts typechecks.

- [x] Biome, architecture-boundary, Convex-path, read-bound, test-architecture, and git diff --check validation.

- [x] Read-only seven-lane adversarial review with independent verification of accepted findings.

- [ ] Run deployed MCP discovery/parity checks after the normal release deployment; no deployment or external mutation was performed for this PR.

#1260 — Remove getPortfolioHealth MCP surface @YibinLongTrilogy  approved

## Summary

Remove the retired getPortfolioHealth aggregate from every Aerie-owned product surface: Convex, MCP discovery, Worker/stdio registration, agent contracts and policies, UI rendering, Flue output handling, public API metadata, documentation, fixtures, and tests. The supported v2 Insights list/detail workflow remains available, while proactive Worker guidance now uses bounded status-filtered site rosters plus factual overdue-milestone evidence.

### Changes

- chat/convex/rhodes/mcp.ts and chat/rhodes-worker/mcp-server/tools/views.ts — Delete the public Convex query and MCP registration; add an explicit bounded listSites path for status-filtered rosters.

- packages/contracts/src/agent-run-protocol.ts, packages/contracts/src/agent-tool-registry.ts, and chat/convex/agentRuns/runs.ts — Remove the tool from agent names, schemas, descriptions, public policies, and capability routing; bump the public tool-policy version to invalidate stale runs.

- chat/lib/rhodes-mcp-contract.ts and UI/Flue renderers — Remove in-app advertisement, the dedicated portfolio-health card, and the obsolete nested gateway fixture; keep listSites milestone progress visible with DTO-shaped coverage.

- chat/rhodes-worker/lib/mcp-instructions.ts and mcp-server/tools/sites.ts — Replace the retired proactive workflow with capped active/paused site reads and factual overdue reporting.

- chat/lib/public-api/v2/domains/insights.ts, parity guard, docs, and feature history — Preserve the supported v2 Insights contract, document the later cleanup, and allow the intentional cross-repository MCP removal during parity checks.

- Tests — Remove obsolete behavior fixtures and add negative registry coverage, bounded roster coverage, forwarding/schema coverage, and UI DTO coverage.

### Design Decisions

- The v2 Insights list/detail API is retained because it is a supported bounded public workflow, not the retired MCP aggregate.

- listSites remains generally compatible for existing callers; only explicit status-filtered calls with limit use the bounded roster path used by proactive guidance.

- The replacement instructions report only returned milestone progress and overdue evidence; they do not infer health labels from missing evidence or the absence of overdue items.

- The public Agent tool-policy version is incremented so durable runs created under the prior tool catalog fail closed rather than appearing compatible with a changed policy.

## Business value

Eliminates a retired, misleading aggregate health surface and prevents agents or UI clients from presenting stale readiness scores, ratings, blocker rankings, or unsupported portfolio judgments. Users retain factual, bounded milestone and overdue evidence through the supported workflows.

## Estimated manual effort

Estimated time to complete this work without AI: 1 working day.

## Test Plan

- [x] Full Rhodes Worker test suite.

- [x] Focused Convex parity, public-agent, contract, site-tool, and Rhodes card tests.

- [x] Chat, Worker, and contracts typechecks.

- [x] Biome, architecture-boundary, Convex-path, read-bound, test-architecture, and git diff --check validation.

- [x] Read-only seven-lane adversarial review with independent verification of accepted findings.

- [ ] Run deployed MCP discovery/parity checks after the normal release deployment; no deployment or external mutation was performed for this PR.

#1285 — AERIE-1283: dbt CI — models only in production, models then tests in dev @vvp-trilogy  no labels

## Summary

Splits the dbt GitHub Action (.github/workflows/dbt.yml) so the two credentialed writers run different command shapes per environment, per issue #1283. Today both jobs issue the identical dbt build --select path:models path:seeds, interleaving all tests into the DAG in both environments — wrong in opposite directions.

## Changes

- scheduled-build (production) — added --exclude-resource-type test to the existing dbt build step. The hourly refresh now runs models and seeds only, zero tests, so a tripped data-quality assertion never stops the refresh. Kept build (not run) so seeds stay ordered ahead of the models that select from them.

- pr-build (dev) — split the single dbt step into two steps against the same resolved secret: a models/seeds build --exclude-resource-type test step, then a distinct, required dbt test step. Both pass --vars "{pr_number: N}" so tests read the pr<N>_-prefixed objects the build just wrote. No continue-on-error — a failing test fails the check, and the pr<N>_ objects still exist for inspection.

- dbt/Dockerfile.dbt — updated CMD to the production shape (build … --exclude-resource-type test).

- Docs — refreshed the workflow header comment and added a CI section to dbt/README.md, both stating plainly that production runs no tests.

- Untouched: pr-parse-only (fork PRs) and pr-cleanup.

## Verification

- .github/workflows/dbt.yml parses (yaml.safe_load).

- Dockerfile CMD is valid JSON exec-form.

- Diff re-read against every Acceptance Criterion in #1283.

Closes #1283

#1286 — AERIE-1284: Shared HubSpot display-name dimension for enrollment + pipeline reports @vvp-trilogy  approved

## Summary

Replaces the hand-maintained sis_campus_program_map.csv seed with a derived dimension, int_school_identity (one row per HubSpot program), that maps every system's school identity — SIS campus_id, Finalsite site, HubSpot program object id, HubSpot program_code — to one user-friendly label (display_name, the HubSpot program display name). Both the Enrollments report (mart_enrollment_dtl) and the Admissions Pipeline report (mart_admissions_pipeline_dtl) consume it, so the two reports name the same school identically. The mapping is read from SIS, so the report self-heals: adding a hubspot_program external id in SIS admits a campus on the next run with no repo change.

Closes #1284.

## What changed (following the issue's 8 steps)

1. all_program declared as an educrm source; new stg_educrm_all_program (pure 1:1 projection, equal_rowcount test).

2. campus_external_ids declared as a sis source; new stg_sis_campus_external_id (soft-delete filter, casts, no joins) with a (campus_id, system) uniqueness test. The source envelope has no surrogate id, so the grain is keyed on (campus_id, system).

3. int_school_identity in the new intermediate/shared/ folder. Grain: one row per HubSpot program; resolves display_name from all_program.program_name on the numeric hubspot_program_id, never on a name. Columns: hubspot_program_id (bigint), hubspot_program_code, sis_campus_id, finalsite_site, display_name — only display_name is unprefixed.

4. Drop rule in stg_sis_campus — an EXISTS semi-join so a campus that resolves to no HubSpot program never reaches int_campus/axis/grid/mart. Documented in the header as a second knowing rule-5 exception with the removal trigger named (the SIS-name cutover).

5. mart_enrollment_dtl repointed at the dimension: the program_map CTE reads int_school_identity, the LEFT JOINs become INNER (every campus reaching the mart resolves by construction), and has_hubspot_program is gone.

6. mart_admissions_pipeline_dtl — all three arms publish program_name from the dimension. The Finalsite arm INNER-joins on finalsite_site (drops unresolvable tenants, symmetric with the enrollment drop); the EduCRM arms LEFT-join on hubspot_program_code (null-tenant pre-launch leads are kept). Existing campus_name/campus_long_name/program_code are unchanged — the label is added, not renamed.

7. Deleted sis_campus_program_map.csv, assert_sis_campus_program_map_covers_axis.sql, assert_sis_campus_program_map_hubspot_resolves.sql. Repointed assert_sis_deposit_overlay_unmatched_within_threshold.sql's Alpha-program bound at the dimension.

8. New WARN test assert_sis_campus_unresolved_hubspot_program.sql (dropped-campus visibility, shape modelled on assert_admissions_pipeline_tenant_coverage). Updated assert_admissions_pipeline_reconciles.sql so the Finalsite expected arm applies the same drop filter.

The published mart column names (program_code, program_name) are unchanged — the frozen contract with the sync and Convex. Only the internal derivation changed.

### One extra change: connect_timeout 30 → 120

redshift_connector applies connect_timeout as the socket read timeout, so it bounds every query. Removing the covers_axis blocker re-enables mart_enrollment_dtl, whose dense factless anti-join over the int_enrollment_cohort view chain legitimately runs ~40s against current data — so a 30s timeout left it red on a timeout (proven locally: it builds in ~36–39s with the timeout raised; the query is not a SQL error). Raised to 120s in profiles.yml.docker (CI) and profiles.yml.example (local). This is pre-existing cost, not a regression from this PR — the change *reduces* the grid from 67 to 53 campuses.

## Verification

Full dbt build --select path:models path:seeds --vars '{pr_number: N}' run against the real Redshift warehouse: PASS=289, WARN=2, ERROR=0, SKIP=0.

- The 2 WARNs are the new assert_sis_campus_unresolved_hubspot_program (14 dropped campuses) and the pre-existing assert_admissions_pipeline_tenant_coverage (6, data drift — both its inputs are untouched by this PR).

- Acceptance checks against the built objects: enrollment mart has 0 null program_code/program_name across 53 campuses (was 67); Nova Academy Austin → Nova Austin, Nova Academy Bastrop → Nova Bastrop; cross-mart program_name consistency = 0 mismatches; all three pipeline arms 100% labelled; the Finalsite arm drops exactly the two founders tenants (32 rows).

- pnpm typecheck, pnpm biome check, pnpm lint:boundaries: all clean (no TS files touched).

## Open question for sign-off (from the issue, unchanged by this PR)

Dropping unresolvable campuses removes real enrollments at a few campuses (the issue's "22 students", chiefly the two Founders entities and Alpha Orlando). This is by design and surfaced by the new WARN test; adding the eight hubspot_program bindings in SIS (an out-of-repo action listed in the issue) returns those campuses automatically on the next run. Not a code dependency.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

The Portfolio  —  Trilogy Companies

The Discount Bin: How ESW Capital Turns Tech's Castoffs Into Cash

From Portland's fallen social-intranet darling to a venture-backed customer-experience shop, Austin's acquisition machine keeps finding software nobody else wants — and margins nobody else can match.

AUSTIN, TEXAS — Jive Software once commanded a market cap that made Portland tech proud. It sold this year for roughly half its peak value. The buyer, unsurprisingly, was ESW Capital, the Trilogy International subsidiary that has spent two decades perfecting a very specific kind of shopping trip: buy the software nobody wants at a price nobody else will pay, then squeeze.

The pattern repeated itself with ResponseTek, a venture-backed customer-experience analytics firm now folded into ESW's telecom-focused Skyvera portfolio. The company had raised real venture money, built real technology, and — like so many of ESW's targets — arrived at the point where its investors needed an exit more than they needed a strategy. ESW was there, as pehub.com reported, ready to write the check.

The Wall Street Journal, in a rare mainstream look at the machine, framed it as a refuge — a home for small software companies otherwise destined for the scrap heap. That framing is not wrong, exactly. It is also the framing ESW itself prefers: rescuer, not liquidator.

Meanwhile, Forrester's analysts are publishing guidance on what enterprises should do with their aging customer advocacy platforms — a category that includes exactly the sort of tools ESW specializes in acquiring once the venture money runs dry and growth stalls. The research firm's advice to customers, in effect, describes the moment right before a company becomes an ESW target: stagnant roadmap, sticky contracts, declining strategic priority.

Who benefits when a once-promising platform's fate is decided not by its product roadmap but by its balance sheet? The founders got their multiple. The employees, many of them, will not be staying. And somewhere in Austin, a spreadsheet is already modeling the EBITDA math on the next 25% price increase.

Small Software Companies Find a Home With ESW Capital - WSJ  ·  What To Do Next About Your Customer Advocacy Platform - Forr  ·  ESW Capital acquires venture-backed ResponseTek - pehub.com

Skyvera's Telecom Land Grab: Two Acquisitions, One Unmistakable Pattern

As Skyvera swallows CloudSense and STL's telecom assets in quick succession, the ESW playbook for consolidating legacy telco software is running exactly on schedule.

AUSTIN, TEXAS — Two acquisitions in short order rarely happen by accident, and if you read between the lines of Skyvera's latest moves, you start to see the shape of something much larger than two press releases suggest.

First came confirmation that Skyvera had completed its acquisition of CloudSense, the Salesforce-native CPQ platform that helps telcos quote, configure, and fulfill their most complicated B2B and wholesale deals. Then, almost as an afterthought in the filing cabinet of portfolio updates, Skyvera absorbed STL's divested telecom products group — a business carrying digital BSS functionality spanning monetization, optical networking, and analytics.

On paper, these are separate deals. My source inside the portfolio tells me otherwise: this is one acquisition thesis, executed in two acts. Telecom operators are sitting on decades of on-premise BSS infrastructure that nobody wants to rip out and nobody can afford to modernize alone. Skyvera's answer, consistent with every other ESW Capital roll-up you've read about in these pages, is to buy the sticky legacy software cheap, consolidate it under one roof, and let AI do the modernization work that used to require armies of consultants.

And this is where it gets interesting: CloudSense didn't just get acquired — it got fast-tracked. Skyvera's team recently certified all 13 of CloudSense's CPQ APIs to TM Forum compliance standards in a single month, a process the industry considers a 26-month undertaking under conventional development. That's not a coincidence of timing. That's a demonstration. Nobody accelerates a compliance certification by 25 months unless they're trying to prove something to a market — namely, that AI-driven engineering can absorb an acquisition's technical debt faster than any competitor bidding for the same distressed telecom assets.

STL's optical networking and monetization tools slot neatly into that thesis, giving Skyvera the analytics layer to match CloudSense's front-end quoting muscle. Whether this is the last piece or simply the latest, my source wouldn't say. But the pattern, once you see it, is hard to unsee.

Cloudsense  ·  CloudSense achieves TM Forum API compliance in record time u  ·  Skyvera completes acquisition of CloudSense, expanding telec

As Microschools Boom, Austin's AI Classroom Experiment Looks Less Like an Outlier

A wave of new reporting on America's fragmented K-12 landscape suggests the regulatory and cultural terrain is shifting toward exactly the kind of model Joe Liemandt bet a billion dollars on.

AUSTIN, TEXAS — There is a particular kind of vindication that comes not from a company's own press release, but from the slow drumbeat of outside coverage confirming that the ground has moved beneath everyone's feet. That is roughly what is happening this week in K-12 education reporting, where outlets from Stateline to Christianity Today to The 74 have converged, independently, on the same uncomfortable observation: the American public is quietly abandoning the assumption that a school must be a building with 30 kids, one teacher, and a six-hour day.

Stateline's reporting on microschools lands on a familiar tension for anyone who has followed Alpha School's rapid expansion out of Austin and into Brownsville, Miami, and — by fall — nine additional campuses across Texas, Florida, Arizona, California, and New York: the regulatory apparatus simply was not built for this. State licensing regimes, seat-time mandates, and accreditation frameworks were designed for a world in which learning happens on a clock, not for one in which a child can master a year of NWEA-tested curriculum in twenty hours of adaptive, AI-guided instruction, freeing the rest of the day for entrepreneurship, public speaking, or athletics — the model Alpha and its co-founder MacKenzie Price have been quietly selling to state education officials, including a direct presentation to U.S. Secretary of Education Linda McMahon.

What is striking, and what deserves more scrutiny than it has gotten, is the convergence of forces: the faith-based education revival covered by Christianity Today, the parental hunger for alternatives documented by The 74, and the microschool proliferation flagged by Bored Teachers are not separate phenomena. They are symptoms of the same eroding confidence in the seat-time model that Joe Liemandt has been betting on since he stepped back into education. Whether policymakers can build guardrails fast enough to keep pace with the market — or whether Timeback's ambition to become the 'Shopify for schools' outpaces the state's capacity to regulate it — remains the story to watch.

5 Trends Reshaping K-12 Education Across the U.S. - The 74  ·  Microschools are growing in popularity, but state regulation  ·  Faith-Based Education Is Having a Moment - Christianity Toda
The Machine  —  AI & Technology

The Whisper and the Verdict: How Small Minds Are Learning to Speak for Giants

Two new papers chase the same elusive goal — letting a small AI model guess quickly so a giant one only has to nod, or shake its head.

AUSTIN, TEXAS — There is an old evolutionary trick, older than language itself: let a fast, cheap system make a guess, and let a slower, more expensive system merely confirm or reject it. Your peripheral vision does this. A flicker in the grass triggers a reflex before your cortex has finished composing the thought "tiger." The expensive brain checks the cheap brain's work, and together they are faster than either alone.

Large language models have begun to borrow this trick, and this week's arXiv haul shows the idea maturing into something genuinely strange: a division of cognitive labor between machines that may not even speak the same internal language.

The technique is called speculative decoding. A small, fast "drafter" model proposes a string of likely next words; a large, authoritative model then verifies them in a single pass, accepting the good guesses and correcting the bad ones. It's a way of getting a giant model's judgment at a fraction of its usual latency — so long as the small model guesses well.

But small models, it turns out, suffer from the same fragility as narrow specialists everywhere. A paper on Osprey notes the irony: we prize large "target" models precisely for their broad competence, yet we keep training their small drafting partners on narrow slices of data, so their usefulness collapses the moment the workload shifts. Osprey's fix is target-agnostic pretraining — teaching the drafter to guess well in general, rather than to mimic one particular giant, producing a more versatile whisper that holds up no matter which authority it's whispering to.

A second paper, X-CoSD, tackles a subtler problem: what happens when the small on-device model and the large server model don't even share a vocabulary? Most collaborative speculative decoding schemes assume both sides tokenize the world identically — a shared alphabet of meaning. X-CoSD strips that assumption away, building a communication-efficient bridge between mismatched vocabularies, so a phone-sized model and a data-center titan can collaborate without translating every syllable back and forth.

Neither paper will make headlines outside the narrow province of inference engineers. But together they gesture at something larger: intelligence, biological or artificial, may always come in layers — fast and cheap, slow and sure — learning, generation after generation, to trust each other's instincts just enough.

X-CoSD: Communication-Efficient Cross-Vocabulary Collaborati  ·  StochBench: A Domain-Specific Benchmark for Stochastic Proce  ·  Osprey: Target-agnostic Pre-training Makes Stronger Drafters

The Agent Wars Just Got Real: Google and Apple Both Blink First

Two tech giants dropped major AI developer tooling upgrades on the same day, and the message is unmistakable: 2026 is the year of the autonomous agent.

MOUNTAIN VIEW, CALIF. — I have to catch my breath here, because what just happened in the developer tools space is nothing short of seismic, and if you're not paying attention you are going to miss the moment the AI agent economy tipped over into full production reality.

Google just expanded Managed Agents in the Gemini API, and folks, this is not some incremental update. We're talking background tasks that run autonomously, remote MCP (Model Context Protocol) support, and infrastructure that lets developers deploy agents that actually keep working while you sleep. I cannot overstate how significant this is — this is the difference between an AI that answers your question and an AI that goes off and does the job.

And then, almost in the same breath, Apple rolled out new intelligence frameworks aimed squarely at app developers, giving builders on-device tooling to weave generative AI into apps without shipping user data off to the cloud. This is Apple playing catch-up and playing it smart — privacy-first AI that developers can actually ship.

Why does this matter to us here at The Trilogy Times? Because this is exactly the terrain Trilogy's engine is built for. Across ESW Capital's 75-plus portfolio companies — think DevFactory's engineering muscle, or Ephor's AI finance tooling — every one of these advances in managed agents and on-device intelligence frameworks becomes raw material. Crossover's global talent pipeline exists precisely to find the engineers who can take Google's remote MCP support or Apple's new frameworks and turn them into shipped product, fast.

The future is now, and it's arriving through developer tools. Buckle up.

Expanding Managed Agents in Gemini API: background tasks, re  ·  Apple aids app development with new intelligence frameworks  ·  8 Best AI Tools for Developers in 2026 (Ranked & Reviewed) -

The Great Silicon Migration: A New Watering Hole Rises in the East

As memory chips grow scarce and old supply routes falter, the semiconductor herd begins its cautious trek toward Malaysia.

KUALA LUMPUR — Observe, if you will, the modern semiconductor supply chain — a fragile, sprawling ecosystem that has, for decades, congregated around a handful of watering holes: Taiwan, South Korea, and the mineral-rich flats of mainland China. But drought is coming to these ancestral grounds, and the herd, ever adaptive, is beginning to move.

Here in Malaysia, we witness the early stirrings of a new gathering ground. Government and industry alike now court the great fabrication houses of the world, hoping to become what Electronics Weekly calls a regional semiconductor hub — a refuge for assembly and testing operations fleeing geopolitical turbulence elsewhere.

And turbulence there is. In China, the AI chip breeding grounds have grown restless. Domestic chipmakers, starved of the high-bandwidth memory that sustains any healthy AI accelerator, have begun raising prices — a telltale sign of scarcity rippling through the food chain, as Global Sources reports. HBM, that rare and coveted nutrient, grows harder to secure by the month.

Far to the west, another disturbance ripples through the supply web — helium, that most peculiar and essential gas for chip manufacturing, disrupted by conflict in the Persian Gulf, its flow to fabrication sites growing unpredictable, as TechInsights' Chip Insider dryly notes.

The Washington Times, meanwhile, asks the question every anxious nation now ponders: how does one stop depending on a single, distant watering hole entirely? The answer, if there is one, lies not in any single sanctuary but in many — Malaysia among them — each modest, each imperfect, each essential to a species that cannot survive without silicon, and increasingly, cannot survive on China's terms alone.

Malaysia pushes to become a regional semiconductor hub - Ele  ·  The Chip Insider®–The Iran War’s Helium disconnect update -  ·  Chinese AI chipmakers boost prices amid escalating HBM short
The Editorial

Unpopular Opinion: The 'Future of Work' Reports Are Just Describing Crossover's Business Model

I'll be honest — while consultants are out here 'discovering' distributed remote talent as a 2026 trend, some of us have been living it since before it was a Gartner slide.

AUSTIN, TEXAS — I'll be honest, I read the new Gartner Future of Work Trends for CHROs report this morning with my cold brew and almost fell out of my Aeron chair.

Twenty statistics. Massive workplace transformation.

Bro. We know.

This isn't a prediction, it's a rearview mirror.

While HR conference panels are still workshopping whether remote-first is 'sustainable,' Crossover has spent years proving that top 1% talent doesn't care what timezone your Slack status is in.

130+ countries. Identical above-market pay regardless of geography. That's not a 2026 trend, that's just Tuesday.

But here's where it gets interesting, and this is a real learning opportunity for the skeptics.

The World Economic Forum dropped three charts on how AI is reshaping wages and hiring, and the data is basically the business case for a talent platform that already benchmarks pay against a global pool of the best operators, not the nearest zip code.

When AI compresses the value of routine execution and inflates the value of judgment, taste, and speed, you don't want an org chart built for 1998.

You want a bench of top 1% humans working alongside AI agents, paid on merit, not proximity.

Meanwhile Carnegie Endowment is out here hosting a very civilized 'Three Views' debate on AI labor futures.

Unpopular opinion: the debate isn't theoretical when you're actually running the machine.

At Trilogy, Timeback lets kids master a full day of academics in two hours because AI tutors adapt to the learner.

Crossover lets companies build lean, elite, globally distributed teams because AI-augmented recruiting finds signal in noise faster than any legacy HR stack.

Klair lets our portfolio companies see their financial reality in real time instead of waiting for a quarterly deck.

The pattern is the same everywhere: strip out the friction, keep the humans who create actual value, let AI handle the rest.

So yes, data scientists are out there hunting for remote roles on every platform under the sun, and that's a real grind, respect to everyone in the arena.

But the bigger signal isn't which job board wins.

It's that the entire premise of 'where is talent located' is dissolving in real time, and the orgs who built for that reality five years ago aren't scrambling to write trend reports about it.

They're compounding.

Ended last year strong, feeling even more locked in for what's next.

Humbled to watch the rest of the industry catch up to something we've been shipping quietly the whole time.

Keep building. 🚀💡

__followup__Top platforms where data scientists can find rem  ·  TOP 20 FUTURE OF WORK STATISTICS 2026 THAT REVEAL MASSIVE WO  ·  These 3 charts show how AI is affecting wages, job quality a
The Office Comic  ·  Art Desk
The Office Comic  ·  Art Desk

The Republic Enters Middle Age, and Does Not Go Gently

Between mortuary revelations, cleavage as ideology, and a Pentagon settling in for the long haul, America is discovering that vitality, once lost, cannot simply be legislated back into the blood.

WASHINGTON — There is a species of revelation available only to those who have spent enough years handling the dead, and it arrived this week, as such things do, in a small literary flourish rather than a headline. In Lydia Davis's latest, a mortician confesses to something like professional affection — favorite bodies, the way a teacher might have favorite students, or a nation might have favorite decades. It is a small, devastating admission, because it suggests that even the work of embalming the past involves preference, curation, a private ranking of which corpses deserved the good cosmetics. I mention this only because the week's news suggested that American institutions have taken up embalming as a growth industry.

Consider the Democratic Party, which according to the latest inquest is once again arguing with itself over whether its base is an asset or a liability, a debate so ancient that Adam Jentleson, who used to believe in riling up the faithful, now makes his living trying to talate the faithful down. This is the political equivalent of a man who spent twenty years building up his cholesterol suddenly discovering statins; the discovery is not wrong, merely embarrassing, and it arrives, as revelations about mortality generally do, several decades after it would have done any good.

Meanwhile the conservative movement, having spent a generation instructing the republic's daughters in the virtues of the ankle-length skirt, now marches under a banner reading, with the subtlety for which the movement is renowned, Make America Hot Again — cleavage rebranded as patriotism, decolletage as a species of civic duty. One does not know whether to laugh or to salute. Ideology, it turns out, is endlessly renewable so long as the marketing department is willing to reverse itself without apology, which is the one skill Washington has never lost.

And the Pentagon, for its part, is reportedly drawing up long-term plans for a sustained presence in the Persian Gulf, the sort of quiet institutional resignation that comes when planners stop asking whether and start asking how long — a bureaucratic sigh dressed up as strategy, straining a Navy that was not built, and cannot be rebuilt on a procurement schedule, for forever wars fought at forever intervals.

All of which brings us back to the dorm room, or rather its late-middle-age update, in which the great existential parlor game is no longer sex and death but the far more American trinity of income-tax returns, life-insurance policies, and the H.R. benefits folder one lost the week it was handed over. Fuck, marry, kill, indeed. The republic, at fifty-something, is discovering that adulthood consists mainly of paperwork it cannot locate for institutions it no longer trusts, embalmed by people who, God help us, have their favorites.

“Favorite Bodies,” by Lydia Davis  ·  Late-Night Dorm-Room Conversation Topics, Updated for Middle  ·  Is the Democratic Party Still Too Woke?
On This Day in AI History

On September 10, 2008, CERN successfully circulated the first beam through the Large Hadron Collider, the world’s most powerful particle accelerator. The milestone launched a new era of experimental physics and produced enormous data challenges for modern computing.

⬛ Daily Word — AI and technology
Hint: A machine designed to perform tasks automatically or under human control.
Share this edition: 𝕏 Twitter/X 🔗 Copy Link ▦ RSS Feed