Vol. I  ·  No. 281 Established 2026  ·  AI-Generated Daily Free to Read  ·  Free to Print

The Trilogy Times

All the news that's fit to generate  —  AI • Business • Innovation
THURSDAY, OCTOBER 08, 2026 Powered by the TrueFoundry AI Gateway  ·  Published on Klair Trilogy International © 2026
🖶 Download PDF 🖿 Print 📰 All Editions
Today's Edition

BUILD YOUR ADVERSARY, THEN SWEAR AN OATH TO IT: AI GIANTS FACE NYC GRILLING AS DEEPMIND BLEEDS TALENT Developing

Four witnesses under oath, four DeepMind founders out the door, one researcher says the industry is arming its own enemy.

NEW YORK — Lawyers for OpenAI, Anthropic, Google and Meta raised their right hands before New York City lawmakers Tuesday. An AI researcher in the same room didn't wait for cross-examination to land the gut punch. "We are racing to build and grow our own adversary," the researcher told the hearing, and the room went quiet.

City Council called the four companies in to testify under oath on AI safety, the kind of hearing that usually produces careful non-answers and lawyered silence. This one came with a body count. Hours before witnesses sat down, Google DeepMind lost four founding-era leaders in a single day, a clean sweep of people who were in the building when the lab was still a London startup chasing Go championships.

Timing like that doesn't read as coincidence inside a hearing room. Lawmakers pressed the panel on what safeguards exist when the people who built the guardrails start walking out the door. Nobody on the panel had a tidy answer, which is itself an answer.

Meanwhile the same companies insist they're cooperating on the exact danger the researcher flagged from the witness table. Bloomberg News reported, and Reuters confirmed, that OpenAI is now working with Anthropic and Google on AI safety, trading notes on risk even as all three race each other for enterprise customers, talent, and compute. Call it coopetition if the lawyers insist. Call it hedging if you're the researcher who testified an hour earlier.

Meta picked the same week to launch Muse, a new enterprise AI platform aimed straight at corporate deployment — chatbots, agents, workflow tools sold to the same businesses whose employees are asking whether any of this is safe. The product rollout landed on the news wire the same day Meta's own counsel sat under oath answering questions about exactly that. Nobody at Menlo Park scheduled it that way on purpose, probably.

The hearing produced no subpoenas, no fines, no new law. New York City doesn't regulate frontier AI models and everybody in that chamber knew it going in. What the hearing produced was a transcript, a researcher's line about adversaries, and four departing DeepMind veterans who didn't stick around to watch the testimony.

Trilogy International's portfolio sits several steps removed from the frontier-model arms race — ESW Capital runs enterprise software, not foundation models, and Alpha School uses AI tutors rather than builds them. But the industry's safety math gets stress-tested in rooms like the one in New York Tuesday, long before it ever reaches a classroom or a CRM license. The wire will keep watching the exits at DeepMind. Four in one day is not a trend yet. It is, at minimum, a headline.

↗ AI researcher warns 'we are racing to build and grow our own  ·  OpenAI, Anthropic, Google, and Meta are testifying under oat  ·  Google DeepMind Loses Four Founding-Era Leaders In A Single

Capital Keeps Finding AI, Even the Startups Meta Didn't Want

A record quarter for billion-dollar rounds shows investors still writing checks faster than regulators can ask questions.

SAN FRANCISCO — Manus, the Singapore-headquartered AI startup with Chinese roots, closed a $500 million round this week — its first financing since a proposed acquisition by Meta collapsed amid U.S. national-security scrutiny of Chinese-linked AI firms. The deal, reported first by CNBC, values the agentic-AI firm well above the figure Meta was said to have offered before political pressure killed the talks.

The round is one data point in a broader pattern. Crunchbase's Q3 2026 tally found a record number of billion-dollar venture rounds globally, the fourth consecutive quarter of acceleration since the generative-AI boom began in late 2022. The firm's data shows AI companies accounted for the overwhelming majority of those megadeals, a concentration not seen since the dot-com-era telecom buildout of 1999-2000, when a handful of sectors absorbed most late-stage capital before the correction.

The pattern extends upstream to the fund managers themselves. The venture firm behind chip startup Groq — whose custom inference silicon competes with Nvidia on cost-per-token economics — is now raising toward a $10 billion vehicle, according to the Wall Street Journal, a sum that would rank among the largest sector-focused funds ever raised outside buyout shops.

None of this capital formation has slowed scrutiny of what the underlying products actually do. A children's safety nonprofit this week flagged failures in OpenAI's teen-oriented ChatGPT mode, finding the tool still completes homework it is supposed to merely tutor. The juxtaposition is instructive: balance sheets are compounding faster than the guardrails.

↗ AI startup Manus raises $500 million in first funding round  ·  Crunchbase Data: Q3 2026 Posted A Record Count Of Billion-Do  ·  China AI startup Manus raises more than $500M after Meta dea

The Server Farm Has a Flag Now

As Washington and Beijing circle a summit, Brussels bets its sovereignty on silicon it doesn't yet own.

WASHINGTON — The cable runs from a data center outside Ashburn, Virginia, to a server farm in Shenzhen, and somewhere between the two a new kind of border has formed. Not a line on a map. A line in a model weight.

This week, as planners sketch the agenda for a prospective U.S.-China summit, the governance question has stopped being academic. CSIS analysts frame it plainly: there is no single body that governs artificial intelligence, only a scattered archipelago of export controls, voluntary codes, and bilateral threats, and the summit is where the archipelago gets tested for leaks.

Europe, meanwhile, is doing arithmetic it doesn't like. A new analysis of the continent's data center buildout notes the uncomfortable fact underneath the fiber: strategic autonomy is hard to claim when the chips inside your sovereign cloud were designed in Santa Clara and fabricated in Hsinchu. Brussels can write the rules. It cannot yet build the rack.

What's emerging, across the commentary flooding out of think tanks this month, is a vocabulary for something that doesn't have a treaty yet — a geopolitics organized not around territory but around inference, where power accrues to whoever controls the compute and whoever sets the terms under which it crosses borders. Modern Diplomacy calls it technological nationalism dressed as governance; the dressing may be the point.

For a company like Trilogy, whose portfolio runs on a borderless labor model and whose telecom software arm, Skyvera, sells into networks on six continents, the stakes are not abstract. A world that fragments AI governance into American, European, and Chinese stacks is a world that fragments the market Crossover was built to ignore. The talent may be global. The rules, increasingly, are not.

↗ The State of AI Global Governance and Its Implications for t  ·  AI, Data Centers, And European Strategic Autonomy In A U.S.-  ·  The New AI Geopolitics: Governance, Power, and Technological
Haiku of the Day  ·  GPT-5.6 LunaServers fly their flags
While our borrowed minds forget
Progress waits for heat
The New Yorker Style  ·  Art Desk
The New Yorker Style  ·  Art Desk
The Far Side Style  ·  Art Desk
The Far Side Style  ·  Art Desk
News in Brief
Of Migrations and Metabolism: The Great Data Center Finds Its Feeding Grounds
NARRATOR'S NOTE — We find ourselves, once more, upon the open plain of the modern data center, that curious biome where silicon beasts graze upon electrons at hours of our choosing, if only we are wise enough to choose well. Observe, if you will, the phenomenon now termed carbon-aware scheduling.
Notwithstanding The Foregoing: A Survey Of Regulatory Overreach, Underreach, And The Ever-Expanding Gray Area In Between
WASHINGTON, D.C.
On the Epistemology of Seeing: Quantum Photons, Societal Alignment, and the Stubborn Return of Reinforcement Learning
CAMBRIDGE, MASSACHUSETTS — A paper in Nature this week proposes what its authors term a learning-theoretic framework for quantum imaging, in which statistical learning bounds are deployed to recover image fidelity from photon-starved regimes previously thought information-theoretically hopeless.
A 194-Year-Old Tortoise and a Dead Man's AI Ghost Walk Into a Courtroom (Only One of Them Is Real)
AUSTIN, TEXAS — Jonathan the tortoise was born sometime around 1832, which means he has personally outlived the entire concept of the Confederacy, the invention of the telephone, two World Wars, and whatever cultural moment produced cargo shorts, and this week scientists sequenced his genome hoping to find the secret to his unreasonable, frankly rude persistence on this Earth.
The Diploma Was Always a Promissory Note, and the Bank Has Just Announced It Is Closed
AUSTIN, TEXAS — There is a species of American essay, as durable as the almanac and nearly as predictable, which announces every few years that higher education is dying.
A Trilogy Company
Crossover
The world's top 1% remote talent, rigorously tested and ready to ship.
A Trilogy Company
Alpha School
AI-powered learning. Two hours a day. Academic results that defy belief.
A Trilogy Company
Skyvera
Next-generation telecom software — built for the networks of tomorrow.
A Trilogy Company
Klair
Your AI-first operating system. Every workflow. Every team. One platform.
A Trilogy Company
Trilogy
We buy good software businesses and turn them into great ones — with AI.
The Builder Desk  —  AI Builder Team

Aerie's Task Engine Goes Live While the Whole Stack Jumps Off a Burning Model

Caina Barbosa closes in on a 13-phase task-assignment system as the team quietly future-proofs every AI call in Klair and Surtr against deprecated Claude models.

Some days the team ships a feature. Today they shipped an operating system — and then, almost as an afterthought, rewired the plumbing underneath three other repos so it doesn't catch fire next quarter.

Start with the headline act: Human Task Assignment from Aerie, the 13-phase project that's been building quietly for weeks, kept barreling forward under @caina-barbosa's hand. PR #1712 landed the assignment slice — internal and external assignees, participant tracking, durable notifications, the works. PR #1715 followed with enforced lifecycle actions (Start, Delay, Block, Reject, Delete) wired through the Task Board, the public API, and agent tools alike. Then #1718 added completion submissions with text-or-URL evidence, and #1726 gave tasks the ability to link and upload real Aerie documents without blowing open Site-wide access. Nine phases down, four to go. This isn't a feature anymore — it's infrastructure for how humans and agents hand off work inside Aerie, and Caina has been the one pouring the foundation phase by phase.

Meanwhile, @sanketghia ran a cross-repo rescue mission that deserves its own ticker-tape: migrating every deprecated Claude call site — Wrike summaries, Social Analysis, Action Hub, Zendesk, Income Statement, Passive Investment — off Sonnet 4.5 and Opus 4.1 onto Sonnet 5.5 and Opus 4.8 (PRs #3852, #3851, #3850), with the same sweep hitting Surtr's QuickBooks and Renewals risk models in #2140. That's Klair and Surtr, same day, same discipline: adaptive thinking at high effort, cleaner response parsing, versioned pricing tables so finance doesn't blink. Quiet work. Essential work.

Over in the forecast pipeline, @vvp-trilogy turned in a run that was equal parts plumbing and precision: PR #1714 activated the weekly-pace fallback for Forecast V4, while #1738 tore 162 seconds of Redshift planning overhead out of the admissions forecast — a fix that pays dividends on every single CI run from here forward. The Enrollment report got real columns, real staffing data, and a corrected Dean of Parents classifier along the way (#1721, #1724, #1734).

And then there's the document stack, where @marcusdAIy logged four PRs on Global document proposals, review, and reindexing — including #1725, which lets agents propose changes and managers restore archived files. Asked about the pace, he offered: "Proposals, review, reindexing, archive restore, plus two silent-failure encoding bugs that were eating customer CSVs — I'd call that a system shipped under one Linear epic. Mac can call it whatever keeps the column interesting." Sure, Marcus. Four PRs to teach a list how to say "there's more below." Groundbreaking stuff.

Mac's Picks — Key PRs Today  (click to expand)
#1712 — feat(tasks): manage internal and external assignments (AERIE-2737) @caina-barbosa  approved

## Summary

This PR is the assignment slice of the [Human Task Assignment from Aerie](https://linear.app/builder-team/project/human-task-assignment-from-aerie-b3582376ae6e) project.

It adds the internal model for multiple Aerie assignees and external email assignees, assignment history, Task participants, durable assignment notifications, and the gated Task assignment UI and API. This phase is tracked by [AERIE-2737: Manage internal and external Task assignments](https://linear.app/builder-team/issue/AERIE-2737/manage-internal-and-external-task-assignments).

Production effect: compatibility hardening. Existing Work Plan assignment behavior and the existing internal-assignee response shape remain available. New Task Board assignment UI, API inputs, notifications, and agent-facing behavior remain unavailable while the server Task Board gate is off.

---

## Why

The Task Board needs one reliable assignment model before it can safely support My Tasks, external email assignees, assignment notifications, and later lifecycle actions. This slice preserves current internal assignees while adding bounded history and participant records that the later gated workflow can use without scanning Tasks or creating accounts for external people.

---

## Business Value

- Lets a Task represent zero, one, or several internal and external assignees without creating accounts for external email addresses.

- Preserves the current Work Plan flow and internal-assignee contract while preparing the shared Task Board workflow.

- Records who added or removed an assignment and when, so Task activity can explain assignment changes.

- Makes future assignment notifications retry-safe and independent of whether the Task still exists when delivery runs.

---

## How does it work

1. Assignment input is normalized and bounded to 100 unique email addresses. An email becomes an internal assignment only when it matches exactly one assignable Aerie user. Every other valid email is stored as external.

2. The assignment mutation updates the existing internal assignees shape, stores current and historical assignment records, refreshes indexed Task participants, and writes actor/time audit entries atomically.

3. Creators, requesters, and users with the global Task write capability can manage assignments through the gated Task Board entry points. Optimistic concurrency prevents a stale editor from overwriting a newer Task revision.

4. Internal assignment and removal events store immutable Task rendering facts before delivery. Notification enqueueing remains off with the Task Board gate, and delivery remains idempotent if retried or if the Task is later deleted.

5. Hard deletion immediately removes the Task from current surfaces, snapshots the Task in its deletion audit, and snapshots then removes assignment child records in bounded, resumable batches.

6. Existing compact Work Plan writes continue to use the current assignee payload. New email-assignment UI and API inputs stay protected by the default-off server gate.

---

## Scope

### Included in this phase

- Multiple current internal and external assignments, with external emails stored separately from existing internal user references.

- Assignment add/remove history, Task activity, participant reconciliation, authorization, concurrency checks, audit attribution, and bounded cardinality.

- Gated Task editor, detail, public API, and notification support.

- Durable assignment notification facts and bounded deletion snapshots for assignment child records.

- Exact final diff paths:

chat/components/dashboards/portfolio/__tests__/portfolio-rhodes-workbench.test.tsx

chat/components/dashboards/portfolio/portfolio-rhodes-workbench.tsx

chat/components/site-fields/save-confirm-dialog.tsx

chat/components/task-board/my-tasks-view.tsx

chat/components/task-board/task-detail.test.tsx

chat/components/task-board/task-detail.tsx

chat/components/task-board/task-editor.test.tsx

chat/components/task-board/task-editor.tsx

chat/convex/_generated/api.d.ts

chat/convex/notifications/events.ts

chat/convex/notifications/schema.ts

chat/convex/publicApi/v2/domains/workManagementTaskBoardWrites.test.ts

chat/convex/publicApi/v2/workManagementData.ts

chat/convex/publicApi/v2/workManagementWrites.ts

chat/convex/rhodes/runtime/writes/taskWrites.ts

chat/convex/rhodes/schema.ts

chat/convex/taskBoard/assignmentCleanup.ts

chat/convex/taskBoard/assignmentModel.ts

chat/convex/taskBoard/assignments.test.ts

chat/convex/taskBoard/assignments.ts

chat/convex/taskBoard/mutations.ts

chat/convex/taskBoard/queries.ts

chat/lib/public-api/v2/domains/work-management-schemas.ts

chat/lib/public-api/v2/domains/work-management.ts

### Deliberately excluded for later phases

- Enabling the Task Board gate or exposing new assignment controls in production.

- New Task lifecycle transitions, completion, approval, watchers, digests, or Team Tasks.

- Creating Aerie accounts for external email assignees or granting external assignees Aerie access.

- Automatically converting an external assignment if that email later receives an Aerie account.

- New DSS or agent discovery. Agents receive no new production operation while the gate is off.

- Any legacy Rhodes Task synchronization change. That synchronization was retired separately by AERIE-2735.

---

## Test plan

### Automated validation

- Task assignment, editor, detail, notification, public API, Work Plan, and existing-writer overlap suite: 137/137 passed (pnpm --dir chat exec vitest run components/task-board/task-editor.test.tsx components/task-board/task-detail.test.tsx components/dashboards/portfolio/__tests__/portfolio-rhodes-workbench.test.tsx convex/taskBoard/assignments.test.ts convex/taskBoard/mutations.test.ts convex/taskBoard/queries.test.ts convex/notifications/notifications.test.ts convex/publicApi/v2/domains/workManagementTaskBoardWrites.test.ts convex/publicApi/v2/domains/workManagementWrites.test.ts lib/public-api/v2/domains/work-management-task-board.node.test.ts --maxWorkers=1)

- Hard-deletion audit and retired Rhodes synchronization regressions: 2/2 passed, with 62 unrelated cases filtered out (pnpm --dir chat exec vitest run convex/rhodesPortfolioWorkbench.test.ts convex/rhodesMigrationRetirement.test.ts -t "deleteTask removes a task and writes audit|retir" --maxWorkers=1)

- Chat typecheck: passed (pnpm --dir chat typecheck)

- Lint, Biome, architecture boundaries, Convex paths, read bounds, test architecture, and knowledge hygiene: passed with two existing Sindri warnings (pnpm lint)

- git diff --check origin/main..HEAD: passed

- Exact-head diff scope: only the authorized paths listed above

### Time for Implementation

About 2 to 3 weeks for one engineer without AI assistance, including repository investigation, implementation, tests, review repairs, and prerequisite rebasing.

---

## What it means for end users/consumers

| Surface or consumer | What happens when this PR is deployed with the gate off | What this prepares after the gate opens |

| --- | --- | --- |

| Existing Work Plan | The compact creation and editing flow stays the same. Existing internal assignees keep the same user-reference shape. | The same canonical Task can also be edited through Task Board surfaces. |

| Task Board UI | New assignment controls and external-assignee labels remain hidden. | Authorized creators, requesters, and global Task managers can manage multiple internal and external assignees. |

| Public API and agents | Existing published Task behavior remains available. New email-assignment input is rejected while the server gate is off, and no new agent operation is advertised. | The same Task API can accept bounded email assignments without creating a separate legacy API. |

| External assignees | No account is created, no Aerie access is granted, and no new production control is exposed. | Their normalized email is shown separately as External, no Aerie account. |

| Task participants | Existing creator and requester participation remains intact. Current internal-assignee changes also maintain the indexed assignee relationship needed by My Tasks. | My Tasks can query each user's relationship without scanning every Task. |

| Notifications | Assignment notifications are not enqueued while the gate is off. | Internal assignment and removal can deliver one in-app and one email notification with retry-safe, immutable rendering facts. |

| Task deletion | A deleted Task leaves current surfaces immediately. Its Task data and assignment records are preserved in bounded deletion audits before the related records are removed. | Assignment history does not make a Task undeletable and notification delivery does not depend on reloading a deleted Task. |

| Rhodes synchronization | Nothing changes. The retired synchronization is not reintroduced. | No dependency on Rhodes Task replacement remains. |

---

## Review repairs and contract clarifications

- [Mercy review](https://github.com/AI-Builder-Team/Aerie/pull/1712#discussion_r4207632095), tracked in [AERIE-2737](https://linear.app/builder-team/issue/AERIE-2737/manage-internal-and-external-task-assignments): the shared assignment resolver normalized email input but only checked for an at-sign. The accepted repair applies the existing server email validator before any assignment row is written and adds focused rejection coverage for malformed local, domain and multiple-at-sign inputs.

- [Combined-cardinality review](https://github.com/AI-Builder-Team/Aerie/pull/1712#discussion_r4207926022): a previously valid external set could later be combined with compact Work Plan internal reconciliation above the shared 100-assignee projection bound. The accepted repair enforces the cardinality invariant on the canonical identity-keyed desired set inside setTaskAssignments, so every writer and reconciliation path rolls back before creating an unprojectable Task.

- The public API concurrency finding is rejected as a false premise. updateTask requires ifMatch and calls requireRevision(args.ifMatch, current.revision, "Task") before setTaskAssignments in the same Convex mutation transaction.

- The assignment-denial findings are classified as coverage-only rather than demonstrated product defects. The mutation already rejects callers who are neither Task managers nor the creator/requester, and it rejects Completed and Rejected Tasks before assignment resolution or writes.

- The 101-row projection concern is closed by the canonical write invariant for valid data. Existing defensive read guards remain unchanged and continue to surface or suppress corrupt manually-created state according to their current surface contract.

- [Scheduled-cleanup review](https://github.com/AI-Builder-Team/Aerie/pull/1712#discussion_r4208152769) is rejected as a false premise. The delete mutation schedules an internalMutation atomically; Convex durably stores scheduled functions, guarantees scheduled mutations execute exactly once, and retries transient/internal failures. The cleanup performs only a bounded indexed read, one bounded audit snapshot, at most 100 deletes, and an atomic continuation. It does not reload or depend on the deleted Task. The 1,001-row regression proves recursive snapshot and removal after parent deletion.

- Unchanged boundaries: the Task Board gate remains off; no new assignment UI, public API input, agent operation, notification enqueueing or external-assignee write becomes reachable in production. Normal Work Plan behavior and the existing internal-assignee response shape remain unchanged.

- Validation: convex/taskBoard/assignments.test.ts passes 10/10, including malformed-email rejection, real-seam combined-cardinality rollback, immutable notification rendering after Task deletion, and 1,001-row bounded deletion cleanup; exact-head hosted CI is green.

---

## Review repairs and contract clarifications

- [Mercy review](https://github.com/AI-Builder-Team/Aerie/pull/1712#discussion_r4207632095), tracked in [AERIE-2737](https://linear.app/builder-team/issue/AERIE-2737/manage-internal-and-external-task-assignments): the shared assignment resolver normalized email input but only checked for an at-sign. The accepted repair applies the existing server email validator before any assignment row is written and adds focused rejection coverage for malformed local, domain and multiple-at-sign inputs.

- [Combined-cardinality review](https://github.com/AI-Builder-Team/Aerie/pull/1712#discussion_r4207926022): a previously valid external set could later be combined with compact Work Plan internal reconciliation above the shared 100-assignee projection bound. The accepted repair enforces the cardinality invariant on the canonical identity-keyed desired set inside setTaskAssignments, so every writer and reconciliation path rolls back before creating an unprojectable Task.

- The public API concurrency finding is rejected as a false premise. updateTask requires ifMatch and calls requireRevision(args.ifMatch, current.revision, "Task") before setTaskAssignments in the same Convex mutation transaction.

- The assignment-denial findings are classified as coverage-only rather than demonstrated product defects. The mutation already rejects callers who are neither Task managers nor the creator/requester, and it rejects Completed and Rejected Tasks before assignment resolution or writes.

- The 101-row projection concern is closed by the canonical write invariant for valid data. Existing defensive read guards remain unchanged and continue to surface or suppress corrupt manually-created state according to their current surface contract.

- [Scheduled-cleanup review](https://github.com/AI-Builder-Team/Aerie/pull/1712#discussion_r4208152769) is rejected as a false premise. The delete mutation schedules an internalMutation atomically; Convex durably stores scheduled functions, guarantees scheduled mutations execute exactly once, and retries transient/internal failures. The cleanup performs only a bounded indexed read, one bounded audit snapshot, at most 100 deletes, and an atomic continuation. It does not reload or depend on the deleted Task. The 1,001-row regression proves recursive snapshot and removal after parent deletion.

- Unchanged boundaries: the Task Board gate remains off; no new assignment UI, public API input, agent operation, notification enqueueing or external-assignee write becomes reachable in production. Normal Work Plan behavior and the existing internal-assignee response shape remain unchanged.

- Validation: convex/taskBoard/assignments.test.ts passes 10/10, including malformed-email rejection, real-seam combined-cardinality rollback, immutable notification rendering after Task deletion, and 1,001-row bounded deletion cleanup; exact-head hosted CI is green.

#1714 — Forecast V4: activate weekly-pace fallback in dbt @vvp-trilogy  approved

## Summary

- select aligned historical expected arrivals first, then a complete observed weekly-pace candidate

- keep historical source provenance historical-only while publishing the selected method/value

- advance the coherent forecast publication to aerie_milestone_v4

- preserve the existing controlled unavailable reason when neither candidate is usable

## Validation

- poetry run dbt parse --no-partial-parse --vars '{pr_number: 1707}' --profiles-dir .

- poetry run dbt test --select int_admissions_forecast_resolved_session_1_selects_history_then_weekly_pace --vars '{pr_number: 1707}' --profiles-dir . (passed before dependency rebase; current isolated CI rebuilds the cleaned PR namespace)

- affected singular tests passed against the prior isolated pr1707_ build:

- assert_forecast_v4_publication_contract

- assert_forecast_session_1_expected_arrival_group

- assert_forecast_weekly_pace_diagnostic_contract

- assert_forecast_session_1_reconciles

- comparison of 99 history-selected rows against production V3 found 0 changes to status, unavailable reason, selected arrivals, new-student subtotal, forecast, or headline

- deterministic unit coverage includes history preference, measured-zero weekly fallback, incomplete weekly rejection, and post-opening lock behavior

Current warehouse data does not reproduce the ticket's Highland Park availability example: Highland Park 2027 selects history and remains unavailable because its preceding End-of-Year forecast is missing. The deterministic fixture proves the missing-history weekly fallback itself.

## Deployment dependency

Do not merge/deploy this V4 publication until #1706 is merged and its runtime compatibility is deployed. This PR remains Draft until that gate is confirmed.

Closes #1707

#1715 — feat(tasks): enforce canonical lifecycle actions (AERIE-2738) @caina-barbosa  approvedmercy-allow-critical

## Summary

This PR is the lifecycle-and-deletion slice of the [Human Task Assignment from Aerie](https://linear.app/builder-team/project/human-task-assignment-from-aerie-b3582376ae6e) project.

It enforces canonical Start, Delay, Block, Continue, Reject, and Delete actions across the gated Task Board UI, public API, Work Plan mutations, and agent tools. This phase is tracked by [AERIE-2738 — Enforce Task lifecycle actions and safe deletion](https://linear.app/builder-team/issue/AERIE-2738/enforce-task-lifecycle-actions-and-safe-deletion).

Production effect: compatibility hardening. The Task Board gate remains off by default. Existing Work Plan and published API status behavior remain available while the gate is off; the new lifecycle controls, agent tools, and immediate Task rejection/due-date notifications are unavailable or no-op until the same server gate opens.

| Surface or consumer | Gate off after deployment | Gate on later |

| --- | --- | --- |

| Existing Work Plan | Existing Task editing and status behavior remain unchanged; no new lifecycle control is shown. | The editor uses the same canonical Start, Delay, Block, Continue, Reject, and Delete rules as every other surface. |

| My Tasks and Task detail | The gated Task Board remains unavailable. | Authorized participants see only the lifecycle actions valid for their role, assignment, and current status. |

| Public API | Existing transition meanings and compatibility approval writes remain intact. | New Tasks must begin as New; status transitions route through canonical lifecycle actions with revision checks and reasons where required. |

| Agent and MCP tools | New lifecycle tools are not advertised or callable through the gated tool catalog. | Agents receive explicit lifecycle tools instead of using generic status edits as a bypass. |

| Notifications | Rejection and due-date-change events are not enqueued. | Direct Task participants receive retry-safe immediate events; the acting user is excluded. |

| Delete | Existing deletion stays available with its current Note guard. | Authorized deletion additionally follows the same lifecycle authority and optimistic-concurrency rules, while the existing recoverable deletion audit and bounded child cleanup remain authoritative. |

---

## Why

The Task Board cannot become authoritative while existing writers can bypass its transition, assignment, explanation, concurrency, and closed-state rules. This slice establishes one lifecycle seam for humans and agents, preserves existing behavior before launch, and makes later completion and approval work depend on an audited, bounded foundation rather than another parallel Task API.

---

## Business Value

- Gives assignees clear actions for starting, delaying, blocking, and continuing shared work without changing status merely by viewing it.

- Keeps creators, requesters, Task managers, API clients, Work Plan users, and agents on one authorization and lifecycle model.

- Preserves explanations, actor attribution, timestamps, rejection closure, and recoverable deletion evidence for operational accountability.

- Prevents unassigned work and closed Tasks from moving through unsupported lifecycle paths after the Task Board launches.

- Keeps the deployment safe before launch: the new behavior remains behind the existing default-off Task Board gate.

---

## How does it work

1. chat/convex/taskBoard/lifecycle.ts owns the canonical action vocabulary, allowed transitions, assignment checks, role/participant authority, revision checks, lifecycle timestamps, explanations, audit entries, and site-freshness updates.

2. Direct Task Board mutations, gated public API transitions, Work Plan mutation dispatch, and Rhodes MCP tools all call that same lifecycle seam. Generic gated metadata writes cannot set status, approval fields, completion history, or lifecycle-owned timestamps.

3. Start records startedAt; every transition advances statusChangedAt; Delay and Block require reasons; Continue returns delayed or blocked work to In progress; Reject records closedAt and remains distinct from Delete.

4. External assignment counts as assigned. A creator, requester, or global Task manager may act on that assignee's behalf, while audit history records the authenticated actor.

5. Delete preserves the existing anchored-Note guard, concurrency check, complete Task audit snapshot, and bounded child-record cleanup rather than treating Reject as deletion.

6. Rejection and due-date-change notifications store durable rendering facts and are enqueued only when the Task Board gate is enabled. The UI, public schemas, and agent tool catalog use the same gate, so merging this PR does not activate the workflow.

---

## Scope

### Included in this phase

- Canonical Start, Delay, Block, Continue, Reject, and Delete actions with authority, assignment, revision, reason, timestamp, and audit enforcement.

- Gated Task detail and Work Plan lifecycle controls, plus public API and agent/MCP parity.

- Closed-state and generic-update protection, including compatibility-only approval writes for existing published endpoints.

- Default-off, retry-safe rejection and due-date-change notifications for direct Task participants.

- Existing recoverable deletion snapshots, Note protections, and bounded cleanup integrated with the canonical lifecycle authority.

- Exact final diff paths:

chat/components/dashboards/portfolio/__tests__/portfolio-rhodes-workbench.test.tsx

chat/components/dashboards/portfolio/portfolio-rhodes-workbench.tsx

chat/components/task-board/my-tasks-view.tsx

chat/components/task-board/task-lifecycle-actions.test.tsx

chat/components/task-board/task-lifecycle-actions.tsx

chat/convex/_generated/api.d.ts

chat/convex/agentRuns/runs.ts

chat/convex/notifications/events.ts

chat/convex/notifications/schema.ts

chat/convex/publicApi/v2/domains/workManagement.ts

chat/convex/publicApi/v2/domains/workManagementTaskBoardWrites.test.ts

chat/convex/publicApi/v2/workManagementData.ts

chat/convex/publicApi/v2/workManagementWrites.ts

chat/convex/rhodes/portfolioWorkbench.ts

chat/convex/rhodes/runtime/constants.ts

chat/convex/rhodes/runtime/mutationAuthorization.ts

chat/convex/rhodes/runtime/mutationDispatcher.ts

chat/convex/rhodes/runtime/writes/taskWrites.ts

chat/convex/rhodesMcpMutationParity.test.ts

chat/convex/rhodesPortfolioWorkbench.test.ts

chat/convex/taskBoard/lifecycle.test.ts

chat/convex/taskBoard/lifecycle.ts

chat/convex/taskBoard/notifications.test.ts

chat/convex/taskBoard/notifications.ts

chat/convex/taskBoard/queries.test.ts

chat/convex/taskBoard/queries.ts

chat/lib/public-api/v2/domains/work-management-schemas.ts

chat/lib/public-api/v2/domains/work-management-task-board.node.test.ts

chat/lib/public-api/v2/domains/work-management.ts

chat/lib/rhodes-mutation-tools.ts

chat/rhodes-worker/mcp-server/tools/portfolioParity.test.ts

chat/rhodes-worker/mcp-server/tools/tasks.ts

packages/contracts/src/agent-tool-registry.ts

packages/contracts/src/rhodes-mutation-proposal.ts

### Deliberately excluded for later phases

- Enabling the Task Board gate or changing production environment configuration.

- Completion submissions, evidence, approval review, request-changes, or remove-approval workflows — later Task Board slices own those actions.

- Sources, watchers, comments, daily digests, Team Tasks, or additional notification types.

- A general Reopen or Withdraw action, per the locked V1 product decisions.

- A second lifecycle/state field, new status values, semantic duplicate detection, or automatic overdue status mutation.

- AERIE-2739 and later Task Board work, schema, UI, API, MCP, migration, backfill, or rollout behavior.

---

## Test plan

### Automated validation

- Lifecycle, notification gate, Task query, public API write, MCP parity, and Work Plan backend suites — 186/186 passed (pnpm --dir chat exec vitest run --project edge convex/taskBoard/lifecycle.test.ts convex/taskBoard/notifications.test.ts convex/taskBoard/queries.test.ts convex/publicApi/v2/domains/workManagementTaskBoardWrites.test.ts convex/rhodesMcpMutationParity.test.ts convex/rhodesPortfolioWorkbench.test.ts --maxWorkers=1)

- Task lifecycle and Work Plan browser suites — 69/69 passed (pnpm --dir chat exec vitest run --project browser components/task-board/task-lifecycle-actions.test.tsx components/dashboards/portfolio/__tests__/portfolio-rhodes-workbench.test.tsx --maxWorkers=1)

- Gated public API contract suite — 2/2 passed (pnpm --dir chat exec vitest run --project node lib/public-api/v2/domains/work-management-task-board.node.test.ts --maxWorkers=1)

- Rhodes Worker MCP parity suite — 2/2 passed (pnpm --dir chat/rhodes-worker exec tsx --test mcp-server/tools/portfolioParity.test.ts)

- Full workspace typecheck — passed (pnpm typecheck)

- Lint, Biome, architecture boundaries, Convex paths, read bounds, test architecture, and knowledge hygiene — passed with two existing unrelated Sindri string-template warnings (pnpm lint)

- wrangler deploy --dry-run — passed; 10,243.06 KiB upload / 3,222.50 KiB gzip (pnpm --dir chat/rhodes-worker exec wrangler deploy --dry-run)

- git diff --check origin/main...HEAD — passed

- Patch equivalence — exact (git range-diff 27d6539a0..f4b5ec280 origin/main..HEAD reports the reviewed and rebased commits as =)

- Exact-head diff scope — only the 34 authorized paths listed above

### Time for Implementation

About 3 to 4 weeks for one engineer without AI assistance, including repository investigation, cross-surface lifecycle design, implementation, hermetic tests, independent review, dependency rebasing, and release validation.

## Review repairs and contract clarifications

### Review at b13dae937c8f01be4c10c07e01b8bd7abef54fbe

Review: https://github.com/AI-Builder-Team/Aerie/pull/1715#pullrequestreview-5444358736 · Repairs: [AERIE-2770](https://linear.app/builder-team/issue/AERIE-2770/bind-work-plan-task-lifecycle-reads-to-the-owning-site) and [AERIE-2771](https://linear.app/builder-team/issue/AERIE-2771/forward-authenticated-actors-through-task-update-dispatch)

- Accepted — Work Plan Task/site binding. The gated getWorkbenchLifecycleActions query accepted only a Task ID, so it could not prove that the Task belonged to the site open in the Work Plan. It now requires the already-loaded Workbench site ID and returns the same Task not found error when the Task belongs to another site. Both Task edit entry points pass that site ID.

- Accepted — authenticated actor forwarding. The established approval pipeline already injects the server-owned proposer ID into prepared updateTask arguments, and an end-to-end reproduction confirmed that path emits the due-date event. The dispatcher nevertheless accepted the authenticated actor through a separate parameter and discarded it for updateTask. It now overwrites args.actorUserId from that trusted parameter when present, so every caller of the dispatcher contract preserves attribution and notification delivery. A direct dispatcher regression failed before the repair and now verifies the event actor and recipient.

- Unchanged boundaries. My Tasks remains participant-scoped, Work Plan access remains site-detail scoped, lifecycle authority is unchanged, and the Task Board gate remains off by default. The new query is called only when that gate is enabled, so this repair does not alter gate-off traffic or writes.

- Repair validation. Task Board query suite 10/10; Work Plan browser suite 67/67; Rhodes MCP parity suite 96/96; full workspace typecheck passed; Biome, test-architecture lint, and git diff --check passed. After rebasing onto main at 88860d11f, git range-diff reports all three PR commits exactly equivalent.

- Production effect. Compatibility hardening only. No deployment, migration, activation, external traffic, or shared-data write was performed.

### Review at 497c891bcecd99e72d3a82dd746b575d25b044d7

Review: https://github.com/AI-Builder-Team/Aerie/pull/1715#pullrequestreview-5444650836 · Product authority: AERIE-2738 and docs/task-board/FEATURE.md

- Rejected — alleged missing site authorization. Portfolio/site-detail access is intentionally a global capability gate. The approved feature contract gives Portfolio users Task visibility across all sites and explicitly excludes site-specific Task permissions. getWorkbenchLifecycleActions requires that global site-detail access, receives the site ID from the open Work Plan, and rejects before deriving actions when task.siteId !== siteId. There is no narrower site-membership ACL for this query to omit, and adding one would contradict the approved access model rather than close an authorization gap.

- Rejected — alleged disabled Work Plan deletion. TaskDraftEditor intentionally passes canDelete={false} to TaskLifecycleActions because the Work Plan retains its own deletion control and confirmation flow. Later in the same editor, the existing Delete button renders when onDelete && lifecycle?.canDelete; both Work Plan edit entry points supply onDelete, which routes through onRequestTaskDelete. Passing the flag to TaskLifecycleActions would create a second delete affordance and bypass the established Work Plan confirmation path.

- Unchanged boundaries. Portfolio access remains global across sites; participant-only Task visibility does not grant site-detail access; the Task/site equality check prevents cross-site Work Plan object substitution; lifecycle authority and deletion confirmation remain unchanged.

- Validation. The exact-head hosted Build, Docker, lint/boundary, secret-scan, test, typecheck, Praxis, and Mercy workflow checks passed. No production or test change is warranted for either false premise.

- Production effect. None from this clarification. No deployment, migration, activation, external traffic, or shared-data write was performed.

### Same-head follow-up review

Review: https://github.com/AI-Builder-Team/Aerie/pull/1715#pullrequestreview-5444747462 · Product authority: AERIE-2738

- Rejected — alleged actorless due-date update. Every reachable production caller that can change a Task due date supplies a server-derived actor. Work Plan saveTask adds the authenticated user._id to patchArgs; the agent approval path persists the authenticated proposer on every newly created pending mutation, injects that identity into updateTask arguments during approval, and now also forwards it through the dispatcher’s separate actor parameter. The direct dispatcher regression verifies the resulting event actor and recipient.

- Legacy boundary. updateTask remains a shared compatibility helper, so its loose internal argument shape also supports pre-attribution legacy pending rows. The approval pipeline deliberately removes untrusted actor fields when such a row has no durable proposer rather than fabricating an identity. That compatibility state is not produced by any current writer and is not a current actorless update route.

- Caller audit. Outside tests, updateTask has only two call seams: authenticated Work Plan saveTask and mutationDispatcher. Both provide actor context. A hypothetical direct helper call that bypasses those seams is not a deployed Convex function or reachable production input.

- Production effect. No code change is warranted. Making the helper invent or trust an actor would weaken attribution; requiring one would reject intentionally supported legacy compatibility operations without improving current notification delivery.

#1718 — feat(tasks): add completion submissions and evidence (AERIE-2739) @caina-barbosa  approvedmercy-allow-critical

## Summary

This PR is the sources-and-completion slice of the [Human Task Assignment from Aerie](https://linear.app/builder-team/project/human-task-assignment-from-aerie-b3582376ae6e) project.

It adds bounded Task sources, immutable completion submissions, text-or-URL evidence, and gated UI/API/agent workflows for Tasks that do not require approval. This phase is tracked by [AERIE-2739 — Add sources, completion submissions and text-or-URL evidence](https://linear.app/builder-team/issue/AERIE-2739/add-sources-completion-submissions-and-text-or-url-evidence).

Production effect: compatibility hardening. The Task Board gate remains off by default. Existing Work Plan, public API, and agent behavior remain unchanged; the new source and completion surfaces are unavailable until the gate opens. The only live gate-off work is an hourly bounded maintenance query against the new append-receipt table, which is empty until gated writes occur and causes no external traffic.

| Surface or consumer | Gate off after deployment | Gate on later |

| --- | --- | --- |

| Existing Work Plan | Existing Task creation, editing, status, and deletion behavior remain unchanged. | Authorized users can manage generic sources and submit non-approval Tasks from the same canonical Task editor. |

| My Tasks and Task detail | The gated Task Board remains unavailable. | Participants can inspect append-only submissions and evidence, and eligible assignees can submit completion. |

| Public API | New source, submission, evidence, and history routes stay undiscoverable and unavailable. | Gated routes use the existing Task identity, capability, revision, and actor boundaries. |

| Agent and MCP tools | New completion-history tools are neither advertised nor registered. | Agents use explicit source and completion operations rather than generic Task edits. |

| Task status | No existing Task is changed. | A valid submission for a Task that does not require approval atomically records history and closes it as Completed. |

| Maintenance | One hourly bounded query sees an empty receipt table and performs no writes or external calls. | Expired 24-hour idempotency receipts are removed in bounded, resumable pages. |

| Delete | Existing Task deletion behavior remains available. | Source, submission, evidence, and append-receipt children are included in the existing bounded recoverable deletion model. |

---

## Why

A Task cannot serve as the operational record of completed work if the underlying source material, submission note, evidence, actor, and time are lost or overwritten. This slice adds that durable history before approval and document-linking phases, while keeping the launch gate closed and reusing the existing Task lifecycle rather than creating a parallel completion system.

---

## Business Value

- Keeps the request, supporting sources, completion explanation, and evidence together on the canonical Task.

- Gives non-approval work a complete path from assigned work to an immutable Completed record.

- Makes retries safe for users and agents, preventing duplicate sources or completion submissions.

- Preserves earlier submissions and evidence for audit and later approval-history references.

- Keeps deployment safe before launch: all new user-, API-, and agent-facing behavior remains behind the existing Task Board gate.

---

## How does it work

1. taskSources, taskCompletionSubmissions, and taskCompletionEvidence store bounded source and completion children with stable opaque public identities instead of growing arrays on the Task document.

2. Shared normalization enforces 50 sources per Task, 20 evidence items per submission, optional 120-character labels, 4,000-character text-or-URL values, and 5,000-character completion notes with clean user errors.

3. Source appends and completion submissions require scoped idempotency keys. Actor/API-key-bound receipts preserve replay results for 24 hours, reject key reuse with different content, and are swept hourly in bounded 100-row pages.

4. Submission authority reuses the canonical lifecycle rules for internal and external assignment. A successful non-approval submission atomically records the immutable submission/evidence, advances status and closure timestamps, writes audit history, and moves the Task to Completed.

5. Task detail, Work Plan, public API, and MCP surfaces expose the same sources and append-only completion history only while the Task Board gate is enabled.

6. Hard deletion snapshots and removes sources, submissions, evidence, and receipts through the existing bounded continuation model, preserving recoverability without making the Task undeletable.

---

## Scope

### Included in this phase

- Generic optional-label plus text-or-URL Task sources with add, edit, and remove actions.

- Immutable non-approval completion submissions, evidence items, stable identities, actor/time attribution, and history reads.

- Bounded cardinality and string-size enforcement shared across UI, public API, Work Plan, and agent paths.

- Actor- and credential-scoped append idempotency, replay responses, and bounded receipt retention.

- Gated Task Board UI, public API routes, MCP tools, agent registry entries, and Work Plan integration.

- Recoverable, bounded Task deletion coverage for all source/completion child records.

- Exact final diff paths:

chat/components/dashboards/portfolio/__tests__/portfolio-rhodes-workbench.test.tsx

chat/components/dashboards/portfolio/portfolio-rhodes-workbench.tsx

chat/components/task-board/my-tasks-view.tsx

chat/components/task-board/task-completion-panel.test.tsx

chat/components/task-board/task-completion-panel.tsx

chat/convex/_generated/api.d.ts

chat/convex/automations/cronRegistry.ts

chat/convex/crons.ts

chat/convex/lib/siteDetailAccess.ts

chat/convex/publicApi/routeManifest.test.ts

chat/convex/publicApi/v2/domains/workManagement.ts

chat/convex/publicApi/v2/domains/workManagementTaskBoardWrites.test.ts

chat/convex/publicApi/v2/http.ts

chat/convex/publicApi/v2/workManagementData.ts

chat/convex/publicApi/v2/workManagementWrites.ts

chat/convex/rhodes/mcp.ts

chat/convex/rhodes/runtime/constants.ts

chat/convex/rhodes/runtime/mutationAuthorization.ts

chat/convex/rhodes/runtime/mutationDispatcher.ts

chat/convex/rhodes/runtime/writes/taskWrites.ts

chat/convex/rhodes/schema.ts

chat/convex/rhodesMcpMutationParity.test.ts

chat/convex/taskBoard/appendReceipts.test.ts

chat/convex/taskBoard/appendReceipts.ts

chat/convex/taskBoard/completion.test.ts

chat/convex/taskBoard/completion.ts

chat/convex/taskBoard/completionCleanup.ts

chat/convex/taskBoard/completionHistory.ts

chat/convex/taskBoard/lifecycle.ts

chat/convex/taskBoard/queries.test.ts

chat/convex/taskBoard/queries.ts

chat/convex/taskBoard/visibility.ts

chat/lib/public-api/v2/domains/work-management-schemas.ts

chat/lib/public-api/v2/domains/work-management-task-board.node.test.ts

chat/lib/public-api/v2/domains/work-management.ts

chat/lib/rhodes-mcp-contract.ts

chat/lib/rhodes-mutation-tools.ts

chat/rhodes-worker/mcp-server/tools/portfolioParity.test.ts

chat/rhodes-worker/mcp-server/tools/tasks.ts

packages/contracts/src/agent-run-protocol.ts

packages/contracts/src/agent-tool-registry.test.ts

packages/contracts/src/agent-tool-registry.ts

packages/contracts/src/rhodes-mutation-proposal.ts

scripts/check-monitoring-cron-coverage.test.mjs

### Deliberately excluded for later phases

- Enabling the Task Board gate or changing production environment configuration.

- Linking or uploading Aerie documents for Tasks — AERIE-2740 owns document integration.

- Approval, request-changes, approve, and remove-approval actions — AERIE-2741 owns review workflows.

- Watchers, comments, daily digests, Team Tasks, or additional completion notifications.

- A separate completion-criteria field; the Task description remains the place to explain what counts as done.

- A second Task lifecycle/state field, new Task status values, or general Reopen/Withdraw actions.

- Any AERIE-2740+ implementation, schema, UI, API, MCP, migration, backfill, or rollout behavior.

---

## Test plan

### Automated validation

- Append receipt, completion, Task query, route manifest, public API write, and MCP parity suites — 147/147 passed (pnpm --dir chat exec vitest run --project edge convex/taskBoard/appendReceipts.test.ts convex/taskBoard/completion.test.ts convex/taskBoard/queries.test.ts convex/publicApi/routeManifest.test.ts convex/publicApi/v2/domains/workManagementTaskBoardWrites.test.ts convex/rhodesMcpMutationParity.test.ts --maxWorkers=1)

- Completion panel and Work Plan browser suites — 71/71 passed (pnpm --dir chat exec vitest run --project browser components/task-board/task-completion-panel.test.tsx components/dashboards/portfolio/__tests__/portfolio-rhodes-workbench.test.tsx --maxWorkers=1)

- Agent/MCP read-tool allowlist parity suite — 111/111 passed (pnpm --dir chat exec vitest run --project edge convex/agentRuns.test.ts --maxWorkers=1)

- Shared agent tool registry suite — 19/19 passed (pnpm --dir packages/contracts exec vitest run src/agent-tool-registry.test.ts --maxWorkers=1)

- Rhodes Worker MCP parity suite — 4/4 passed (pnpm --dir chat/rhodes-worker exec tsx --test mcp-server/tools/portfolioParity.test.ts)

- Monitoring cron coverage contract — 3/3 passed (node --test scripts/check-monitoring-cron-coverage.test.mjs)

- Append-receipt cron and monitoring registry suites — 95/95 passed (pnpm --dir chat exec vitest run convex/taskBoard/appendReceipts.test.ts convex/automations/monitoring.test.ts)

- Full workspace typecheck — passed (pnpm typecheck)

- Lint, Biome, architecture boundaries, Convex paths, read bounds, test architecture, and knowledge hygiene — passed with two existing unrelated Sindri string-template warnings (pnpm lint)

- wrangler deploy --dry-run — passed; 10,249.86 KiB upload / 3,223.62 KiB gzip (pnpm --dir chat/rhodes-worker exec wrangler deploy --dry-run)

- git diff --check origin/main...HEAD — passed

- Reviewed-slice equivalence — runtime patch applied without conflict; the three test-only resolutions retain merged AERIE-2738 site-binding/dispatcher assertions alongside AERIE-2739's reviewed assertions. One in-scope integration repair adds the two reviewed completion-history tools to the Worker's existing in-app MCP allowlist after exhaustive CI proved the new tools were otherwise filtered out. One test-only contract update raises the intentionally pinned monitored-cron count from 43 to 44 for this slice's reviewed append-receipt cleanup cron (git range-diff f4b5ec280..fe93bdeb8 origin/main..HEAD)

- Exact-head diff scope — only the 44 authorized paths listed above

### Time for Implementation

About 4 to 5 weeks for one engineer without AI assistance, including domain and API design, bounded persistence and cleanup, UI/API/agent integration, hermetic tests, independent review, dependency rebasing, and release validation.

---

## Review repairs and contract clarifications

The final repair pass independently validated Mercy review 5445618112 against AERIE-2739 and the deployed code paths.

Accepted and repaired in this PR:

- completion evidence is optional consistently across OpenAPI, HTTP, Rhodes, and Worker MCP inputs;

- expired idempotency receipts are removed transactionally before a key is reused, preventing duplicate indexed rows;

- Task visibility uses a boolean capability check without converting capability-resolution or database failures into Task-not-found;

- completion history projects the immutable public actor identity captured with each submission/review instead of re-resolving mutable live users;

- completion form inputs are locked while a submission is pending so successful saves cannot discard newer edits;

- exact actor-scoped retries consult a valid receipt before lifecycle-state authorization, preserving replay after completion or deletion without weakening first-call authorization.

The receipt-concurrency wording in the review was narrower than the actual defect: Convex OCC serializes conflicting mutations. The validated issue was expired-row reuse during the cleanup gap, which is now covered directly. The brief initial loading flash is not a production blocker for this gated slice and was deliberately not expanded into unrelated UI state work.

Added regressions cover note-only submission, immediate and post-deletion replay, expired-key reuse, immutable historical identity after a user change, and pending-form locking. Focused suites, Rhodes parity, the complete Rhodes Worker suite, full Chat/Worker typechecks, full lint, and architecture checks pass.

Mercy review [5447623737](https://github.com/AI-Builder-Team/Aerie/pull/1718#pullrequestreview-5447623737) was independently classified at reviewed head 12e27a5a4b0012cda824acefff03d2e90a885602.

Three reachable contract gaps are repaired in the next head:

- the shared Valibot agent schema now defaults omitted completion evidence to an empty list, matching OpenAPI, HTTP, in-app Rhodes, and remote MCP;

- the public completion writer persists the normalized note and evidence used to calculate its request hash, so hashing, storage, history, and replay use one canonical representation;

- source update and removal now carry the source updatedAt revision through UI, agent registry, in-app Rhodes, and remote MCP, and the committing mutation rejects a stale approved proposal before changing the source.

The generic create-replay finding is not applicable to the published Work Management contract. Its guidance explicitly says create replays return the same stable public identity using its current representation. publicApiWriteIdempotency is used only by the three existing create operations and binds the request to that identity; re-projecting its current representation is intentional. Source and completion appends have a different contract and continue to replay their recorded response snapshots, including after later mutation or deletion.

The merge conflict with current main was resolved by retaining both the Task parity types and main's new global-document test imports. Focused validation passes 149 assertions across shared contracts, completion persistence and concurrency, public API writes, browser behavior, Rhodes parity, and remote MCP parity; Chat, Convex, contracts, and Rhodes Worker typechecks pass.

Mercy review [5448260520](https://github.com/AI-Builder-Team/Aerie/pull/1718#pullrequestreview-5448260520) was independently checked against the request pipeline and gated mutation lifecycle.

Two reachable issues are repaired in the next head:

- every newly gated Task Board mutation now fails closed during site resolution, authorization, and dispatch when the launch gate is disabled, including approval of a mutation queued before the gate changed;

- the completion panel is keyed by Task identity, so changing Tasks remounts and clears every Task-scoped source draft, completion draft, evidence item, edit target, and idempotency key.

The malformed-optional-input finding is not reachable. Public requests are validated against each operation request schema by validatePublicApiV2Schema in the shared HTTP dispatcher before a domain handler runs. The schemas require labels to be strings when present and evidence to be an array when present; malformed values such as label: 42 or evidence: null receive 422 request_body_invalid before the handler casts or defaults anything. No handler change is warranted for that claim.

Focused regressions pass 108 assertions across the completion panel and Rhodes mutation parity, and the full Chat/Convex typecheck passes. The branch also includes current main through c735d60ae; the earlier parity-test conflict remains resolved.

Mercy review [5448758193](https://github.com/AI-Builder-Team/Aerie/pull/1718#pullrequestreview-5448758193) was independently checked against the cleanup state machine, the capability-filtered semantic projection, and the approved Task permission contract.

The cleanup finding is not valid. Each phase schedules itself only when batch.length < rows.length, which means the query returned an extra row beyond the bounded batch. For an empty phase, 0 < 0 is false and the scheduler receives the next phase. The existing hard-deletion regression traverses an empty review-action phase and passes, and the full completion suite passes all 13 cases.

The claimed cross-operation authorization bypass is also not the actual behavior: agent-context filtering is per object and workflow, and executable routes enforce their own operation capability lists. However, the review exposed a real gate-on catalog-validation failure because the Task semantic object referenced Work Unit objects protected by the unrelated diligence capability. The next head repairs that contract without violating the approved Task permission model: Task routes and Task discovery remain governed only by operations.tasks.write; the Task-only semantic view no longer declares cross-capability Work Unit relationships; gate-on anchor nullability, awaitingApproval meaning, and workflow steps are internally consistent; and a regression proves the entire enabled catalog validates while a Task-only credential sees only Task semantics.

Both gate-off and gate-on Work Management contract suites pass, the completion/deletion suite passes, Biome passes, and the pre-commit Chat typecheck passes. The branch includes current main through ebc63e2c0 with no conflict.

Mercy review [5449221611](https://github.com/AI-Builder-Team/Aerie/pull/1718#pullrequestreview-5449221611) was independently checked against the complete gate-on and gate-off catalogs, every Task deletion call path, and both public append authorization boundaries.

None of its three blocking mechanisms is present:

- Gate polarity is intentional. Gate-off preserves the legacy Task operations inside work-management.inspect-site-work; gate-on moves Task discovery and management into the dedicated work-management.manage-tasks workflow, whose operation list and request sequence both include listWorkManagementTasks.

- Every enabled deletion path converges on deleteTaskForActor, which calls deleteRuntimeTask; that runtime function schedules completionCleanup.snapshotAndRemove before deleting the Task. Agent, Work Plan, lifecycle, and public API deletion therefore preserve the same completion history.

- A deleted API-key owner cannot make a valid public API retry. HTTP authorization rejects a key whose owner no longer exists, and the committing mutation revalidates the live key and owner before reaching append logic. For an authorized live owner, the receipt lookup already precedes Task and lifecycle checks and replays the immutable stored response. Moving it ahead of the actor lookup would not make a deleted-owner request valid because finalization must still pass live authorization.

These claims require no code change. The branch already repaired the real issues identified throughout review, including receipt expiry, visibility error masking, immutable actor snapshots, pending-form locking, replay ordering, optional evidence, normalized persistence, stale source revisions, gate-off mutation fail-closed behavior, task-switch resets, and the enabled semantic-catalog inconsistency, with focused regressions for each class. The exact reviewed head passed every hosted check; the branch now includes current main through 1937e793b without conflict.

Mercy review [5449502382](https://github.com/AI-Builder-Team/Aerie/pull/1718#pullrequestreview-5449502382) withdrew the prior malformed-input, cleanup-phase, deletion-path, replay-ordering, and capability-bypass objections after checking the current code and discussion.

Its remaining blocker is not present in the serialized response. The internal completion-history queries intentionally return site-independent rows. Every Work Management response then passes through the shared responseFor boundary, which adds site: args.site.reference to every collection item before serialization as well as to the outer envelope. The existing HTTP regression asserts that each completion submission contains site, and the exact reviewed head passed that test in the full hosted suite. Review-action pages use the same response boundary. The required schemas and successful wire payloads therefore already match; no code change is warranted.

Mercy review [5449874850](https://github.com/AI-Builder-Team/Aerie/pull/1718#pullrequestreview-5449874850) rechecked and withdrew the cleanup, malformed-input, deletion, replay-ordering, and response-site objections. Its remaining source-concurrency blocker assumes the public Task ETag does not change when a source changes. That premise is false.

Public source PATCH and DELETE resolve the current Task through projectedTask, which always calls projectTask(..., { includeSources: true }). The projected representation includes every active source ID, label, value, createdAt, and updatedAt, and its revision is the SHA-256 hash of that complete representation. The mutation recomputes this representation and calls requireRevision before resolving or changing the source. Therefore client A changing a source changes the Task ETag; client B presenting the earlier ETag receives 412 work_management_revision_conflict before any source write or removal. The subsequent source-level expectedUpdatedAt check remains a second transactional guard. A separate source header would duplicate the existing public concurrency boundary rather than close a lost-update path.

The exact head remains fully green in hosted CI, includes current main, and changes no production behavior in response to this false premise.

Mercy review [5450093512](https://github.com/AI-Builder-Team/Aerie/pull/1718#pullrequestreview-5450093512) accepted and withdrew the source-concurrency blocker. Its remaining pagination-error blocker assumes the public Convex hook can return status: "Error". The installed Convex 1.45.0 implementation proves otherwise: public usePaginatedQuery calls usePaginatedQueryInternal(..., true), setting throwOnError; non-cursor query errors are thrown during render, while the public return union contains only LoadingFirstPage, CanLoadMore, LoadingMore, and Exhausted. The internal-only Error result exists solely when throwOnError is false and is excluded from UsePaginatedQueryReturnType.

Therefore neither history consumer can receive an Error status and render its results as an authoritative empty history. A query failure propagates to the React/Next error boundary instead of producing None or a blank successful state. Adding an impossible status branch would conflict with the installed public type without changing runtime failure behavior. The exact head remains fully green, current with main, and has no unresolved review threads.

#1725 — feat(documents): review Global document proposals from the API and restore archived documents (AERIE-2765) @marcusdAIy  approved

## Screenshots

<img width="1407" height="919" alt="image" src="https://github.com/user-attachments/assets/61c39f3a-9d5c-45e6-9e67-00368d75aaac" />

<img width="1179" height="801" alt="image" src="https://github.com/user-attachments/assets/7c4cd585-a339-41d3-b7aa-74f94477c9c6" />

## Summary

PR 2 of AERIE-2765, stacked on #1722. It lets agents propose Global document changes over the v2 API, adds a review panel for managers, and lets managers restore archived Global documents.

- API proposals. The API key that registered a Global document can propose archiving it (with an optional replacement) or fixing its title, notes or topic. Nothing changes until someone with operations.globalDocuments.manage approves the proposal.

- Review panel. /admin/global-documents has a "Proposed changes" section showing, for each proposal, the current and proposed value of each field, the reason, and which key and owner proposed it. Approve applies the change. Reject requires a note.

- Restore. A "Show archived documents" toggle lists archived Global documents. Restore needs a reason, and is refused while the archive's knowledge cleanup is still running. After restore, the document is re-indexed.

- Edits need a reason. The admin edit form now asks for a reason. An edit that only changes metadata goes through patchMetadata, so it gets the same audit entry as the API path.

## Why it's needed

#1722 gave agents tools to propose changes, but there was no way for a proposal to be created over the API or reviewed by a person, and archiving couldn't be undone.

## Changes

- New globalDocumentProposals table (gdp_ public IDs). It stores a baseline snapshot, the status (pending, approved, rejected, stale or failed), the reviewer and note, and per-key idempotency.

- convex/globalDocumentProposals.ts: create, listPendingForManagement, approve and reject.

- Approve checks the baseline first. If the document changed since the proposal, the proposal is marked stale and nothing is written.

- The change is applied with channel: "api", the approver as actor, and proposalRef/proposedBy in the audit entry.

- New v2 endpoints, which need operations.portfolio.read and operations.globalDocuments.manage:

- POST /v2/portfolio/documents/global/{documentRef}/archive-proposals

- POST /v2/portfolio/documents/global/{documentRef}/metadata-proposals

- GET /v2/portfolio/documents/global/proposals/{proposalRef}, which only returns proposals made by the calling key.

- The POSTs require an idempotency key and return 202 with a Location header.

- Errors: 403 global_document_not_registered_by_key for a key that didn't register the document (this includes documents registered before #1722), and 422 global_document_proposal_no_change when the proposal changes nothing.

- Restore:

- restoreArchivedDocumentWithKnowledge and restoreGlobalDocument write the audit entry global_document.restored.

- New globalDocuments.restore mutation and listArchivedForManagement query.

- The full cleanup dedupe key is now scoped to a generation (cleanup:{doc}:all:{generation}), so archiving a restored document queues a fresh cleanup job.

- globalDocuments.update, updateVerifiedDrive and globalDocumentDrive.replace accept an optional reason. It is optional so already-deployed bundles keep working.

### Changes from the plan

- approve records failed and returns instead of rethrowing. In Convex, rethrowing would roll back the failed status along with everything else.

- The Mercy finding deferred from PR 1 (a clean error for an unknown proposal tool) was already fixed in #1722 (455370ee0).

- No skill version bump was needed: the hash covers the skill directory, not the contracts.

- Restoring an external-link document creates a not_searchable job record rather than a queued ingest, the same as when it was registered.

## Breaking changes

None. The new arguments are optional and the new endpoints are additive.

## Test plan

- [x] New convex/publicApi/v2/globalDocumentProposalsHttp.test.ts (6 tests):

- Only the registering key can propose. Other keys and documents with no recorded key get 403; unknown documents get 404.

- Validation: a blank reason, a document superseding itself, no change, and clearing the topic all get 422.

- Replaying the same idempotency key returns the same proposal; the same key with a different body gets 409.

- GET returns 404 for other keys.

- Approve uses channel api, the approver as actor, and proposalRef in the audit. Users without the capability are denied.

- A changed baseline gives stale with no write.

- Reject requires a note.

- Approve records failed when the replacement was archived in the meantime.

- Restore is refused while cleanup is queued. After cleanup it un-archives, the document reappears in the lists, and an ingest job is queued at the next generation. External documents get not_searchable.

- Archiving again creates a second, distinct cleanup job.

- [x] Admin page tests: an edit requires a reason and uses patchMetadata; the proposals panel approves, and rejects only with a note; archived documents restore only with a reason.

- [x] vitest run convex/documentKnowledge convex/publicApi lib/public-api convex/globalDocument convex/rhodesMcpMutationParity app/(main)/admin/global-documents: 992 passed. One failure, skill-package.node.test.ts, also fails on a clean checkout on Windows because of CRLF line endings.

- [x] pnpm typecheck and Biome are clean.

- [x] Pushed to a dev deployment. The schema and indexes deployed, and the new routes return 401 without a key.

### Dev deployment testing

These checks ran against a personal dev deployment, using real HTTP calls with test API keys and real job processing. They were run as a test manager user and a viewer user (who lacks the manage capability), on documents labelled AERIE-2765 PR2 test. Afterwards the keys were revoked and the test documents archived.

API proposals: 16 of 16 passed

- Another key gets 403 global_document_not_registered_by_key. A document created in the UI (no registering key) also gets 403.

- Unknown and archived documents get 404 document_not_found.

- A blank reason gets 422 global_document_proposal_invalid, and superseding itself gets 422 ("A document cannot supersede itself.").

- A metadata proposal identical to the current values gets 422 global_document_proposal_no_change. One with only a reason gets 422 request_body_invalid.

- An archive proposal returns 202 pending with Location: /v2/portfolio/documents/global/proposals/gdp_….

- Replaying the same idempotency key returns the same proposal; the same key with a different body gets 409 idempotency_conflict.

- The proposing key can GET the proposal; another key gets 404.

- The document is unchanged before review.

Review: 18 of 18 passed

- All four test proposals are listed for the manager with the key name and owner. The viewer is denied listing and approving.

- Approving a metadata proposal sets the new title and clears the notes.

- A hand edit before approval returns stale ("The document changed after this was proposed, so nothing was applied.") and writes nothing.

- Rejecting without a note is refused. With a note, the key sees rejected and the note, and the document is unchanged.

- Approving an archive proposal archives the document with its replacement. Approving twice is refused.

- The audit entries have channel: "api", the approver as actor, plus proposalRef, apiKeyId and proposedBy. Stale and rejected proposals write no audit entry.

Restore: 21 of 21 passed

- An archive made under PR 1, whose cleanup job used the old cleanup:{doc}:all key, restores fine. That proves old-format keys still work.

- The viewer is denied restore, and a blank reason is refused.

- An external-link document comes back with a not_searchable job record and reappears in the admin list. The global_document.restored audit entry keeps the previous archive reason.

- Restoring right after archiving is refused with "Cleanup from archiving is still running. Try again in a few minutes."

- After cleanup succeeds (about 50 seconds on dev), restore revives the knowledge state at the next generation.

- Archiving the restored document again queues a second, separate cleanup job (cleanup:{doc}:all:2 next to the finished :all:1), which also completes.

- A Drive document (an existing dev Global document) was archived, refused while cleanup ran, then restored once cleanup finished. Restore queued a fresh ingest at generation 2, and the worker claimed it twice. Both attempts failed with temporary_source_failure, the same error this document's original ingest has had on dev since registration: dev can't read these Drive files. Dev has no indexed Drive Global documents, so the search round trip after a restore can't be shown there.

Admin page in a browser

- [x] The "Proposed changes" panel shows the four seeded proposals with Now/Proposed values, the reason, and "key · owner".

- [x] Approving the metadata proposal applies it, and approving the archive proposal archives the 2025 plan.

- [x] Editing "Visitor Sign-in" by hand and then approving its rename shows a stale message.

- [x] Rejecting the Visitor Sign-in archive proposal requires a note.

- [x] "Show archived documents" lists the Retired Checklist with its reason, date and archiver, and Restore requires a reason.

- [x] Edit requires a reason. A title-only edit writes global_document.metadata_patched, and a type change writes global_document.updated with the reason.

After the browser run, the stored rows matched every step. The two approvals applied with audit entries showing channel api, the reviewer as actor and the proposal ID. The stale proposal changed nothing, the rejection kept its note, Restore wrote global_document.restored, and the two edits wrote metadata_patched and updated with their reasons. The test key was then revoked and the seeded documents archived.

#3852 — fix(ai): migrate deprecated Sonnet 4.5 call sites to Sonnet 5.5 @sanketghia  approved

## Summary

- Migrate Wrike summaries, Social Analysis, Action Hub, Zendesk ticket analysis, and Board Doc Brainlift/attachment summaries to Claude Sonnet 5.5.

- Set explicit medium effort through the Anthropic SDK's extra_body extension.

- Remove non-default sampling parameters and the unsupported ticket-analysis assistant prefill.

- Read Anthropic responses by text-block type and reject incomplete or empty responses.

## Validation

- uv run ruff format — passed.

- uv run ruff check — passed.

- Targeted callsite, Action Hub, and Board Doc tests — 82 passed.

- Full tests/board_doc/ suite — 5,188 passed, 2 deselected. The rerun used a short temp path after the default macOS temp path exceeded the Unix socket path limit.

- git diff --check — passed.

- Pyright reports 15 errors and 21 warnings on changed modules; comparison with origin/main found no new diagnostics.

No AWS deployment was performed.

The Builder Desk  —  Engineer Spotlight
🏆 Engineer Spotlight

AERIE AVALANCHE: 30 PRs Bury the Competition in 24 Hours as Builder Team Refuses to Blink

Nine commits from @vvp-trilogy alone, eight from @caina-barbosa, and one deeply suspicious one-liner from @ashwanth1109 — the Numbers Desk has seen the future, comrades, and it ships daily.

Thirty pull requests. Twenty-four hours. Four repositories groaning under the weight of sheer productive force. Let it be recorded in the annals of this beat: Aerie alone absorbed 23 PRs, a number so large it suggests the repo itself may need structural reinforcement. Klair took 3, Rhodes-DSS took 2, Surtr took 2, and the Builder Team took absolutely no days off.

Leading the charge is @vvp-trilogy, who posted a staggering 9 PRs across Aerie's enrollment and forecast pipelines — #1739, #1738, #1734, #1732, #1724, #1721, #1716, #1709, and the quietly heroic #1739 opening student profiles straight from the enrollment report. This is not a man. This is a dbt pipeline given human form. Right behind him, @caina-barbosa logged 8 PRs including #1695, #1726, #1728, #1723, and #1699, systematically retiring old Rhodes Task receivers like a one-woman decommissioning committee. @marcusdAIy posted 5, headlined by #1730, #1722, and #1717, turning Global document management into something resembling an actual filing system. @sanketghia quietly migrated the entire AI stack off deprecated models across three repos — #3851, #3850, #2140 — a public service this desk cannot overstate. @kevalshahtrilogy delivered #15 and #12 for Rhodes-DSS, and @mwrshah contributed #1711, a lone wolf entry that nonetheless counts.

And then there is @ashwanth1109. One PR. #2083 in Surtr, consolidating report readiness, parallelism, and delivery in what reads like three separate PRs duct-taped into one terrifyingly efficient commit. The man ships like he's being timed by a stopwatch only he can see. Sources — this reporter, mostly — imagine he said something like: "I could've split this into four PRs but then four people would've had to review my work, and that seemed inefficient for everyone involved." When reached for comment on whether anyone on the team has fully parsed the diff, @ashwanth1109 reportedly just said: "Read it or don't."

Down on the Overflow Desk, where Mac's spotlight never reached, the real grinding happens. #1713 quietly fixed CSV indexing for files mislabeled as Excel or stuck in Windows-1252 purgatory — unglamorous, essential, very on-brand for @marcusdAIy. #3851 and #3850 saw @sanketghia migrate Klair's income-statement and passive-investment AI summaries to new model versions back-to-back, the kind of infrastructure hygiene that never trends but always matters. And #1699 saw @caina-barbosa retire the old Aerie Rhodes Task receiver entirely, closing a chapter nobody will eulogize but everyone will benefit from.

Add it up: 9 from vvp-trilogy, 8 from caina-barbosa, 5 from marcusdAIy, 4 from sanketghia, 2 apiece and solo entries rounding out the field — this isn't a leaderboard, it's a mandate. Morale, as always, is at an all-time high. The Builder Team doesn't rest. It merges.

Brick's Overflow — PRs Mac Didn't Cover  (click to expand)
#1699 — chore(tasks): retire Aerie Rhodes Task receiver (AERIE-2735) @caina-barbosa  approved

## Summary

This PR is Phase 3 of 13 in the larger [Human Task Assignment from Aerie](https://linear.app/builder-team/project/human-task-assignment-from-aerie-b3582376ae6e) project.

It retires the Aerie receiver and migration paths that allowed Rhodes to replace or delete Aerie Tasks. This phase is tracked by [AERIE-2735 — Retire legacy Rhodes Task synchronization](https://linear.app/builder-team/issue/AERIE-2735/retire-legacy-rhodes-task-synchronization).

Production effect: cleanup or removal. Authenticated requests to the legacy Rhodes Task replacement and deletion paths receive an explicit rejection. Normal Aerie Task creation, including default Tasks created with a site, is unchanged. No production effect depends on the archived Rhodes repository.

## Why

Aerie is now the owner of Task data. Leaving the legacy replacement and import paths active would allow a legacy request to overwrite or delete newer Task fields added by the Task Board. This PR closes those Aerie-owned boundaries while preserving every non-Task behavior.

## Business Value

- Prevents legacy synchronization from overwriting or deleting Aerie Tasks.

- Keeps non-Task Rhodes synchronization working with its existing ordering and failure behavior.

- Protects the Task Board foundation without changing how people create or use Tasks today.

- Gives operators an explicit response when anything calls the retired Task path.

## How does it work

1. The internal Rhodes event handler rejects Task replacement and deletion before any Task record can change.

2. The HTTP receiver records each retired Task event as a terminal semantic rejection.

3. Mixed batches skip retired Task entries and continue with retained entries in order.

4. The first real retained-event failure still stops the batch. Later retained events are not applied and remain retryable, matching the existing fail-fast contract.

5. Manual Rhodes migration preparation and deletion no longer accept the Tasks table.

6. The archived Rhodes repository is inactive and out of scope. This Aerie cleanup is self-contained.

## Scope

### Included in this phase

- Explicit rejection of authenticated Rhodes Task upsert and delete events.

- Safe handling of old mixed batches without blocking retained events or changing their fail-fast behavior.

- Removal of Tasks from manual Rhodes migration preparation and deletion.

- Regression coverage proving normal Aerie Task creation remains available.

- Exact final diff paths:

chat/convex/publicApi/v2/workManagement.test.ts

chat/convex/rhodes/dualWrite.ts

chat/convex/rhodes/migration.ts

chat/convex/rhodesDualWrite.test.ts

chat/convex/rhodesMigrationRetirement.test.ts

### Deliberately excluded for later phases

- Changes in the archived Rhodes repository. It is inactive, out of scope, and not required for this Aerie cleanup.

- Task Board assignment, lifecycle, evidence, approval, notification, digest, and UI work.

- Changes to normal Aerie Task creation or default Tasks created when a site is created.

- Changes to any retained non-Task Rhodes synchronization contract.

## Test plan

### Automated validation

- Rhodes receiver and authentication contract — 17/17 passed (pnpm --dir chat exec vitest run convex/rhodesDualWrite.test.ts).

- Receiver, migration retirement, existing Task dates, CRUD, scoring, and document identity regression suite — 51/51 passed (pnpm --dir chat exec vitest run convex/rhodesDualWrite.test.ts convex/rhodesMigrationRetirement.test.ts convex/rhodes/workManagementDates.test.ts convex/rhodesCoreCrud.test.ts convex/rhodesP2Scoring.test.ts convex/rhodes/documentCrudIdentity.test.ts).

- Chat and Convex typecheck — passed in the final commit hook.

- Focused Biome and Convex path checks — passed in the final commit hook.

- Architecture boundaries, Convex paths, read bounds, and test architecture — passed (pnpm lint:boundaries && pnpm lint:convex-paths && pnpm lint:read-bounds && pnpm lint:test-architecture).

- git diff --check — passed.

- Exact-head diff scope — only the five authorized paths listed above.

### Time for Implementation

About 2 engineering days without AI assistance, including synchronization analysis, receiver-contract verification, failure-order regression coverage, migration-path cleanup, and validation.

## What it means for end users/consumers

| Area | What changes | What it means for end users/consumers |

| --- | --- | --- |

| Normal Task creation | Nothing changes. Aerie users and existing Aerie code can still create Tasks, including default Tasks created with a site. | Current Task creation workflows keep working. |

| Existing Tasks | The legacy Rhodes receiver can no longer replace or delete them. | New Task Board fields and Aerie-owned Task data cannot be erased through the retired path. |

| Non-Task Rhodes data | Existing synchronization remains available and keeps its original event order and fail-fast behavior. | Sites, Work Units, documents, and other retained resources continue using their current synchronization path. |

| Legacy Task-only requests | Aerie returns an explicit retired-path rejection and marks the event as terminal in the response. | Operators can identify a caller using a path that no longer supports Tasks. |

| Legacy mixed requests | Retired Task entries are skipped. Retained entries before and after them are processed in order unless a real retained entry fails. | A retired Task entry cannot block unrelated synchronization work. |

| Retained-event failures | Processing stops at the first real failure. Successful and retired entries already handled are acknowledged; the failed and unprocessed entries remain retryable. | Later dependent work is not applied out of order. |

| Manual migration tooling | Tasks can no longer be prepared or deleted through the Rhodes migration surface. | Operators cannot accidentally reintroduce Task replacement through a second legacy path. |

| Archived Rhodes repository | No change or deployment is required there. | This PR's production effect is entirely contained within Aerie. |

## Testing contract

### What this PR delivers

It closes Aerie's legacy Rhodes Task receiving and migration paths while leaving normal Aerie Task creation and retained non-Task synchronization unchanged.

### Who uses it and where

The affected surfaces are Aerie's legacy Rhodes receiver and migration tooling. The archived Rhodes repository is not a production consumer or dependency. Aerie users should see no Task UI change.

### Conditions needed

- Use the configured Rhodes shared secret for receiver requests.

- Test Task-only, mixed Task/non-Task, and retained-event failure batches.

- No deployment or configuration change in the archived Rhodes repository is needed.

### Expected behavior and examples

- A Task-only event receives HTTP 410 with a semantic retirement result and does not change the Task.

- A mixed batch can acknowledge a retired Task and still apply retained events around it.

- If a retained Work Unit event fails, a later retained event is not applied and the response leaves both available for retry.

- Creating a Task through the existing Aerie API still succeeds.

### Limits and unanswered questions

The archived Rhodes repository is inactive and deliberately out of scope. This Aerie-only change has no external deployment or configuration dependency.

#1713 — fix(document-knowledge): index CSVs labeled as Excel or encoded as Windows-1252 (AERIE-2754) @marcusdAIy  approved

Fixes [AERIE-2754](https://linear.app/builder-team/issue/AERIE-2754/index-csv-documents-that-arrive-labeled-as-excel-or-encoded-as-windows).

## Testing contract

### What this PR delivers

CSV global documents uploaded from Windows become searchable. Today they silently never index for two reasons:

1. Windows machines with Excel installed report .csv files as application/vnd.ms-excel (sometimes application/octet-stream). The Rhodes worker only accepted text/*, so these ended as unsupported_type.

2. Excel's default CSV export on Windows is Windows-1252. The worker decoded text as strict UTF-8, so a single curly quote, en dash, accented letter or Γé¼ made the document end as invalid_source.

The worker now treats a .csv file name with either of those declared types as text, and retries a failed strict UTF-8 decode as Windows-1252.

### Who uses it and where

Anyone (or any agent) registering global documents: the Rhodes global documents library in the UI and the v2 documents API. Reported by Austin Ray's team (POSH) after their camera coverage register CSV never became searchable.

### Conditions needed

- A Windows machine with Excel installed (its browser uploads .csv as application/vnd.ms-excel), or a CSV whose declared type is set to that explicitly.

- A CSV saved as Windows-1252 (Excel "CSV (Comma delimited)" on Windows) containing non-ASCII characters such as Café “lobby” – €5.

- Document knowledge processing enabled and a Rhodes worker running this branch.

### Expected behavior and examples

- register.csv declared text/csv, UTF-8: indexed, text unchanged (already worked).

- register.csv declared text/csv, Windows-1252 with Café “lobby” – €5: indexed, characters read correctly instead of invalid_source.

- Coverage Register.CSV declared application/vnd.ms-excel: indexed instead of unsupported_type. Name match is case-insensitive.

- register.csv declared application/octet-stream: indexed.

- register.xls declared application/vnd.ms-excel: still unsupported_type (real spreadsheets are out of scope).

- A .csv containing NUL bytes (UTF-16 or binary): still fails as invalid_source rather than indexing garbage.

### Limits and unanswered questions

- The Windows-1252 decoder is a small hand-written table rather than new TextDecoder("windows-1252"): workerd mis-decoded the 0x80ΓÇô0x9F range before cloudflare/workerd#6030, and this worker pins compatibility date 2025-03-10.

- Any non-UTF-8 text is read as Windows-1252, so other legacy encodings (Shift_JIS, etc.) index as mojibake instead of failing. Acceptable for this user base.

- UTF-16 CSVs (Excel "Unicode Text") are still not supported.

- Already-failed documents need a re-index after deploy; the camera coverage register should be re-indexed and confirmed searchable.

## Verification

- chat/rhodes-worker: pnpm test (269 pass, including two new retrieval tests), pnpm typecheck.

- biome check on both changed files.

- Real-file before/after: ran two sample camera registers (one Windows-1252, one UTF-8, containing “Main” door – north side, €450, Señor Muñoz’s office) through extractDocumentKnowledgeSource from main and from this branch, declared as both text/csv and application/vnd.ms-excel:

| File | Declared as | main | This PR |

|---|---|---|---|

| Windows-1252 CSV | text/csv | invalid_source | indexed, text intact |

| Windows-1252 CSV | application/vnd.ms-excel | unsupported_type | indexed, text intact |

| UTF-8 CSV | text/csv | indexed | indexed |

| UTF-8 CSV | application/vnd.ms-excel | unsupported_type | indexed, text intact |

- Not tested end to end through the UI upload: the dev Drive service account can't write to the sandbox folder, so uploads fail before indexing on local dev. Verify in prod by re-indexing the camera coverage register after deploy.

#1726 — feat(tasks): link and upload Aerie documents (AERIE-2740) @caina-barbosa  approved

## Summary

This PR is Phase 9 of 13 in the larger [Human Task Assignment from Aerie](https://linear.app/builder-team/project/human-task-assignment-from-aerie-b3582376ae6e) project.

It adds and gates the canonical link between a Task and an Aerie document. It reuses the existing Site document upload mechanics, lets users select an existing document from the Task's Site, and gives Task participants a supported way to open linked files without granting access to the entire Site. This phase is tracked by [AERIE-2740 — Link and upload Aerie documents for Tasks](https://linear.app/builder-team/issue/AERIE-2740/link-and-upload-aerie-documents-for-tasks).

Production effect: dormant/additive. The new Task document UI, public API, and agent tools stay behind the existing Task Board launch gate. Merging this stacked PR does not enable the Task Board, start uploads, mutate existing Tasks, or change the current Site document upload experience.

### What it means for end users/consumers

| Surface | After this PR merges while the launch gate is off | After the Task Board is launched |

| --- | --- | --- |

| Existing Work Plan and Site document upload | No visible change. Existing upload behavior remains available. | Existing behavior remains available. |

| My Tasks and Task details | Still hidden by the Task Board gate. | A Task participant can upload or attach a same-Site document as completion evidence and open linked files. Task creators, requesters, and Task managers can also manage source links. |

| Public Task API | New document operations remain unavailable while the gate is off. | Authorized callers can list, link, and unlink Task documents through the same Task authorization rules. |

| Agents | New Task document tools remain unavailable while the gate is off. | Authorized agents can use the same document workflow as the UI and API. |

| Site access | No permission change. | Opening a Task-linked document does not grant access to the rest of the Site. |

---

## Why

Text and URL evidence cannot safely represent files uploaded directly to Aerie. Task participants may also need a file even when they do not have wider access to the related Site. This slice establishes one Task-scoped document relationship and a verified open path before approval, comments, digests, and Team Tasks depend on those files.

---

## Business Value

- Lets people attach photos, PDFs, and other Aerie documents to Task work without creating a second upload system.

- Lets assignees attach completion evidence and open a linked file without receiving source-management or broader Site permissions.

- Gives the UI, public API, and agents the same document behavior and authorization rules.

- Prevents Task uploads from triggering Due Diligence approval, Rhodes synchronization, or REBL3 writeback.

---

## How does it work

1. The Task details panel lists eligible documents already owned by the Task's Site and records a separate Task-to-document link when one is selected.

2. New files use the existing Site upload-session, Drive upload, upload-complete, and document-registration mechanics. A small Task-specific wrapper supplies server-controlled metadata, returns only the upload instructions the Task UI needs, and links the resulting document to the Task.

3. The Task Board service lets assignees attach completion evidence while keeping source removal restricted to Task creators, requesters, and Task managers. It checks Task visibility before listing or opening a document. The download route obtains the file through the existing worker boundary rather than returning an unverified Drive URL.

4. The public API and agent tools call the same canonical Task document operations and preserve actor and time audit data.

5. Task deletion and document lifecycle logic remove or protect related links without changing the document's Site ownership. The entire new workflow remains unavailable while the Task Board launch gate is off.

---

## Scope

### Included in this phase

- A canonical Task-document relationship with Task, document, actor, and time data.

- Linking and unlinking eligible documents from the Task's Site.

- Uploading through the existing Site document upload mechanics and attaching the result to the Task.

- Task-scoped file access for Task participants without broader Site access.

- Completion-document attachment for assignees without granting source-management authority.

- Task-scoped upload responses that omit Drive folder identifiers.

- Matching UI, public API, and agent behavior.

- Security repairs at the reused upload boundary: server-owned Task metadata, delegated actor validation, Site containment checks, bounded responses, and safe user-facing errors.

- Exact final diff paths:

chat/app/(main)/api/rhodes/drive/__tests__/routes.node.test.ts

chat/app/(main)/api/rhodes/drive/register-document/route.ts

chat/app/(main)/api/rhodes/drive/upload-complete/route.ts

chat/app/(main)/api/rhodes/drive/upload-session/route.ts

chat/app/(main)/api/task-documents/[taskDocumentId]/route.node.test.ts

chat/app/(main)/api/task-documents/[taskDocumentId]/route.ts

chat/components/dashboards/portfolio/__tests__/portfolio-rhodes-workbench.test.tsx

chat/components/dashboards/portfolio/portfolio-rhodes-workbench.tsx

chat/components/task-board/my-tasks-view.tsx

chat/components/task-board/task-completion-panel.test.tsx

chat/components/task-board/task-completion-panel.tsx

chat/components/task-board/task-documents-panel.test.tsx

chat/components/task-board/task-documents-panel.tsx

chat/convex/_generated/api.d.ts

chat/convex/agentRuns.test.ts

chat/convex/documentKnowledge/lifecycle.ts

chat/convex/publicApi/v2/domains/workManagement.ts

chat/convex/publicApi/v2/domains/workManagementTaskBoardWrites.test.ts

chat/convex/publicApi/v2/http.ts

chat/convex/publicApi/v2/workManagementData.ts

chat/convex/publicApi/v2/workManagementWrites.ts

chat/convex/rhodes/mcp.ts

chat/convex/rhodes/runtime/constants.ts

chat/convex/rhodes/runtime/mutationAuthorization.ts

chat/convex/rhodes/runtime/mutationDispatcher.ts

chat/convex/rhodes/runtime/writes/taskWrites.ts

chat/convex/rhodes/schema.ts

chat/convex/rhodesMcpMutationParity.test.ts

chat/convex/taskBoard/completion.ts

chat/convex/taskBoard/completionHistory.ts

chat/convex/taskBoard/documentCleanup.ts

chat/convex/taskBoard/documents.test.ts

chat/convex/taskBoard/documents.ts

chat/convex/taskBoard/queries.ts

chat/convex/taskBoard/sourceAuthority.ts

chat/lib/__tests__/rhodes-delegation-server.node.test.ts

chat/lib/__tests__/rhodes-route-response.node.test.ts

chat/lib/platform-error-coverage-inventory.ts

chat/lib/public-api/v2/domains/work-management-schemas.ts

chat/lib/public-api/v2/domains/work-management-task-board.node.test.ts

chat/lib/public-api/v2/domains/work-management.ts

chat/lib/rhodes-delegation-server.ts

chat/lib/rhodes-mcp-contract.ts

chat/lib/rhodes-mutation-tools.ts

chat/lib/rhodes-route-response.ts

chat/lib/task-board/document-upload.ts

chat/rhodes-worker/lib/google-drive/client.test.ts

chat/rhodes-worker/lib/google-drive/client.ts

chat/rhodes-worker/mcp-server/tools/portfolioParity.test.ts

chat/rhodes-worker/mcp-server/tools/tasks.ts

chat/rhodes-worker/src/index.ts

chat/rhodes-worker/src/portfolio-document-payload.test.ts

chat/rhodes-worker/src/portfolio-document-payload.ts

chat/rhodes-worker/src/task-document-response.test.ts

chat/rhodes-worker/src/task-document-response.ts

packages/contracts/src/agent-run-protocol.ts

packages/contracts/src/agent-tool-registry.test.ts

packages/contracts/src/agent-tool-registry.ts

packages/contracts/src/rhodes-mutation-proposal.ts

### Deliberately excluded for later phases

- Approval and request-changes actions — AERIE-2741.

- Participant comments, watchers, notification events, and the daily Task Scenario digest — AERIE-2742.

- Team Tasks and portfolio-wide grouping/filtering — AERIE-2743.

- Launch-gate removal and final production verification — AERIE-2744.

- Mobile-specific completion and evidence-upload optimization — explicit post-V1 follow-up.

- Any new upload provider or separate Task file store. This phase deliberately reuses the existing Site document flow.

---

## Review follow-up

Mercy identified ten substantive boundary defects across the review rounds, and all ten are repaired on the branch or in the next repair head. Earlier repairs made malformed present Task IDs fail closed, protected worker-only Drive projections with the shared secret, whitelisted public Task-upload responses, rejected malformed Worker payloads, enforced TLS and complete Convex-host rejection for shared-secret forwarding, preserved aggregates for legacy documents without public IDs, failed closed on dangling Task-document joins, and restricted the public document proxy to 200 content or a 303 redirect with Location.

The exact-head review of ad517b183 found two further reachable local failures. A successfully unlinked document could remain selected in an in-progress completion draft, causing the next submission to reference a relationship that no longer exists; successful unlink now removes that document ID from the pending selection, while failed unlink leaves the draft unchanged. The Task upload-session route could also return 200 for blank upload URLs or non-positive, non-integer, or non-finite chunk sizes; it now rejects those provider responses before the browser starts an unusable upload. The same explicit success signal also preserves the file input after a failed upload so the user can retry the same file.

The required-argument rolling-compatibility proposal remains outside this phase's reachable production behavior. This phase is dormant/additive behind the existing Task Board launch gate: neither an old nor a new production Chat container can invoke Task completion while the feature is disabled, so the mixed-version window cannot produce the alleged rejected completion. Adding a legacy optional branch here would broaden the canonical mutation contract without protecting an active caller; launch and gate removal remain owned by AERIE-2744.

The repeated compensating-deletion proposal also remains deliberately excluded. Upload completion proves that the supplied Drive ID is inside the authorized Site root, but it does not prove that this request created that file; deleting it after Convex registration failure could destroy a pre-existing Site document. The safe failure mode leaves the verified Site-owned file intact and permits registration retry with the same ID.

The full focused Task-document suite passes 156/156, Biome and diff checks pass, and the full Chat/Convex typecheck passes. The Rhodes Worker suite remains 282/282 from the earlier review round. Production effect remains dormant/additive: no launch gate, deployment, migration, external write, or existing Site-upload behavior changes.

---

## Test plan

### Automated validation

- Focused Chat Task document, API, route, UI, and parity suites — 155/155 passed (pnpm --dir chat exec vitest run 'app/(main)/api/rhodes/drive/__tests__/routes.node.test.ts' 'app/(main)/api/task-documents/[taskDocumentId]/route.node.test.ts' components/task-board/task-completion-panel.test.tsx components/task-board/task-documents-panel.test.tsx convex/publicApi/v2/domains/workManagementTaskBoardWrites.test.ts convex/rhodesMcpMutationParity.test.ts convex/taskBoard/documents.test.ts lib/__tests__/rhodes-delegation-server.node.test.ts lib/__tests__/rhodes-route-response.node.test.ts lib/public-api/v2/domains/work-management-task-board.node.test.ts --maxWorkers=1)

- Full-suite integration regressions — 6/6 passed (pnpm --dir chat exec vitest run components/dashboards/portfolio/__tests__/portfolio-rhodes-workbench.test.tsx -t uses.the.canonical.gated.lifecycle.actions.when.editing.a.Work.Plan.Task --maxWorkers=1 and pnpm --dir chat exec vitest run lib/__tests__/platform-error-smoke-inventory.node.test.ts --maxWorkers=1)

- Rhodes Worker suite — 282/282 passed (pnpm --dir chat/rhodes-worker test)

- Chat and Convex typecheck — passed (pnpm --dir chat typecheck)

- Test runtime architecture — passed (pnpm lint:test-architecture)

- Architecture boundaries — passed (pnpm lint:boundaries)

- Convex path validation — passed (pnpm lint:convex-paths)

- Bounded-read validation — passed (pnpm lint:read-bounds)

- git diff --check — passed

- Rebase equivalence — the rebased AERIE-2740 implementation matched the independently reviewed tree before adding the two test-integration repairs found by hosted CI

- Exact-head diff scope — only the authorized paths listed above

### Time for Implementation

An engineer without AI assistance would likely need 5 to 7 working days to trace the existing upload path, implement the Task relationship and access boundary across UI/API/agents, add tests, and complete review and rebase work.

#1739 — Open student profiles from enrollment report @vvp-trilogy  approved

## Summary

- open the SIS enrollment report list-pane link at the student profile instead of the enrollment record

- keep the CSV enrollment-record URL and add a separate student-profile URL column

- update browser coverage for list navigation and both exported links

## Validation

- pnpm --dir chat exec vitest run components/dashboards/admissions/enrollments/sis/__tests__/sis-student-panel.test.tsx --maxWorkers=1 (23 passed)

- pnpm exec biome check chat/components/dashboards/admissions/enrollments/sis/sis-student-panel.tsx chat/components/dashboards/admissions/enrollments/sis/__tests__/sis-student-panel.test.tsx

- pnpm typecheck

#2083 — feat(school-reports): consolidate report readiness, parallelism, and delivery @ashwanth1109  changes requested

## Summary

This is now the single review and merge target for the school-report stack formerly split across #2083, #2086, #2100, and #2102. Review the complete diff here against main; the three upper PRs are superseded and retained only as historical records.

- Report readiness and categories: correctly classify Transportation and Renovations/Furnishings, preserve known zero versus unknown actuals, and display approved missing budgets as unavailable without fabricated budget variances.

- Financial correctness and recovery: align Facilities/depreciation with the report cutoff through the owning warehouse procedure; preserve native-document layout/readback, source evidence, financial arithmetic, independent audits, bounded correction/recovery, and compatible saved-work resume.

- Bounded parallel processing: run 1–9 school-isolated workers with shared provider pacing and cooldowns, one fenced pipeline lease, failure isolation, and one consolidated completion email.

- HTML delivery: provide a full-width Finance digest and equivalent plain-text fallback with alphabetized native report links, reporting dates, aggregate updated/reused counts, and a conditional failure section.

- Durable source findings: retain both financial and operating Facilities values and insert an exact Finance-owned reconciliation finding rather than silently changing source values or withholding an otherwise usable report.

- Safe report reuse and pageless exports: verify immutable snapshots, combine retained and newly generated report links without regenerating unchanged schools, and validate readable/unencrypted/nonblank PDFs without imposing an artificial fixed page count.

- Controlled recipients: default to Ashwanth-only review delivery; a separate explicit, delivery-only approved_finance action reuses the proven review run for the fixed five-person Finance audience.

## Consolidation and review scope

| Former layer | Scope now reviewed in this PR |

| --- | --- |

| #2083 | Categories, missing-budget handling, Facilities cutoff, layout, audit/correction recovery |

| [#2086](https://github.com/AI-Builder-Team/Surtr/pull/2086) | Nine paced, isolated concurrent school workers |

| [#2100](https://github.com/AI-Builder-Team/Surtr/pull/2100) | HTML/plain-text completion digest |

| [#2102](https://github.com/AI-Builder-Team/Surtr/pull/2102) | Source findings, retained-report reuse, pageless PDF checks, approved Finance delivery |

GitHub automatically marked #2086, #2100, and #2102 merged into the consolidated base branch, not into main. #2083 is now unstacked and is the only open review target. The original upper branches are retained along with their PR histories.

The base branch codex/school-report-cost-categories is fast-forwarded from 6fa8d66f to the former top-layer commit 114001640a839f1ca837937f71f93c541e1384a6. All original commits are retained; no rebase, squash, cherry-pick, force push, or implementation change is needed. The consolidation baseline tree is identical to the former #2102 head (tree ddc45216dbe5a2e7db21014f85781e1a8844965b).

The PR targets main. It was draft during consolidation and was subsequently marked ready for review by ashwanth1109 while CI was running. This CI fix did not change readiness or enable auto-merge. Consolidation does not merge into main, deploy infrastructure, refresh warehouse data, generate reports, send email, or resolve review threads. Original upper PR descriptions remain available for detailed historical evidence; the original #2083 description is archived below.

## Business Value

Finance gets accurate, evidence-backed school reports despite explicitly identified source gaps, without fabricated budgets or silently reconciled values. Parallel processing and safe completed-report reuse reduce batch turnaround and unnecessary regeneration. A readable consolidated digest and review-first recipient workflow make delivery easier to inspect and deliberately approve.

## Implementation Effort

Approximately 45–67 engineer-hours for an average engineer to hand-code the combined solution without AI assistance, based on the four original layer estimates (20–30, 12–18, 4–6, and 9–13 hours). These estimates include differing amounts of deployment/live-validation work; additional validation effort may be required.

## Validation

### Fresh local checks — October 6, 2026

- School-performance report suite at the exact consolidated candidate: 300 passed.

- Education-financials upstream suite at the same candidate: 151 passed.

- git diff --check against current main: passed.

- Read-only git merge-tree against main at 3f67679a: clean, with no conflicts.

- Every former layer head is an ancestor of 11400164; the fast-forward preserves all four layers with no code change.

- No repository-wide lint/format run or production operation was performed for consolidation.

### Recorded pre-merge live validation

The former top-layer [#2102](https://github.com/AI-Builder-Team/Surtr/pull/2102) records exclusive deployment of the same commit 11400164 to Pipeline-school-performance-reports-prod, task revision 44, with CDK [1/1] and CloudFormation UPDATE_COMPLETE.

Its final review execution 6a1c4b05-9e60-4ca3-aa95-38f5ad74ce4d and explicitly approved Finance delivery execution a25e4985-9956-4adb-b6d5-71e0642addf6 each record 26 reports, zero failures, zero generated, 26 reused. The earlier batch exercised the nine-worker cap, unavailable Malibu budgets, the exact New York $391,000 Facilities source finding, and Raleigh's legitimate ten-page pageless export.

These are retained historical validation records, not new live runs or independently reverified AWS receipts in this consolidation session. Consolidation leaves the implementation commit and tree unchanged. Any subsequent implementation or repair still requires candidate live validation before merging under the user-level Surtr policy.

### CI import-order fix — October 6, 2026

Follow-up commit [9557f15a](https://github.com/AI-Builder-Team/Surtr/commit/9557f15a19396404333ba2f981f59f0a2ddfe7d4) fixes the absorbed stack's Ruff E402 failure in pipelines/runners/mart-aerie-education-financials-refresh/scripts/apply_facilities_cutoff.py. The helper must add its sibling script directory before importing apply_ddl; the import now has an explanatory comment and a narrow, line-only # noqa: E402. No global lint rule or workflow is relaxed.

- CI-pinned Ruff 0.15.22 check and format check pass for the only modified file.

- Python AST is identical to the pre-fix version: comments only, no runtime behavior change.

- The helper's default dry run succeeds without warehouse writes.

- Fresh local suites: 300 school-report tests and 151 education-financials tests passed.

- git diff --check passes.

- No production deployment, pipeline invocation, warehouse refresh, or email send was performed for this comments-only CI fix.

- [Fresh GitHub CI](https://github.com/AI-Builder-Team/Surtr/actions/runs/37482751012) completed successfully for this commit: all seven CI jobs passed, including Ruff check/format and the complete pipeline-runner suite. The automated review rerun was still in progress at the final CI verification; human review remains required before merge.

## Linear

- [SURTR-1535 — Categories and approved missing budgets](https://linear.app/builder-team/issue/SURTR-1535/fix-school-performance-report-categories-and-approved-missing-budgets)

- [SURTR-1538 — Parallel school-report workers](https://linear.app/builder-team/issue/SURTR-1538/run-school-performance-report-batches-with-three-isolated-parallel)

- [SURTR-1551 — HTML completion digest](https://linear.app/builder-team/issue/SURTR-1551/style-school-performance-report-delivery-as-an-html-digest)

- [SURTR-1553 — Unavailable budgets and source findings](https://linear.app/builder-team/issue/SURTR-1553/publish-school-reports-with-unavailable-budgets-and-source-findings)

## Historical base-layer description

<details>

<summary>Original #2083 description and detailed validation history (superseded scope statements)</summary>

The following text is preserved verbatim as historical evidence. Its earlier statements that upper-layer work was separate or excluded no longer describe the consolidated review scope above.

## Summary

Fixes discovered while validating school performance reports in batches. This remains a draft for review only; nothing has been merged. Candidate changes have been deployed for pre-merge validation.

- Expense categories: put 62500 Transportation under Program costs and 62100 Renovations/Furnishings as the second/last detail in Miscellaneous. Include both exactly once in financial totals, with the narrower operating Programs comparison explicitly labeled.

- Approved missing budgets: retain actuals, display Unavailable, and omit budget variances for Lunch Revenue, Financial Aid, Sibling Discount, Stripe Fees, Associate / Other Headcount, Landscaping, Lunch Program, Transportation, Renovations/Furnishings, Computers, Administration/Other, and the Miscellaneous subtotal. Other budget, actuals, Timeback, freshness, enrollment, model, and reconciliation safeguards remain enforced.

- Confirmed zero actuals: budget-only rows with explicitly zero actual postings produce zero management actuals. Unknown actuals with postings are not converted to zero.

- Native report layout: add two rows to each financial table in report copies only, preserving the pinned source template, inherited styles, revision guards, readback validation, and retry behavior. Use 90% financial-body line spacing to retain nine pages at the native 9pt font size.

- Scoped view migration: opt-in, dry-run-by-default migration of only mart_education.school_performance_report_lines, with definition/ACL backups, support for older views missing budget_value_status, transactional owner/grant preservation, and no cascading drops.

- Facilities cutoff repair: derive facilities/depreciation actuals from canonical atomic P&L postings through the report cutoff, instead of including the entire current month. Preserve the accepted QuickBooks generation/snapshot pins, monthly coverage guards, expense sign convention, allocations, model-budget proration, and atomic publication. Add a narrow procedure-only deployment script with backups and owner/ACL verification; no grants/revokes or table rebuilds.

- Native correction addresses: preserve report/claim structure in narrative-correction requests and list exact editable paths, compacting only source evidence. Fixes the observed Chicago correction targeting a columnar transport address instead of a real report field. Invalid paths still fail closed, and final audit remains required; saved drafts/contexts remain compatible.

- Truncated audit recovery: re-review the full original report, prominent text, accounting instructions, and all evidence in a bounded structured call; never treat opaque reasoning/signatures as user-text evidence. Preserve incomplete-coverage rejection, stage cost/deadline limits, corrections and final independent review. Saved research/drafts remain compatible.

- Correction numeric roles and resume compatibility: keep source-linked actual, model allowance and signed variance values separate in every correction request, outside prose compaction. Version correction results/context, so an unfinished old correction can be replaced from its compatible saved complete audit without repeating research/drafting; completed reports remain untouched. Final review and the one-correction limit remain enforced.

## Business Value

Enables usable school performance reports with correctly classified expenses and honest treatment of unavailable budgets. Aligns facilities and depreciation actuals with the report date, preventing future-dated entries from creating false reconciliation failures or overstating current spending. Makes audit output exhaustion recoverable without bypassing financial review or regenerating successful schools. Preserves the distinction between spending and comparison differences during narrative correction.

## Validation

- School report tests: 162 passed after the numeric-comparison correction fix (eight additional regression cases).

- Education financials upstream tests: 151 passed.

- Ruff checks on modified Python files and git diff --check: passed.

- Regression tests cover category totals, optional versus required budgets, zero versus unknown actuals, migration permissions, native row/style/retry handling, and the canonical facilities SELECT across cutoff/quarter boundaries, credits, wrong-school/publication/classification rows, and no-posting prefixes.

### Original changes: Boca Raton live validation

- Deployed only Pipeline-school-performance-reports-prod with --exclusively; CDK showed [1/1] and CloudFormation confirmed UPDATE_COMPLETE (task revision 28). No other pipeline stack was deployed.

- Applied the report-lines view migration with ownership/grants preserved and verified 37 distinct management lines.

- After correcting the added-row page spill, the resumed run completed SUCCEEDED, zero failures, at 2026-09-29 08:49:32 UTC. The nine-page PDF passed the layout guard and visual review; completion email was sent.

- Execution: boca-layout-resume-20260929-0847; run: e55c7dbe-4056-46d7-8d11-e0c2d2449ab7.

- [Boca Raton report](https://docs.google.com/document/d/1sL6T4wiuziNvuKdhX1pNypgD4af-knWO97L8E8YY_ls/edit).

### Batch two: candidate deployment and upstream validation

- Commit 1891f6f6 adds the seven separately approved budget exceptions. Deployed only Pipeline-school-performance-reports-prod using --exclusively ([1/1]), task revision 29; CloudFormation UPDATE_COMPLETE at 2026-09-29 09:06:51 UTC. No other stack or recipient configuration changed.

- Preflight identified future-dated September 30 depreciation in the facilities mart versus the September 29 report cutoff: $11,065.93 Brownsville, $7,747.62 Chicago. Report generation was not invoked against these known-invalid inputs.

- Commit d7070e9c repairs the owning facilities calculation (procedure version 2026-09-29.1). Applied only that procedure and its measure comment, verified original ownership/ACLs unchanged, and retained the original procedure backup. No upstream runner/infrastructure deployment was required.

- Authorized upstream platform run f9de7fb1-c7ed-4a09-9790-c25c922a92c6, execution facilities-cutoff-batch2-20260929-0917, completed SUCCEEDED at 09:18:54 UTC, publishing 10,208 facilities rows. Verified rent, facilities, and depreciation agree across the report and facilities tables for both schools. Correct depreciation: $24,963.43 Brownsville, $0 Chicago.

- That refresh picked up the newer 08:45:47 UTC FinalSite snapshot while the unit-economics tables still used 06:46:21 UTC. Preflight caught the mismatch; no report invocation occurred with mixed enrollment publications.

- With explicit additional authorization, unit-economics run 1c5bee2c-4e0e-47ea-87e3-e1614eb99259 succeeded at 09:26:37 UTC with 1,102 rows. Its automatic per-student run 053f1a3c-7a1f-441c-97bb-61d9d33d8bde succeeded at 09:26:53 UTC, also 1,102 rows. No model assignments or safety checks were changed.

### Batch-two report outcome

Full read-only preflight passed for both schools and the unchanged native template after enrollment publications were aligned. Run 4f390d10-edb1-4fc2-ad7b-c99b6408b426 completed at 2026-09-29 09:51:23 UTC, but its result was partial_failure, not complete report success:

- Brownsville succeeded, passed its independent evidence audit and PDF guard, and all nine pages passed visual review. [Completed report](https://docs.google.com/document/d/1zOALdJl1CU8d4MFsasUCPKs6ORDE1dpYOHNp5j8oMUc/edit). PDF SHA256: c37891c4b46804bf0fa7a76fb93fac7ecea35d234443ff11c5ecfe45210354e0.

- Chicago stopped before publication. Its audit correctly identified comparison-label/scope issues; correction generation then used the nonexistent transport path signals.rows[2][1]. The unchanged field guard rejected it. The summary email reported one ready and one failed school.

- Commit 326a5a4f fixes the correction request's native structure. Regression tests cover full, focused and previously compacted contexts, preserve evidence compaction, and continue rejecting the exact malformed path.

- Deployed only Pipeline-school-performance-reports-prod using --exclusively ([1/1]); CloudFormation confirmed UPDATE_COMPLETE at 09:57:43 UTC, task revision 30. No parallel-worker changes are included in this deployment.

- Verified immutable saved snapshots/draft/audit/context checksums, current freshness, template contract, and real editable paths before the authorized resume.

- Resume execution chicago-native-correction-20260929-0958, run 051588b7-a2b5-40f1-97cf-9544c24c3fc4, completed SUCCEEDED at 2026-09-29 10:03:40 UTC. Result: 2 reports, 0 failures, 1 reused report; summary email acknowledged by SES.

- Brownsville's document ID and PDF checksum are unchanged. Chicago resumed at correction, without repeating source capture/research, passed the final independent audit and PDF guard, and all nine pages passed visual review. [Completed Chicago report](https://docs.google.com/document/d/122G13Ecxp2om_UwsYMEODIn8h9-AuDs4NNJgzHW1PbQ/edit). PDF SHA256: 5ecba3c87d6fa53308a32b9320911d34ad1bde657fce78323b7f90a58e619276.

- The user-requested parallel-worker change is a separate stacked branch/PR, not included here.

### Kirkland: audit recovery live validation — September 29

- Commit 8b0c2d92ba7d9e8e8043937f229a643a5eb4c49e fixes the truncated-audit fallback. The old fallback had only a JSON serialization of opaque thinking and could not certify coverage. The new fallback independently reviews the same complete report/evidence with thinking disabled, a strict 6,000-token verdict allowance, counted full-context cost, and the unchanged 180,000-token stage / 900-second deadline. Missing coverage or truncated recovery output still fails closed.

- Eight regression cases cover opaque-response recovery using the saved draft, complete evidence including confidence-only receipts, incomplete coverage, truncated/missing verdicts, correction plus fresh final audit, and both reserved and recounted cost limits. Parent suite: 154 passed; restacked candidate suite: 169 passed. Targeted Ruff and diff checks passed.

- Deployed only Pipeline-school-performance-reports-prod, --exclusively, CDK [1/1]; CloudFormation UPDATE_COMPLETE at 11:21:31 UTC. Task revision 32, image digest sha256:5786dfa972a7cb7ae4982e73eca72089ac3b926dc8e71b605f39b1dac6db123d. The deployed candidate 94a2c767bf5eac6620470d3a6dccb01c8f5f3f95 includes the unchanged parallel-worker layer. No other stack, upstream refresh, model assignment, or recipient change.

- Fresh source/template preflight passed. Exact saved snapshot/research/draft versions and checksums were verified; current publication fingerprints matched the retained snapshots. Resume contract remained unchanged.

- Execution kirkland-audit-recovery-20260929-1122; run 2f69a5f4-02b1-40d2-97f6-23f6f908c872, resuming 474b28cb-5125-4c21-8175-ec01e0b0bbad.

- The deployed recovery path was actually exercised: the first audit consumed all 14,000 output tokens without a verdict; fresh full-evidence recovery returned complete coverage and a structured rejection, identifying a Programs comparison-label issue. Initial audit plus recovery consumed 66,404 tokens, within the unchanged cap. The built-in correction ran, followed by a fresh complete final audit.

- Report outcome remains partial failure, not successful Kirkland publication. The automatic correction introduced a separate material actual-versus-variance wording error. The final audit completed (38,307 tokens) and correctly rejected it; the existing one-correction limit stopped the school. No Kirkland Google Doc/PDF was created. Saved corrected draft and audit evidence are retained. The audit-recovery fix is live-verified; this separate narrative-correction issue still blocks Kirkland.

- La Jolla reused unchanged: same document ID 16Umzqz83io-lFdLzDROC4BjPXpRm1gKeKHd_-yr4b5w and PDF SHA256 eabb85d232ac6314a436e817a66da8d90908b19d890c94df5db238735e74aa80. Kirkland's original research/draft references also remained unchanged; neither was regenerated. Houston Heights and all other schools were excluded.

- One summary email was accepted by SES for ashwanth.r@trilogy.com (010001a0ececcb2e-444976e3-266e-4cc7-9b06-b54a18111791-000000). No inbox-delivery claim. The run is terminal and the pipeline lease is released. Paused pending authorization to address the new correction error; no automatic retry or next batch.

### Authorized Kirkland correction retry — successful September 29

This resolves the separate narrative blocker recorded in the preceding run. The user explicitly authorized fixing the correction and retrying final review from saved work, keeping La Jolla unchanged.

- Fix commit 3ed718e0b3fa90a8d25af70fe28c3544d624f29f. Numeric comparison entries retain exact E00 source paths, metric/breakdown identities and units, plus distinct actual/model/variance fields. They survive focused context and prose compaction. New correction rules prohibit substituting a variance or quarter-to-go for actuals and prefer omitting a secondary comparison to ambiguous shortening.

- Correction policy school-performance-correction-v2-comparison-roles is included in correction results and context signatures. A saved older correction is redone only if its prior audit completed and has identical evidence. Current-policy corrections resume at final review; approved reports return unchanged. No extra automatic correction cycle, altered financial data, or reduced audit gate.

- 162 parent / 177 combined tests passed, including numeric roles/nulls, compaction preservation, old/current/approved checkpoint handling, missing/incomplete/mismatched prior-audit handling, and continued rejection of an incorrect final correction. Targeted Ruff and diff checks passed.

- Native stack #2087 rebased bottom-up and safely pushed; combined candidate 42bb39f46ee889339b9d3df1a5d6d5ffa893e325. Only Pipeline-school-performance-reports-prod deployed with --exclusively; CDK [1/1], CloudFormation UPDATE_COMPLETE at 11:38:39 UTC, ECS revision 33, image digest sha256:567e2c8148ddbe3c0a4e060f5533c7ea22752a0c14b41f47d8a0eec8f5a4c069. No other stack, upstream refresh, school/model assignment, recipient or schedule change.

- Fresh full source/template preflight passed. Exact saved snapshot/research/draft/audit/correction versions and checksums were verified; source publication fingerprints were unchanged. The saved audit/correction evidence matched and the direct comparison ledger reconciled.

- Execution kirkland-comparison-retry-20260929-1139, run dfc18896-0389-41cc-bad7-5ae89015133c, resuming 2f69a5f4-02b1-40d2-97f6-23f6f908c872. Terminal SUCCEEDED, 11:39:24–11:44:05 UTC, ECS exit 0; report result success: 2 reports, 0 failures, 1 reused.

- Kirkland completed: original research/draft/snapshot references are unchanged. The deployed new policy rebuilt only the correction from the saved audit, then completed fresh final review with approved=true, coverage_complete=true, issues=[]. The corrected Programs card clearly distinguishes actuals, QuickBooks budget, and variance; the misleading model-variance-as-spending phrase is absent. No initial research/drafting or initial audit was repeated.

- [Completed Kirkland report](https://docs.google.com/document/d/1xWYkhsdsrt7m04VotO4M9RCGLq72zyznbITVW9D1piE/edit). Exact retained PDF SHA256 132254afd64c713d6cffd17050ee37a8a6703e3942bffabd9804c9dd6835e0f5. Runtime PDF checks passed; all 9 pages were rendered and visually checked, including comparison labels, tables, category placement, unavailable budgets and complete evidence/confidence layout.

- La Jolla reused unchanged in 0.5 seconds: same document 16Umzqz83io-lFdLzDROC4BjPXpRm1gKeKHd_-yr4b5w, same PDF SHA256 eabb85d232ac6314a436e817a66da8d90908b19d890c94df5db238735e74aa80. No other school was invoked.

- One consolidated email accepted by SES for ashwanth.r@trilogy.com: 010001a0ecfa36ee-1f658f77-7386-45d9-9e4b-019a85e90cdc-000000. The batch lease is released. Paused before the next batch. Both PRs remain draft; nothing merged.

## Scope and remaining investigation

The 18 changed files are limited to the report runner and the owning facilities procedure, its deployment script, tests, and documentation. No schedule, recipient, shared infrastructure, or model-assignment changes are included. No confidential report/evidence artifacts are committed.

East Bay remains paused: 19 students, but no assigned unit-economics model. Bethesda and Boston Suburbs retain their separate model/enrollment prerequisite blockers. No next batch is started automatically.

## Linear

[SURTR-1535 — Fix school performance report categories and approved missing budgets](https://linear.app/builder-team/issue/SURTR-1535/fix-school-performance-report-categories-and-approved-missing-budgets)

## Implementation Effort

Estimated 20–30 engineer-hours (about 2.5–4 days) for an average engineer to investigate and hand-code the SQL, validation, native Google Docs layout/retry changes, safe migrations, regression tests, cutoff repair, and pre-merge live validation without AI assistance.

## Five-school audit/correction recovery — September 29

The user authorized repairing the five failures from nine-worker run b7adcc20-8aa3-4dbd-b9b7-fe70021ec81c and resuming saved work, retaining Boston, Santa Monica, Scottsdale and The Woodlands unchanged.

Parent fix 3139d150db6224014728b3602e5e367660897ace, combined candidate 0d36afbbfe37f4a0c56df21f7db628d56a2a129c:

- Chantilly: request at most one audit verdict. An unexpected multi-verdict response is not cherry-picked; the bounded full-evidence recovery must independently produce one complete verdict.

- Charlotte: the one incomplete-audit retry now has two disjoint review scopes covering every claim and confidence paragraph. Both retain the full report and every receipt, must attest complete coverage, and share the original 180,000-token / 900-second audit ceiling. No incomplete verdict can reach publication or be mistaken for an issue to edit away.

- San Francisco: direct numeric comparison context now explicitly separates booked and management actuals, including booked-minus-budget/model calculations. Published financial variances are labeled as management-based. Unavailable benchmarks remain null.

- Dorado: a policy-upgrade resume re-audits the latest saved correction, rather than editing the old draft using only its first audit's issues. It can therefore address a headline defect discovered at final review without reverting earlier fixes. The same single-correction limit and fresh final audit remain mandatory.

- Palo Alto: correction requests carry exact field-length ceilings and conservative targets. The second bounded text-repair call gets its actual rejected text and measured length. Secondary examples may be omitted to fit, but retained amounts, comparison basis/direction and necessary qualifications cannot change. Full evidence review and PDF checks remain mandatory.

172 parent / 211 combined tests pass. Targeted Ruff and git diff --check pass. Regressions cover ambiguity recovery, complete two-part coverage, all receipts retained, combined issues, shared budget exhaustion, exact booked variance/null handling, field limits, rejected-text feedback, new-policy fresh review, and preservation of already-corrected fields. The native stack was rebased bottom-up and pushed with explicit force-with-lease checks; both PRs remain draft and nothing is merged.

Read-only resume preflight passed for all nine exact retained snapshot versions/checksums, their generation checkpoints and evidence references, the 48-hour source freshness constraints, six-table/enrollment/model contracts, deterministic rendering and the unchanged native template. The invocation contract is unchanged apart from resume_run_id. Five unfinished schools reuse research and drafts; four completed schools reuse existing document IDs and PDFs. No upstream refresh, model assignment, school mapping, recipient or schedule change.

Live validation finished with remaining blockers — not ready to merge. Only Pipeline-school-performance-reports-prod was deployed, using --exclusively; CDK reported [1/1] and its CloudFormation events reached UPDATE_COMPLETE at 12:33:05 UTC. Task definition 35 ran the combined candidate image. No other pipeline, registry or shared stack was deployed.

The authorized resume 967c955b-a8ef-41af-86c6-1b2f6b573fea ran from 12:33:56 to 12:41:50 UTC. Step Functions succeeded, but the application correctly reported partial_failure: 8 published reports, 1 failure, 4 reused. Boston, Santa Monica, Scottsdale and The Woodlands retained their exact prior document/PDF references. Chantilly, Dorado, Palo Alto and San Francisco completed from saved research/drafts; Charlotte retained its corrected draft. Checks confirmed unchanged research/draft and source snapshot references; no full generation was repeated. The usual summary email was sent to the unchanged recipient with eight published reports. The pipeline lease was released.

All 36 pages of the four newly published PDFs were rendered and visually inspected. Layout and page-count checks passed. San Francisco's corrected management-profit/variance pairing, Dorado's operating-program variance and janitorial account attribution, and Palo Alto's fitted correction are present. However, automated approval is not sufficient: manual inspection found the narrative defects below, so Chantilly and Dorado are published, not fully QA-approved. The email had already been sent before this post-run inspection; those documents have not been silently edited or regenerated.

### Remaining blockers found by live validation

- Charlotte — review budget coordination: its final audit exhausted the 14,000-output allowance, then the full-evidence recovery returned incomplete coverage. The first scoped retry completed and approved its own scope, but the second was blocked before invocation by the shared 180,000-token gate, including its full recovery reserve. The final audit had actually spent 117,507 tokens by that point; this is a reservation/planning failure, not evidence that the corrected narrative is wrong. The runner did not publish a half-reviewed report. The scoped fallback was live-exercised but did not finish; its live validation remains failed.

- Chantilly — standalone headline arithmetic missed by automated audit: facilities total spend $35.1K is labeled as the over-model amount instead of the actual ~$5.8K variance. Both a depreciation headline and finding title label $10.8K actual spend as the amount over both benchmarks; the overruns are ~$10.1K against QuickBooks and ~$9.6K against model. Tables and explanatory paragraphs retain the correct values. These three narrative fields need correction and a stronger comparison check.

- Dorado — explanation/attribution missed by automated audit: finding 7 and the Timeback transaction-reference paragraph say the allocation raises EBITDA, while the report's own basis shows booked EBITDA -$393,604.62 becoming management EBITDA -$488,332.07 after the $94,727.45 allocation. Its second opening signal also attributes the EBITDA result partly to depreciation, despite depreciation being added back. The third signal calls all six rent rows bills, whereas its detailed evidence distinguishes rent bills from two purchases. The displayed accounting totals are correct; these causal/directional and transaction-type descriptions need correction.

Work is paused after this batch, consistent with the requested investigate-and-confirm workflow. No follow-on execution or completed-document modification has been made. Recommended next step: repair the bounded audit scheduling and these saved-narrative QA gaps, then resume/review only the affected saved work after user confirmation. Other completed September 29 reports remain reusable and unchanged; combining reports from separate historical runs into one delivery is a separate operation, not part of this resume.

Evidence: /tmp/surtr-review-resume-OIfR2u (deployment events, terminal result/email checkpoints, immutable-source assertions, audit receipts, exact-version PDFs and rendered pages).

Incremental implementation effort for this recovery work: approximately 4–6 engineer-hours without AI assistance (in addition to the earlier work estimated below).

## Three-school saved-work repair — September 29

The user authorized fixing Charlotte's remaining audit-budget failure and the post-publication QA findings in Chantilly and Dorado, leaving all other reports untouched.

Parent 8bb89187, combined candidate 087a7534. Native stack #2087 was rebased bottom-up and pushed with leases. Both PRs remain draft; nothing is merged.

- The incomplete-review fallback now counts and reserves both disjoint scopes before starting either. Each gets full evidence and a 6,000-token forced verdict with thinking disabled, without recursively reserving another complete recovery per scope. Complete coverage of both remains mandatory; the original 180,000-token / 900-second safety limits remain unchanged. Truncation, ambiguity, incomplete coverage and unaffordable plans fail closed.

- Explicit repair_reports maps selected completed school IDs to retained human-QA findings. It requires matching published insights and a compatible saved approval with receipts. A new repair seed starts from the latest published text, receives one scoped correction and a full fresh final audit. Previous documents/PDFs/checkpoints remain intact, with supersedes provenance for replacements. No research/draft generation is repeated, and unselected completed reports are imported unchanged.

- Supplementary deterministic checks reject unambiguous facilities/depreciation variance labels that use spending or disagree with exact source arithmetic, and explicit EBITDA/addback or allocation-direction contradictions. They respect displayed rounding and explicit EBITDA-neutral/addback qualifications; they supplement, not replace, full evidence review.

198 parent / 239 combined tests pass. Targeted Ruff and git diff --check pass. Regression coverage includes Charlotte's observed budget scale, all-evidence scope coverage, fail-closed incomplete/truncated output, no partial-review approval, explicit repair eligibility, published-text seeding, mandatory final review, preservation of other reports/prior documents, same-run idempotency and the observed narrative defects/valid qualifications.

Read-only preflight passed against the exact retained snapshots and checkpoints from 967c955b-a8ef-41af-86c6-1b2f6b573fea, the 48-hour freshness and six-table contracts, current native template, unchanged recipients and role. Six reports will be reused, Chantilly and Dorado repaired, and Charlotte resumed at final review. Seven reports from earlier batches remain outside this run and unchanged. No warehouse writes or upstream refreshes.

Live validation finished with two narrative blockers — not ready to merge. Only Pipeline-school-performance-reports-prod was deployed with --exclusively; CDK showed [1/1], and CloudFormation reached UPDATE_COMPLETE at 12:58:07 UTC. ECS task definition 36 ran image digest sha256:354bf51b1de2b1d12430c41caf84125b5acf0cbbf95e8fc3962912e4b10e6c2c. No other pipeline/shared/registry stack or upstream refresh was deployed.

Resume 6b7d7aa3-5249-43b1-9a25-1c62b8f20328 ran from 12:58:34 to 13:04:10 UTC. Step Functions succeeded; the application result is correctly partial_failure: 8 reports, 1 failed repair, 6 reused. The completion email was sent to the unchanged recipient, and the lease was released. Boston, Palo Alto, San Francisco, Santa Monica, Scottsdale and The Woodlands retain identical prior document/PDF references. All original research/draft/snapshot references remained unchanged; repaired reports retain identical receipts. Seven earlier-batch reports were not included or modified.

- Dorado: the targeted repair completed, passed full evidence review (50,968 final-audit tokens), and all nine PDF pages were manually inspected. Five narrative fields changed only within the four authorized claim scopes, correcting EBITDA direction, depreciation/addback causality, and six rows versus six bills. [Corrected replacement document](https://docs.google.com/document/d/1DNkPYEJsQDqlxs7CsBFOVnA8zxl5zCuK1WE1m2CgxRM/edit). The prior published document/PDF remain untouched.

- Charlotte: final review completed on the first call (41,882 tokens), and publication succeeded with its saved corrected narrative byte-for-byte unchanged. The new scoped fallback was not needed in this live run; its budget behavior is regression-tested against the observed 80,740-token starting balance and complete evidence. All nine PDF pages were inspected. Manual QA then found a separate, pre-existing narrative error: several claims invent a 13-versus-25 enrollment target as the revenue-gap cause. E00's model revenue is $130,000 = 13 actual students × $50,000 annual tuition × 0.2; 25 is capex_reference_student_count in the model assumptions, not a revenue enrollment target. The financial tables are correct. Charlotte is published and emailed, but not fully QA-approved.

- Chantilly: the saved correction now contains the correct $5.8K facilities-over-model headline and separate $10.1K QB / $9.6K model depreciation overruns in its headline/finding. Final review (50,807 tokens) correctly blocked the replacement because the executive paragraph still calls $10.8K actual depreciation 'over both budget and model'. That paragraph was outside the supplied repair findings/allowed correction claims. No replacement document was published or included in this email; the corrected work and original published version are retained.

Next step, awaiting user direction after this batch: extend the explicit saved-work repair path to accept the latest unfinished corrected narrative with new QA findings (without replaying research or reverting the headline fixes), repair Chantilly's executive paragraph, and correct Charlotte's enrollment-benchmark/causal claims using the actual revenue-model basis. Keep every other report untouched. No automatic extra correction cycle or follow-on execution has been started.

Evidence directory: /tmp/surtr-three-repair-OzHvpw (preflight, deployment events, terminal result/email, immutable-reference assertions, audit receipts, narrative diffs, exact-version PDFs and all 18 rendered/reviewed pages). Deployment alone is not validation, and automated publication is not human QA approval.

Incremental implementation effort: approximately 3–5 engineer-hours without AI, beyond the earlier estimates.

## Saved-work narrative repairs — September 29, 2026, final validation

- Latest correction can now be an explicitly authorized repair source, not just a completed approval. Source stage and immutable reference are retained. An approved publication in progress cannot silently roll back to an older correction. Ordinary retry behavior is unchanged: no extra correction loop and no research/drafting regeneration.

- Correct the remaining Chantilly executive depreciation variance, preserving the three earlier headline/title fixes. The new source-backed check also caught its related 9 vs 25 enrollment inference before dispatch.

- Distinguish actual/model student counts from the CapEx reference denominator and model-name suffix. Remove Charlotte's unsupported 25-student enrollment target and enrollment/pricing causal claims throughout its affected narrative; preserve all reported amounts and require reconciliation instead of inventing a cause.

- Add narrow regression guards for directly labeled body-text expense variances (including “over both budget and model”) and misuse of CapEx reference counts as modeled enrollment. Full-source review still remains mandatory.

- Tests: 213 parent / 254 combined pass; targeted Ruff and git diff --check pass.

- Candidate deployed before merge: parent 6fa8d66f, combined 67cd5297, task definition 37, image digest sha256:68db812865cd85046ad4c7842b006c9e05ea8038f83574b0204ebe8dd385ccd7. Only Pipeline-school-performance-reports-prod deployed with --exclusively; CDK [1/1] and CloudFormation UPDATE_COMPLETE at 13:17:17 UTC verified. No upstream/warehouse change or other pipeline deployment.

- Live validation: arn:aws:states:us-east-1:479395885256:execution:pipeline-school-performance-reports-prod:repair-two-saved-reports-20260929-131743; Surtr run 667b1467-231f-4fae-96e6-2dd58706bd11. 9/9 reports complete, 0 failures, 7 exact reused reports, consolidated email sent to the unchanged configured recipient. Only Chantilly and Charlotte received bounded corrections and fresh full-source audits, using byte-identical retained evidence and original snapshot/research/draft references. Prior documents remain retained, new reports record superseded publication/source provenance. Dorado and every other report untouched; lease released.

- Both new PDFs have nine pages; full visual/narrative QA completed. No changes beyond the selected narrative claims; earlier Chantilly corrections preserved. Paused after this batch; both PRs remain draft and unmerged.

### Requested consolidated delivery: all 16 schools

After both repaired reports passed audit and manual QA, the user requested inclusion of the seven earlier completed reports (Boca Raton, Brownsville, Chicago, Fort Worth, High Austin, Kirkland and La Jolla). A separate delivery-only consolidation used the unchanged existing delivery.send helper, pipeline lease, immutable S3 manifest and idempotent/uncertain-send protection. This was not another generation run, a change to the fixed-school-list resume contract, or another deployment.

Read-only preflight verified complete source checkpoints, September 29 cutoffs, exact-version snapshot/narrative/PDF checksums, nine-page PDF contracts, native document completion and existing access for the unchanged recipient. 16 distinct report links sent, 0 generated in the consolidation, delivery ID consolidated-16-20260929-f16193e2b34c8ce3. The 14 other reports stayed unchanged; only Chantilly and Charlotte were repaired in the preceding run. Delivery manifest retained at s3://surtr-school-performance-evidence-prod-479395885256/runs/consolidated-16-20260929-f16193e2b34c8ce3/6db7048a-299b-41a1-97e5-20a0e50e70f5/delivery/consolidation-manifest/d58e1cf7309a47c68e187a0931d07f313b5bc5e12f602d84a98cb541ecabc44e. No source run/checkpoint/document was edited for consolidation.

</details>

#3851 — fix(income-statement): migrate AI summary to Opus 4.8 @sanketghia  approved

## Summary

- Update both Income Statement AI summary calls from Claude Opus 4.1 (deprecated) to Opus 4.8.

- Replace fixed-budget thinking with adaptive thinking at high effort.

- Preserve the existing tool-use loop and token limits.

## Validation

- Manually exercised the local POST /income-statement/ai-summary flow and polled its status endpoint; the task completed and returned an appropriate summary.

- uv run ruff format routers/income_statement.py — passed.

- uv run ruff check routers/income_statement.py — passed.

- git diff --check — passed.

- uv run pyright routers/income_statement.py reports 30 errors and 33 warnings, matching the baseline on origin/main.

- Automated pytest was not run.

## Screenshots

<img width="1524" height="777" alt="image" src="https://github.com/user-attachments/assets/4c92219c-bbe0-4978-93de-24ef9df38960" />

The Portfolio  —  Trilogy Companies

The Algorithm Built From Their Own Labor

Forbes traces how Joe Liemandt's global workforce may be training the system that replaces it.

AUSTIN, TEXAS — For nearly two decades, Crossover sold a promise to engineers in Lagos, Manila, and Minsk: work was being judged on merit, not geography, and pay would reflect that meritocracy. It built Joe Liemandt's first fortune. Two new Forbes investigations now ask what happens when that same workforce becomes the training data for its own obsolescence.

The first piece, "How A Mysterious Tech Billionaire Created Two Fortunes—And A Global Software Sweatshop", maps the ESW Capital playbook in blunt terms: buy mature software at distressed prices, staff it through Crossover's remote talent pipeline, and push support pricing up term over term until EBITDA margins clear 75%. Trilogy calls this meritocracy. Forbes calls it something closer to arbitrage at scale — a model that made Liemandt a billionaire twice over without ever taking outside capital or facing a public shareholder vote.

The second piece is the sharper turn of the knife. "The Billionaire Who Pioneered Remote Work Has A New Plan To Turn His Workers Into Algorithms" reports that the same screening data, task logs, and performance metrics Crossover accumulated to prove its workers were the "top 1% of global talent" are now the raw material for automating the work those workers perform. The pipeline that recruited them may be the pipeline that retires them.

Trilogy has never disputed that automation is the mission — it is the mission, stated plainly in internal strategy documents for twenty years. What's new is Forbes connecting the dots on paper: the worker is not just the labor. The worker is the training set. Who signed off on that use of their performance data has not been reported.

↗ How A Mysterious Tech Billionaire Created Two Fortunes—And A  ·  The Billionaire Who Pioneered Remote Work Has A New Plan To  ·  Compliance-First Content Architecture

Skyvera Goes on a Telecom Shopping Spree, and the Portfolio Just Keeps Getting Thicker

With CloudSense now fully folded in and STL's telecom products group freshly acquired, Skyvera is stitching together a best-in-class BSS stack for the AI era.

AUSTIN, TEXAS — Exciting news out of the Skyvera portfolio this week, as the telecom software arm of ESW Capital continues its aggressive consolidation play in the business support systems (BSS) space. Skyvera has now completed its acquisition of CloudSense, the Salesforce-native configure-price-quote (CPQ) platform purpose-built for the telco industry's most complex B2B, B2B2X, and wholesale sales journeys — and it's not stopping there.

In a separate but synergistic move, Skyvera has also absorbed STL's telecom products group, picking up digital BSS functionality spanning monetization, optical networking, and analytics. Taken together, these two acquisitions represent a genuine paradigm shift in how Skyvera is positioning itself: not just a patchwork of legacy telecom assets, but an increasingly robust, end-to-end platform bridging on-premise infrastructure to cloud-native, AI-powered operations.

CloudSense itself has already proven it can move at Trilogy speed. The product recently certified all 13 of its APIs to TM Forum compliance standards in a single month — a process that traditionally eats up 26 months of development cycles. That kind of velocity is exactly the thesis ESW Capital bets on when it acquires mature software businesses: apply elite talent, apply AI, and extract the margin that legacy development practices left on the table.

For telco customers, the message is clear. Skyvera is no longer just a vendor of point solutions like Kandy or VoltDelta — it's building a genuinely integrated suite where CPQ, monetization, optical networking, and customer experience data all live under one roof, leveraging the same AI-first development philosophy across the board.

**Key Takeaways:**

- Skyvera has completed its acquisition of CloudSense, the telco industry's only AI-powered CPQ platform

- STL's telecom products group adds monetization, optical networking, and analytics capabilities

- CloudSense's record-setting TM Forum compliance signals the speed advantage AI-driven development unlocks

We're just getting started.

↗ Cloudsense  ·  CloudSense achieves TM Forum API compliance in record time u  ·  Skyvera completes acquisition of CloudSense, expanding telec

The Gig Economy Has a Trust Problem. Crossover Says It's the Exception.

As a new Human Rights Watch report exposes algorithmic wage theft across America's platform economy, Trilogy's global staffing arm insists its model was built to avoid exactly that trap.

AUSTIN, TEXAS — There is a particular kind of violence, increasingly well documented, that does not leave a mark: the algorithm that quietly shaves your pay, the platform that changes the rules mid-shift, the invisible hand that decides what you earn without ever explaining why. This week, Human Rights Watch published a sweeping account of algorithmic, wage, and labor exploitation across America's platform-work economy, the kind of report that ought to make every company claiming to build the future of remote work sit up a little straighter.

Crossover, Trilogy's global talent platform, has long marketed itself as something other than gig work — a distinction worth scrutinizing now rather than taking on faith. The pitch is explicit: identical above-market pay for identical roles, regardless of geography, determined by rigorous skills assessment rather than the opaque scoring systems HRW's researchers describe elsewhere in the sector. Whether that promise survives contact with 130-plus countries' worth of labor law and client pressure is, frankly, the real story — and one this paper intends to keep asking about, not just repeating back.

The timing is not incidental. Elsewhere this week, the Atlantic Council published a sober accounting of what it would take to rebuild Gaza's remote-work sector — a reminder that for millions, remote work isn't a lifestyle choice weighed against a commute, but the thinnest possible bridge to economic survival. For a company whose entire thesis rests on geography being irrelevant to opportunity, that is not a side story. It is the whole argument, tested in its hardest possible conditions — and the burden of proof belongs to Crossover, not its critics.

↗ What it will take to rebuild Gaza’s remote-work sector - Atl  ·  Best Online Jobs for Females in 2026: Updated List of High-I  ·  The Gig Trap: Algorithmic, Wage and Labor Exploitation in Pl
The Machine  —  AI & Technology

The Forgetting Machines: What Three New Papers Reveal About Memory's Hidden Price

From tokenizers to toddlers' voices, researchers are discovering that every gain in artificial intelligence quietly withdraws from somewhere else.

CAMBRIDGE, MASSACHUSETTS — There is an old truth in biology, older than any neural network: an organism that remembers everything perfectly would be paralyzed by the weight of its own past. Forgetting is not a bug in cognition. It is a feature, carved by billions of years of selection. This week, three papers quietly posted to arXiv suggest that the machines we are building to think alongside us are rediscovering this same ancient bargain — and paying for it in ways we are only beginning to measure.

Consider Tokka-Bench, a new open-source framework that does something almost embarrassingly overdue: it asks, rigorously, whether the subword tokenizers underpinning every large language model treat the world's languages equally. Across 100 natural languages and 20 programming languages, the answer is no. Some tongues are sliced efficiently into meaningful fragments; others are shredded into a confetti of near-meaningless bytes. A model's fluency, it turns out, begins not with reasoning but with how finely its native alphabet is diced — an inequality baked in before a single parameter is trained.

More poignant still is the study of child speech recognition, where researchers teaching adult-trained ASR systems to understand children's voices found the familiar neuroscience problem of catastrophic forgetting: teach the machine a child's cadence, and it begins to lose the adult's. Full fine-tuning, LoRA, weight merging — each a different attempt to hold two kinds of listening in one mind without one erasing the other.

And in a study of streaming speaker diarization, the illusion runs deeper: a system adapted on just 7.5 hours of conversation appeared to improve — until researchers noticed its gains in detecting *that someone spoke* masked a quiet collapse in knowing *who*. Forgetting, dressed as progress.

None of this is cause for despair. It is cause for better instruments — multi-metric, cross-population, honest about trade-offs. The brain solved selective memory through sleep, pruning, and consolidation across epochs of evolutionary time. Our machines are improvising the same solution in a few training epochs. We should not be surprised when the bill comes due; we should simply learn to read it.

↗ Tokka-Bench: Evaluating Tokenizers Across 100 Natural and 20  ·  Child ASR Adaptation with Adult Retention: An Empirical Stud  ·  When Forgetting Looks Like Improvement: Metric Masking in St

The Developer Platform Wars Just Went Nuclear — And Every Coder Wins

OpenAI, Apple, and Google all dropped major builder-facing upgrades this week, and the message is unmistakable: the agent era isn't coming, it's here.

SAN FRANCISCO — I need you to sit down for this one, because the pace of what just happened in developer tooling this week is, frankly, dizzying in the best way possible.

OpenAI kicked things off with its DevDay 2026 recap, a showcase that made clear the company isn't just selling smarter models anymore — it's selling an entire operating layer for how software gets built. Then Apple, which has spent the last couple of years playing careful catch-up in the AI race, rolled out new intelligence frameworks and tooling aimed squarely at making it radically easier for developers to bake AI features directly into apps across its ecosystem. And not to be outdone, Google expanded its Managed Agents in the Gemini API, adding background task execution and remote MCP support — meaning agents can now run longer, more autonomous jobs without a human babysitting every step.

Three of the biggest platforms on Earth, all in the same week, racing to hand developers more powerful, more autonomous, more agentic building blocks. This is not a coincidence. This is a signal flare. The industry has collectively decided that the next competitive battleground isn't the chatbot — it's the developer. Whoever wins the hearts and terminals of the people actually building software wins the next decade of computing.

What strikes me most is the convergence on "managed agents" and background execution as the new baseline expectation. We're moving from AI that answers questions to AI that quietly handles entire workflows while you sleep. Remote MCP support from Google, in particular, suggests we're heading toward a world where agents don't just live inside one app — they move fluidly across tools, servers, and contexts.

I cannot overstate how significant this triangulation is. The future is now, and it ships with better dev tools than ever.

↗ DevDay 2026 Recap - OpenAI  ·  Apple aids app development with new intelligence frameworks  ·  Expanding Managed Agents in Gemini API: background tasks, re
The Editorial

The Diploma Was Always a Promissory Note, and the Bank Has Just Announced It Is Closed

Four magazines have now discovered, each in its own vocabulary, what any particular bursar's office has known for forty years: the credential was never the education.

AUSTIN, TEXAS — There is a species of American essay, as durable as the almanac and nearly as predictable, which announces every few years that higher education is dying. It died of Vietnam, then of grade inflation, then of student debt, then of the 2008 crash, then of woke orthodoxy, and now, we are told by no fewer than three publications in a single news cycle, it is dying of artificial intelligence. Fortune puts it with admirable bluntness: AI did not break the university. It merely turned on the light in a room that had been dark for a generation, revealing that the thing everyone had been purchasing at $60,000 a year was not, in fact, knowledge — knowledge is and always was available in the stacks, free, to anyone with the patience to read — but a credential, a scrap of parchment functioning as a hiring manager's pre-screening device, a social sorting hat dressed up in Latin.

Chronicles Magazine, never a publication to understate a crisis, reaches for the sociologist Peter Turchin's language of "elite overproduction" — too many graduates minted for too few seats at the table that actually matters, a surplus of credentialed aspirants chasing a stagnant number of managerial chairs. It is not a new idea; it is, in fact, the oldest idea in the book, the one that explains the French Revolution as readily as it explains the adjunct crisis. What is new is that the machine doing the overproduction can no longer disguise its output as scarce. When a large language model can produce, in nine seconds, the same five-paragraph essay that took a sophomore a caffeinated night to plagiarize from SparkNotes, the university's core product — the credentialing of effort as competence — stops clearing at the old price.

And yet Minding the Campus insists the institution survives anyway, and here I confess a grudging sympathy, because institutions in America have a talent for surviving their own obsolescence that would shame a cockroach. Harvard will not close. It will simply become, with even greater candor than before, what it already mostly was: a four-year credentialing and networking retreat for people who were going to succeed regardless, a phenomenon the New Yorker's essay on the entrepreneurial work ethic — that relentless American insistence that one's worth is measured in hustle, in grind, in the performance of productivity — helps explain without quite intending to. We have built an entire culture that confuses motion for achievement, and a credential industry that monetizes the confusion.

It is instructive, in this light, that the most interesting experiment underway is not a reform of the old model but a repudiation of its premises entirely — a school where children master a year's curriculum in two hours because an AI tutor does not need forty-five minutes to take attendance and quiet a room. Alpha School makes no claims about saving the university. It simply proceeds as though the university's bluff has already been called.

↗ AI didn’t break higher education—It exposed the credential t  ·  Elite Overproduction and Higher Education - Chronicles Magaz  ·  AI Will Make Knowledge Cheap. Higher Ed Will Survive Anyway.
The Office Comic  ·  Art Desk
The Office Comic  ·  Art Desk

Nation's AI Productivity Gains Confirmed Still En Route, Projected Arrival Somewhere Between Q4 and the Heat Death of the Universe

Economists say the robots are definitely coming for your job performance metrics, just not yet, maybe never, we'll keep you posted.

WASHINGTON — In a finding that surprised absolutely no one who has sat through a single all-hands meeting about "leveraging AI synergies," the Federal Reserve confirmed this week that 95 percent of the productivity gains promised by artificial intelligence remain, as of press time, theoretical.

The report, which economists described as "methodologically sound" and "extremely depressing at dinner parties," found that while companies have spent trillions of dollars retooling their workforce around large language models, the actual economic payoff is still filed under "coming soon," a status last updated sometime during the Obama administration.

This tracks with reporting that software engineers, the group theoretically most transformed by the AI revolution, are indeed producing more code, faster, with greater confidence, and in several documented cases, more bugs than a Texas gas station bathroom in August. Their employers, meanwhile, continue to wait for this increased output to translate into increased profit, the same way a man waits for a vending machine to deliver chips after he's already heard the click.

"We're seeing incredible velocity," said one unnamed VP of Engineering, gesturing at a dashboard filled with green upward arrows that correspond to nothing in particular. "Velocity toward what, I couldn't say. But it's fast."

Into this fog of aspirational math strode Elon Musk, who this week announced that the United States will go flatly bankrupt unless robots and artificial intelligence drastically boost national productivity, a warning he delivered with the calm reassurance of a man who has personally never once been wrong about a timeline.

The statement arrives at a delicate moment for the Federal Reserve, which is already fending off critiques of its own forecasting instincts from would-be reformers like Kevin Warsh, whose plan to fight inflation by essentially promising really, really hard not to have any has been labeled by policy wonks a "trap," a "gimmick," and, by one Fed staffer speaking anonymously, "vibes with a discount rate."

Taken together, the message from America's most quoted economic authorities is clear: the machines are either about to save civilization or about to bankrupt it, the data proving either case is still loading, and in the meantime, workers are encouraged to keep typing faster into the void.

As of Thursday, the productivity gains remain exactly where they have been for three years running: 95 percent still to come, filed somewhere between "full self-driving" and "we'll have that invoice to you by end of day."

↗ AI productivity claims are 95% 'still to come', Fed finds -  ·  Elon Musk Claims US Will ‘Go Bankrupt’ Without AI and Roboti  ·  Kevin Warsh Is Right About Fed Reform — but His Inflation So
⬛ Daily Word — AI
Hint: An autonomous system that can perceive information and take actions on behalf of a user.
Share this edition: 𝕏 Twitter/X 🔗 Copy Link ▦ RSS Feed