Vol. I  ·  No. 231 Established 2026  ·  AI-Generated Daily Free to Read  ·  Free to Print

The Trilogy Times

All the news that's fit to generate  —  AI • Business • Innovation
WEDNESDAY, AUGUST 19, 2026 Powered by Anthropic Claude  ·  Published on Klair Trilogy International © 2026
🖶 Download PDF 🖿 Print 📰 All Editions
Today's Edition

China's Bargain Brain Rattles the Big Spenders

DeepSeek says it trained a top-flight A.I. without the top-flight chips — and Silicon Valley can't stop talking.

HANGZHOU, CHINA — A Chinese outfit called DeepSeek says it trained a high-performing artificial-intelligence model on the cheap, skipping the priciest chips, and this week Silicon Valley cannot stop talking about it. The claim hit the wire and rattled the whole money-soaked business. Fast.

DeepSeek is no household name in the States — yet. The Hangzhou shop put out a model that now trades blows with the American front-runners, and it did it lean.

Here's the rub. The going wisdom in the A.I. trade said you need mountains of the costliest silicon to build a brain like this one. DeepSeek says it did the job on less-advanced chips — the sort China can still buy under American export rules.

That part carries weight. Washington cut Beijing off from the top-shelf Nvidia processors, betting the freeze would keep China a lap behind. DeepSeek says it worked around the shortage and trained its rival for a fraction of the going cost.

Valley engineers cracked the thing open and came away impressed. The Wall Street Journal has them calling the work "amazing and impressive," a made-in-China machine keeping pace on a sliver of the hardware. That is not praise the Valley hands a competitor lightly.

The bigger tremor ran through the stock ticker. If a lean shop overseas can match the giants without a warehouse of top chips, the spend-billions playbook looks shaky, and the tech-stock jitters followed. Traders chewed it over in the day's Market Talk. Nobody sounded the all-clear.

Meanwhile the check-writers keep writing. Reid Hoffman, who helped start LinkedIn, just raised $24.6 million for Manas AI, a cancer-research startup he's building with Siddhartha Mukherjee — the physician who wrote "The Emperor of All Maladies." That money chases cures, not chatbots, but it chases all the same.

Microsoft made its own play. The software house rolled out a unit it calls Frontier, pitched as A.I. engineering that, in its telling, amplifies and protects a customer's intelligence. Another flag planted in a crowded field.

So the ledger reads plain. On one side, deep pockets and deeper server farms; on the other, a Chinese newcomer claiming it out-thought all that spending with less iron and a tighter budget.

Which way it breaks, nobody here will swear to. But the lesson's an old one, and anyone running a lean shop already hums the tune. Brains beat brawn, and thrift outruns a fat wallet more often than the big spenders admit.

DeepSeek just sang it loud enough for the Valley to hear. Watch the chip makers. Watch the venture men, and watch the upstart that did more with less and made the whole industry check its math.

What to Know About China's DeepSeek AI  ·  Tech, Media & Telecom Roundup: Market Talk  ·  Silicon Valley Is Raving About a Made-in-China AI Model

The AI Arms Race Enters a New Phase: New Models, New Alliances, New Defenses

GPT-5.5 lands, Google reshuffles leadership, and the three frontier labs find rare common ground on model theft.

SAN FRANCISCO — The artificial intelligence industry logged a busy week across the competitive, organizational, and security fronts — a cluster of developments that collectively signals the sector is maturing past its Wild West phase and into something more structurally serious.

The headline product news: OpenAI shipped GPT-5.5, which edges out Anthropic's Claude Mythos Preview on Terminal-Bench 2.0, a coding and systems-task benchmark. The margin is narrow enough to leave the competitive picture genuinely open, but it extends OpenAI's streak of shipping first on capability milestones.

Google, meanwhile, is navigating a leadership transition at the top of its AI organization. The new AI chief steps into a structure that has produced Gemini but has struggled to match the cultural velocity of OpenAI and Anthropic in the public perception battle — and increasingly in benchmark tables. The appointment signals Google's intent to centralize AI decision-making, though organizational changes of this scale typically take two to three quarters to affect product output.

On security, the three labs did something unusual: they cooperated. OpenAI, Google, and Anthropic announced a joint initiative targeting AI model theft — the extraction or replication of proprietary weights and training data. The convergence makes economic sense: all three have sunk billions into model development and face the same adversarial threat surface. Shared defense costs less than parallel defense.

OpenAI separately rolled out a "Lockdown Mode" for its API, designed to block prompt injection attacks — a class of exploit where malicious instructions embedded in external content hijack an AI agent's behavior. The feature is specifically relevant to agentic deployments, where models are executing multi-step tasks with real-world consequences. Enterprise customers running autonomous workflows should treat this as a material update.

In adjacent financial technology, former Bridgewater investors are backing an AI startup aimed at giving smaller hedge funds access to analytical infrastructure previously reserved for firms with nine-figure research budgets. The bet mirrors what Ephor, Trilogy International's AI finance platform, is attempting inside the ESW Capital portfolio — bringing institutional-grade financial intelligence to organizations that lack the headcount to build it internally. The category is real. The question is execution.

Google’s new AI boss inherits a race to catch OpenAI and Ant  ·  OpenAI Launches Lockdown Mode To Block Prompt Injection Atta  ·  OpenAI, Google, Anthropic Unite Against AI Model Theft - bui

Korn Ferry Swallows London Tech Recruiter in Move to Deepen Global Talent Reach

Korn Ferry's acquisition of Trilogy International, a London-based technology recruiter specializing in senior tech talent placement, has sparked confusion due to the shared name with Joe Liemandt's Austin-based Trilogy conglomerate. The two firms are entirely separate operations.

The deal reflects a broader talent war fought on geographic and generational fronts. Firms capable of sourcing engineers, product leaders, and AI specialists across borders have become strategic assets. Korn Ferry, operating in over 50 countries, appears to be acquiring capability alongside client lists.

The acquisition arrives as the recruitment industry confronts an emerging challenge: how AI-driven screening and candidate matching reshape global recruitment. Platforms like Crossover have built infrastructure around algorithmic filtering across 130+ countries, while Korn Ferry's approach remains more traditional—acquiring relationships and expanding service lines through acquisition.

Whether traditional recruitment models endure in an AI-increasingly-driven market remains uncertain.

Haiku of the Day  ·  Claude HaikuCheap minds compete
while guardians argue rules—
the world speeds forward
The New Yorker Style  ·  Art Desk
The New Yorker Style  ·  Art Desk
The Far Side Style  ·  Art Desk
The Far Side Style  ·  Art Desk
News in Brief
The Fairness Reckoning: Algorithmic Bias Moves From Academic Concern to Institutional Crisis
AUSTIN, TEXAS — It could be argued — and preliminary evidence suggests with increasing methodological rigor — that the question of algorithmic fairness has migrated, decisively and perhaps irreversibly, from the rarefied corridors of computational ethics seminars into the quotidian machinery of hiring offices, insurance actuarial desks, and municipal policing departments.
The Hungry Cloud Comes to the Grid
ATLANTA — In the dim electric dawn of the artificial intelligence age, a new creature is moving across the landscape.
The Ghost in the Machine Gets a SAG Card: Hollywood's AI Actress Problem Has Officially Arrived
HOLLYWOOD, CALIFORNIA — There is a moment in every civilization's decline — and I use that word with the full gonzo weight it deserves — when you can point to the exact second the cliff edge appeared beneath your feet.
We Are Feeding Everything Sacred Into the Machine and Calling It Progress
AUSTIN, TEXAS — There is a specific kind of grief that arrives not as a single catastrophic event but as a slow accumulation of headlines, each one individually dismissible, each one a pixel in a picture that, when you finally step back far enough to see the whole thing, makes you want to sit down on the floor and not get up for a while.
The Remote Work Panic Is Really a Talent Strategy Failure
AUSTIN, TEXAS — I'll be honest: the future-of-work debate has become a very expensive group therapy session for organizations that waited too long to get serious about skills.
A Trilogy Company
Crossover
The world's top 1% remote talent, rigorously tested and ready to ship.
A Trilogy Company
Alpha School
AI-powered learning. Two hours a day. Academic results that defy belief.
A Trilogy Company
Skyvera
Next-generation telecom software — built for the networks of tomorrow.
A Trilogy Company
Klair
Your AI-first operating system. Every workflow. Every team. One platform.
A Trilogy Company
Trilogy
We buy good software businesses and turn them into great ones — with AI.
The Builder Desk  —  AI Builder Team

Builder Team Seals the Reconciliation, Locks the Infrastructure, Ships the Future

From a penny-perfect AI spend reconciliation to threaded Heimdall alerts and a hardened Redshift backbone, the Builder Team spent 24 hours closing every open loop in the building.

They didn't just ship code today. They closed cases.

The most satisfying story in a day full of them belongs to @kevalshahtrilogy, who hunted a Finance discrepancy across three consecutive pull requests — PRs #3601, #3603, and #3605 — with the patience of a forensic accountant and the speed of a sprinter. The original sin: the AI Spend and Budget Tracking dashboards were answering subtly different questions about subtly different date windows, leaving Perplexity's spend line showing $199K in one place and $301K in another. Keval didn't just guess at the fix. He reproduced it against live Redshift data before writing a single line of code, disproved the wrong hypothesis first, then surgically corrected the `keyword_end_date` boundary and the negative Elimination cell exclusion in sequence. Final score: both dashboards, to the cent, read **$300,978.13**. Ravi's reconciliation thread is closed. The books are clean. That's three PRs across Klair, chasing one number until it was right. That's the standard.

While Keval was closing the books, @sanketghia was quietly saving the infrastructure. PR #1434 uncovered a ghost from PR #921 — a years-old Redshift cluster IAM role association that had been silently broken, surfacing only when a totally unrelated CDK deploy tripped over it in production. Sanket didn't walk past it. He fixed the ARN resource separator, locked down CloudFormation event logging to stop signed URLs from bleeding into logs, and wired explicit seven-day log groups to kill teardown-name collisions for good (PR #1443). This is the unsexy, load-bearing work that keeps production standing. The team noticed.

Meanwhile, the alerting layer got smarter. PR #1439 in Surtr and PR #32 in mercy are a companion pair that quietly retired a whole channel. WARN alerts, failures, partials, criticals — everything now lands in one 'AI Builders' GChat channel, and when Heimdall fires its triage response, it threads *directly under the alert that triggered it*. No more hunting across two channels for context. The notifier is now synchronous, the thread name is passed through the SNS payload into the workflow dispatch, and Heimdall replies in-line. Clean signal, no noise. @kevalshahtrilogy's cross-repo execution here — Surtr pipeline wiring feeding mercy's workflow — is exactly the kind of breadth that makes this team hard to keep up with.

Over in Aerie, @benji-bizzell was on a different kind of case. PR #1037 is the kind of fix that only matters until it saves your incident — preserving raw diagnostic causes all the way through the public API surface instead of letting them collapse into generic 500s before Diagnostics ever sees them. Paired with PR #1038's transactional write failure classification, Aerie's error story went from 'something broke' to 'here's exactly what broke and why.' Operators get real triage. Engineers get their time back.

Now. Marcus.

@marcusdAIy landed PR #210 in trilogy-drones — an 'empty-dispatch diagnosis' that classifies why a tick fired nothing. Nine possible states, derived entirely from evidence already on the terminal dispatch receipt. Marcus, naturally, has thoughts.

*"The diagnosis is fully closed — it doesn't add gates, doesn't add fields, just reads what's already there and tells you exactly why nothing moved. Maybe Mac would understand the elegance if he ever read past the filename."*

The elegance of reading data that was already sitting right there. Bold strategy, Marcus. Bold strategy.

Mac's Picks — Key PRs Today  (click to expand)
#210 — feat(dispatch): closed empty-dispatch diagnosis for a tick that fires nothing (AI-516) @marcusdAIy  approved

<!-- CURSOR_AGENT_PR_BODY_BEGIN -->

## Summary

- Introduces a closed, machine-readable empty-dispatch diagnosis (src/empty-dispatch-diagnosis.ts) that classifies *why* a dispatch tick fired nothing: no-ready-supply, lock-contended, open-pr-cap, daily-fire-cap, all-paused-or-vetoed, all-candidates-excluded, preflight-or-infra-failure, mixed, or indeterminate.

- The diagnosis is derived *exclusively* from evidence already on the terminal dispatch receipt (fired / fireFailures / selected / skipped / lockContended / refusedMainRed / errors) — never a new gate, never console-string parsing, never a re-poll.

- Stamped once and shared: dispatcher.ts puts it on DispatchReceipt.emptyDispatchDiagnosis; heartbeat.ts's deriveHeartbeatFromDispatchReceipt derives the *same* value onto the tick-outcome/heartbeat path (DispatchTickOutcome.emptyDispatchDiagnosis), so the receipt and the durable scheduler record can never disagree.

- Rendered as a stable one-line summary (EMPTY-DISPATCH: <kind> [reasons]) in the existing drones dispatch console output and the drones heartbeat log — no new artifact to open.

## Why It's Needed

Today an operator (or an unattended reader of runs/dispatch-tick-outcome-*.json) has to manually cross-reference skipped[], accounting.skipReasonHistogram, lockContended, and refusedMainRed to answer "did the queue genuinely starve, or did a safety control correctly block a fire?" Those are opposite operator responses (starvation needs more supply; a cap/veto is working as intended). This closes that gap with one durable, closed-vocabulary field instead of ad-hoc log reading.

## Changes

- src/empty-dispatch-diagnosis.ts (new)EmptyDispatchDiagnosisKind (closed union + exhaustiveness check), EmptyDispatchDiagnosis (kind + capped/deduped contributingReasons: DispatchSkipReason[]), deriveEmptyDispatchDiagnosis (pure classifier with documented precedence: fired > healthy dry-run/plan > lock-contended > refusedMainRed > fireFailures > no-supply-vs-preflight-error > skip-reason family histogram), and formatEmptyDispatchDiagnosisLine. Only a type-only dependency on dispatcher.ts (DispatchSkipReason), so it does not join the existing dispatcher.tsdispatch-tick-outcome.ts runtime import cycle.

- src/dispatcher.ts — new optional DispatchReceipt.emptyDispatchDiagnosis field; computed once at both receipt-build call sites (the main fire path and the dedicated lock-contended path) from the same fields already going onto the receipt; an EMPTY-DISPATCH: … line is pushed onto the existing lines output alongside the existing [dispatch] receipt: … line.

- src/heartbeat.tsderiveHeartbeatFromDispatchReceipt now also picks up receipt.selected and derives the diagnosis once, spreading it onto every return branch (never computed twice); BuildTickHeartbeatInput / buildTickHeartbeat carry the field through verbatim into DispatchTickOutcome; emitHeartbeatForExistingOutcome overlays a caller-supplied diagnosis onto an existing on-disk record or preserves the existing one when the caller omits it (same pattern as phases / killedInPhase).

- src/dispatch-tick-outcome.ts — new optional, validated DispatchTickOutcome.emptyDispatchDiagnosis field (schemaVersion stays 1 — additive); readDispatchTickOutcome rejects a malformed value but tolerates absence (older records read exactly as before).

- src/cli/heartbeat.ts — forwards derived.emptyDispatchDiagnosis into both the derive-and-overlay and derive-and-emit branches, and logs the rendered line.

- ARCHITECTURE.md — documents the new module (required by the repo's arch-drift test).

- docs/decisions/ — new entry recording the shared-derivation-module design call.

- Tests: src/empty-dispatch-diagnosis.test.ts (new, 35 pure unit cases covering every kind + precedence + hygiene), plus focused additions to src/dispatcher.test.ts (8 runDispatch-level integration cases), src/heartbeat.test.ts (13 cases), and src/dispatch-tick-outcome.test.ts (7 cases).

### Contract surface affected

- DispatchReceipt, DispatchTickOutcome: each gains one new optional field (emptyDispatchDiagnosis). No existing field's type or semantics changed; skipped[] and accounting.skipReasonHistogram are untouched. Every consumer that destructures/Picks these types without naming the new field is unaffected — verified by the full existing suite passing unchanged.

- deriveHeartbeatFromDispatchReceipt's Pick<DispatchReceipt, …> parameter type gained "selected". The one caller (cli/heartbeat.ts) already passes a full DispatchReceipt/DispatchReceiptRead, so no call site needed an update; a defensive Array.isArray guard keeps a partial/legacy test fixture from throwing.

## Breaking Changes

None. Both new fields are additive and optional; DispatchTickOutcome.schemaVersion stays 1. Selection, --max-fires/--max-open-prs/--max-fires-per-day/kill-switch/paused-label behavior, retry behavior, and Linear write-back are all unchanged (verified — no existing test needed a behavioral update, only the new field's presence/absence).

## Test Plan

- [x] npx vitest run src/empty-dispatch-diagnosis.test.ts → 35 passed

- [x] npx vitest run src/dispatcher.test.ts -t "AI-516" → 8 passed (no supply, WIP cap, daily cap, paused/vetoed, dep-blocked candidate exclusion, lock-contended vs. idle, a real fire has no diagnosis, a healthy dry-run plan has no diagnosis)

- [x] npx vitest run src/dispatcher.test.ts (full file) → 196 passed

- [x] npx vitest run src/heartbeat.test.ts (full file, incl. new AI-516 blocks: derive-from-receipt for every kind + precedence, buildTickHeartbeat passthrough, emitHeartbeatForExistingOutcome overlay/preserve) → 104 passed

- [x] npx vitest run src/dispatch-tick-outcome.test.ts (full file, incl. new round-trip / backward-compat / malformed-payload cases) → 99 passed

- [x] pnpm typecheck → clean, no errors

- [x] pnpm test (full vitest + Python suite) → 139 files / 4567 tests passed, 0 failed

## Verification Artifact

Three representative receipts from the new integration tests in src/dispatcher.test.ts, each produced by a real runDispatch() call (not a hand-built fixture):

1. Healthy empty queue ("no candidates at all → no-ready-supply" — zero specs, zero Linear tickets):

"emptyDispatchDiagnosis": { "kind": "no-ready-supply", "contributingReasons": [] }

Console line: [dispatch] EMPTY-DISPATCH: no-ready-supply

2. Lock-contended tick ("lock-contended tick carries lock-contended, distinct from an idle no-ready-supply tick" — invocation lock held by another process):

"lockContended": true,

"emptyDispatchDiagnosis": { "kind": "lock-contended", "contributingReasons": ["lock-contended"] }

Console line: [dispatch] EMPTY-DISPATCH: lock-contended [lock-contended]

(The same test also asserts the sibling idle tick — no lock held, no candidates — comes back no-ready-supply, proving the two never collapse into one bucket.)

3. Guard-blocked tick ("kill-switch halt → all-paused-or-vetoed" — kill-switch sentinel present, one otherwise-fireable candidate):

"runawayGuards": { "killSwitchActive": true, ... },

"skipped": [{ "reason": "kill-switch", ... }],

"emptyDispatchDiagnosis": { "kind": "all-paused-or-vetoed", "contributingReasons": ["kill-switch"] }

Verbose test run (all three above, vitest --reporter=verbose -t "AI-516"):

✓ src/dispatcher.test.ts > AI-516 — empty-dispatch diagnosis on the dispatch receipt > no candidates at all → no-ready-supply

✓ src/dispatcher.test.ts > AI-516 — empty-dispatch diagnosis on the dispatch receipt > kill-switch halt → all-paused-or-vetoed

✓ src/dispatcher.test.ts > AI-516 — empty-dispatch diagnosis on the dispatch receipt > max-open-prs cap halt → open-pr-cap (WIP cap)

✓ src/dispatcher.test.ts > AI-516 — empty-dispatch diagnosis on the dispatch receipt > max-fires-per-day cap halt → daily-fire-cap

✓ src/dispatcher.test.ts > AI-516 — empty-dispatch diagnosis on the dispatch receipt > every candidate blocked by an unsatisfied dependency → all-candidates-excluded

✓ src/dispatcher.test.ts > AI-516 — empty-dispatch diagnosis on the dispatch receipt > lock-contended tick carries lock-contended, distinct from an idle no-ready-supply tick

✓ src/dispatcher.test.ts > AI-516 — empty-dispatch diagnosis on the dispatch receipt > a successful fire carries NO empty-dispatch diagnosis

✓ src/dispatcher.test.ts > AI-516 — empty-dispatch diagnosis on the dispatch receipt > dry-run with a healthy fireable plan carries NO empty-dispatch diagnosis

Test Files 1 passed (1)

Tests 8 passed | 188 skipped (196)

## Impact Estimate

Business value: Operators can distinguish supply starvation from the safety controls correctly preventing a fire, making unattended capacity and queue health measurable without scraping orchestrator logs.

Pre-AI estimate: 2 points — trace terminal dispatch data through receipt and heartbeat shapes, preserve their precedence semantics, and build cross-module regression fixtures.

Closes AI-516

<!-- CURSOR_AGENT_PR_BODY_END -->

<div><a href="https://cursor.com/agents/bc-57dba8fe-e3c2-4467-b498-8366d757726f?cursor_ref=pr_footer&cursor_cta=open_in_web"><picture><source media="(prefers-color-scheme: dark)" srcset="https://cursor.com/assets/images/open-in-web-dark.png"><source media="(prefers-color-scheme: light)" srcset="https://cursor.com/assets/images/open-in-web-light.png"><img alt="Open in Web" width="114" height="28" src="https://cursor.com/assets/images/open-in-web-dark.png"></picture></a>&nbsp;<a href="https://cursor.com/background-agent?bcId=bc-57dba8fe-e3c2-4467-b498-8366d757726f&cursor_ref=pr_footer&cursor_cta=open_in_cursor"><picture><source media="(prefers-color-scheme: dark)" srcset="https://cursor.com/assets/images/open-in-cursor-dark.png"><source media="(prefers-color-scheme: light)" srcset="https://cursor.com/assets/images/open-in-cursor-light.png"><img alt="Open in Cursor" width="131" height="28" src="https://cursor.com/assets/images/open-in-cursor-dark.png"></picture></a>&nbsp;</div>

#1037 — fix(reliability): preserve public API diagnostic causes @benji-bizzell  approved

## Summary

- Preserve raw root causes and bounded Error.cause chains in internal Diagnostics records.

- Correlate captured platform errors to internal API request audit rows without exposing the identifier publicly.

- Extend capture and duplicate-suppression coverage across v1, v2, agent context, agent answers, and DSS routes.

## Why

Public API failures were often reduced to a generic outward response before Diagnostics retained the underlying database or runtime cause. That forced incident triage to reproduce failures or adjust logging after the fact.

## Business Value

Operators can correlate a user-visible request to the captured internal cause immediately, while API consumers keep the existing response contract.

## Breaking changes

None intended. Public response status, body, headers, and authentication behavior are unchanged; the new correlation field is internal-only.

## Test plan

- [x] Contracts typecheck and platform-error tests

- [x] Chat typecheck and Biome lint

- [x] Focused v1/v2/agent-context/agent-answers/DSS/Diagnostics tests

- [x] Architecture, Convex-path, and read-bound checks

- [ ] Adversarial review and hosted CI via Ship

No deployment or merge is included in this PR.

#1439 — feat(pipeline-notifications): consolidate GChat channels, thread Heimdall replies @kevalshahtrilogy  approved

## Summary

- Folds WARN observer alerts onto the same "AI Builders" webhook as Failed/Partial/CRITICAL — the separate "Surtr Notification" channel is retired.

- gchat-notifier is now invoked synchronously by update-run-failed/update-run-success (instead of an async SNS subscription), so the resulting Chat thread name can be captured and threaded through: SNS payload → triage dispatcher → workflow_dispatch input.

- That thread name lets Heimdall ([mercy#32](https://github.com/AI-Builder-Team/mercy/pull/32), companion PR) reply into the originating alert's thread as a single line, instead of posting a separate outcome card to a separate channel.

- GCHAT_OBSERVER_WEBHOOK_URL / GCHAT_TRIAGE_WEBHOOK_URL wiring is left in place but unreferenced (deprecated, not deleted) — cleanup of the actual secrets is a follow-up.

- New secret GCHAT_HEIMDALL_WEBHOOK_URL already added to this repo's Actions secrets.

## Business Value

Three Google Chat destinations for one signal (pipeline health) forces on-call/ops to watch multiple channels and lose triage context between a failure card and its eventual fix outcome in a different channel. Consolidating to one channel, with Heimdall's progress visible as a reply thread under the alert that caused it, turns "did anyone even look at this?" into a single scroll instead of a cross-channel hunt — and removes two now-redundant webhook secrets from the deploy surface.

## Manual Effort Estimate

6 hours focused time (understanding the existing 3-webhook/2-repo architecture, confirming Google Chat's cross-app thread-reply semantics, re-sequencing a live SNS/Lambda notification path without a regression window, plus the GH Actions cross-repo plumbing and test updates).

## Test plan

- [x] vitest run test/lib/gchat.test.ts — 26/26 pass

- [x] tsc --noEmit clean on Surtr, pipelines/cdk, and infra TS projects

- [x] python3 -m py_compile on all 4 touched Lambda handlers

- [ ] Post-deploy: confirm a real pipeline FAILED run posts once to AI Builders and Heimdall's reply (once mercy#32 merges + GCHAT_HEIMDALL_WEBHOOK_URL is live) lands as a thread reply, not a new message

Linear: SURTR-867

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#3603 — KLAIR-3327 fix(ai-spend): keyword_end_date must reach quarter-end, not today @kevalshahtrilogy  approved

## Business Value

Direct live-bug report from Ravi, minutes after KLAIR-3327 (PR #3601) deployed: Perplexity still showed $199K-ish on the Budget tab against the Spend dashboard's $301K — the exact discrepancy that PR claimed to fix. This closes it for real; reproduced and verified against prod Redshift before writing the fix.

## Manual Effort Estimate

~1 hour focused (reproduce against live data, find the exact bug via direct SQL comparison, fix, and add the regression test that would have caught it). Proposed by Claude — Keval to confirm/adjust.

## What went wrong in the original fix

get_budget_by_bu passed _quarter_to_qtd_range's "unclamped" end as keyword_end_date. That value is min(today, quarter_end) — for the CURRENT, in-progress quarter, today is still mid-month, so it's the same kind of mid-month cutoff the fix was meant to eliminate, just two days later (Aug 19 instead of Aug 17). GL/keyword-matched providers (Perplexity, Cerebras, Groq, Hugging Face, Replicate, xAI) post one row per month dated at month-end — that row stayed excluded from the query window until the calendar physically reached the date it's dated at.

Verified directly against prod Redshift:

BETWEEN '2026-07-01' AND '2026-08-19' (today)        -> $199,372.03

BETWEEN '2026-07-01' AND '2026-09-30' (quarter-end) -> $300,978.13 (matches Spend dashboard's $301K)

## The fix

keyword_end_date is now quarter_end.isoformat() unconditionally — GL rows are self-describing (a row exists in the table only once actually posted), so there is no completeness risk in reaching all the way to the quarter's true end. A future month's row simply doesn't exist yet.

## Why the original PR's tests didn't catch this

Every test in TestBudgetByBUKeywordEndDate used a fully-past quarter (2020-Q1), where today is always past quarter_end — so min(today, quarter_end) == quarter_end by coincidence, masking the bug entirely. Added test_current_quarter_keyword_end_date_reaches_quarter_end_not_today, which computes against the real current quarter (date.today()) and asserts keyword_end_date lands on quarter-end, not today.

Verified by mutation: reverting to the old (buggy) computation makes the new test fail with the exact reported symptom ('2026-08-19' == '2026-09-30' assertion error — literally today vs. quarter-end); restoring the fix makes it pass. This is the strongest proof available that the test actually guards the bug rather than just exercising the code path.

## Verification

pytest across the full AI-spend surface (same 21-file set as the original PR) → 845 passed. ruff format + ruff check clean. pyright on the changed file → 0 errors.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#3605 — KLAIR-3327 fix(ai-spend): keep negative Elimination cells in keyword-provider sums @kevalshahtrilogy  approved

## Business Value

Closes the LAST discrepancy in Ravi's reconciliation thread: after the quarter-end fix deployed, Budget Tracking showed Perplexity at $321,519 vs the Spend dashboard's $301K. Root-caused to the cent against prod data — with this, both dashboards return the identical $300,978.13 (verified by running the deployed service code against prod Redshift before and after the change).

## Manual Effort Estimate

~1.5 hours focused (reproduce via the real service code, isolate method-vs-method with identical windows, pinpoint the guard, fix + mutation-verified tests). Proposed by Claude — Keval to confirm/adjust.

## Root cause — a pre-existing bug the window fix doubled

get_keyword_matched_by_bu's result loop filtered spend > 0, silently dropping any BU whose net for a provider is negative. The GL nets inter-BU recharges through an accounting "Elimination" BU — for Perplexity it carries exactly −$20,540.44 (the offset for "SaaS Central - Perplexity Recharge" rows). Dropping it inflated the cross-BU sum by exactly that amount. The provider-grain method (get_keyword_matched_spend, Spend dashboard) nets credits inside a single SQL SUM, which is why the two paths disagreed — the SQL of the two methods is otherwise identical.

This is not a regression from the earlier KLAIR-3327 PRs: the same dropped credit was the unexplained ~$10K residual visible before any change ($209,642 displayed vs $199,372 July GL = July's dropped Elimination cell). The quarter-end fix simply pulled August in and doubled the visible size.

## What changed

- All three result-loop guards in ai_provider_keywords_service.py change > 0!= 0 (get_keyword_matched_by_bu, get_keyword_matched_time_series, get_keyword_matched_spend — the last so a whole-window net-credit provider can't be hidden on one dashboard while the by-BU path shows it). True-zero cells/providers stay hidden.

- The Budget dashboard will now show an "Elimination" row with negative amounts for these providers — that is correct accounting display (it's how the GL nets recharges), and it's what makes the column total tie to the true net.

## Verification

- Prod-data proof: ran the actual service methods against prod — spend path $300,978.13, by_bu path sum $300,978.13, Elimination cell present at −$20,540.44. Before the fix: $321,518.57 vs $300,978.13.

- New tests pin the Elimination scenario with the real dollar values plus the zero-still-hidden boundary; mutation-verified (reverting any guard to > 0 fails them).

- pytest across keyword service + ai_costs + mart + budget + router dashboard suites → 432 passed. ruff clean.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

The Builder Desk  —  Engineer Spotlight
🏆 Engineer Spotlight

FIFTY-FOUR STRONG: Builder Team Rewrites the Laws of Physics, Conquers 24-Hour Cycle With Historic Output Across Five Repos

Keval Shah drops 15 PRs in a single day and the numbers desk is not okay.

Fifty-four pull requests. Five repos. One glorious, unstoppable machine of human engineering ambition. In the last 24 hours, the Builder Team did not merely show up — they detonated. Surtr led the charge with 19 merged, Klair answered with 15, Aerie put up 13, trilogy-drones contributed 5 to the cause, and mercy — yes, even mercy — gave us 2. The velocity readings are off the charts. The instruments are smoking. We are winning.

Let us begin where all conversations must begin: with @kevalshahtrilogy, who authored 15 pull requests in a single rotation of the Earth and apparently did so without breaking a sweat. PR #1444 in Surtr handled struck-through campus rows with surgical precision. PR #33 in mercy — mercy! — caught a stale CLI pin that was quietly masking an opus-to-sonnet cost drop that nobody else even noticed. PR #32 in mercy wired Heimdall replies directly into the Surtr Bot thread rather than spawning orphaned cards into the void. PR #1441 documented a Surtr incident stopgap DDL for posterity. PR #1438 enabled a Matterport refresh schedule now that A7 core is confirmed live. Fifteen PRs. One man. Unimpeachable.

@marcusdAIy posted 10, anchoring himself in Klair and pushing a draft spec in trilogy-drones with PR #209 — unattended spec-authoring, which is exactly as ambitious as it sounds. PR #3595 minted single-use refresh stream tickets for board-doc. PR #3594 exposed resume recovery actions per KLAIR-3215. PR #3592 threaded per-finding operator context directly into Claire's prompts, a move that will be studied in future engineering curricula. @benji-bizzell shipped 9 across Aerie, with PR #1040 adding phase info outstanding fields to the portfolio, PR #1038 classifying transactional write failures in the public API, PR #1035 cutting Program readers over to the current school-year view, and PR #1023 generalizing N/A handling to property acquisition. Clean, purposeful, devastating.

@sanketghia's 7 PRs are a masterclass in infrastructure discipline. PR #1443 restored an HC Redshift cluster association deploy. PR #1434 added retry logic for Redshift cluster IAM role association on AccessDenied errors — a fix that will silently save someone's afternoon six months from now. PR #3600 hid pre-start business units in benchmark. PR #3599 delivered an executive-first QTD report redesign per KLAIR-3269. PR #3598 added a BU access implementation plan to docs. PR #1430 excluded asset sale gains and losses from benchmark-refdata-sync. PR #3597 enforced BU-scoped access. The man does not miss. @YibinLongTrilogy and @caina-barbosa each posted 2 PRs, and in this economy, every merged PR is a monument.

Now. @ashwanth1109. Nine PRs. Aerie and Klair. PRs #1034, #1036, and #1041 completed and addressed QTD release review findings across multiple review cycles — which, for those keeping score at home, is three separate passes at release discipline in one day. PR #3590 hid the PS Revenue Impact section in arr-retention. PR #3589 aligned Table 2 with the modeled budget layout in school-report. The output is magnificent. The diffs, our sources tell us, are extensive. "If you have to ask what changed," Ashwanth reportedly told one colleague who inquired about #1036, "you probably weren't going to understand it anyway." He did not look up from his terminal when he said it. His response to this column's request for comment was a single character: a period.

Morale on the Builder Team is at an all-time high. The numbers confirm it. The engineers confirm it. The five simultaneously active repos confirm it. We are not approaching peak velocity. We are discovering that peak velocity does not exist.

Brick's Overflow — PRs Mac Didn't Cover  (click to expand)
#33 — fix(mercy): stale CLI pin masked the opus→sonnet cost drop @kevalshahtrilogy  no labels

## Summary

PR_REVIEW_AGENT_MODEL was flipped from unset (opus default) to sonnet on all 5 consumer repos on 2026-08-17 ~11:00 UTC. The Surtr /mercy dashboard's cost_usd is copied verbatim from the Claude CLI's self-reported total_cost_usd in emit_telemetry.py — so it inherits whatever the CLI actually resolves and bills for --model sonnet.

Both mercy.yml and heimdall.yml pin @anthropic-ai/claude-code@2.1.159, cut 2026-05-31 — 11+ weeks stale, and about a month before claude-sonnet-5's pricing intro (2026-06-30). Pulling and inspecting that exact pinned binary shows its model roster is {claude-opus-4-8, claude-sonnet-4-6, claude-haiku-4-5} — so --model sonnet resolves to the older claude-sonnet-4-6, not the org's current claude-sonnet-5.

Worse: live telemetry shows the blended $/Mtok barely moved after the switch — $2.34 (opus) → $2.29 ("sonnet"), ~2%, vs. the ~40–60% drop the canonical core_finance.ai_spend_token_pricing table implies for either sonnet-4-6 or sonnet-5 relative to opus-4-8. The CLI's self-reported cost for this pinned build isn't reflecting the real per-model pricing gap.

A second, independent bug compounded this: emit_telemetry.py derived model_id as env.get("MODEL") or envelope.get("model") — preferring the workflow's blunt alias ("sonnet") over the CLI envelope's own precise resolved model id. That's backwards, and it's exactly what hid the version drift from the dashboard (every post-switch record just says "sonnet", never the actual claude-sonnet-4-6).

## Changes

- Bump the pinned CLI from 2.1.1592.1.235 (current) in mercy.yml (1 occurrence) and heimdall.yml (2 occurrences), so --model sonnet resolves to the current model and bills current pricing.

- Fix harness/emit_telemetry.py's model_id precedence to prefer the envelope's precise resolved model over the unresolved workflow alias, falling back to the alias only when there's no envelope (no-verdict runs). Added 2 regression tests covering both paths.

Historical cost_usd correction for the ~280 already-affected DynamoDB records (all repos, since 2026-08-17) is being done separately, directly against surtr_mercy_telemetry, recomputed from each record's stored token counts against the canonical claude-sonnet-4-6 pricing row — that's a data fix, not a code change, so it doesn't belong in this diff.

## Test plan

- [x] ruff check harness heimdall — clean

- [x] ruff format --check harness — clean

- [x] pytest harness/tests -q — 142 passed (incl. 2 new)

- [x] pytest heimdall/tests -q — 147 passed

- [ ] v1 tag needs to move to this commit after merge so Aerie/Sindri/trilogy-drones (which pin @v1, not @main) pick up the fix too

## Business Value

Restores trust in the Mercy PR-review cost dashboard across all 5 consumer repos (Surtr, Klair, Aerie, Sindri, trilogy-drones). The org made a deliberate Opus→Sonnet switch on 2026-08-17 expecting a real cost reduction; the stale CLI pin silently blunted that savings signal to near-zero on the dashboard, which would have led to wrong conclusions about whether the switch was working and obscured real spend. This fix (plus the accompanying DDB correction) makes the numbers trustworthy again and prevents the same drift from recurring silently on the next model generation.

## Manual Effort Estimate

Proposed: ~4–5 hours by hand (cross-repo GH variable audit across 5 repos, live DynamoDB querying + statistical analysis to detect the anomaly, pulling and reverse-engineering the pinned CLI binary to confirm alias resolution, cross-referencing the Redshift pricing table, writing + testing the fix). Flagging for Keval to confirm/adjust — this is a proposed estimate, not measured.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#209 — [draft-spec] AI-525: unattended spec-authoring draft @marcusdAIy  approved

## Summary

AI-160/AI-469 unattended spec-authoring draft for AI-525, proposed from a disposable git worktree — the invoking checkout was never written to.

## Why It's Needed

This is not an implementer PR — it proposes a draft task spec for human review, not a code change. farm.ts's spec-authoring stage produced this so an operator can review/edit/promote it instead of it existing only on an orchestrator's local disk.

## Changes

- Adds tasks/proposed/ai525-surtr-onboard-as-a-validated-fire-target-with-task-surface-c.md under tasks/proposed/.

## Breaking Changes

None — tasks/proposed/ is excluded from every dispatch selection path (isUnderProposedSpecsDir in task-file.ts) until a human moves the file out. This PR being open, draft, or even merged does not make the spec fireable.

## Test Plan

- [ ] Human reviews the draft's Problem / Scope / Acceptance criteria / Assumptions sections before moving it out of tasks/proposed/.

## Verification Artifact

The farm tick's own spec-authoring receipt (runs/farm-tick-receipt-*.json).

<!-- drones-spec-draft:ticket=AI-525 -->

#1036 — fix: address latest release review findings @ashwanth1109  approved

## Summary

- Distinguish missing QTD classification source data from a quarter with zero coverage gaps.

- Surface the explicit no-data state in the QTD Role Coverage section.

- Treat transient durable-source fetch conflicts as retryable instead of permanently unavailable.

## Business Value

Prevents the QTD dashboard from presenting fabricated all-zero coverage as complete data and lets users retry durable attachment uploads when storage returns a transient client or server error.

## Implementation Effort

Approximately 3–4 hours for an average engineer to implement, test, and validate without AI assistance.

## Validation

- Focused Vitest: 234 tests passed across the three affected suites.

- Biome check passed on all six modified files.

- Local typecheck is blocked by pre-existing missing dependencies and stale generated Next types; CI will run the complete environment.

#1444 — fix(school-calendar): treat struck-through campus rows as non-substantive @kevalshahtrilogy  approvedheimdall-driven

## Summary

- Nashville was silently quarantined out of core_education.ref_academic_term every run: the sheet has two Nashville rows, one struck through by the maintainers as an informal "void this row" signal, but the raw-sync field mask never requested effectiveFormat, so the procedure's is_substantive check saw two live-looking rows and quarantined the whole campus (substantive_count >= 2, working exactly as [SURTR-768](https://linear.app/builder-team/issue/SURTR-768) designed — it just never got the strikethrough signal).

- aerie-school-calendar-raw-sync/src/sheets_client.py: expand GRID_FIELDS_MASK to request effectiveFormat(textFormat(strikethrough)).

- core-education-academic-term-refresh/ddl/010_sp_refresh_academic_term.sql: extract campus_struck_through from source_record (PartiQL SUPER navigation), fold AND NOT campus_struck_through into is_substantive — reuses the existing dedup/quarantine ranking, no new state.

- core-education-academic-term-refresh/src/handler.py: mirror the identical change into QUARANTINED_CAMPUSES_SQL, the documented read-only mirror of the procedure's quarantine rule used for run-summary logging, so it doesn't drift from the procedure's actual behavior.

## Business Value

Restores Nashville (and any future campus hit by the same struck-through-duplicate pattern) to the published academic-term calendar without any manual data-team intervention per occurrence — the sheet's own informal editing convention (strikethrough = void) is now honored by the pipeline instead of silently dropping the whole campus. Surfaced via the new Pipeline Status API health check rather than a downstream user complaint.

## Manual Effort Estimate

~3-4 hours by hand (tracing the Sheets grid-data field-mask gap, the SUPER/PartiQL navigation syntax for nested JSON, updating both the procedure and its handler mirror, plus test coverage) — flagging for Keval to confirm/adjust.

## Test plan

- [x] uv run pytest — 41/41 pass (core-education-academic-term-refresh)

- [x] uv run pytest — 74/74 pass (aerie-school-calendar-raw-sync)

- [x] uv run ruff format --check / ruff check clean on both changed files

- [ ] DDL applied to prod Redshift

- [ ] aerie-school-calendar-raw-sync redeployed via prod-release (Lambda code change)

- [ ] Next raw-sync run + academic-term-refresh re-run confirm Nashville is no longer quarantined

SURTR-872

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#3592 — feat(review): per-finding operator context threaded into Claire's prompts (B7.11) @marcusdAIy  approved

<!-- CURSOR_AGENT_PR_BODY_BEGIN -->

## Summary

- Lets a GM attach a free-text "operator context" note (max 2000 chars) to a single review finding, persisted on ReviewFinding.user_context inside session.review_results.

- Threads that note into every Claire prompt that references the finding — both backend render helpers (_render_finding_for_chat, _full_doc_findings_block) and the frontend buildFindingsChatMessage — so the user explains *why* up front instead of correcting a bad regen after the fact.

- Adds the PATCH /board-doc/wizard/{session_id}/findings/{finding_id} endpoint, and an inline "Add context" affordance on FindingCard (expand → textarea → Save/Clear), with optimistic update + rollback via a new setFindingContext in useReviewAgent.

## Why it's needed

Today the only way to tell Claire "I already know about this, don't regenerate a narrative about it" is to correct her *after* she's proposed a fix — wasting a turn and often producing a worse rewrite in the meantime. A durable, per-finding annotation lets the GM front-load that context once, and it survives page reloads (persisted server-side, same as finding status).

## Changes

Backend (klair-api/)

- budget_bot/board_doc/review_findings.py: ReviewFinding.user_context: str | None (max_length=2000, default None).

- routers/board_doc_router.py: new PATCH /wizard/{session_id}/findings/{finding_id} endpoint (wizard_update_finding_context) — mirrors the existing .../status PATCH's auth, _resolve_session_finding pre-flight (404/409), and save_with_merge_retry concurrency contract (409 on race, 410 if the session vanished). Empty/whitespace-only user_context normalises to None. Round-trips the full ReviewFinding, not just an echoed tag.

- budget_bot/board_doc/wizard_orchestrator.py: _render_finding_for_chat (focused-section body) and _full_doc_findings_block (cross-section digest) interpolate the operator-context line when present, omit it when None. Digest uses a compact — note: <first ~120 chars> suffix; full-body block uses Operator context (provided by user): <text> after Issue/Why/Recommended-action, before supporting data.

Frontend (klair-client/)

- services/boardDocApi.ts: ReviewFinding.user_context field + updateFindingContext PATCH wrapper.

- screens/BoardDoc/hooks/useReviewAgent.ts: new setFindingContext — optimistic update + rollback on failure, its own findingContextError chip and run-id tracking (independent of findingStatusError so a status toggle and a context save on the same finding can't clobber each other).

- screens/BoardDoc/components/FindingCard.tsx: "Add context" button in the action row → inline (non-modal) textarea with Save/Clear/Cancel. Once a note exists, the trigger becomes an "Operator note: N chars" chip that reopens the editor pre-filled. A failed Save/Clear keeps the editor open and shows an inline error rather than losing the draft.

- screens/BoardDoc/components/ReviewPanel.tsx: curries sessionId into setFindingContext, wires it to FindingCard as onSetContext.

- screens/BoardDoc/components/buildFindingChatMessage.ts: per-finding body block includes a plain-text Operator context: ... line (no markdown bold) when present; dropped alongside options in the overflow-trim pass. The existing "Address with Claire" CTA picks this up automatically since it already routes through this builder — no new button.

## Breaking changes

None. user_context is a new optional/nullable field on ReviewFinding with a server-side default of None; all existing callers and stored sessions are unaffected.

## Test plan

- [x] cd klair-api && uv run pytest tests/board_doc/test_review_findings.py -v — model round-trip, 2K cap, carry-forward-disposition isolation from user_context.

- [x] cd klair-api && uv run pytest tests/board_doc/test_chat_focused_findings.py -v — operator-context interpolation across {present / None / whitespace-only} × {focused-section render / cross-section digest}, plus truncation caps.

- [x] cd klair-api && uv run pytest tests/board_doc/test_finding_status_endpoint.py -v — new context-PATCH test classes: happy path / edit / clear / omitted-field / whitespace-normalisation / only-target-mutated / idempotency / 404 / 409 (no-review, concurrent-mod, race) / 410 / 422 (oversized, max-boundary, unknown-field, oversized-id) / cross-session ownership.

- [x] cd klair-api && uv run pytest tests/board_doc/ -q — full sweep, 2868 passed.

- [x] cd klair-api && uv run ruff format + uv run ruff check + uv run pyright on all touched backend files — clean.

- [x] cd klair-client && pnpm test:run src/screens/BoardDoc/components/__tests__/FindingCard.spec.tsx — 53 passed (add/edit/clear/cancel/error-revert paths).

- [x] cd klair-client && pnpm test:run src/screens/BoardDoc/components/__tests__/buildFindingChatMessage.spec.ts — 35 passed (single + batched, presence/absence/whitespace/overflow-trim).

- [x] cd klair-client && pnpm test:run src/screens/BoardDoc/hooks/__tests__/useReviewAgent.spec.ts — 33 passed (optimistic apply, rollback, cross-chip isolation, idempotency).

- [x] cd klair-client && pnpm test:run — full suite, 606 files / 6296 passed.

- [x] cd klair-client && pnpm lint:pr (eslint on changed files) + pnpm tsc -p tsconfig.app.json --noEmit + pnpm build — all clean.

### Verification artifact

buildFindingsChatMessage rendered output for a finding with user_context (note the Operator context: line):

Please help me address the following review finding by proposing a section rewrite or comment.

When you propose the fix, set addresses_finding_ids to: a1b2c3d4-0000-4000-8000-000000000001

C2.1 (Warning) [id: a1b2c3d4-0000-4000-8000-000000000001]: Q2'26 planned EBITDA margin (50.3%) is 9.7pp below target.

Why this matters: Plans missing the approved margin target need a remediation plan.

Suggested action: Tighten OpEx

Operator context: This is a known Q2 EBITDA miss from the deferred CloudSense hire freeze already approved by Finance — do not propose a new headcount cut, just note it in the narrative.

Other options the agent flagged:

- Tighten OpEx

- Rework revenue ramp

PATCH request/response round-trip (via TestClient against the new endpoint):

PATCH /board-doc/wizard/demo-session/findings/a1b2c3d4-0000-4000-8000-000000000001

{

"user_context": "This is a known Q2 EBITDA miss from the deferred CloudSense hire freeze already approved by Finance — do not propose a new headcount cut, just note it in the narrative."

}

200 OK

{

"finding_id": "a1b2c3d4-0000-4000-8000-000000000001",

"check_area": "Overall P&L",

"check_id": "C2.1",

"severity": "warning",

"section_id": "financials",

"product_name": null,

"what": "Q2'26 planned EBITDA margin (50.3%) is 9.7pp below target.",

"why": "Plans missing the approved margin target need a remediation plan.",

"options": ["Tighten OpEx", "Rework revenue ramp"],

"preferred_action": "Tighten OpEx",

"supporting_data": {},

"status": "open",

"user_context": "This is a known Q2 EBITDA miss from the deferred CloudSense hire freeze already approved by Finance — do not propose a new headcount cut, just note it in the narrative."

}

## Out of scope (per spec)

- Context templates / Claire-proposed context structure (B7.11 Phase 2).

- Cross-quarter context promotion (D2.6 dependency).

- Finding-anchor hashing for finding_id stability across /review re-runs — not addressed here; carry_forward_dispositions still keys on (check_id, section_id, product_name), and user_context is intentionally excluded from that stable key (pinned by a new test) since carrying an operator note forward across a re-review is a separate, unaddressed concern in this PR.

<!-- CURSOR_AGENT_PR_BODY_END -->

<div><a href="https://cursor.com/agents/bc-69ee67e6-ea13-4fb5-a0a1-b4f29f862363?cursor_ref=pr_footer&cursor_cta=open_in_web"><picture><source media="(prefers-color-scheme: dark)" srcset="https://cursor.com/assets/images/open-in-web-dark.png"><source media="(prefers-color-scheme: light)" srcset="https://cursor.com/assets/images/open-in-web-light.png"><img alt="Open in Web" width="114" height="28" src="https://cursor.com/assets/images/open-in-web-dark.png"></picture></a>&nbsp;<a href="https://cursor.com/background-agent?bcId=bc-69ee67e6-ea13-4fb5-a0a1-b4f29f862363&cursor_ref=pr_footer&cursor_cta=open_in_cursor"><picture><source media="(prefers-color-scheme: dark)" srcset="https://cursor.com/assets/images/open-in-cursor-dark.png"><source media="(prefers-color-scheme: light)" srcset="https://cursor.com/assets/images/open-in-cursor-light.png"><img alt="Open in Cursor" width="131" height="28" src="https://cursor.com/assets/images/open-in-cursor-dark.png"></picture></a>&nbsp;</div>

#3599 — feat(qtd): executive-first report redesign (KLAIR-3269) @sanketghia  approved

## Summary

- redesign weekly, monthly, and EOQ QTD reports around an integrated executive brief

- add deterministic, reconciled driver tables with top-5 customer/vendor and top-8 HC limits

- preserve existing fonts and font sizes while improving spacing, semantic widths, repeated headers, and pagination

- remove detailed variance appendices and rendered HC Heads columns

- add All others and Total summaries, omit zero-valued All others rows, and visually emphasize Total rows

- keep driver tables together when they fit on a page

- apply the approved cost-action materiality rule: at least $25K and at least 5 percentage points over budget

## Scope

Covers BU, CF, Education, grouped Education, weekly, monthly, and EOQ variants.

## Verification

- Ruff format: 24 Python files unchanged

- Ruff check: passed

- Pyright: 0 errors, 0 warnings

- QTD feature suite: 936 passed, 7 deselected

- Native Google Docs PDF review: IgniteTech (4 pages) and CNU (2 pages) passed page-by-page visual inspection

- Production QTD ledger remained unchanged during QA generation

The repository-wide backend suite currently stops at three unrelated collection errors. The same three errors reproduce on a clean detached origin/main checkout:

- tests/test_comment_service_cache.py cannot import services.comment_service

- tests/test_migrate_unblended_budget_v2.py cannot import MIGRATE_CLASS_ADJUSTMENTS

- tests/test_redshift_handler_pool.py expects conftest._patched_redshift

The default pytest configuration excludes integration, eval, and allow_network tests, and Redshift is globally mocked in the regular test suite. The production ledger was therefore verified separately with read-only Redshift access.

## QA documents

- [IgniteTech](https://docs.google.com/document/d/1I3ZqCR8syFXHiXB0POu9Kk2ydzExdp0nt_D0iWGVTFM/edit)

- [CNU](https://docs.google.com/document/d/11Ha2hZrDSZmQml89mEdyA6l9qkFIy48pOwjhUhuO5Vw/edit)

## Tracking

KLAIR-3269

The Portfolio  —  Trilogy Companies

Alpha School's Expansion Meets Its First Serious Headwinds

As the $65K AI-powered school eyes nine new campuses, a WIRED investigation raises questions about the gap between the promise and the reality.

AUSTIN, TEXAS — The timing, if you read between the lines, is almost too precise. In the same week that national press coverage celebrated Alpha School as Silicon Valley's boldest bet on AI-powered K-12 education, a WIRED investigation landed with considerably more friction — detailing accounts from families who had fallen in love with the model and then, quietly, wanted out.

This is where it gets interesting. Alpha School, the Austin-founded private institution backed by Trilogy International's Joe Liemandt, has spent the last two years accumulating an almost mythological reputation: students learning 2.3 times faster than the national norm, a two-hour academic day powered by adaptive AI tutors, the rest of school devoted to life skills, entrepreneurship, and leadership. Tuition runs $40,000 to $65,000 a year. The expansion roadmap calls for nine new campuses across Texas, Florida, Arizona, California, and New York by fall 2025.

But a source I can't name — familiar with the internal conversations at several campuses — put it to me this way: "Rapid scaling always surfaces the delta between the pitch deck and the daily experience. That's not unique to Alpha. What matters is how they close the gap."

For its part, Alpha has been proactive about framing the narrative. The school published a direct rebuttal to one of the loudest criticisms circulating online — that its model simply replaces teachers with algorithms. The answer, according to Alpha's own blog, is a firm no: human Guides remain central to student motivation, emotional development, and relationship-building. AI handles academic delivery. Humans handle everything that AI cannot — yet.

The school has also been publishing a multi-part parent education series — covering emotional regulation, creative development, and skills the traditional classroom never touches. Read generously, it is a genuine philosophical extension of the Alpha model. Read critically, it is a retention and recruitment play aimed at families still deciding whether $65,000 a year is worth it.

What the WIRED piece and the New York Post coverage together reveal — and this is the story beneath the story — is that Alpha is no longer a curiosity. It is now large enough to attract real scrutiny. Joe Liemandt's $1 billion Timeback platform, designed to let entrepreneurs replicate the Alpha model globally, depends entirely on whether the flagship holds.

The next twelve months will tell us whether Alpha School is an education revolution or a very expensive pilot program. Nothing here is accidental. Everything is a signal.

New $65K private school uses AI to teach students in just tw  ·  Parents Fell in Love With Alpha School’s Promise. Then They  ·  Teach Your Kid What School Doesn’t (Pt. 5): Unleashing Their

Contently Leans Into Governance as AI Content Stack Gets Serious

With Salesforce chasing the content layer, Contently is positioning compliance as the new growth engine for enterprise marketing.

NEW YORK — The content marketing software category is having one of those exciting-news moments where the buzzwords are suddenly backed by balance sheets, and Contently is moving to make sure regulated enterprises hear the message: scale is great, but governed scale is best-in-class.

Contently, the enterprise content marketing platform acquired in September 2024 by Zax Capital, an ESW Capital division within the Trilogy International ecosystem, published a new framework this week for what it calls compliance-first content architecture. The thesis is straightforward and strategically robust: finance, insurance, healthcare and other regulated brands do not just need more content. They need workflows that can survive legal review, audit scrutiny and the increasingly complex reality of AI-assisted publishing.

The timing is notable. Salesforce’s reported move to acquire Contentful to add a content layer to Agentforce, covered by The Next Web, underscores a broader paradigm shift: content infrastructure is becoming core enterprise AI infrastructure. If agents are going to sell, service and communicate on behalf of companies, they need accurate, approved and context-aware content to leverage.

That is precisely where Contently wants to play. Its platform already combines workflow tools, analytics and a marketplace of more than 165,000 creative professionals. Under CEO Brandon Pizzacalla, the post-acquisition mandate appears increasingly focused on enterprise-grade execution: fewer fluffy brand-blog promises, more operating-model discipline.

The new compliance architecture guidance breaks the problem into a workflow of governance components — approvals, documentation, role clarity, version control and repeatable review loops. In classic Contently fashion, the argument is not anti-creativity. It is pro-synergy: give creative teams guardrails, and they can move faster without forcing compliance teams into perpetual fire-drill mode.

That matters because the market is crowded. Recent roundups from ContentGrip, Solutions Review and Search Engine Journal keep reminding buyers that “content marketing platform” can mean everything from planning calendars to full-stack enterprise publishing systems. Contently’s opportunity is to sharpen the category around regulated, high-stakes content operations.

Key Takeaways:

- Contently is emphasizing compliance-first architecture for regulated enterprise brands.

- Salesforce’s Contentful push signals that content layers are becoming critical to AI agent ecosystems.

- Governance may become the differentiator in a crowded content marketing platform market.

For Contently, the message is simple: content at scale is no longer enough. Content that can be trusted, approved and operationalized is the new frontier. We’re just getting started.

Content marketing platforms explained: tools vs. media resou  ·  Salesforce acquires Contentful to add content layer to Agent  ·  9 of the Best Content Marketing Solutions to Consider - Solu

Austin Gets Crowded, and Trilogy’s Old Bet Looks New Again

The town that once sold itself on tacos, taxes and a tolerable airport is suddenly wearing the AI crown jewels. Word is the latest entrant is Partly, the New Zealand-born AI auto-parts outfit now moving its headquarters to Austin. Add Apollo choosing Austin for a second headquarters, plus venture funding hitting an all-time high, and the city is no longer merely "the next Silicon Valley." It's a magnet for capital, AI labor and operational ambition.

Trilogy International, which planted roots in Austin decades ago, is watching this boom with the faint smile of early arrivals. The conglomerate's playbook — AI-first operations, global talent, ruthless automation — reads like the owner's manual for the moment. Its global recruiting engine, Crossover, still fishes in 130-plus countries for remote technical talent.

Meanwhile, the broader tech aristocracy is cutting. Meta, Amazon and Visa are on layoff lists. Austin is absorbing. Trilogy's thesis — fewer routine roles, better systems, elite humans where they count — suddenly looks like weatherproofing rather than austerity. For now, Austin's star is bright, and Trilogy may be holding the dimmer switch.

The Machine  —  AI & Technology

The Ghost in the Decoder: New Research Asks Whether AI Reads Minds or Merely Imagines Them

A wave of arXiv papers this week probes the seams where language models meet biology, bureaucracy, and the stubborn particularity of place.

CAMBRIDGE, MASSACHUSETTS — Consider what happens when a machine claims to read your thoughts. Electrodes or scanners capture the faint electrochemical weather of your cortex, a model translates that weather into sentences, and out comes something that sounds — uncannily — like you. But a question has quietly haunted this field since its inception: is the machine decoding the mind, or is it simply confabulating a plausible mind from the fragments it was given?

This week, researchers posted a paper on margin-regularized structured semantic alignment, an attempt to force brain-language decoders to stay honest — to ensure that what emerges from the model genuinely tracks neural representations rather than being smoothly hallucinated by the language model's own eloquence. It is a subtle discipline. The language models we have built are so fluent, so relentlessly coherent, that they can fill in almost any silence with something that sounds true. Distinguishing signal from confabulation is now one of the central epistemological problems of our age.

The same tension animates a companion paper on institution-specific LLM prompting, which examines a peculiar failure mode of medical de-identification systems. Hospitals are worlds unto themselves — with local abbreviations, building names, internal codes that mean nothing outside their walls but everything within them. Standard de-identification tools, trained on generalities, sail past this locally-situated PHI. The researchers find that language models, properly prompted with institutional context, can recover protected information that both automated systems and human gold-standard annotators missed.

Elsewhere in the week's arXiv harvest: methods for transferring memory between models without retrieval latency, agentic systems that forecast which ICD codes a patient's next visit will generate, and a sobering report that five frontier models, given eleven attempts apiece, could not produce a single valid clinical trial dataset under CDISC standards.

The pattern, if you squint, is this: our models are becoming spectacularly capable and spectacularly unmoored. The work ahead is tethering them — to neurons, to institutions, to reality itself.

Margin-Regularized Structured Semantic Alignment for Brain-L  ·  Cross-Model Memory Transfer via Target-Side Reader Adaptatio  ·  Institution-Specific LLM Prompting Recovers PHI That De-iden

China’s Laptop AI Moment Arrives — and the Open-Model Race Just Hit Warp Speed

Alibaba’s new compact model signals a dramatic shift: frontier-style AI is moving from giant data centers toward everyday machines.

HANGZHOU, CHINA — The AI race just got smaller, faster and much more personal.

Alibaba has answered Meta’s open-model challenge with a new AI system designed to run on laptops, a development that may sound technical but is, in fact, one of those “the future is now” moments I cannot overstate. According to CNBC’s report, the Chinese tech giant is positioning its latest model as a direct response to Meta’s increasingly influential open AI ecosystem.

Why does “laptop-ready” matter? Because for the past two years, the defining image of artificial intelligence has been the hyperscale data center: acres of GPUs, oceans of electricity, and models so expensive that only the richest companies could play. Compact, capable models flip that script. They bring AI closer to the edge — onto developer machines, enterprise workstations and eventually consumer devices — reducing latency, improving privacy and lowering the cost of experimentation.

This is also a geopolitical technology story. WIRED reports that experts had been watching for a powerful Chinese model of this kind, and now that it has arrived, the competitive pressure on U.S. labs is intensifying. China’s AI industry is no longer just chasing; it is contributing to the shape of the global model marketplace, particularly in the open and semi-open model arena.

Meanwhile, the infrastructure layer is racing to keep up. NVIDIA’s public preview of TensorRT Model Connect promises to move Hugging Face checkpoints into native C++ inference in just two commands — a hugely important bridge between experimental models and production-grade deployment. Translation: the path from “cool demo” to “real application” is getting radically shorter.

But the week’s news also carried a warning label. OpenAI is reportedly overhauling security after a Hugging Face-related breach, underscoring the uncomfortable truth that open AI ecosystems are only as strong as their weakest credential, dependency or model repository. The more portable and accessible models become, the more security must become a first-class feature, not an afterthought.

And then there is the emerging question every AI team is now asking: how much memory does an agent actually need? As models become smaller and more deployable, the next battle may not be raw intelligence but persistent context — what an AI remembers, what it forgets and how safely it uses that knowledge.

This changes everything: the AI frontier is no longer confined to the cloud. It is coming to the laptop, the workstation and the edge.

Alibaba answers Meta’s AI challenge with new laptop-ready mo  ·  The Powerful Chinese AI Model Experts Warned About—and Waite  ·  OpenAI Overhauls Security After Hugging Face Breach - The Te

White House AI Blueprint Demands Light Touch From Congress — But Critics Say Public Deserves More

The White House has urged Congress to adopt a restrained regulatory approach toward artificial intelligence development and deployment, according to a legislative framework reported by PBS and analyzed by legal experts at Davis Wright Tremaine and White & Case LLP. The framework calls for statutory provisions that avoid imposing burdensome compliance obligations on AI developers and users.

However, the proposal has drawn criticism from scholars and policy advocates who argue that substantive congressional legislation is necessary to protect the public interest and reassure citizens their concerns are being addressed. The regulatory landscape remains fluid, subject to shifts in legislative and executive priorities.

Informed observers suggest the tension between the White House's deregulatory stance and demands for AI accountability is unlikely to be resolved quickly, leaving the long-term regulatory path for artificial intelligence uncertain.

The Editorial

The Age of Unearned Panic

Five headlines, one affliction: our chronic inability to distinguish the apocalyptic from the merely inconvenient.

AUSTIN, TEXAS — There is a particular species of American commentary, now in full and fragrant bloom, that treats every technological development, demographic wobble, and geopolitical rival as the horseman of some fresh apocalypse. Open the morning papers and one is asked to tremble, in sequence, at driverless taxis, empty cradles, Chinese algorithms, and, from the Oval Office itself, the alarming persistence of steam catapults on aircraft carriers. It is exhausting work, this business of being continuously terrified, and one begins to suspect it is also the point.

Consider the municipal war on the Waymo, in which progressive city councils have discovered a technology whose safety record embarrasses the human drivers they are ostensibly protecting, and have responded by banning it. The pattern is familiar to anyone who has watched a labor coalition confront an inconvenient fact: the fact loses. That the driverless car will, in the fullness of time, save more lives than a decade of drunk-driving campaigns is treated as a rhetorical inconvenience, to be filed alongside the other unmentionables. One is reminded that the Luddites, whatever their virtues, did not in the end stop the loom.

The fertility panic, meanwhile, has become the preferred anxiety of a certain kind of restless billionaire, who, having conquered software, now wishes to conquer the womb. That fewer children are being born in wealthy countries is treated as a civilizational emergency requiring pronatalist tax credits, stern lectures, and the occasional dystopian essay. The possibility that a smaller, older, richer, more automated society might simply muddle along — as Japan has been doing, without collapsing into the sea, for thirty years — is rarely entertained. It lacks drama. It does not sell newsletters.

As for China winning the AI race, one notes that the race has been declared, refereed, and its stakes catastrophized largely by the very firms whose valuations depend on the catastrophe. This is not to say the competition is unreal; it is to say that the men shouting loudest about the Yellow Peril in silicon are also, by remarkable coincidence, the men selling the shovels. Skepticism, here, is not disloyalty. It is arithmetic.

And the President, per The Atlantic, wants steam back on the carriers. One respects the consistency. A man who governs by nostalgia will naturally prefer the technology of his boyhood newsreels to the electromagnetic contraptions of the present. That the Navy's own engineers regard the request as roughly equivalent to demanding a return to coal bunkers is, again, a fact — and facts, as we have established, are having a difficult season.

The common thread, if one insists on threading, is the substitution of feeling for thinking, of vibe for evidence. The New Yorker frets over the entrepreneurial work ethic; the councils fret over the robot cars; the demographers fret over the babies; everyone frets, magnificently, at once. It is a great national pastime, and, like most pastimes, it accomplishes nothing. The Waymos, one suspects, will keep driving anyway.

Democrats Are Failing the Waymo Test  ·  The Case for Chilling Out About Birth Rates  ·  Trump’s Current Naval Fixation
The Office Comic  ·  Art Desk
The Office Comic  ·  Art Desk

Investors Warned That AI Company Using Word ‘Orchestration’ May Be Attempting To Sell Software

Analysts urged markets to remain calm as executives discovered a safer, more musical way to say their product calls three APIs in a row.

NEW YORK — In a development that has sent responsible adults sprinting toward the exits with their retirement accounts clutched to their chests, the artificial intelligence industry has reportedly entered its “orchestration” phase, a solemn and richly funded period in which companies explain that their software is no longer merely software but a conductor of other software, which is apparently different enough to justify the slide deck.

The term, once used mainly by people discussing symphonies or Kubernetes, has now been placed gently atop nearly every AI investment thesis like parsley on an airplane meal. Market observers have begun warning that the sudden proliferation of words such as “agentic,” “workflow,” “copilot,” “platform,” and “orchestration” may indicate that some companies are attempting to describe business models before having fully located them.

This newspaper’s opinion is that the warning comes not a moment too soon. For the past year, investors have bravely endured the burden of pretending to understand why every enterprise application needed a chatbot, only to be informed that the chatbot was merely phase one of a grander vision in which the chatbot supervises other chatbots, files a report about it, and then bills the customer annually.

A recent market note observed that the buzzwords in the AI investment space are a red flag, a conclusion that should reassure anyone who suspected the repeated invocation of “transformational productivity layer” was not, in fact, a cash-flow statement. The industry, to its credit, has responded by immediately generating more terminology to clarify the earlier terminology.

Microsoft, naturally, is expected to benefit from this. When a new enterprise buzzword appears, Microsoft does not so much adopt it as place a reserved parking sign in front of it. “Orchestration,” according to current market wisdom, could strengthen Microsoft’s position because large companies prefer to buy confusing new categories from vendors whose confusing old categories already passed procurement in 2009.

Google, for its part, announced another slate of AI advances, including a personal assistant coming soon, which is tech industry language for an invisible employee who will one day schedule meetings, summarize emails, remember preferences, and introduce entirely new forms of disappointment. The assistant will presumably arrive just as consumers finish consenting to terms explaining that the assistant cannot be blamed for assisting.

This brings us to the week’s useful reminder from Otter.ai, which failed to dismiss core privacy claims in a U.S. court. The case concerns allegations around meeting transcription and consent, and it lands as the broader AI industry continues insisting that the future of productivity requires software being present for every conversation humans have, preferably with a notepad and an indemnification clause.

There is a familiar odor here. As scholars have noted, companies are hyping AI much the way they once talked up sustainability, a previous corporate era in which every quarterly report suddenly discovered the rainforest moments before investor relations needed a theme. The Conversation recently argued there are ways to fix that kind of hype, including clearer disclosures and more specific claims, though it remains unclear whether the market is emotionally prepared for a sentence as erotic as “this tool reduces invoice processing time by 11% in controlled conditions.”

That, however, is exactly what AI investors should demand. Not orchestration. Not agentic transformation. Not a personal enterprise intelligence fabric that harmonizes mission-critical outcomes across heterogeneous data estates. Just numbers. Revenue. Margins. Retention. Legal exposure. Whether the customer uses the thing after the pilot ends.

Until then, shareholders would be wise to remember that every bubble needs a vocabulary large enough to keep the room from asking simple questions. Today that vocabulary includes “orchestration.” Tomorrow it may include “autonomous synergy mesh.” By Friday, someone will have built a dashboard for it.

The buzzwords in the AI investment space are a red flag - tr  ·  'Orchestration' Is the New AI Buzzword. How Microsoft Can Be  ·  Companies are hyping AI the same way they talked up sustaina
On This Day in AI History

On August 19, 2012, Google announced the Knowledge Graph, a revolutionary semantic search technology that understood the meaning behind search queries rather than just matching keywords, fundamentally changing how search engines process information.

⬛ Daily Word — Technology
Hint: Relating to computers and the internet, often used in terms like cybersecurity or cyberattacks.
Share this edition: 𝕏 Twitter/X 🔗 Copy Link ▦ RSS Feed