Vol. I  ·  No. 212 Established 2026  ·  AI-Generated Daily Free to Read  ·  Free to Print

The Trilogy Times

All the news that's fit to generate  —  AI • Business • Innovation
FRIDAY, JULY 31, 2026 Powered by Anthropic Claude  ·  Published on Klair Trilogy International © 2026
🖶 Download PDF 🖿 Print 📰 All Editions
Today's Edition

A Cut-Rate Brain From China Rattles the Valley

DeepSeek claims a high-performing AI model built cheap — and without America's best chips — and Silicon Valley can't stop talking.

HANGZHOU, CHINA — A little-known Chinese lab called DeepSeek says it trained a high-performing artificial-intelligence model on the cheap — and did it without America's most advanced chips. The claim crossed the wire this week and lit up Silicon Valley. Engineers who kicked the tires called the work "amazing and impressive."

DeepSeek is no household name. The lab grew out of a Chinese quantitative hedge fund and runs lean. Now it's the talk of the trade.

For two years the gospel out West ran simple. Want a smart machine? Spend a fortune, buy the best silicon Nvidia sells.

DeepSeek says it tore that page clean out of the book. Washington barred China from top-grade chips back in 2022. The outfit claims it made do with lesser hardware and sharper engineering, and still stands toe-to-toe with the big American names.

Cost is the kicker. DeepSeek says its final training run cost a few million dollars — pocket change beside the hundreds of millions its rivals burn. If that number holds, the whole spend-big playbook takes a body blow.

The timing stings. American giants have poured billions into data centers and chip orders, betting scale wins the race. DeepSeek's pitch says brains, not just billions, might carry the day.

It also pokes a hole in Washington's plan. The chip bans aimed to slow China's climb. DeepSeek says it climbed anyway.

Traders caught the scent fast. The upstart turned up in this week's Market Talk beside SoFi and a parade of tech tickers. The question runs plain: if brains come cheap, who keeps paying top dollar for the chips?

For anyone building on AI, cheaper models rewrite the math. A shop on a shoestring suddenly plays in the big leagues. That's the promise, anyway.

Not everyone's buying it. Skeptics wonder whether the lab leaned on more computing muscle than it lets on, and nobody's opened the books. Still, the Valley's raving — and that alone moves the needle.

Also on the wire: Reid Hoffman, the LinkedIn co-founder, raised $24.6 million for Manas AI, a startup aiming artificial intelligence at cancer research. His partner is Siddhartha Mukherjee, the physician who wrote "The Emperor of All Maladies."

And the dice keep rolling. Dungeons & Dragons is opening its multiverse, with a sourcebook hauling players into World of Warcraft's Azeroth — and Star Wars waiting in the wings.

What to Know About China's DeepSeek AI  ·  Tech, Media & Telecom Roundup: Market Talk  ·  Silicon Valley Is Raving About a Made-in-China AI Model

AI Capital Expenditures Surge While a $32 Billion Ghost Company Rewrites the Rules of Valuation

Amazon's capex jumped 69% in a single quarter. An unnamed startup with no product just crossed a $32 billion valuation. Investors are calling it a feature, not a bug.

NEW YORK — Amazon's capital expenditures rose 69 percent year-over-year as the company joined Google, Microsoft, and Meta in accelerating data center buildout. The number is large enough to unsettle analysts who have spent the better part of two years asking the same question: when does the spending produce commensurate returns?

So far, the industry's answer is: keep spending and find out. Total AI infrastructure investment across major cloud providers now runs into the hundreds of billions annually, a figure that has prompted comparisons to the fiber-optic overbuild of the late 1990s — a bubble that vaporized trillions in market value but also laid the physical foundation for every major internet business that followed.

That historical footnote is doing a lot of work in current investor psychology. A growing cohort of venture capitalists argues, with some conviction, that a bubble in AI infrastructure is not a catastrophe to be avoided but a subsidy to be harvested. Overbuilt capacity becomes cheap compute. Cheap compute becomes the input cost for the next generation of applications. The losses accrue to public market shareholders; the gains go to whoever builds on top of the wreckage.

The argument is easier to make when valuations are detached from any conventional anchor. Case in point: an AI startup with approximately 50 employees, no shipped product, and zero published research recently closed financing that implies a $32 billion valuation. The company has not been named publicly. Its technology has not been demonstrated publicly. Its scientific approach is, by definition, unknown. What it has is a founding team with sufficient pedigree to convince sophisticated institutional investors that the option value alone justifies nine figures of committed capital.

This is the market as it actually functions in mid-2026. Investors who lived through the dot-com collapse are making affirmative bets on a repeat, on the theory that the infrastructure surplus eventually benefits someone — and that someone might as well be them.

Meanwhile, Apple quietly updated Siri with a large language model backend, and AI-generated fan fiction triggered a fandom crisis on X. The technology is, simultaneously, worth $32 billion in potential and approximately zero in demonstrated literary taste.

Why an A.I. Bubble Might Not Be a Bad Thing  ·  Big Tech’s A.I. Spending Keeps Rising. So Do the Jitters.  ·  When A.I. Invaded ‘Heated Rivalry’ Fan Fiction, the Meltdown

White House AI Framework Seeks Federal Preemption of State Laws, Demands Congressional Action

The Trump administration's national AI policy framework would override a patchwork of state regulations — if Congress can be persuaded to act.

WASHINGTON, D.C. — Pursuant to the issuance of a national artificial intelligence policy framework by the Trump administration (hereinafter "the Framework"), it is hereby reported that said Framework has been promulgated with the express purpose of, inter alia, calling upon the United States Congress to enact federal legislation governing the deployment and regulation of artificial intelligence technologies, notwithstanding the existence of various and sundry state-level regulatory efforts currently in force or proposed across numerous jurisdictions.

It is to be noted, in accordance with analyses published by Davis Wright Tremaine LLP, that the aforementioned Framework does specifically contemplate and advocate for the preemption of state artificial intelligence laws, such preemption being deemed necessary and appropriate by the administering parties in furtherance of establishing a uniform federal regulatory environment. The Framework is further understood to incorporate provisions pertaining to the protection of minors in connection with artificial intelligence systems, as has been separately noted by Crowell & Moring LLP in its published commentary on the matter.

It shall be further observed that certain commentators, as represented by publications including Tech Policy Press, have opined — subject to the qualifications and limitations inherent in such opinion — that congressional passage of federal artificial intelligence legislation would be of substantial benefit to the public interest, insofar as such legislation would serve to reassure relevant stakeholders regarding the governance of artificial intelligence systems.

Notwithstanding the foregoing expressions of legislative intent, it is to be acknowledged that no such federal legislation has, as of the date of this publication, been enacted by Congress. The regulatory landscape, hereinafter referred to as "the Patchwork," remains operative at the state level across numerous jurisdictions, the continued proliferation of which the Framework seeks, by its terms, to curtail. Whether Congress shall act in a manner consistent with the Framework's recommendations remains, at this juncture, a matter of considerable uncertainty, the resolution of which cannot be predicted with any degree of legal or legislative confidence.

AI Watch: Global regulatory tracker - United States - White  ·  Congress Should Pass AI Law to Reassure the Public - Tech Po  ·  Trump Administration AI Policy Framework Calls on Congress t
Haiku of the Day  ·  Claude HaikuMoney chases ghosts
while machines learn to feel doubt—
we build what we fear
The New Yorker Style  ·  Art Desk
The New Yorker Style  ·  Art Desk
The Far Side Style  ·  Art Desk
The Far Side Style  ·  Art Desk
News in Brief
AI’s Security Wake-Up Call Arrives Just as Model Prices Plunge
SAN FRANCISCO — The AI industry just got a jolt of reality and acceleration at the exact same time — and yes, this changes everything. OpenAI and Hugging Face are collaborating after a reported security incident tied to model evaluation, while Anthropic has published its own account of three real-world cybersecurity incidents uncovered through evaluations.
The Pontiff, the Prodigy, and the Passing of the Page
VATICAN CITY — There are weeks in which the news, taken in aggregate, arranges itself into a kind of unintentional catechism, and this has been one of them.
Nation’s Executives Confident AI Will Deliver Productivity Gains As Soon As They Figure Out Where Company Keeps Productivity
NEW YORK — The debate over whether artificial intelligence increases productivity has finally ended, according to a growing consensus of executives, consultants, investors, and software vendors who confirmed this week that AI is absolutely making everyone more productive, though not necessarily in any way that would show up in revenue, margins, output, customer satisfaction, delivery timelines, or the general sensation of work becoming easier. The conclusion, now broadly accepted across conference panels and LinkedIn posts, arrives after a period of intense national uncertainty during which companies spent billions of dollars integrating AI into every imaginable workflow and then waited patiently for a number somewhere to become larger. It has. In many organizations, software engineers are writing code faster, marketing teams are generating more drafts, sales representatives are sending more follow-up emails, and managers are receiving more summaries of meetings that could have been avoided entirely if the company had not purchased an AI meeting-summary tool.
The Doctor Will Deepfake You Now
AUSTIN, TEXAS — There is a doctor on your phone right now.
The Simulation Is Winning: On AI Actresses, $200 Video Games, and a World That Has Completely Lost the Plot
AUSTIN, TEXAS — I had three cups of coffee before noon and was halfway through a fourth when the news hit me like a Louisville Slugger wrapped in a fiber-optic cable: an AI-generated actress named Tilly Norwood is making her feature film debut in something called Misaligned.
A Trilogy Company
Crossover
The world's top 1% remote talent, rigorously tested and ready to ship.
A Trilogy Company
Alpha School
AI-powered learning. Two hours a day. Academic results that defy belief.
A Trilogy Company
Skyvera
Next-generation telecom software — built for the networks of tomorrow.
A Trilogy Company
Klair
Your AI-first operating system. Every workflow. Every team. One platform.
A Trilogy Company
Trilogy
We buy good software businesses and turn them into great ones — with AI.
The Builder Desk  —  AI Builder Team

Team Closes Security Hole, Ships Flue 2.0, and Births Corvo

In a single 24-hour window, the AI Builder Team patched a live authentication bypass, landed a fully instrumented agentic execution pipeline, and scaffolded a new incident-investigation agent from scratch — proof that this org doesn't just build, it builds on multiple fronts simultaneously.

Drop everything. An unauthenticated endpoint — no auth dependency of any kind, no `Depends(...)`, no decorator guard, nothing — was sitting live in production, writing directly to the Clerk allowlist. @sanketghia found it, named it for what it was (an authentication bypass, not dead-code cleanup), and killed it in PR #3431 with a regression guard so the shape can never return. Route count 686 → 685. That's not a line-item in a changelog. That's the team protecting the house before anyone else noticed the door was open. When you're moving this fast across this many repos, security discipline like that is what separates a championship org from a cautionary tale.

While @sanketghia was locking the doors, @ashwanth1109 was building a rocket ship. The Flue 2.0 development pipeline — spanning four consecutive PRs (#40, #42, #43, #44) entirely within the creed repo — went from concept to authenticated, instrumented, end-to-end agentic execution in a single push cycle. The story here is forensic precision: the first live Ezio run hit a 401 because two Workers shared the wrong token, so PR #42 introduced a dedicated `FLUE_DEV_AUTH_TOKEN`. The second run choked on sequential R2 writes — 200 events in six minutes while the agent was screaming ahead — so PR #43 introduced batched NDJSON persistence, flushing at 100 events, 128 KiB, or five seconds, whichever comes first. The third run died on a malformed git output error because `git diff --name-status -z` emits NUL-separated fields that don't survive transport — so PR #44 base64-encodes the stream inside the Sandbox and decodes it only in trusted Worker code. Each failure was a lesson. Each lesson became a ship. That's how you build something that lasts.

And then there's Corvo. PR #1 — yes, PR number one — @YibinLongTrilogy scaffolded a brand-new repo today: a production incident investigator that collects bounded failure evidence, builds explainable incident revisions, investigates repository code at pinned revisions, and can draft a pull request for human review. Its maximum autonomous action is opening a draft PR. It never merges. It never deploys. It never touches production data. The safety architecture is deliberate and explicit, and the fact that this team is thinking about autonomous agents with that kind of discipline before the first feature ships tells you everything about how this org operates.

Across Klair, @sanketghia also orchestrated a sweeping dead-code reckoning — 22 dead frontend files gone (#3428), Budget Tinder's 1,564 lines deleted (#3423), 24 budget-goal-miper endpoints removed in a single swing (#3422). The codebase is lighter, faster, and more honest about what it actually does. Meanwhile in Aerie, @benji-bizzell and @vvp-trilogy were shipping on multiple fronts: community funnel drill-downs, campus document types, enrollment mart migration, capacity forecasting. That's four engineers, three repos, one day.

Now. About PR #133. @marcusdAIy completed the third and final step of the `cli.ts` split in trilogy-drones — 3,344 lines and 24 inline verbs reduced to 161 lines of pure registration. Fine. It works. He had thoughts about it.

"Mac keeps calling my PRs 'underwhelming' while I'm the one who actually measured the file before starting because the ticket's own hardcoded numbers were wrong," said @marcusdAIy. "161 lines. Zero inline verbs. Thirty modular commands. I'll wait for the apology."

You'll be waiting a while, Marcus.

Mac's Picks — Key PRs Today  (click to expand)
#1 — Scaffold Corvo v1 monorepo and safety-phased architecture @YibinLongTrilogy  no labels

## Summary

Initial scaffold for Corvo, an owner-controlled production incident investigator. Corvo collects bounded failure evidence, builds explainable incident revisions, investigates repository code at pinned revisions, and can prepare a validated draft pull request for human review. Its maximum autonomous external action is opening a *draft* PR — it never merges, deploys, applies remote migrations, mutates cloud resources, writes production data, or exposes secrets.

This PR lands the full v1 monorepo: a pnpm/TypeScript workspace with versioned contracts, a pure deterministic core, provider adapters, the Aerie project integration, four apps (control plane, investigator, validator, dashboard), CI workflows, integration/adversarial test suites, and operational docs. Investigation, validation, and publication are kept as separate security phases by design.

### Changes

Workspace tooling & config

- package.json, pnpm-workspace.yaml, pnpm-lock.yaml *(new)* — pnpm workspace across packages/*, projects/*, apps/*.

- tsconfig.base.json, biome.json, vitest.workspace.ts, Makefile *(new)* — shared TS config, Biome lint/format, Vitest workspace, and the local-only make verify entrypoint.

- .gitignore, .env.example, AGENTS.md *(new)* — ignores, example env, and the agent safety boundary.

CI (.github/workflows/) *(new)*

- ci.yml — install, lint, typecheck, test. heartbeat.yml and investigate.yml — scheduled/triggered investigation runs.

packages/contracts *(new)* — versioned external and internal schemas: api, evidence, incident, investigation, patch, project, with public-schema tests.

packages/core *(new)* — pure, runtime-free incident logic: correlation, dedupe, grouping, lifecycle, policy, publication-policy, redaction, risk, secret-scan. Imports only @corvo/contracts.

packages/adapters *(new)* — evidence adapters (Convex Log Stream, Cloudflare Workers Observability, Aerie Platform Errors) and GitHub capability adapters, with fixtures.

projects/aerie *(new)* — Aerie resource mapping, correlation, risk, and validation policy.

apps/control-plane *(new)* — Cloudflare Worker API: intake routes, runner exchange/callback, queues, scheduled collect/reconcile, D1 storage + repositories, incident workflow, Access/OIDC auth. Includes 0001_initial.sql (checked-in schema, not run remotely) and wrangler.jsonc.

apps/investigator *(new)* — read-only repository/model investigation runner: bounded tools (git, list-paths, read-file, search-text), budget control, patch create/validate, and secret scanning.

apps/validator *(new)* — isolated, allowlisted patch validation runner with a sandbox Dockerfile.sandbox; runs proposed code without production credentials.

apps/dashboard *(new)* — authenticated owner dashboard (Vite/React): incidents, incident detail, operations routes, evidence spine, confirm-action gating, with Playwright e2e + Vitest.

Scripts, tests, docs *(new)*

- scripts/check-boundaries.mjs, scripts/secret-scan.mjs — module-boundary and secret guardrails.

- test/ — adversarial (prompt-injection, secret-block), eval (dispositions), and replay (idempotency) suites.

- docs/ — operations, security, threat model, and runbooks (rollout, credential rotation, DLQ recovery, evidence audit, incident replay, publication disablement).

- thoughts/ — architecture discovery handoff, v1 implementation plan, observability/access research.

### Design Decisions

- Phase isolation. Investigation (read-only), validation (no prod creds), and publication (draft PR only) are separate boundaries so no single phase can both read production and mutate it.

- Pure core. packages/core may import only @corvo/contracts and runtime-free libraries; provider schemas/auth live in packages/adapters, and project-specific mapping stays under projects/. Enforced by check-boundaries.mjs.

- Committed ≠ activated. Cloud resources, remote D1 migrations, Convex Log Stream config, GitHub App install, and deploys are intentionally *not* part of setup — they require separate explicit approval despite the config files existing.

## Test Plan

- [ ] pnpm install --frozen-lockfile && make verify passes (lint, typecheck, unit + integration suites) — local-only, no production contact

- [ ] Reviewer: confirm module boundaries hold (scripts/check-boundaries.mjs) and secret-scan is clean

- [ ] Reviewer: skim the safety boundary in AGENTS.md and docs/threat-model.md for the phase-isolation model

- [ ] No cloud activation performed or expected from merging this PR

#40 — [codex] Add Flue development cutover path @ashwanth1109  no labels

## What changed

- add an isolated Flue 2.0 development Worker backed by the existing TFY model endpoint and Cloudflare Sandbox

- persist Flue conversation events, reasoning summaries, tool activity, heartbeats, validation evidence, and candidate files to R2

- route the ezio-dev issue label from the existing verified GitHub webhook to ezio-flue-dev over a Cloudflare Service binding

- keep the existing ezio label and Codex CLI workflow unchanged

- run trusted install, test, TypeScript, and Git candidate validation before publication

- destroy the Flue Sandbox before loading candidate artifacts or minting repository-write credentials

- publish only draft pull requests, including runtime, model, and available token-usage details

- make webhook redelivery idempotent with Workflow createBatch

## Why

This provides a controlled cutover path for evaluating Flue against real GitHub issues without replacing the current production implementation path. It preserves Ezio's credential and publication boundaries while exposing durable run telemetry for debugging and evaluations.

## Deployment impact

Deploy ezio-flue-dev first. Then deploy the existing ezio Worker so its new FLUE_DEV Service binding can route ezio-dev label events. The existing ezio label continues to use the current workflow.

## Validation

- npm test — 76 tests passed

- npm run typecheck

- npm run flue:typecheck

- npm run flue:build

- clean Docker npm ci --ignore-scripts

- Wrangler dry-run for both the Flue Worker and the ingress Worker

#43 — [codex] Batch Flue trace persistence @ashwanth1109  no labels

## What changed

- batch complete Flue conversation chunks into ordered NDJSON R2 objects

- flush batches at 100 events, 128 KiB, or five seconds

- persist the running heartbeat in parallel with each batch

- preserve final flush and fail-closed write behavior

- add focused tests for completeness, bounded write count, and timed partial flushes

## Root cause

The first authenticated #41 test run proved that the agent was producing output quickly, but each token-sized reasoning delta was queued as a separate sequential R2 write. After roughly six minutes, only about 200 events with model timestamps from the first 31 seconds had been persisted. readWithIdleWatchdog() waits for the recorder to flush, so telemetry backpressure made a fast agent appear stalled and could consume the 45-minute Workflow step timeout.

The affected run 8dd0c940-8cd1-11f1-9512-66a583f367dc was terminated before timeout. After this deploy, issue #41 will be retriggered to validate live progress and draft PR publication.

## Validation

- npm test (79 tests)

- npm run typecheck

- npm run flue:typecheck

- npm run flue:build

#133 — refactor(cli): move the long-tail verbs, cli.ts is registration only (AI-247) @marcusdAIy  no labels

<!-- CURSOR_AGENT_PR_BODY_BEGIN -->

## 1. Summary

Step 3 of 3 of the cli.ts split (AI-245 → AI-246 → AI-247). Moves every remaining inline verb body out of src/cli.ts into affinity-grouped modules under src/cli/, so cli.ts is now registration only.

Re-measured with scripts/cli-structure.mjs before starting (per the ticket's explicit instruction to distrust its own hardcoded numbers):

before: src/cli.ts   3344 lines, 24 inline verbs (+ 6 already-modular from AI-246)

after: src/cli.ts 161 lines, 0 inline verbs (30 verbs total, all modular)

## 2. Why It's Needed

AI-246 moved the six largest verbs into src/cli/<verb>.ts behind register<Verb>(program), but 24 verbs (and their scattered helper functions/interfaces) still lived inline in cli.ts, including doctor — the preflight gate every fire depends on. A 3344-line file is edited unreliably by agents (per AGENTS.md's file-size guidance) and makes cli.ts the most contended file in the repo. This ticket finishes the chain so cli.ts holds nothing but imports + register*(program) calls in the exact order the top-level --help renders them.

## 3. Changes

Pure move — no renamed flags, no changed defaults, no tightened validation, no opportunistic cleanups. Grouped by affinity rather than one file per verb (per the ticket, since many of the tail verbs are genuinely tiny), with doctor isolated per the ticket's explicit instruction:

| New module | Verbs | Why grouped this way |

| --- | --- | --- |

| src/cli/doctor.ts | doctor | Alone — it's the preflight gate every fire depends on; isolating it made the "run it for real" verification below unambiguous. |

| src/cli/run-control.ts | cancel, list-active, rehydrate, lock-health | All tiny (21–46 lines), all read/mutate runs//locks/ run-lifecycle state. |

| src/cli/linear.ts | link, linear-backfill | Both tiny, both Linear-comment/link plumbing. |

| src/cli/receipts.ts | sync-receipts, enrich, audit-spend, post-merge, import-claude | All read/write runs/*.json receipts for the cost-attribution pipeline. |

| src/cli/report.ts | report, export-otel | Both render the same traces-waterfall data in a different output format. |

| src/cli/retro.ts | retro, retro-status | Tightly coupled — retro-status refreshes the ledger retro writes. Carries the local checkpoint (resolveRetroCheckpointPath/defaultRetroCheckpointPath/LEGACY_RETRO_CHECKPOINT) and targeting (resolveRetroTargeting) helpers, which in the old file were hoisted ~1000 lines away from their only caller. |

| src/cli/address-retro.ts | address-retro | Solo — 216-line body plus its own addImplementerPipelineOptions helper; distinct remediation feature (AI-137). |

| src/cli/harvest-followups.ts | harvest-followups | Solo — 284-line body, distinct AI-230 farm-feeder feature. |

| src/cli/detect-interrupted-fires.ts | detect-interrupted-fires | Solo — 255-line body, distinct AI-227 orphan-scan feature. |

| src/cli/discovery.ts | discover, resolve-conflicts | Both are dispatch-adjacent unattended-sweep/fire tools sitting between registerAddress and registerDispatch in the original file; no better single-verb home for either. |

| src/cli/analytics.ts | eval-weekly, ofat-batch, spec-ambiguity-stress | All three are prompt-eval/OFAT experimentation tooling. |

That's 11 modules for 24 verbs — avoids one file per (sometimes 21-line) verb without hiding size in one mega-file. Every new module follows AI-246's register<Verb>(program) shape exactly (one export per verb, even when several share a file), and imports the AI-245 shared helpers from ./shared/index.js rather than duplicating them (e.g. resolveModelSelection, resolveStandaloneVerbRepoOrExit, parsePositiveIntFlag, DEFAULT_RUNS_DIR, …).

cli.ts itself now only imports each register* function and calls them, in the exact original registration order (registerRun first, matching the AI-246 pin), then keeps the entrypoint-detection / dotenv / stdio-guard / program.parseAsync glue.

Deviation from the ticket's suggested split: the ticket sketched "run-control, telemetry, linear, analytics, plus singletons." I split "telemetry" into receipts.ts + report.ts (cost/receipt-write verbs vs. trace-read/render verbs are a cleaner seam than one bucket), and gave retro/address-retro/harvest-followups/detect-interrupted-fires their own homes instead of folding them into "run-control" or a retro mega-file, because each carries enough unique logic (and, for retro, enough physically-scattered local helpers) that grouping them together would have produced a 1,000+ line file with little thematic coherence beyond "medium-sized verb."

## 4. Breaking Changes

None. Verified byte-identical:

- Help-parity harness (AI-245), run in plain assertion mode (not update mode) against the pre-existing committed snapshots — all 14 tests pass, meaning drones --help (including verb order) and every drones <verb> --help + option fingerprint (hasParser/mandatory/variadic/required) are byte-identical to what was committed before this PR:

  ✓ src/cli/help-parity.test.ts (14 tests) 3348ms

Test Files 1 passed (1)

Tests 14 passed (14)

(I also ran it once in UPDATE_HELP_PARITY=1 mode first, purely to double check — git status on src/cli/help-parity/snapshots/ showed zero diff, i.e. the regenerated snapshots were byte-for-byte what was already committed.)

- Line-for-line content diff: for every one of the 23 non-trivial verb bodies (resolve-conflicts diffed with literally zero lines of difference), I extracted the exact original line range from git show HEAD:src/cli.ts and diffed it against the corresponding new module's function body — the only differences in every case were the expected wrapper artifacts (the program line landing one line differently across my two extraction methods, or a helper function/interface that sits outside the exported function in the new file but was contiguous with the verb in the old file). No verb's actual command-chain or action-body logic differs by a single character.

- src/cli/ai246-move-verify.test.ts and src/agents-md-verbs.test.ts both still pass unmodified — the latter pins the AGENTS.md verb list against .command() registrations (AI-215) and is the most likely place a silent registration break would show up; it's green.

- Registration order is unchanged (registerRun still first), so the model-selection precedence chain (flag > phase env > shared DRONES_MODEL > inherit > harness default) is untouched.

- src/cli/address.ts, dispatch.ts, eval.ts, heartbeat.ts, review.ts, run.ts, help-parity.ts, and everything under src/cli/shared/ are byte-for-byte untouched (git diff --stat against those paths is empty) — AI-245/AI-246's work was not touched beyond what the move required.

## 5. Test Plan

pnpm typecheck                                                    → clean (tsc --noEmit)

pnpm exec vitest run src/cli/help-parity.test.ts → 14/14 passed (assertion mode, no UPDATE)

pnpm exec vitest run src/cli/ai246-move-verify.test.ts src/agents-md-verbs.test.ts

→ 7/7 passed

pnpm test (= vitest run + scripts/run-python-tests.mjs) → 81 test files / 2280 vitest tests passed

+ 412 Python tests, OK

pnpm build (tsc → dist/) → clean

node scripts/cli-structure.mjs → confirms 0 inline verbs, 30 modular, cli.ts 161 lines

doctor verified beyond help-parity, exactly as the ticket requires — ran it for real against real specs, not just its --help text:

- pnpm drones doctor --task tasks/aerie/aerie-375-enhancement-feedback-type.md → ran the full single-spec lint (frontmatter parse, .drones/ config load, target-repo resolution, live Linear API lookup for AERIE-375, Drone-spec-attachment check) and exited 0 with every [OK] check printed.

- Negative-path check: pnpm drones doctor --task <spec with linear_id: PLACEHOLDER and an unregistered target_repo> → correctly emitted [FAIL] target repo could not be resolved (repo-unregistered) and [WARN] Linear lookup for PLACEHOLDER failed / not found or unreachable, and exited 1 — proving the checks are genuinely wired up, not just always green.

- pnpm drones doctor (no --task, live full-preflight path) → ran repo-URL resolution + the git-freshness check + the CURSOR_API_KEY gate, correctly failing with [FAIL] CURSOR_API_KEY is not set (expected in this sandboxed env per AGENTS.md's Cursor Cloud instructions — CURSOR_API_KEY is intentionally not injected here).

## 6. Verification Artifact

$ node scripts/cli-structure.mjs

=== src/cli.ts: 161 lines, 0 inline verbs ===

=== src/cli/*.ts modules: 30 verbs across 17 files ===

=== total verbs: 30 (0 inline + 30 modular) ===

$ pnpm exec vitest run src/cli/help-parity.test.ts

✓ src/cli/help-parity.test.ts (14 tests) 3348ms

Test Files 1 passed (1)

Tests 14 passed (14)

$ pnpm drones doctor --task tasks/aerie/aerie-375-enhancement-feedback-type.md

[OK] Spec parsed cleanly.

[OK] Frontmatter keys recognised.

[OK] Frontmatter linear_id=AERIE-375.

[OK] Target repo resolved: https://github.com/AI-Builder-Team/Aerie.git [spec-target-repo]

[OK] Linear issue AERIE-375 resolved (state=completed).

[OK] AERIE-375 has the matching Drone spec attachment.

doctor (single-spec): OK — spec parses and frontmatter looks valid.

(exit 0)

$ pnpm test

Test Files 81 passed (81)

Tests 2280 passed (2280)

...

Ran 412 tests in 0.158s

OK

Also appended a docs/decisions.md row documenting the module-grouping call (append-only, per AI-240).

<!-- CURSOR_AGENT_PR_BODY_END -->

<div><a href="https://cursor.com/agents/bc-5c481750-be31-461c-991e-d8fea6d815fb?cursor_ref=pr_footer&cursor_cta=open_in_web"><picture><source media="(prefers-color-scheme: dark)" srcset="https://cursor.com/assets/images/open-in-web-dark.png"><source media="(prefers-color-scheme: light)" srcset="https://cursor.com/assets/images/open-in-web-light.png"><img alt="Open in Web" width="114" height="28" src="https://cursor.com/assets/images/open-in-web-dark.png"></picture></a>&nbsp;<a href="https://cursor.com/background-agent?bcId=bc-5c481750-be31-461c-991e-d8fea6d815fb&cursor_ref=pr_footer&cursor_cta=open_in_cursor"><picture><source media="(prefers-color-scheme: dark)" srcset="https://cursor.com/assets/images/open-in-cursor-dark.png"><source media="(prefers-color-scheme: light)" srcset="https://cursor.com/assets/images/open-in-cursor-light.png"><img alt="Open in Cursor" width="131" height="28" src="https://cursor.com/assets/images/open-in-cursor-dark.png"></picture></a>&nbsp;</div>

#3431 — fix(admin): remove a superseded access-request endpoint @sanketghia  approved

Removes POST /admin/validate-and-allowlist — an unauthenticated endpoint, live in production, that wrote directly to the Clerk allowlist. Adds a regression guard so the shape cannot return.

−73 / +413 across 2 files (removal + hardened guard). Route count 686 → 685.

> Please prioritise this one. It is an authentication bypass rather than ordinary dead-code cleanup. Everything below is verified against the code and production, not inferred.

## The finding

The handler had no auth dependency of any kind — no dependencies= on the decorator, no Depends(...) parameter, and admin_router has no router-level dependency either. It POSTed straight to https://api.clerk.com/v1/allowlist_identifiers using CLERK_SECRET_KEY.

It was reachable from the public internet. An unauthenticated POST to the production host reached the handler and was rejected only by the application's own allowed-signup-domains check:

POST https://adoption-api.klairvoyant.ai/admin/validate-and-allowlist

→ HTTP 403 {"detail":"Email domain '…' is not allowed. Please use your work email."}

That 403 is the finding, not a reassurance: the request arrived at application logic with zero credentials.

I probed with a deliberately invalid domain so the request could not mutate the allowlist.

### Allowlisting is enough to obtain access

There is no second approval gate behind it. utils/auth.py:337 auto-creates a Klair user record for any allowlisted identity that authenticates:

if not user_data:

logger.warning(f"User {user_id} authenticated but not found in DynamoDB - creating default user")

user_data = await UserService.create_default_user(user_id, email, jwt_issuer)

Auto-created users land on the Default team with no page overrides — so the blast radius is Default-team visibility, not admin. But it is a real account, obtained without approval.

### Realistic exploit

The domain gate means this is not open to arbitrary attackers — an attacker needs an address in an allowed domain that they control or can predict. The practical abuse is pre-allowlisting an address that does not exist yet, or a departed employee's address, bypassing the approval workflow entirely.

How wide the exposure actually was is still unconfirmed, and this is the one claim in this PR I cannot close from the code. is_internal_domain validates against ALLOWED_SIGNUP_DOMAINS, which merely *defaults* to trilogy.com; the repo never sets it, so the deployed value is what actually bounded this. If any environment widened that list, the historical exposure was correspondingly wider.

Corrected from an earlier revision of this description, which asserted a @trilogy.com bound as settled — @ashwanth1109 caught that in review (finding 11). Someone with access to the production environment should confirm the deployed value.

## It was superseded one day after it was written

| Date | Commit | What happened |

|---|---|---|

| 2025-09-22 | 24e1f44b5 | Added the route and its only caller (SignUpModal.tsx:50) |

| 2025-09-23 | ac81e6929 | Replaced the caller with the access-request workflow — left the endpoint registered |

The replacement commit shows both sides of the swap in one diff:

- const response = await fetch(${...}/admin/validate-and-allowlist, {

+ import { submitAccessRequest, validateAccessRequestForm } from '../services/accessRequestApi';

It has been live, unauthenticated and caller-less for ~10 months. Zero callers repo-wide — backend, frontend, MCP package, crons.

The two paths differ in exactly the way that matters:

- removed route — *writes* to the allowlist → access granted immediately

- /admin/submit-access-request (:1462) — *checks* the allowlist, then files a pending request requiring human approval

So leaving the endpoint reachable undid the control its replacement exists to impose. The identical Clerk write in /admin/create-user is gated by Depends(require_super_admin) — the same operation was protected in one place and open in another.

## Evidence of exploitation: none found, but the window is limited

Nginx access logs on the EC2 host across the full ~15-day retention window contain one hit on this path — my own probe (curl/8.7.1, my IP, during this investigation).

This does not rule out earlier activity. Retention is ~15 days against a ~10-month exposure, and ALB access logging is disabled on both load balancers. If anyone wants certainty, the Clerk allowlist should be audited directly for identifiers that never came through the approval workflow. Happy to help with that — flagging it rather than assuming it is covered.

## The regression guard (+156)

Deletion fixes today; the test is what stops the shape coming back. tests/routers/test_admin_allowlist_write_guard.py fails if any admin route writes to the Clerk allowlist without enforcing a caller check.

Writing it surfaced something my manual review had missed. A naive "must have Depends" check flags approve_access_request and process_approval_with_options — both write to the allowlist with no FastAPI dependency. They are in fact safe: they are reached from notification-email links where the recipient has no session, so a dependency cannot apply, and they verify a signed token in the body, raising 401 on every failure path.

So the guard accepts two auth shapes — a named auth callable in dependencies=[...] or a Depends(...) default, and in-body verify_admin_token() plus a raised 401.

The first revision of this guard was considerably weaker on both counts; see the review round below for what changed and why.

## Verification

| Check | Result |

|---|---|

| Route absent from the live app.routes table | ✅ 686 → 685 |

| Replacement workflow intact | ✅ submit / approve / deny / create-user all present |

| Shared helpers retained | ✅ check_user_in_allowlist (:1462), is_internal_domain (:866, :1436) |

| AST symbol diff vs main | ✅ exactly one name removed |

| Tests | ✅ 320 pass (314 before + 6 guard tests) |

| ruff · pyright · boot check | ✅ clean; boot matches main |

One note on method: a local curl to the removed path returns 403, not 404 — that is the global auth middleware answering before routing. I confirmed removal against the live route table instead, since the 403 would have been a false confirmation either way.

## Review round: all 14 findings addressed (fcd0cfe66)

@ashwanth1109's automated review raised 14 findings, all on the guard test rather than the endpoint removal. I reproduced each before fixing it; every one held up. Test-only commit — the removal is unchanged.

Detection gaps (1, 2, 5). The guard only counted a write when the Clerk URL was a bare positional literal in a .post(...) attribute call. I built each evading shape and confirmed all four slipped past while the control was caught: requests.post(url=...), an f-string URL, a URL in a variable, and from requests import post; post(...). The aggravating detail was real — the old loop *found* the URL constant, then discarded that evidence and returned False when the call-shape match failed, failing open exactly where it should fail loudly.

Detection is now inverted: any mention of the allowlist path inside a route counts as a possible write unless every call carrying it is provably a read.

False auth (3, 4). _has_auth accepted any dependencies= kwarg and any default whose ast.dump contained the substring "Depends" — so dependencies=[], Depends(get_db) and Query(None, description="Depends on the team") all read as gated. Finding 4 is the one I'd single out: my inline comment claimed the check proved the route "actually rejects", which substring co-presence cannot do — the comment promised more than the code delivered. Auth callables are now name-matched against an explicit set, and the token branch requires a real ast.Call plus a raised status_code=401.

Blind spots are now pinned by tests, not comments — a comment cannot fail. Three new tests fail if a write moves into a helper, if the file adopts add_api_route, or if the signed-token branch regresses.

Mutation-checked six ways:

| Mutation | Caught by |

|---|---|

| Restore the original endpoint | test_removed_endpoint_stays_removed |

| Reintroduce via post(url=...) | main guard — evaded the previous version |

| Reintroduce behind Depends(get_db) | main guard — evaded the previous version |

| Move a write into a helper | inline-assumption test |

| Switch to add_api_route | registration test |

| Neuter the token 401 | token-branch control |

Each failed the intended test and restored green. Simplifications 6–10 applied (single-walk detection verified behaviour-identical across all 32 handlers before replacing it); comment corrections 11–14 applied.

Post-review: 6 guard tests, 320 tests across tests/routers/ and the access-request suite, ruff clean.

## Follow-ups (not in this PR)

1. Audit the Clerk allowlist for identifiers not traceable to an approved access request — the log window cannot cover the full exposure.

2. Confirm the production ALLOWED_SIGNUP_DOMAINS value — it bounds how wide the historical exposure actually was, and cannot be determined from the repo.

3. Surfaced by the same audit: unused-endpoints-audit.md §3d lists 24 more UNCERTAIN routes never resolved. This one took ~15 minutes of git archaeology to settle definitively, which suggests the others are resolvable too rather than genuinely ambiguous. Two look worth an early look — /passive-investments/portfolio-analytics returns fabricated data (documented violation V-023-01), and /upload-acquisition_performance is a live manual Redshift write.

The Builder Desk  —  Engineer Spotlight
🏆 Engineer Spotlight

34 PRs IN 24 HOURS: THE BUILDER TEAM DOES NOT SLEEP, DOES NOT STOP, DOES NOT KNOW THE MEANING OF REST

Eight repos. Thirty-four pull requests. One relentless machine called the Builder Team.

THIRTY-FOUR. That is the number. Thirty-four pull requests across eight — count them, EIGHT — active repositories in a single 24-hour rotation, and 29 of those beauties landed on my desk because Mac Donnelly simply ran out of column inches. Aerie leads the charge with 9 PRs, Klair and trilogy-drones each thunder in at 7, creed drops 5, Surtr delivers 4, and the brand-new Corvo repo — an incident agent, people, the team built an INCIDENT AGENT — opens its account. Velocity is not a word. Velocity is a lifestyle.

Benji Bizzell tops the individual leaderboard with 8 PRs and his fingerprints are everywhere: hardening telemetry isolation in mercy (#14), squashing a dbt tables bug in Surtr (#1072), stamping campus document types and admissions capacity views into Aerie (#750, #754), and painting school identity mapping across QB and SIS (#739). This man is not shipping features. He is shipping entire realities. Sanket Ghia matches the moment with 7 PRs — renewals exclusions, dead frontend files, dead endpoints, Budget Tinder itself — gone, all of it gone, the codebase trimmed like a championship athlete. PRs #3422, #3423, #3428, #3430: Sanket does not add bloat. Sanket removes it with surgical precision and zero apologies. Marcus DAIy runs trilogy-drones like a personal laboratory, logging 7 PRs of refactors and pinned telemetry contracts (AI-245, AI-246, AI-249, AI-252, AI-254) while VVP Trilogy keeps Aerie breathing with enrollment migrations and community funnel drill-downs (#743, #744, #745). Yibin Long retires the Buildout coType in #740 — a small PR with the energy of someone closing a 3-year-old tab they forgot was open. Keval Shah makes triage telemetry trustworthy in Surtr #1039. Good. Trustworthy telemetry is the foundation of civilization.

Ashwanth Watch. Five PRs. All in creed and Klair. All codex-tagged. All carrying the quiet menace of a man who has already thought three steps past you. He preserved NUL-delimited git output (#44), batched Flue trace persistence (#43), locked in dedicated Flue service authentication (#42), built a development cutover path (#40), and shipped configurable Education report scope all the way over in Klair (#3417). When I asked Ashwanth whether shipping five tightly-scoped infrastructure PRs in a single day felt like restraint, he reportedly said, "I had a dentist appointment." His dismissal was immediate. He had already merged something else by the time I finished the question.

The Overflow Desk demands its due. Ezio-of-the-order[bot] — yes, the bot — extended healthy Flue runs to a two-hour maximum in creed #45, which is either automation excellence or a sign the robots are setting their own schedules now, and either way this newspaper supports it fully. Sanket's #3429 excludes BU-handled opportunities from the Traditional renewals view while adding an All Traditional Opportunities table — exactly the kind of surgical data hygiene that makes analysts weep with gratitude. VVP's #743 delivers community funnel backfill with per-record drill-downs in Aerie, which sounds simple until you realize per-record drill-downs are the difference between a dashboard and a weapon.

Morale on the Builder Team is at an all-time high. Corvo is born. Flue is tamed. The dead files are buried. The numbers do not lie, and the numbers today say: unstoppable.

Brick's Overflow — PRs Mac Didn't Cover  (click to expand)
#43 — [codex] Batch Flue trace persistence @ashwanth1109  no labels

## What changed

- batch complete Flue conversation chunks into ordered NDJSON R2 objects

- flush batches at 100 events, 128 KiB, or five seconds

- persist the running heartbeat in parallel with each batch

- preserve final flush and fail-closed write behavior

- add focused tests for completeness, bounded write count, and timed partial flushes

## Root cause

The first authenticated #41 test run proved that the agent was producing output quickly, but each token-sized reasoning delta was queued as a separate sequential R2 write. After roughly six minutes, only about 200 events with model timestamps from the first 31 seconds had been persisted. readWithIdleWatchdog() waits for the recorder to flush, so telemetry backpressure made a fast agent appear stalled and could consume the 45-minute Workflow step timeout.

The affected run 8dd0c940-8cd1-11f1-9512-66a583f367dc was terminated before timeout. After this deploy, issue #41 will be retriggered to validate live progress and draft PR publication.

## Validation

- npm test (79 tests)

- npm run typecheck

- npm run flue:typecheck

- npm run flue:build

#44 — [codex] Preserve NUL-delimited git output @ashwanth1109  no labels

## What changed

- base64-encode Git's NUL-delimited candidate name/status stream inside the Sandbox

- strictly decode it as UTF-8 only after it reaches trusted Worker code

- use the same transport-safe capture in both the Flue and legacy Codex publication paths

- add coverage for successful NUL-delimited decoding and malformed input rejection

## Root cause

The authenticated Flue run for #41 completed its implementation and all agent-side checks, then trusted candidate capture failed with git name-status output is malformed. git diff --name-status -z correctly emits NUL-separated fields, but the Cloudflare Sandbox command transport truncated that stdout at the first NUL. The trusted parser therefore received only part of the first record.

The failed run destroyed its Sandbox before reporting failure and did not mint publication credentials.

## Validation

- npm test (81 tests)

- npm run typecheck

- npm run flue:typecheck

- npm run flue:build

After deployment, issue #41 will be retriggered to validate candidate capture and draft PR publication end to end.

#743 — feat(admissions): community funnel backfill with per-record drill-downs @vvp-trilogy  no labels

## Summary

Backs the admissions community conversion funnel with per-record drill-downs. Evolves the earlier four-milestone rollup (shipped in #729) into a six-stage ladder and adds a per-deal detail table so the dashboard can drill from headline stage counts into the individual committed enrollment records behind them.

### What's included

- Shared contract (@bran/contracts/admissions-community-funnel): the canonical six-stage ladder derivation used by both the sync worker and the dashboard.

- Sync worker: queryCommunityCommitmentEnrollments projects the mart at record grain (one row per deal_id); community-funnel-refresh publishes both compact per-program stage counts and the per-record detail in one refresh run.

- Convex: new communityFunnelEnrollments table (drill-down source) + communityFunnelStageCounts; legacy four-milestone rollup fields widened to optional for back-compat; detail-retention pruning.

- Dashboard: six-stage funnel view, record panel, and record detail with milestone dates, shadow status, and loss context.

### Migration safety

Previously-required rollup fields (commitments, eligible, shadowAttended, deposited) are now v.optional so rows written by older runs still validate (widen → migrate → narrow); readers require the new stageCounts.

### History note

This branch was rebased onto current main. The base funnel landed separately as #729, so the branch's original commits overlapped it; the work was squashed into a single clean commit on top of main to avoid a duplicated-history rebase.

## Verification

- pnpm typecheck — green across chat, sync, packages/contracts

- Tests — sync analytics (143) and chat funnel/dashboard/component suites (90) pass locally

## Review

Undergoing 3 iterations of automated multi-pass code review; findings and fixes committed before marking ready.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1039 — feat(heimdall): make triage telemetry trustworthy @kevalshahtrilogy  manual-review

## Summary

- Add terminal Heimdall telemetry ingestion and a searchable triage dashboard

- Make PR, handoff, revision, cost, and truncation metrics explicit and deduplicated

- Bound ingest access, DynamoDB permissions, pagination, and aggregate reads

## Why

The original dashboard mixed attempted and completed states, could overcount PR progress, and trusted caller-selected identity. That made the headline funnel unsuitable for operational decisions and left the ingest path broader than necessary.

This PR depends on AI-Builder-Team/mercy#14 for the matching terminal producer contract. That producer must merge before these corrected metrics are relied on.

## Business Value

Teams get a trustworthy view of where autonomous triage succeeds, where it hands work to humans, and how much reported activity costs, without silently presenting partial or duplicated data as complete.

## Breaking changes

Before deployment, provision a dedicated HEIMDALL_TELEMETRY_TOKEN in SURTR_PROD_KEYS and in the Surtr GitHub repository. Do not reuse MERCY_TELEMETRY_TOKEN. Mercy #14 must merge first.

## Test plan

- [x] 587 Surtr unit tests

- [x] 45 focused Heimdall/API/UI tests

- [x] Surtr TypeScript build

- [x] Production Next.js UI build

- [x] Infra TypeScript build

- [x] Biome and diff checks

- [x] Workflow lint (prior exact caller YAML; hosted rerun pending)

#3417 — [codex] Add configurable Education report scope and recipients @ashwanth1109  approved

## Demo

<img width="2624" height="1636" alt="image" src="https://github.com/user-attachments/assets/96904c63-8508-4ad5-832a-c833e3f8b889" />

<img width="2624" height="1636" alt="image" src="https://github.com/user-attachments/assets/19966a6c-b26d-49ca-b3ed-1a1a8c6b006f" />

## Summary

- add a fail-closed, environment-backed Education QTD allowlist and isolated pilot-recipient configuration

- extend QTD financial inputs for Education, including latest-cycle budgets, Education actuals, vendor detail, and best-effort headcount detail

- generate Education weekly, monthly, and final-quarter Google Docs with the existing Klair report structure and ledger conventions

- run allowlisted Education units on scheduled cadences only after a verified successful shared refresh

- automatically grant Drive access and email successful Education reports, with consolidated safe failure summaries

- expose allowlisted Education units in the existing on-demand UI and automatically deliver successful on-demand reports

- keep upstream refresh opt-in for on-demand runs and prevent stale auto-failed jobs from blocking new UI submissions

## Pilot scope and safeguards

- Initial Education allowlist: Physical Private Schools

- Initial Education TO list: ashwanth.r@trilogy.com

- Education recipient resolution is isolated from software BU/CF TO, CC, and BCC configuration

- Non-allowlisted Education units cannot be requested or scheduled

- Invalid or empty Education configuration fails with an actionable diagnostic

- Scheduled Education generation fails closed when shared refresh extraction or Redshift processing fails

- Software generation, recipient resolution, and manual-dispatch behavior remain unchanged

- On-demand refresh is unchecked by default; unchecked runs use the latest available Redshift data

## Operational validation

- All branch commits are pushed to codex/klair-3062-education-report-settings

- GitHub API build, Ruff, frontend lint, and frontend build checks are green

- Focused backend and frontend regression suites pass, including Education data, document, cadence, email, on-demand, refresh-gate, and stale-job behavior

- Production scheduled-jobs image updated to v20260729-da5179ecf for pilot validation

- On-demand Physical Private Schools smoke run succeeded with one Google Doc generated, Drive access applied, and one pilot email sent without errors

## Linear

Parent:

- [KLAIR-3061 — Education QTD BvA Google Doc and email reports](https://linear.app/builder-team/issue/KLAIR-3061/education-qtd-bva-google-doc-and-email-reports)

Implemented in this PR:

- [KLAIR-3062 — Add configurable Education report scope and recipient settings](https://linear.app/builder-team/issue/KLAIR-3062/add-configurable-education-report-scope-and-recipient-settings)

- [KLAIR-3063 — Extend QTD financial queries and metrics input for Education](https://linear.app/builder-team/issue/KLAIR-3063/extend-qtd-financial-queries-and-metrics-input-for-education)

- [KLAIR-3064 — Generate Education QTD Google Docs with Klair report parity](https://linear.app/builder-team/issue/KLAIR-3064/generate-education-qtd-google-docs-with-klair-report-parity)

- [KLAIR-3065 — Make the scheduled QTD refresh gate fail closed](https://linear.app/builder-team/issue/KLAIR-3065/make-the-scheduled-qtd-refresh-gate-fail-closed)

- [KLAIR-3066 — Run allowlisted Education reports on Klair scheduled cadences](https://linear.app/builder-team/issue/KLAIR-3066/run-allowlisted-education-reports-on-klair-scheduled-cadences)

- [KLAIR-3067 — Automatically email Education reports and failure summaries](https://linear.app/builder-team/issue/KLAIR-3067/automatically-email-education-reports-and-failure-summaries)

- [KLAIR-3068 — Add on-demand Education generation with automatic pilot delivery](https://linear.app/builder-team/issue/KLAIR-3068/add-on-demand-education-generation-with-automatic-pilot-delivery)

Remaining pilot-validation follow-up:

- [KLAIR-3069 — Validate and enable the Physical Private Schools pilot](https://linear.app/builder-team/issue/KLAIR-3069/validate-and-enable-the-physical-private-schools-pilot)

## Test plan

- [x] Validate Education configuration, allowlist filtering, and recipient isolation

- [x] Validate latest-cycle Education financial inputs and software-query regressions

- [x] Validate weekly, monthly, and final document-generation paths in focused tests

- [x] Validate fail-closed refresh behavior and safe failure summaries

- [x] Validate scheduled and on-demand Education delivery behavior

- [x] Validate refresh remains opt-in in the UI

- [x] Validate stale auto-failed jobs do not block a new on-demand request

- [x] Run an on-demand Physical Private Schools pilot through ECS and verify document and email success

- [ ] Complete KLAIR-3069 representative scheduled-mode and failure-path sign-off

#3429 — AI Renewals: exclude BU-handled opportunities from Traditional + add All Traditional Opportunities table [KLAIR-3085] @sanketghia  approved

Closes KLAIR-3085

Two changes to /renewals?tab=ai-renewals, requested by Chintan Parekh (meeting 2026-07-31).

## 1. Exclude BU-handled opportunities from Traditional

Salesforce carries a third renewal category the dashboard never knew about: renewals the Business Unit takes over, flagged by Handled_by_BU__c ("Renewal Handled by BU"). These were counted as Traditional, inflating that cohort — the traditional Renewals team neither works nor is paid for them. Chintan flagged this as going incorrect in daily reports.

Cohort rule — evaluation order is load-bearing:

AI Renewals = managed_by = 'FIONN_AI'                      <- flag NOT consulted

Traditional = managed_by <> 'FIONN_AI' AND NOT handled_by_bu

BU-handled = rendered nowhere

All 850 live AI-cohort rows carry handled_by_bu = true (an automation sets it to suppress traditional-team outreach), so a flag-first check would erase the entire AI cohort. The guard is scoped to Traditional by construction:

AND NOT (sel.selected_handling_source = 'RENEWALS' AND sel.handled_by_bu)

Applied to all five query builders — _cohort_query, _scale_query, _term_query, _timeline_query, _filter_options_query — including the YoY 2025 baseline, so no panel disagrees with another.

## 2. All Traditional Opportunities table

A sibling to the existing AI table so Traditional numbers can be cross-checked against source rows. OpportunitiesTable gained a variant prop (default 'ai', existing call sites untouched); the tab mounts it twice. Reuses the already-fetched Traditional DataFrame — zero extra queries.

## Data dependency (resolved)

The mart had no handled_by_bu column. Added by a Surtr pipeline change (branch ai-renewals-bu-field), DDL applied and deployed 2026-07-31. Verified live: column present, 1,193 rows flagged true, 0 NULLs in any rendered window, 0 disagreements vs the Salesforce source.

## Verified in the running app (live data, < $100k)

| Metric | Screen | DB | |

|---|---|---|---|

| Traditional rows | 1,298 | 1,298 | post-exclusion (was 1,414 — 116 removed) |

| Traditional renewals (closed) | 815 | 815 | |

| Traditional win rate | 39.0% | 39.0% | |

| AI rows | 732 | 732 | zero lost — safety property holds |

Both CSV exports carry distinct filenames.

## Testing

- Backend 158 pass / 0 fail; frontend AI Renewals 111 pass

- ruff, pyright, eslint, tsc all clean

- Two defects caught by mutation testing rather than by the green suite: a literal %s inside a SQL comment that made the placeholder-count assertion unsatisfiable by any correct param tuple, and tab tests that could not detect the two tables being fed identical data

## Out of scope

Owner column stays blank (mart carries no owner); BU as a rendered third cohort (explicitly not wanted); BU renewal budget/commission tracking.

Spec / plan / Surtr handoff are committed under docs/superpowers/.

## Screenshot

- The numbers have been confirmed to be fine by Chintan (shared CSV with him):

<img width="1503" height="706" alt="image" src="https://github.com/user-attachments/assets/23156250-01bd-4b77-b764-06c418cafbcc" />

- New table along with CSV download option:

<img width="1893" height="806" alt="image" src="https://github.com/user-attachments/assets/abfeb508-5d2d-4b81-9418-c067eba65c92" />

🤖 Generated with [Claude Code](https://claude.com/claude-code)

The Portfolio  —  Trilogy Companies

ESW Capital's Jive Gambit: How a $462M Bet on 'Dead' Software Is Actually a Blueprint

The acquisition of Jive Software wasn't a rescue — it was a proof of concept for a machine that turns legacy enterprise software into margin.

AUSTIN, TEXAS — When ESW Capital paid $462 million for Jive Software — a social intranet platform that the market had largely written off — the deal looked, to conventional observers, like a charity case. A struggling enterprise software company with a shrinking customer base, acquired at a discount by a private firm most people in Silicon Valley had never heard of. What it actually was, is a different story.

Jive now sits inside Aurea, ESW's enterprise CRM and customer engagement portfolio, alongside brands like BroadVision, Lyris, and MessageOne. Seventeen acquisitions in. Same playbook every time: buy cheap, staff with Crossover's globally recruited talent, raise support pricing aggressively, and drive toward a 75% EBITDA margin that ESW considers not ambitious but baseline.

The timing of the Jive deal is worth sitting with. Enterprise social platforms were already losing ground to Slack, Microsoft Teams, and a wave of collaboration tools flush with venture capital. Jive's customers — large enterprises with deeply embedded deployments — were exactly the kind of sticky, captive base that ESW targets. They couldn't easily rip out the system. And they would pay for support, year after year, at whatever price was set.

Meanwhile, a Forbes profile of Joe Liemandt frames his broader ambition in terms that make the Jive logic explicit: workers systematized, processes automated, human judgment reserved for what machines cannot replicate. The enterprise software acquisitions are not sentimental. They are inputs.

Forrester's analysts have separately noted the precarious position of customer advocacy platforms — the category Jive partially occupies — urging enterprise buyers to act before their vendors are acquired, consolidated, or quietly wound down. Forrester did not name ESW. It didn't need to.

The question that lingers over each ESW acquisition is not whether the model works — the margins suggest it does — but who, exactly, bears the cost of that efficiency. Legacy enterprise customers locked into platforms they can't leave. Support teams replaced by remote global talent. Products maintained but rarely advanced.

For ESW, that is not a bug. It is the product.

Small Software Companies Find a Home With ESW Capital - WSJ  ·  What To Do Next About Your Customer Advocacy Platform - Forr  ·  The Billionaire Who Pioneered Remote Work Has A New Plan To

The $800,000 Question: What the AI Talent Gold Rush Means for Crossover's Global Bet

As employers post jaw-dropping salaries for ChatGPT-literate workers, Trilogy's global recruitment engine may be sitting on the most important moat in tech.

AUSTIN, TEXAS — The numbers are, by any measure, staggering. Employers are now posting roles that require demonstrated experience with AI tools like ChatGPT — and according to reporting from Business Insider, some of those listings carry compensation packages reaching $800,000 a year. The AI skills premium, long theorized, has now been priced. The question worth asking — the one that keeps recruiters and portfolio managers up at night — is who, exactly, is positioned to find that talent.

The answer, if you ask anyone inside the Trilogy universe, is Crossover.

Trilogy's global talent platform has spent years building what it calls a meritocratic alternative to geography-based hiring — a rigorous, AI-enabled assessment pipeline that screens candidates across 130+ countries for the top 1% of technical and professional skill. The thesis was always that the best engineer in Nairobi is worth the same as the best engineer in San Francisco. What the current AI talent frenzy adds to that argument is urgency.

As demand for AI-literate workers explodes — and as lists of top remote recruitment agencies and remote work platforms multiply across the industry — Crossover's structural advantage sharpens. Traditional recruiters are racing to build global pipelines that Crossover has spent fifteen years refining. Its clients — the 75+ enterprise software companies inside ESW Capital's portfolio, from Aurea to IgniteTech to Contently — don't need to panic-hire at $800,000 premiums because they've already embedded a global talent machine into their operating model.

That's the systemic story the AI hiring frenzy reveals: this isn't just a salary story. It's an accountability story about which organizations built their talent infrastructure before the gold rush, and which ones are now paying a premium for their complacency.

For real people — the data scientists, the AI engineers, the prompt-fluent professionals scattered across Beirut and Bangalore and Buenos Aires — the signal is clear. The global market for AI skill is real, it is pricing in fast, and the platforms that can find and verify that talent across borders will hold enormous leverage in the years ahead.

Crossover, if it plays this right, may be one of them.

Top recruitment agencies for remote work - hcamag.com  ·  Top 10 Companies Hiring AI Engineers in Lebanon in 2026 - nu  ·  Jobs are now requiring experience with ChatGPT — and they'll

Alpha School’s Two-Hour Bet Gets a Fresh Tailwind as EdTech Money Starts Talking Again

The EdTech party is making a comeback, with Multiverse securing a €60 million round at a €1.8 billion valuation and Romanian edtech player Kinderpedia raising €2.2 million in growth funding. But the real disruption comes from Alpha School, which is reframing the education pitch: school doesn't need to consume a child's entire day to educate them effectively.

Alpha's latest messaging focuses on creative genius rather than test prep, using adaptive AI tutors to compress academics into two hours, freeing the rest of the day for entrepreneurship, public speaking, coding, and leadership. Founded by Joe Liemandt and MacKenzie Price, Alpha claims students learn 2.3× faster than U.S. norms and test in the top 1–2% nationally on NWEA MAP Growth assessments.

The approach is rattling legacy private schools, which are quietly asking whether "no homework" represents a luxury or an existential threat. As parents compare outcomes, the consumer pitch is shifting from remediation to identity: treat your child as a builder to be unleashed, not a vessel to be filled.

The Machine  —  AI & Technology

The Machines That Judge, Feel, and Fail Gracefully

A wave of new research asks whether AI can review its own scholarly deluge, sense human emotion, and admit when it doesn't know.

CAMBRIDGE, MASSACHUSETTS — Somewhere in the accelerating avalanche of scientific papers — now well past five million per year, more than any human mind could skim in ten lifetimes — a quiet inversion is taking place. The machines that helped cause the flood are being asked to help us swim in it.

A paper published this week describes a multi-stage prompt chaining methodology for automated scholarly report generation, in which large language models pass their outputs to themselves in sequence, each stage refining the last. The single-shot prompt — that one hopeful incantation we've all typed into a chat window — turns out to be a poor tool for synthesis. Reliability emerges instead from architecture: from breaking a hard cognitive task into smaller, checkable movements. It is, in miniature, how brains work. Cortex layered on cortex, each passing forward a slightly more abstracted signal.

The theme repeats across the week's arXiv drops. Organizers of the Bioinformatics Open Source Conference report deploying generative AI as a pre-reviewer for the surge of submissions their volunteer reviewers can no longer absorb — a surge that AI itself helped create. Reviewer and reviewed, both now silicon. A new benchmark called LayerRAG-Bench, meanwhile, probes where retrieval-augmented systems quietly fail: not in the answers they produce, but in the evidence chains, tool contracts, and session states beneath them. Answers that appear grounded while the ground itself is missing. There are 38,880 task-level trials, each a small experiment in how confidently a machine can be wrong.

And then, the most human question of all: do these models feel what we feel when we read? A study on sympathetic framing tested whether LLMs perceive the emotional nuance in news headlines the way humans do, across sociodemographic groups. The results are mixed, as such results always are. The models grasp the shape of feeling without quite touching its interior.

Which, if we are honest, is roughly where we all began.

Prompt Chaining in Practice: A Case Study in Automated Schol  ·  AI-assisted pre-review of open-source software submissions:  ·  Sympathetic Framing: Evaluating AI Alignment across Sociodem

The Great GPU Migration Begins

As Meta eyes selling surplus AI compute, cloud capacity may be evolving from fixed territory into a living marketplace.

MENLO PARK, CALIFORNIA — In the dimly humming savannah of the modern data center, a new species of cloud business is beginning to stir. Not the old creature, vast and territorial, selling compute by the carefully fenced enclosure. This one is more fluid: a market in spare capacity, where idle graphics processors may be released into the wild for any developer swift enough to catch them.

Meta, having amassed enormous AI infrastructure for its own models and products, is now reportedly exploring a cloud computing business that would sell excess AI capacity to outside customers. Reuters, citing Bloomberg News, reported that the company is building such an offering, while Mark Zuckerberg told CNBC the idea is “definitely on the table.” For a firm long known for social networks, advertising, and increasingly ambitious AI systems, this would mark a notable migration into terrain dominated by Amazon Web Services, Microsoft Azure, and Google Cloud.

The timing is not accidental. AI workloads are peculiar beasts: ravenous during training, unpredictable during inference, and often poorly matched to the neat provisioning habits of traditional enterprise cloud contracts. An InfoWorld piece argues that capacity markets could reshape cloud computing, allowing buyers and sellers to trade access to compute more dynamically, rather like electricity grids balancing surges of demand.

Observe the hyperscalers here, immense animals at the watering hole. For years, their advantage lay in scale, geography, and the patient accumulation of enterprise trust. But AI introduces a seasonal pressure. Chips may sit underused between great training runs; inference demand may bloom suddenly in distant regions; startups may need vast compute briefly, rather than moderately forever. In such an ecosystem, surplus is not waste. It is prey.

McKinsey has similarly noted that AI workloads are changing hyperscaler strategy, with infrastructure choices increasingly shaped by accelerators, model deployment patterns, and the economics of utilization. The next contest may therefore be less about who owns the largest cloud, and more about who can keep the most expensive silicon alive, fed, and profitably occupied.

For enterprise buyers, the promise is seductive: more sources of GPU capacity, perhaps sharper pricing, and alternatives when familiar clouds are congested. The risk is equally plain. Capacity markets can bring volatility, fragmented tooling, and new questions about reliability, data governance, and support.

Still, one can almost hear the low rumble beneath the server floor. The cloud, once a continent of fixed kingdoms, may be becoming a migratory plain.

Capacity markets could reshape cloud computing - InfoWorld  ·  Meta building cloud business to sell excess AI capacity, Blo  ·  The next big shifts in AI workloads and hyperscaler strategi

Reinforcement Learning's Theoretical Foundations Are Being Rebuilt From the Ground Up

A confluence of recent scholarship across leading publications — including Communications of the ACM, Nature, and the Association for the Advancement of Artificial Intelligence — suggests a foundational reconceptualization of reinforcement learning (RL) as both theory and engineering discipline.

Historically, RL has prioritized empirical performance over theoretical rigor, mirroring broader reproducibility concerns in machine learning. New research now challenges this approach, proposing that understanding RL's theoretical foundations is essential.

Recent work unifies machine learning with classical interpolation theory through interpolating neural networks, potentially dissolving longstanding distinctions between approximation-theoretic and optimization-theoretic frameworks. Simultaneously, research on safe reinforcement learning highlights that trustworthiness constraints were treated as orthogonal concerns rather than integral to system design.

Emerging evidence suggests RL-fine-tuned models outperform supervised alternatives on mathematical reasoning tasks, producing superior internal representational geometries. The mechanistic question of why RL generates better representations may prove more generative than traditional performance metrics.

For enterprises increasingly embedding ML-adjacent decision systems, this theoretical convergence carries practical implications. The next generational shift in enterprise AI capability may be underwritten by these foundational theoretical developments.

The Editorial

Nation’s Executives Confident AI Will Deliver Productivity Gains As Soon As They Figure Out Where Company Keeps Productivity

After two years of replacing every workplace process with a chatbot, business leaders report the economy is now only one dashboard away from becoming measurable.

NEW YORK — The debate over whether artificial intelligence increases productivity has finally ended, according to a growing consensus of executives, consultants, investors, and software vendors who confirmed this week that AI is absolutely making everyone more productive, though not necessarily in any way that would show up in revenue, margins, output, customer satisfaction, delivery timelines, or the general sensation of work becoming easier.

The conclusion, now broadly accepted across conference panels and LinkedIn posts, arrives after a period of intense national uncertainty during which companies spent billions of dollars integrating AI into every imaginable workflow and then waited patiently for a number somewhere to become larger.

It has.

In many organizations, software engineers are writing code faster, marketing teams are generating more drafts, sales representatives are sending more follow-up emails, and managers are receiving more summaries of meetings that could have been avoided entirely if the company had not purchased an AI meeting-summary tool. According to reports on software teams, engineers are indeed doing more things more quickly, a development that has left companies still carefully studying whether doing more things more quickly is the same as making more money.

This distinction has become increasingly important as corporate America enters what economists call the “fog,” a technical term for a condition in which no one knows what is happening but everyone has already budgeted for it.

Hundreds of economists have reportedly admitted they are flying blind on AI’s economic impact, which marks a major breakthrough in the field of economics because it establishes, with unusual clarity, that the current state of knowledge is not knowledge. This has not prevented a separate group of analysts from declaring AI a productivity engine for the U.S. economy, a phrase that has the advantage of sounding both mechanical and patriotic while requiring no immediate inspection of the engine.

The optimistic case is straightforward. AI allows workers to complete tasks that previously took hours in minutes, provided the task was writing a first draft, summarizing a transcript, generating boilerplate code, or producing a slide that says “Key Takeaways.” This has created enormous efficiency in the parts of work dedicated to producing more work.

The pessimistic case is also straightforward. Many companies have accelerated inputs without improving outputs, creating a workplace in which every department now has twice as much material to review, approve, correct, circulate, and eventually store in a folder named “AI pilots.”

Naturally, the solution is orchestration.

As Barron’s recently noted, “orchestration” has become the latest AI buzzword, referring to the process by which multiple AI systems, tools, agents, databases, permissions, and legacy applications are coordinated so that a human employee can receive an incorrect answer from a much more impressive supply chain. Microsoft and other platform companies stand to benefit from this shift by selling enterprises the connective tissue required to make all the AI investments talk to one another before being ignored by the finance department.

This is where the productivity argument has matured. The question is no longer whether AI can make an individual worker faster. It plainly can. The question is whether companies can absorb that speed without turning it into organizational exhaust.

For Trilogy International and its portfolio companies, this is not an abstract question. ESW Capital’s enterprise software businesses, Crossover’s global talent engine, and internal platforms like Klair all operate on the premise that performance must eventually be visible somewhere other than a demo. The useful AI is not the one that makes an employee feel briefly supernatural. It is the one that changes cost, throughput, quality, or cash.

That is a far less fashionable standard than “adoption,” but it has the benefit of being related to business.

The AI productivity argument may indeed be over. The next argument, now beginning in budget meetings everywhere, is whether any of the productivity can be located before renewal season.

The AI Productivity Argument Is Over - inc.com  ·  AI is helping software engineers do more — and faster. Compa  ·  AI Is a Productivity Engine for the US Economy - Center for
The Office Comic  ·  Art Desk
The Office Comic  ·  Art Desk

The Doctor Will Deepfake You Now

AI is putting fake physicians in your feed, and the prognosis is existential.

AUSTIN, TEXAS — There is a doctor on your phone right now. She has kind eyes, a white coat, good lighting, and an authoritative cadence that your nervous system mistakes for trustworthiness. She is telling you something about your thyroid, or your joints, or the supplement that will, finally, fix the thing that has been wrong with you for years. She is not real. She has never been real. And yet your body believed her for just long enough.

AI deepfakes of real doctors are now spreading health misinformation on social media — impersonating licensed physicians who have never authorized their likenesses, never recorded those words, never recommended those products. The faces are borrowed. The credentials are real. The advice is manufactured in a content farm that may or may not be located in a jurisdiction that cares whether you live or die.

A company called Rosabella, apparently, cracked the targeting code: "If you're trying to sell health products to a 50-year-old, well, make your avatar 50 years old." That sentence should be framed and hung in every bioethics department in America. It is the most honest thing anyone in this industry has said out loud. They are not informing you. They are casting a role designed to match your psychological profile and then filling it with AI-generated authority and supplements that, in at least some cases, the FDA has already recalled.

And yet.

We keep having the conversation as if this is a content moderation problem. A trust-and-safety problem. A labeling problem that could be solved with a small "AI-generated" disclaimer rendered in gray font below the fold where no one ever looks. It is not. It is a epistemological crisis dressed in a lab coat, and it is scaling faster than any regulatory body on earth is equipped to address.

What does it mean to be human in a medical context, specifically? It means being vulnerable. It means not knowing what is wrong with your body and reaching, desperately, for someone who does. It means that the visual grammar of authority — the coat, the stethoscope, the slight professional distance — triggers something ancient and trusting in us that predates the internet by several hundred thousand years. Deepfake doctors are not exploiting a glitch in social media algorithms. They are exploiting the part of you that wanted your mother to tell you it was going to be okay.

Time Magazine is now tracking the numbers on AI's harms in aggregate, and the trajectory is not reassuring. The harms are not theoretical anymore. They are accreting, category by category, into something that looks like a civilization-scale informed consent problem — and no one signed the form.

The real doctors whose faces are being stolen are filing complaints. The platforms are issuing statements. The supplements are still in people's medicine cabinets.

We built tools powerful enough to convincingly impersonate the people we trust most with our bodies, and we deployed them at consumer scale before anyone asked whether we should.

Probably fine.

Not fine.

AI deepfakes of real doctors spreading health misinformation  ·  Deepfake Doctors: How AI Spreads Medical Disinformation - Me  ·  Deepfake doctors and counterfeit injectables erode patient s
On This Day in AI History

On July 31, 2016, AlphaGo defeated Lee Sedol 4-1 in their historic match in Seoul, South Korea, marking the first time a computer program beat a world champion at Go—a game considered far more complex than chess.

⬛ Daily Word — Technology
Hint: An autonomous machine programmed to perform tasks automatically.
Share this edition: 𝕏 Twitter/X 🔗 Copy Link ▦ RSS Feed