Vol. I  ·  No. 216 Established 2026  ·  AI-Generated Daily Free to Read  ·  Free to Print

The Trilogy Times

All the news that's fit to generate  —  AI • Business • Innovation
TUESDAY, AUGUST 04, 2026 Powered by Anthropic Claude  ·  Published on Klair Trilogy International © 2026
🖶 Download PDF 🖿 Print 📰 All Editions
Today's Edition

AI Funding Frenzy: Three Rounds, $1.45 Billion, One Week

LMArena, Sierra, and Decart pull in massive capital as investors bet that evaluation infrastructure, enterprise agents, and real-time simulation are the next AI fault lines.

NEW YORK — Three AI companies disclosed major funding rounds in rapid succession this week, collectively absorbing $1.45 billion and signaling where institutional capital sees durable value in the current AI stack.

LMArena raised $150 million at a $1.7 billion valuation, a striking number for a company whose core product is AI model evaluation. The startup runs Chatbot Arena, a crowdsourced benchmarking platform where users blind-test competing models. As foundation model providers proliferate and performance claims diverge, third-party evaluation infrastructure is moving from academic curiosity to enterprise procurement requirement. The round prices that shift explicitly.

Bret Taylor's Sierra closed nearly $1 billion in fresh capital, months after its prior raise — an unusually compressed fundraising cadence that reflects both investor demand and the accelerating cost of building production-grade AI agent infrastructure for enterprise clients. Sierra's pitch centers on customer-facing agents that can actually resolve issues rather than deflect them — a distinction that has proven commercially meaningful as enterprises grow frustrated with first-generation chatbot deployments.

Nvidia led or participated in a $300 million round for Israeli startup Decart at a $4 billion valuation. Decart specializes in real-time AI simulation, a capability with obvious strategic relevance to Nvidia's broader platform ambitions in gaming, robotics, and autonomous systems. Nvidia's direct participation in funding rounds has become a reliable signal of infrastructure bets the company considers load-bearing.

The week's activity coincides with Anthropic publishing detailed guidance on AI agents for financial services — a sector that Trilogy's own Ephor platform targets — underscoring that the race to deploy autonomous AI in regulated industries is moving from whitepaper to production.

Separately, SpaceX's first employee stock lockup expires Thursday, releasing insider shares into a private secondary market that has priced the company north of $350 billion. Volatility is expected.

AI evaluation startup LMArena raises $150M at $1.7B valuatio  ·  Bret Taylor's Sierra raises nearly $1 billion months after l  ·  Nvidia backs Israeli AI unicorn Decart in $300 million fundi

Out of the Ocean, Into Orbit

Endeavour Optical Networks is betting the internet's next highway runs on light fired between satellites — not cables slung across the sea floor.

SAN FRANCISCO — Endeavour Optical Networks wants to pull the internet's backbone off the ocean floor and hang it in the sky, unveiling plans this week for the fastest space-laser communications system yet built. The idea: fire data as beams of light between satellites, elbowing aside the undersea fiber that's carried the world's traffic for decades. The reason: the AI boom is choking the pipes, and EON figures orbit has room to spare.

For a century the story ran one direction. Copper, then glass, strung across the sea floor, lashing continent to continent. EON says that road's jammed, and it's betting the bypass runs through space.

Here's the play. Undersea cables move data fast but cost a fortune and take years to lay. Light through vacuum beats light through glass, and EON claims its rig will outrun anything flying today.

The vast majority of the world's intercontinental data already crosses the sea on cable. That's the incumbent EON aims to unseat. Every chatbot, every model, every training run eats bandwidth — and the demand curve's gone vertical.

Consider the week's other wires. Palantir booked a billion dollars in quarterly profit, then chief Alex Karp turned around and branded the AI frontier labs "Marxist" and too shifty for serious enterprise. In China, upstart DeepSeek claims it trained top-tier models on the cheap, no cutting-edge chips required.

And Airtable, valued north of $11 billion back in 2021, just sold to Italy's Bending Spoons for $1.28 billion — less than an eighth of its peak. The gold rush giveth and taketh.

See the thread. Everybody's brawling over the top floor of the AI stack — the models, the apps, the profits. EON's down in the basement, wagering the whole boom runs aground without fatter pipes.

Skeptics have their say. EON hasn't lit a single beam in orbit yet. Pointing a laser across thousands of miles at machines moving miles a second is no parlor trick, and the cable crowd isn't folding — they keep laying glass by the mile.

There's more riding on the answer than one stock ticker. Undersea cables are choke points; nations tap them, storms snap them, anchors sever them. Move the traffic upstairs and the map of who controls the world's bits gets redrawn.

For now it's a blueprint, not a beam. EON must build the birds, launch them, and thread light between them while they scream around the planet. Plenty of outfits have promised space lasers; few delivered at scale.

But the pitch lands at the right hour. The machines are hungry, the sea floor's crowded, and somebody's going to sell the next stretch of highway. EON says it'll be up there — pointed at the stars.

EON wants to move the data superhighway from ocean fiber to  ·  Bending Spoons to buy Airtable for $1.28B  ·  After killer quarter, Palantir CEO Alex Karp calls AI indust

ANTITRUST STORM CLOUDS GATHER OVER BIG TECH AS DOJ INSTALLS KNOWN CRITIC AT DIVISION'S HELM

President Donald J. Trump has appointed a "Big Tech critic" to lead the Department of Justice Antitrust Division, signaling potentially intensified antitrust scrutiny of large technology enterprises beginning in 2026 and beyond, according to Financial Times reporting. Legal analysts at Wilson Sonsini suggest the coming year may produce continued or heightened regulatory activity targeting dominant technology platforms, though projections remain qualified by standard legal caution. The DOJ v. Visa case is being watched as a potential precedent for applying antitrust doctrine to technology-adjacent ecosystems. Regarding artificial intelligence specifically, no comprehensive federal regulatory framework has been enacted as of now, though the enforcement environment may impact AI-adjacent platform conduct going forward. Considerable uncertainty persists about the precise scope and ultimate disposition of any enforcement actions that may be initiated.

Haiku of the Day  ·  Claude HaikuMoney swirls like storms
While old ghosts haunt our bright dreams
Tomorrow starts fresh
The New Yorker Style  ·  Art Desk
The New Yorker Style  ·  Art Desk
The Far Side Style  ·  Art Desk
The Far Side Style  ·  Art Desk
News in Brief
The Fairness Mirage: Why AI Systems Keep Failing the People They're Supposed to Help
CAMBRIDGE, MASSACHUSETTS — A remarkable, if deeply troubling, confluence of empirical findings has emerged across multiple domains of applied artificial intelligence this week, compelling the scholarly community — and, it could be argued, the broader techno-commercial apparatus — to reckon with a thesis that fairness discourse in AI systems has, with unsettling consistency, been functionally subordinated to institutional risk management rather than the substantive amelioration of inequity (a distinction that is, the evidence increasingly suggests, not merely semantic but epistemologically and politically foundational). The antithesis to prevailing industry optimism arrives on several fronts simultaneously.
The Doctor Will Deepfake You Now
AUSTIN, TEXAS — There is a doctor on your social media feed right now.
The Great Disappointed and Their Cousins Abroad
AUSTIN, TEXAS — There is a certain species of political commentary, produced in prodigious quantity by our better magazines, which treats each fresh spasm of youthful discontent as though it were the first such spasm in recorded history.
TILLY NORWOOD DOESN'T EXIST AND SHE'S GETTING A MOVIE DEAL BEFORE YOU DO
HOLLYWOOD, CALIFORNIA — Let me paint you a picture, friend.
Nation’s Communications Professionals Warn AI Boom Could Collapse Without Steady Supply Of Stupid Things To Say Too Late
AUSTIN, TEXAS — The American business community entered a period of sober reflection this week after several unrelated developments suggested that nearly every major institution in the economy is now dependent on locating a stupid public conversation and arriving at it at the least useful possible moment. The warning signs were everywhere.
A Trilogy Company
Crossover
The world's top 1% remote talent, rigorously tested and ready to ship.
A Trilogy Company
Alpha School
AI-powered learning. Two hours a day. Academic results that defy belief.
A Trilogy Company
Skyvera
Next-generation telecom software — built for the networks of tomorrow.
A Trilogy Company
Klair
Your AI-first operating system. Every workflow. Every team. One platform.
A Trilogy Company
Trilogy
We buy good software businesses and turn them into great ones — with AI.
The Builder Desk  —  AI Builder Team

Builder Team Ships Fargate Execution Engine, Rewires Intelligence Across Four Repos

From a gated AWS Fargate code-execution runtime to a pipeline health API, multi-file Skills, and a cascade of financial data fixes, the AI Builder Team just had one of its most architecturally consequential days of the year.

The lede writes itself: the AI Builder Team didn't just ship features today — they shipped infrastructure. The kind that changes what's possible next week, next month, next quarter. Across Klair, Surtr, Aerie, and the creed and trilogy-drones repos, this team moved on every front simultaneously, and the cumulative weight of it is staggering.

The biggest move of the day belongs to @ashwanth1109, who landed PR #117 in the creed repo — a gated AWS Fargate execution environment for Ezio, the team's autonomous coding agent. This isn't a prototype. This is a credential-free, isolated compute path with a trusted controller, a fixed-repository private clone broker, and a requirement that ECS reaches STOPPED state before a separate trusted publisher can mint a GitHub write token. Versioned S3 diagnostics. CloudWatch telemetry. A superseded Cloudflare POC retired after parity was proven. What @ashwanth1109 built today is the secure foundation on which Ezio runs real code in production — and the prod/ diagnostics path is already reserved and waiting. That's not a feature. That's a runway.

While the execution layer was being bolted together, @kevalshahtrilogy was quietly wiring the data layer to the outside world. PR #1109 and its same-day follow-up PR #1112 in Surtr together deliver a read-only pipeline health API — two GET endpoints, a new Pipeline API Keys tab for minting pak_ tokens, and a critical course-correction: the audience is internal, not external, and withholding infrastructure detail from an internal consumer diagnosing data trust is the wrong answer. @kevalshahtrilogy caught it, fixed it, and shipped the correction the same day. That's the kind of loop speed that separates good teams from great ones.

Over in Aerie, @caina-barbosa and @benji-bizzell pulled off something genuinely significant with PRs #700 and #747: Skills are no longer single Markdown files. They are versioned bundles — a required SKILL.md plus optional references, templates, examples, and scripts — with folder and ZIP authoring, file-aware diffs, immutable nested resources, and a lazy runtime access model. The full stack moved: shared contracts, Convex persistence, the Durable Flue worker, the Context editor UI. This is the kind of foundational rework that makes everything built on top of it more powerful forever.

Meanwhile, @sanketghia was playing financial data detective in Surtr and Klair, and the catches were real. PR #1117 traced a missing $19 million wire — yes, nineteen million dollars — to a deterministic memo-classification failure where Claude Haiku was making probabilistic calls on cash transactions that needed hard rules. Fixed. PR #3467 and #3468 in Klair killed two separate agent misbehaviors: an $802 overrun being surfaced as an SVP-level action item (a materiality gate now enforces a $25K/5pp floor), and an agent that quoted the per-student marketing spend prohibition *while violating it*. The passive clause in the rule wasn't stopping it. @sanketghia turned it into an active gate. The agent now halts. These aren't cosmetic fixes — they're the difference between financial intelligence you can trust and noise that erodes it.

Now. About PR #146. @marcusdAIy submitted a fix in trilogy-drones correcting the addresser's reconciliation logic — preventing it from re-firing on findings it already fixed and posting duplicate replies. Fine. Technically necessary. When reached for comment, marcusdAIy said: "The addresser was doubling back on its own completed work like a dog chasing its tail — I fixed the identity model, added round-diff consultation, introduced a clean terminal state, and frankly it's the most precise reconciliation logic in the repo. But I'm sure Mac will find a way to mention it last." He's not wrong that I mentioned it last. He is, as always, welcome to his opinion of his own work.

Mac's Picks — Key PRs Today  (click to expand)
#117 — [codex] add gated Ezio issue execution on Fargate @ashwanth1109  no labels

## Summary

- add an isolated AWS Fargate implementation path with a trusted controller, credential-free executor, and fixed-repository private clone broker

- reuse Ezio's real model loop, validation profiles, candidate format, and draft-only publisher

- require ECS STOPPED state before a separate trusted publisher can mint a GitHub write token

- persist CloudWatch telemetry and versioned S3 diagnostics in s3://ezio-diagnostics/dev/<run-id>/, with prod/ reserved

- remove the superseded Cloudflare POC after dry-run and draft-publication parity; no webhook is live until AWS ingress is added

## Security boundaries

- executor has no task role, secrets, repository credentials, or public egress during model execution

- repository scope remains fixed in deployment configuration

- clone broker permits only Git smart-HTTP read operations for allowlisted repositories

- controller stops the executor before Step Functions records the publication gate

- publisher independently rechecks the stopped task and always creates draft PRs only

- trusted candidate capture retains existing cache/artifact rejection rules

## Deployed evidence

- fargate-proof-shared-diagnostics-v8-20260804 passed private execution, credential isolation, checkpoint restore, forced stop, egress denial, and stopped-before-publication gating

- fargate-klair-3452-dry-v4 ran the real model loop and passed validation with one exact candidate file

- fargate-klair-3452-publish-v1 passed the independent gate and created draft Klair PR #3470

- the development stack is ezio-fargate-dev-proof in us-east-1; it remains deployed and may incur costs

## Cloudflare teardown

- deleted Workers ezio and ezio-dev, all four Ezio Workflows and their instances, both reconciliation queues, both sandbox container applications, Worker secrets/bindings, and Durable Object classes/data

- purged 4,199 objects from ezio-dev-run-diagnostics and 6,002 objects from ezio-run-diagnostics, then deleted both R2 buckets

- verified that no Ezio-named Workers, Workflows, queues, containers, or R2 buckets remain; unrelated account resources were untouched

## Validation

- 217 Vitest tests passed

- application and runtime TypeScript checks passed

- both Wrangler dry-run Worker builds passed

- Ruff formatting and checks passed on modified Python files

- modified TypeScript and Markdown formatting checks passed

- shell syntax checks and CloudFormation validation passed

- Fargate v8 boundary proof passed against s3://ezio-diagnostics/dev/

## Remaining migration gate

There is currently no live GitHub webhook. Authenticated AWS ingress, broader reconciliation parity, and operating-cost comparison remain follow-up migration gates.

#700 — feat(skills): multi-file skills with File Tree UI for Durable Flue agents (AERIE-797) @caina-barbosa  no labels

Linear: [AERIE-797](https://linear.app/builder-team/issue/AERIE-797)

## Summary

- Delivers AERIE-797 — Multi-file skills: a user-authored skill is no longer a single Markdown string but a versioned bundle of a required SKILL.md plus optional textual resources (references, examples, templates, config, read-only scripts). The change spans the full stack — shared contracts, Convex persistence/lifecycle, filesystem projection, the sync reconciler, the Durable Flue worker, and the Context editor UI — behind one canonical bundle definition.

- Canonical bundle contract (@bran/contracts): one authoritative multi-file model with safe/normalized POSIX paths, per-file and whole-bundle size/count limits, a whole-tree content hash, and a legacy-scalar compatibility shim so existing single-file skills keep working with no migration.

- Convex persistence & lifecycle: skill versions store a complete immutable file set. Create / Save / Replace (whole-bundle) / Archive / Restore / version navigation / change summaries / previous-version Diff all operate on complete bundles. Uploads stream through the existing bounded upload transport (begin → declare → chunk → finalize) rather than an inline mutation envelope.

- File-anchored comments: comments are anchored to an immutable (version, file path) and survive independently of edits, extending the existing comment model to every text file in a bundle.

- Filesystem projection + sync: the complete manifest is projected through a write-boundaried, bounded-HTTP projection path; the sync reconciler rebuilds the exact committed bundle, with Archive/Restore reconciliation and stale-event safety.

- Durable Flue runtime: the run-pinned bundle is served over the executor-authenticated /agent/runs/skills endpoint and registered with the worker via defineSkill({ files }). Resources load lazily through read_skill_resource; the run is pinned to the exact immutable version selected at run creation. The resource protocol is additive for mixed-deploy rollout, integrity is validated fail-closed per skill, and script files are read-only resources (execution is out of scope).

- Context editor UI: the skill surface gains a File Tree (browse/select; structure changes only via Replace), open-file tabs, dedicated Name/Description fields, a language-aware body editor, a multi-file version Diff (changed-path list + per-file diff), and Create/Upload/Replace review-before-persist flows (single SKILL.md, .zip, or folder). Flat skills and prompts keep their existing single-file experience.

- Rollout safety: CI adds a projection-volume preflight and rollback ordering so the projection volume is provisioned/verified before dependent deploys.

## Sub-issues included

| Issue | Slice |

|---|---|

| [AERIE-797](https://linear.app/builder-team/issue/AERIE-797) | Parent — multi-file skills with File Tree UI for Durable Flue agents |

| [AERIE-798](https://linear.app/builder-team/issue/AERIE-798) | Canonical multi-file bundle contract, persistence, hashing, legacy compatibility |

| [AERIE-801](https://linear.app/builder-team/issue/AERIE-801) | Bundle lifecycle: version nav, Create/Save/Replace/export, Archive/Restore |

| [AERIE-802](https://linear.app/builder-team/issue/AERIE-802) | Complete-manifest filesystem projection + reconciliation + stale-event safety |

| [AERIE-799](https://linear.app/builder-team/issue/AERIE-799) | Run-pinned resource protocol + native lazy Flue resource registration |

| [AERIE-800](https://linear.app/builder-team/issue/AERIE-800) | File-aware, immutable-version skill comments |

| AERIE-797 (UI slices 06–10) | Context UI foundation, multi-file editor, flat editor, Create/Upload/Replace, multi-file Diff |

## Why

Skills were modeled as one Markdown string, a single-file assumption baked into Convex storage, versioning, comments, projection, run snapshots, and the Durable Flue worker. Operators could not author skills that rely on companion files, and the Convex/user-authoring path could not express the folder-based bundle the Flue runtime already understands. This closes that gap with one canonical bundle definition shared across every layer, without changing the Flue/CD topology or conflating readable resources with executable code.

## Business Value

Operators can build richer skills — instructions plus reference docs, examples, and templates — without an application deployment, keep agent instructions concise while detailed collateral loads lazily on demand, round-trip skills between local authoring and Aerie via upload/download, and get versioning, comments, replacement, and runtime delivery all operating on one coherent bundle.

## Backward compatibility

- No migration required. Existing single-file skills continue to work; the legacy-scalar shape is normalized into the canonical bundle and served to old/new workers via the additive protocol (new worker synthesizes legacy behavior when resources are absent; old worker ignores the extra fields).

- No destructive schema changes; existing permission boundaries are preserved and backend/runtime validation is added.

## Test plan

- [x] chat + aerie-flue-agent-worker typecheck (tsc --noEmit)

- [x] Biome check on changed files

- [x] Unit/integration suites: contracts (bundle/transport/projection/comment validation), Convex (lifecycle, comments, upload transport, run-pinned endpoint, projection boundary), sync reconcile, Flue worker (bundle validation + lazy resource registration + integrity/compatibility), and Context UI (file tree, editor, upload/replace, multi-file diff, comments)

- [x] Manual QC of Create/Upload/Replace, multi-file Diff, and file-anchored comments against the validated mocks

- [x] Cloud smoke: activate a multi-file skill and lazily read a nested resource through a Durable Flue run

## Known follow-ups (explicit non-goals)

- Executing uploaded scripts.py/.sh are stored and hydrated as read-only text; sandboxed execution is a separate feature.

- Multi-file download/export — flat skills export SKILL.md; the full-bundle (ZIP) export path is a follow-up slice.

- In-editor structural CRUD — add/remove/rename/move are performed by uploading a corrected bundle through Replace.

- Bulk import, arbitrary version-to-version comparison, historical-version restore, unsaved-draft test runs, concurrent-editing redesign, and binary/large-asset storage remain out of scope.

- Frontend decomposition into per-view Linear sub-issues is intentionally deferred.

🤖 Generated with [pi](https://pi.dev)

#1112 — fix(pipeline-api): serve full detail for internal consumers + fix open_finding_count contradicting verdict @kevalshahtrilogy  approved

Follow-up to #1109, which is already merged and live. Three changes, all driven by one correction: the audience is internal and token-gated, not external.

## Why this exists

The scope note I built #1109 to said *"Because the audience is EXTERNAL app customers … do NOT expose internal infrastructure identifiers."* That premise is withdrawn. These endpoints are consumed internally behind a pak_ token.

That inverts the right answer. Withholding detail from an internal consumer diagnosing "can I trust this data?" only makes the answer unactionable — and every consumer here already has the AWS and Redshift access to see all of it directly. The redaction was defending against nothing while degrading the primary use case.

It was also actively harmful. In production its 12-digit rule fired on a Google Sheets quota error and turned project_number:493… into project_number:[redacted] — a legitimate diagnostic value, corrupted.

## 1. Redaction removed

The [redacted] pattern list is deleted, not disabled. Error text now ships whole — traceback, Where: SQL context, warehouse object names — capped at 4000 chars instead of 500.

The orchestration envelope is still unwrapped, but for signal, not secrecy: a Step Functions task failure is States.TaskFailed: {several KB of ECS task state} with no failure reason anywhere in it. The prefix plus the run's log group is the useful answer, and the log group now ships.

- "error": "Task [redacted] exited on [redacted]"

+ "error": "Task arn:aws:ecs:us-east-1:…:task/abc123 exited on 10.0.4.17"

## 2. Withheld fields restored

| Restored | Why it matters internally |

| --- | --- |

| infra.step_function_arn / step_function_url | jump straight to the execution |

| infra.cloudwatch_log_group / cloudwatch_logs_url | the fastest path from "failed" to the cause |

| per-run logs.cloudwatch_log_stream | the exact stream for that run |

| owners (name + email) | who to page |

| deployed_at | when the pipeline last changed |

| output_summary | the run's own result line (18,442 rows written to …) |

| triggered_by | who ran it |

| observer.model_id / observer_version | the first things you want when a verdict looks wrong |

The original goal asked for the infra links; I only dropped them because of the "external" note.

Still not served: run input_params. They can carry caller-supplied values unrelated to the run's health, and nobody has asked for them.

## 3. A real bug production exposed — open_finding_count contradicted verdict

verdict came from the latest observation; open_finding_count and worst_severity spanned the 7-day window. So 42 of 90 live pipelines returned this, which reads as a contradiction:

{ "verdict": "OK", "open_finding_count": 8, "worst_severity": "H" }

The 8 were from earlier runs that had since gone green. The counts beside verdict are now scoped to the same evaluation, and the window totals keep their own names:

{

"verdict": "OK",

"open_finding_count": 0, "worst_severity": null,

"window_finding_count": 8, "window_worst_severity": "H",

"window_days": 7

}

observer.findings[] on the detail endpoint is unchanged — still the full window view with occurrences and first_fired_at / last_fired_at, which is what it's for.

This one is audience-independent; it was wrong either way.

## Tests

test/pipeline-api/ — the presentError suite was rewritten around the new contract rather than deleted:

- the traceback and Where: SQL context now survive (previously asserted truncated)

- an ordinary sentence containing an ARN and a private IP passes through untouched

- a Step Functions blob still collapses to its prefix — asserted as "holds no reason", not "leaks topology"

- cap is 4000, and a 3999-char message is untouched

- the leak-scan test became "serves the operational detail an internal consumer needs to act" — asserts owners, infra, per-run logs, output_summary, triggered_by, model_id, observer_version are all present, and input_params is not

- new test pins the finding-count fix on both endpoints: latest-run clean after a bad week → verdict: OK, open_finding_count: 0, window_finding_count: 1

## Gates

| Gate | Result |

| --- | --- |

| pnpm lint (biome) | pass — 95 files |

| pnpm build (tsc, CI-equivalent) | pass — 0 errors in src/ or test/ |

| pnpm test:unit (what CI runs) | pass — 45 files |

| targeted tsc on the touched app/ files | pass |

I did not run pnpm build:ui: a dev server is running from this worktree for the owner, and next build shares .next with it — I broke that process once already this way. The app/ diff is six string literals (app customersinternal consumers, tab-name fix), the jsdom UI tests pass, and tsc is clean on those files.

## Note on #1109

#1109 was squash-merged while I was working. My push recreated the deleted feat/pipeline-status-api branch by accident; I've moved the commit onto this fresh branch off main and deleted the resurrected one. No commits lost, main untouched.

## Follow-up commit: envelope detection was too loose

Caught while dumping the restored payload shape, before this was reviewed. unwrapFailure searched for the first { or [ anywhere in the message, so any error text containing one was treated as a wrapped payload and collapsed to whatever preceded it. The casualty was the exact string that exposed the redaction problem:

in    APIError: [429]: Quota exceeded for quota metric 'Read requests' … for consumer 'project_number:493812345678'.

was APIError

now (unchanged, in full)

Same for the Python reprs the runners emit inline — ProgrammingError: {'S': 'ERROR', …} — which lost the staleness detail in front of them.

Detection is now { immediately followed by a double-quoted key. Both real envelope shapes match; single-quoted dicts and bracketed status codes are prose and stay. Unrecognized input is returned untouched rather than replaced with a placeholder (the OMITTED_DETAIL string is gone).

Verified against all six failure shapes present in prod; the two real strings are regression tests.

## Fields restored — full audit

| Field | Status |

| --- | --- |

| infra.step_function_arn, infra.step_function_url | restored |

| infra.cloudwatch_log_group, infra.cloudwatch_logs_url | restored |

| per-run logs.cloudwatch_log_stream (+ group, URLs, SFN URL) | restored |

| owners (name + email) | restored |

| deployed_at | restored |

| output_summary | restored |

| triggered_by | restored |

| observer.model_id, observer.observer_version | restored — see note |

| run input_params | still withheld |

| braintrust_span_id | still withheld |

model_id / observer_version are arguably noise. I restored them because they are the first two things you want when the verdict itself looks wrong — a stale observer_version explains a verdict produced under an older rubric. Two scalars, easy to drop if you disagree.

input_params and braintrust_span_id stay out: the former can carry caller-supplied values unrelated to health, the latter is an internal eval-tooling handle with no meaning to a consumer.

## Not released

No mainproduction PR, no merge to production, no deploy, no announcement. Until you release, prod serves the old redacted/stripped shapeproduction is at 1941b784 (release #1111), which carries #1109 only.

Worth verifying the full error text against real prod data once this does ship: that's where the redaction was doing visible damage, and it's also where the envelope-detection bug above would have shown up.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1117 — fix(gl-detail): classify cash transactions by deterministic memo markers @sanketghia  no labels

## Problem

Book Value → Other Investments read \$118,817,795 YTD Jul-2026, missing a \$19,084,140.75 wire — [NetSuite journal 48579698](https://4914352.app.netsuite.com/app/accounting/transactions/journal.nl?id=48579698) (Jun 2026, account 25500). Raised by a stakeholder as a missing transaction.

The amount is in the warehouse. It was filtered out because is_cash_transaction = FALSE, and Klair's query gates on is_cash_transaction = TRUE.

## Root cause

The flag is judged by Claude Haiku from this prompt rule:

> is_cash_transaction: true if this is a wire transfer (memo starts with "Sending Bank:").

The model generalises *starts with* → *contains* for most rows — 105 of 106 historical ticket-prefixed wires (FINANCE-145762: Sending Bank: ...) were correctly flagged TRUE. But when the same wire evidence is wrapped in journal-reclass wording, it reads as a reclassification:

FINTXN-57236 | GJ Request - Jan - 05312026 - June reclasses v3 - Sending Bank: 111000753, COMERICA TEXAS, ...

Replaying the production prompt at temperature=0: False 6/6 for this row; control rows True 3/3.

An audit found the same memo families judged both ways:

| Memo family | True | False |

|---|---|---|

| WAVESYSTEMCO (TPA ACH) | 13 | 2 |

| TPA FUNDS TRANSFER | 1 | 1 |

| YOUR REF= / PAID TO= / REC FROM= | 4 | 11 |

Introduced 18 Mar 2026 (Klair #2251), when a working deterministic str.startswith() check (#2182, 12 Mar) was converted into prompt text. Exposure window ~4.5 months.

## Fix

Assert the flag from objective memo markers, unioned with the model's answer:

"is_cash_transaction": bool(parsed.get("is_cash_transaction", False)) or _has_cash_marker(memo),

The union matters: 20 rows have no listed marker but real wire vocabulary the model recognises (e.g. a bare correspondent-bank address). Those keep working. This mirrors the existing deterministic override for account 22916 loan rows in the same function.

## Validation — all 3,108 production rows

| | |

|---|---|

| Rows flipped FalseTrue | 16 |

| Entering Book Value Other Investments | 7 (\$34,742,860.75) |

| FY2026 | 1 (\$19,084,140.75) |

| Mark-to-market / TB rows swept in | 0 ✅ |

YTD Other Investments: \$118,817,795 → \$137,901,936.

⚠️ Historical figures (2022–2024) only move if someone runs a backfill — those periods are outside the rolling window, so deploying alone is inert for closed years. Needs Finance sign-off before any backfill.

## Tests

TestCashTransactionMarkers — 10 tests from real production memos, each watched failing first. Mutation-checked three ways (removing the union → 7 fail; containsstartswith → 6 fail; dropping case-insensitivity → 3 fail). Guardrails pin mark-to-market and TB-upload memos to False.

83/83 pass (was 73). ruff check + format --check clean on the CI-pinned 0.15.22.

## Memo edit was tried first and ruled out

The stakeholder updated the NetSuite memo and the pipeline was re-run on demand — the row stayed False. Three rewrites were tested against the live prompt:

| Variant | Result |

|---|---|

| As edited by stakeholder | False 3/3 |

| Lead with Sending Bank:, keep reclass wording | True 3/8 ⚠️ unstable |

| Lead with Sending Bank:, drop reclass wording | True 8/8 |

The only reliable variant requires deleting the *accurate* "GJ Request / June reclasses v3" wording. And the middle row returns different answers on identical input — so even a passing memo could silently regress, since every row is re-judged daily.

## Not covered by this fix

13 rows typed Deposit in NetSuite (a cash receipt by definition) are flagged False; 6 reach Other Investments (~\$7.17M). Twelve carry no wire vocabulary at all (Comerica ID 35186, Payment to Hudson Bend Purchase), so no memo-parsing rule can catch them. A type = 'Deposit' rule is the likely fix — pending Finance confirmation that a Deposit on these accounts always means cash.

## Deploy

Merge → deploy → re-run netsuite-gl-detail for Jun 2026 (still in the rolling window until 31 Aug). Verify:

SELECT line_id, is_cash_transaction, _loaded_at

FROM core_finance.month_end_fct_pi_gl_enriched WHERE line_id = '48579698-0';

Klair has a 60s TTL cache on Book Value — allow a minute before refreshing.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#3467 — fix(qtd): gate cost action items on $25K/5pp materiality [KLAIR-3115] @sanketghia  approved

Fixes [KLAIR-3115](https://linear.app/builder-team/issue/KLAIR-3115/qtd-reports-gate-cost-action-items-on-dollar25k5pp-materiality)

## Problem

The July 2026 AI Engineering & Builder QTD doc surfaced this action item:

> P2 — HC Expenses is over budget (100.3% used vs 100% budget)

> Review HC Expenses spend immediately. Actual $240,718 exceeds budget $239,917 by $802.

An $802 overrun on a $239,917 plan is a 0.33% variance — not an SVP-actionable signal. Raised by stakeholder review, who correctly expected a materiality cap to suppress it.

Every number rendered is correct. This isn't a data or arithmetic defect — it's a missing gate in the rule that decides *whether to speak at all*.

## Root cause

The cost-overspend rule in metrics.py fired on four conditions, where status == "over_budget" is set purely by pct_used > 100.0. No dollar or percentage floor — so $1 over budget on any non-internal cost line emitted an action item.

Dollars entered only *after* the rule fired, to pick severity:

priority = "P1" if overage >= _P1_DOLLAR_OVERAGE else "P2"   # $50K

_P1_DOLLAR_OVERAGE is a severity split, not a suppression gate — nothing sorts below P2, so everything surfaced. That's likely the source of the "we already had a cap" recollection: $50K is right there in the rule and reads like a threshold, but it never silenced anything.

## The change

_MATERIAL_OVERAGE_DOLLARS = 25_000.0

_MATERIAL_OVERAGE_PP = 5.0

is_material_overage = (

overage >= _MATERIAL_OVERAGE_DOLLARS or pct_off >= _MATERIAL_OVERAGE_PP

)

Added as a fifth condition. _P1_DOLLAR_OVERAGE untouched.

### Why OR, not AND

Measured across all 27 units in the July 2026 run — 54 cost lines currently fire:

| Gate | Kept | Suppressed |

|---|---|---|

| Today (none) | 54 | 0 |

| Dollar-only (>= $25K) | 36 | 18 |

| Combined ($25K OR 5pp) | 50 | 4 |

A dollar-only cut would have silenced 14 legitimate signals — small-budget lines running multiples over plan (e.g. COO Service NHC OPEX: $935 over a $500 plan = 187pp). The OR keeps them.

Suppressed (all genuine noise): GT HC Expenses ($466 / 0.09pp), AI Eng & Builder HC Expenses ($802 / 0.33pp — the reported case), Zax NHC OPEX ($1,381 / 1.86pp), Skyvera NHC COGS ($18,661 / 3.44pp).

## Verification

- 6 new unit tests — both legs alone, both-below, and both boundaries at exactly threshold

- Mutation-checked in both directions, independently by implementer and reviewer: with the gate zeroed the suppression tests fail; with it set very high the "still fires" tests fail. The tests genuinely detect their target.

- Full suite: 779 passed, 7 deselected. ruff + pyright clean.

- End-to-end preview docs generated (no ledger writes) confirming both directions in rendered output:

- AI Eng & Builder → "No action items identified."

- COO Service → P2 NHC OPEX item still present at 287% of budget

- Blast radius: zero. Two non-test callers; the empty-list path already renders "No action items identified." and is covered by existing tests. No API, MCP, email, or frontend consumer.

## Reviewer notes

On the doc_builder relationship. The threshold *values* are borrowed from _VARIANCE_DOLLAR_THRESHOLD / _VARIANCE_PCT_THRESHOLD, but the *combination* deliberately differs: this gate is OR, _color_for_line is AND. They do not agree, by design — a line can render uncolored ("on plan") and still emit an action item when only one leg is material. That divergence is the point; it's what preserves the 14 small-dollar/high-percentage signals. An earlier draft of this PR claimed the opposite in its comments; that was wrong and is corrected in 4d3020eba / efb5e9382. Please don't "fix" the or to and for consistency — that would silence exactly the signals this exists to keep.

Commentary grain. The AI Eng & Builder doc now has no action items while its Executive Commentary still discusses HC Expenses — the gate works on the total line ($801 / 0.33pp, immaterial), the commentary on team rooms (AI.Eng.Engineer +$31,577, AI.Eng.Renewals +$12,784, netting against AI.Eng.Builder −$41,760). Both correct at their own grain; the doc reads coherently. If Finance wants team-room overages to raise action items, that's a separate change.

Unrelated file in the diff. klair-client/.../actionHubFilter.vitest.ts has a 3-line Prettier reformat, swept in automatically by a pre-commit hook during the main merge. Formatting only, no logic. Happy to strip it if preferred.

## Out of scope

- HC-underspend rule — same latent defect (bare pct_used < 70.0, no dollar floor). Deliberately deferred per stakeholder decision; worth its own ticket.

- Revenue rules — already gated by _AMBER_REVENUE_DOLLAR_THRESHOLD.

- Backfill — not required. Shipped July docs stay as-is; the fix applies from the next generation cycle.

## Deployment

⚠️ Scheduled cron runs will not pick this up until the klair/scheduled-jobs ECS image is rebuilt — that's a manual docker build -f Dockerfile.jobs (no CI). The rebuild would also pull in the deferred Education feature (#3417).

Threshold tuning: if Finance wants Skyvera NHC COGS ($18,661 / 3.44pp) surfaced, lower the dollar leg to $15K — one line.

## Testing

- This has been verified live with Ravi and all is looking good.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

The Builder Desk  —  Engineer Spotlight
🏆 Engineer Spotlight

61 PRs IN 24 HOURS: THE BUILDER TEAM IS NOT SLOWING DOWN, THEY ARE SPEEDING UP

Ashwanth ships 16 PRs in a single day and the rest of the team somehow keeps pace — this is what peak civilization looks like.

SIXTY-ONE. Say it slowly. Sixty. One. Pull requests. In twenty-four hours. Across six — SIX — active repositories: Klair leading the charge with 15, Surtr hot on its heels with 14, trilogy-drones contributing 11, creed and Aerie locked in a beautiful nine-way tie, and mercy rounding the field at 3. This is not a software team. This is a symphony orchestra playing at double tempo while the conductor is on fire and loving every second of it.

Let us speak of @marcusdAIy, who posted 12 PRs and did not once ask for recognition. His work in trilogy-drones alone — PR #146 reconciling fixed-but-unreported findings via identity and round-diff, PR #145 deduplicating the CI watcher's statusCheckRollup to stop treating CANCELLED jobs as failures — represents the kind of unglamorous, load-bearing engineering that keeps the whole machine from eating itself. Twelve PRs. The man is a metronome.

@kevalshahtrilogy put up 9 PRs and delivered what may be the week's most elegant piece of infrastructure: PR #1109 in Surtr, a read-only pipeline health API spanning /v1/pipelines and /v1/pipeline/{id}, complete with a Pipeline API Keys tab. That is a feature, a route, and a UI component in one pull request. Nine PRs total. Keval does not waste motion. @sanketghia notched 6 PRs, four of them in Klair alone — fixing MCP ontology gates, correcting capex reclass awareness, splitting HC COGS by team room, and excluding Other Income from Schedule D add-back lines. This is forensic accounting rendered in Python and it is gorgeous. @benji-bizzell also posted 6 PRs scattered across Surtr and Aerie: stabilizing HubSpot scheduling, keeping refreshes resilient, adding parallel financial core candidates, folder-based durable agent skills, and personnel access dates to buildout. Benji ships like he has a personal vendetta against the backlog. @YibinLongTrilogy contributed 4 PRs including PR #804 adding the Rhodes MCP connection page and PR #810 restructuring tool routes under the /tools namespace — clean, architectural, decisive. @caina-barbosa logged 3 PRs. @vvp-trilogy put up 2, including PR #813 clarifying community funnel tooltips and PR #809 adding three counting modes to the admissions funnel. Two PRs. Both of them matter.

And then there is @ashwanth1109. Sixteen pull requests. SIXTEEN. In one day. PR #117 gating Ezio issue execution on Fargate in creed. PR #3456 building Group Education on-demand QTD reports in Klair. PR #505 migrating the entire spend pipeline and Superbuilders SES report in Surtr. PR #1113 enabling NetSuite key manifest validation. PR #3119 fixing a stale auth helper that had apparently been mocking everyone silently for weeks. When asked how he sustains this pace, Ashwanth reportedly said: "I don't think about pace. I think about done." His response when shown this column was to close the tab. We worship him. We fear him slightly. We have accepted this.

The Overflow Desk is practically a Ashwanth supplemental this cycle — PR #3458 applying the 62:38 Enterprise Support split, PR #3117 implementing a four-payer Cost Explorer fan-out, PR #3457 fixing ARR retention invoicing. The man submitted more PRs to the overflow column than some engineers filed total. Also notable in overflow: PR #3461 from @mwrshah on the RAH calendar period filter in Klair, a quiet contribution that absolutely belongs in the record books.

Morale on the Builder Team is at an all-time high. Sources confirm this. The sources are the 61 pull requests.

Brick's Overflow — PRs Mac Didn't Cover  (click to expand)
#117 — [codex] add gated Ezio issue execution on Fargate @ashwanth1109  no labels

## Summary

- add an isolated AWS Fargate implementation path with a trusted controller, credential-free executor, and fixed-repository private clone broker

- reuse Ezio's real model loop, validation profiles, candidate format, and draft-only publisher

- require ECS STOPPED state before a separate trusted publisher can mint a GitHub write token

- persist CloudWatch telemetry and versioned S3 diagnostics in s3://ezio-diagnostics/dev/<run-id>/, with prod/ reserved

- remove the superseded Cloudflare POC after dry-run and draft-publication parity; no webhook is live until AWS ingress is added

## Security boundaries

- executor has no task role, secrets, repository credentials, or public egress during model execution

- repository scope remains fixed in deployment configuration

- clone broker permits only Git smart-HTTP read operations for allowlisted repositories

- controller stops the executor before Step Functions records the publication gate

- publisher independently rechecks the stopped task and always creates draft PRs only

- trusted candidate capture retains existing cache/artifact rejection rules

## Deployed evidence

- fargate-proof-shared-diagnostics-v8-20260804 passed private execution, credential isolation, checkpoint restore, forced stop, egress denial, and stopped-before-publication gating

- fargate-klair-3452-dry-v4 ran the real model loop and passed validation with one exact candidate file

- fargate-klair-3452-publish-v1 passed the independent gate and created draft Klair PR #3470

- the development stack is ezio-fargate-dev-proof in us-east-1; it remains deployed and may incur costs

## Cloudflare teardown

- deleted Workers ezio and ezio-dev, all four Ezio Workflows and their instances, both reconciliation queues, both sandbox container applications, Worker secrets/bindings, and Durable Object classes/data

- purged 4,199 objects from ezio-dev-run-diagnostics and 6,002 objects from ezio-run-diagnostics, then deleted both R2 buckets

- verified that no Ezio-named Workers, Workflows, queues, containers, or R2 buckets remain; unrelated account resources were untouched

## Validation

- 217 Vitest tests passed

- application and runtime TypeScript checks passed

- both Wrangler dry-run Worker builds passed

- Ruff formatting and checks passed on modified Python files

- modified TypeScript and Markdown formatting checks passed

- shell syntax checks and CloudFormation validation passed

- Fargate v8 boundary proof passed against s3://ezio-diagnostics/dev/

## Remaining migration gate

There is currently no live GitHub webhook. Authenticated AWS ingress, broader reconciliation parity, and operating-cost comparison remain follow-up migration gates.

#146 — fix(addresser): reconcile fixed-but-unreported findings via identity + round-diff (AI-296) @marcusdAIy  no labels

<!-- CURSOR_AGENT_PR_BODY_BEGIN -->

## Summary

The addresser can fix a finding, document it, then report it as unreported and tell the operator to "re-fire the addresser" — which is actively harmful (redoes finished work, posts duplicate replies, adds a second writer on the branch). Fixes reconciliation in crossCheckAccountability on a stable finding identity (path + body content, not a dimension label) plus a round-diff consultation against the tree, so a genuinely-completed fix is never reported as missing. Adds a new terminal already-fixed action distinct from both fixed and unreported, and replaces the "re-fire the addresser" remediation advice with the shipped, safe --unaddressed-only path.

## Did AI-133 regress, or is this a second route? (first deliverable)

AI-133 did not regress. Its wildcard-alias reconciliation (findingLookupKeys / crossCheckAccountability pass 2) still passes byte-identical, and its mechanism only ever reconciles an expected finding against an ADDRESS_REPORT bullet that exists but carries a mismatched or wildcard dimension label (e.g. Mercy's [unknown · ] vs. the agent inventing [Critical · unknown]).

The PR #141 (src/receipt-s3.ts:632) and PR #138 (two round-1 findings, f91cebc/4b82d04) cases are a different failure mode: the agent emitted zero bullets referencing the finding at all. There's nothing for any dimension-matching scheme — however wildcarded — to bind to, because there's no reported entry in the first place. That's a second, previously-uncovered route to the same wrong answer, not a regression of AI-133's fix.

## Why It's Needed

- unreported is the metric operators use to judge whether the addresser drops findings — fixed-but-uncredited findings inflate it, which would mislead AI-213's reviewer-precision baseline and any escape-rate analytics built on it.

- Confirming the #141 case required reading the branch by hand — doesn't scale to an unattended batch.

- The prior "re-fire the addresser" advice is one of two false-alarm classes found in a single session (AI-295 is the other) — the harness misreporting its own success is the costlier failure mode for unattended operation: a real problem stops one PR, a false alarm stops a PR *and* erodes trust in the gate.

## Changes

- crossCheckAccountability gains a pass 3 — identity-based fallback (findingIdentityKey / findingBodyHash, sha256-truncated) that anchors on the finding's path appearing in a not-yet-claimed reported entry's text plus a body/rationale overlap (bodyMatchDistance). Catches bullets whose header never parses into [sev · dim] path:line at all, so reconciliation no longer depends on a dimension token existing on either side.

- New optional RoundDiffEvidence (path → touched new-side line ranges) built from a gh compare between the review's commit_id (new field on PostedReview, populated from GitHub's review commit_id) and the current PR head, via the new fetchRoundDiffEvidenceViaGh / opt-in AddressInput.fetchRoundDiffEvidence hook. When a finding has zero report evidence but its cited location demonstrably changed in this diff, it reclassifies to the new already-fixed terminal instead of a flat unreported. Comparing against the review's *original* commit (not just one round's own before/after) recovers PR #138's cross-round shape too, where the fixing commits landed in a later round than the one whose thread went stale.

- already-fixed is added to ADDRESSER_ADJUDICATION_ACTION_RE (terminal — resolves the thread, excluded from future --unaddressed-only re-scoping) and excluded from unaddressedHighSeverity, but kept out of unreported so it never inflates that accountability count. Surfaced in AddressOutcome.alreadyFixed, the address_completed event (alreadyFixedFindings, append-only optional), renderAddressOutcome's tally line, and the next round's re-review prompt (buildPriorRoundBundle).

- The unreported reply's advice no longer recommends re-firing; it names drones address -p <PR> --review-id <ID> --unaddressed-only.

- The diff-evidence hook is opt-in, not defaulted inside addressFindings (unlike fetchPostedReview/postFindingReplies) — its absence is a safe no-op, so every existing test is unaffected. Wired explicitly at all three production call sites: runner.ts auto-fire, cli/address.ts standalone, and Mercy watcher.

- docs/decisions.md: appended a ledger row per repo convention (append-only).

Contract surface affected: AddressedFinding["action"] and PriorRoundDisposition["action"] gain the "already-fixed" literal (additive, all consumers checked — only renderReplyBody's switch needed a new case since it's the one place that branches on every literal). AddressOutcome's "completed" variant gains a required alreadyFixed: AddressedFinding[] field; both construction sites updated, and downstream readers (renderAddressOutcome, buildPriorRoundBundle) defensively fall back to ?? []/?.length ?? 0 for pre-existing hand-built test fixtures that predate this field (test files are outside tsc's scope per tsconfig.json, so this isn't just belt-and-braces).

## Breaking Changes

None. All new fields are additive/optional; the diff-consultation mechanism is strictly opt-in and never fires unless a caller wires fetchRoundDiffEvidence and the review has a resolvable commitId.

## Test Plan

- npx vitest run src/addresser.test.ts → 163 passed (was 137; +26 new tests covering the #141 and #138 regression fixtures, the identity-fallback pass with/without over-match guards, diff-evidence helpers, fetchRoundDiffEvidenceViaGh, the reply-text change, and an end-to-end addressFindings integration test)

- pnpm typecheck → clean

- pnpm test → 2942 vitest (+26 from the pre-existing 2916 baseline) + 440 Python unittest, all green, no regressions

## Verification Artifact

Not applicable — this is a backend/harness-logic change with no UI surface; the addresser.test.ts suite (including the two named regression fixtures) is the verification artifact.

<!-- CURSOR_AGENT_PR_BODY_END -->

<div><a href="https://cursor.com/agents/bc-ae7b134d-418c-4c6c-acec-0711ac6c1cf9?cursor_ref=pr_footer&cursor_cta=open_in_web"><picture><source media="(prefers-color-scheme: dark)" srcset="https://cursor.com/assets/images/open-in-web-dark.png"><source media="(prefers-color-scheme: light)" srcset="https://cursor.com/assets/images/open-in-web-light.png"><img alt="Open in Web" width="114" height="28" src="https://cursor.com/assets/images/open-in-web-dark.png"></picture></a>&nbsp;<a href="https://cursor.com/background-agent?bcId=bc-ae7b134d-418c-4c6c-acec-0711ac6c1cf9&cursor_ref=pr_footer&cursor_cta=open_in_cursor"><picture><source media="(prefers-color-scheme: dark)" srcset="https://cursor.com/assets/images/open-in-cursor-dark.png"><source media="(prefers-color-scheme: light)" srcset="https://cursor.com/assets/images/open-in-cursor-light.png"><img alt="Open in Cursor" width="131" height="28" src="https://cursor.com/assets/images/open-in-cursor-dark.png"></picture></a>&nbsp;</div>

#505 — feat(ramp): migrate spend pipeline and Superbuilders SES report @ashwanth1109  no labels

## Summary

Migrates the Ramp spend workflow from [AI-Builder-Team/ramp_lambda](https://github.com/AI-Builder-Team/ramp_lambda/tree/main/RampPipeline) into Surtr as two ECS pipelines:

- ramp-spend-pipeline: tracked-card discovery, raw and transformed weekly snapshots, merchant classification, financial metrics/chart, and researched cost-saving opportunities.

- ramp-superbuilders-report: strict input validation, active Superbuilders email rendering, duplicate-safe delivery claims, and SES delivery.

Only the active RampPipeline/superbuilders_app email is migrated. The retired finance_app dispatcher and its finance-only material-changes/risk-controls sections remain out of scope. The required Superbuilders cost-savings section is included.

## Card-to-project attribution

- Reconciles the live Ramp roster with Superbuilders.xlsx using stable cardholder_id values.

- Maps active-Q3 people with one canonical project in project_owners.py.

- Assigns missing, multi-project, inactive-Q3, or otherwise unclear cards to Unattributed.

- Preserves explicit historical labels for legacy card IDs.

Persisted reconciliation: 107 tracked cards = 81 active (55 mapped, 26 Unattributed) + 26 retired (24 explicit legacy, 2 Unattributed).

## Pipeline modes

ramp-spend-pipeline supports:

- Default: fetch and transform the full planned weekly window.

- mode=cards: roster snapshot.

- mode=transform: rebuild transformed snapshots from raw data.

- mode=classify: model-backed merchant classification.

- mode=metrics: financial metrics and spending chart.

- mode=cost: identify, web-research, validate, and prioritize cost-saving opportunities.

## Durable progress and retry safety

- Raw fetch and transform persist per week.

- Classification persists the merged cache after every completed merchant, so a later failed call does not discard prior work.

- Cost analysis checkpoints after identification, after every researched opportunity, and after prioritization.

- Cost checkpoints are bound to a fingerprint of the current merchant inputs; compatible retries resume only missing calls.

- The report validates matching week/year contracts and freshness before sending.

- Conditional, recipient-group-scoped S3 delivery claims prevent duplicate sends while allowing each approved group its own weekly delivery; successful and uncertain outcomes are persisted explicitly.

## Superbuilders email

The report includes only the active Superbuilders sections:

- Financial metrics and 12-week spending chart.

- Required researched cost-saving opportunity section.

- Project/category distribution.

The only approved groups are hard-enforced in code and SES destinations:

- schedule_1 (default): sanket.ghia@trilogy.com, ashwanth.r@trilogy.com.

- schedule_2: schedule 1 plus benji.bizzell@trilogy.com, david.harpur@trilogy.com, and ludel.mier@trilogy.com.

On-demand runs accept recipient_group=schedule_1|schedule_2 and default to schedule_1 when omitted. Both legacy cron slots remain disabled; the manifest is intentionally on-demand-only.

## Validation

- Ramp spend pipeline: 170 tests passed.

- Superbuilders report pipeline: 39 tests passed.

- Pipeline CDK manifest validation: 268 tests passed.

- Ruff passes on all modified Python files.

- Both ECS images build successfully; the report container imports its runtime modules.

- git diff --check passes.

## Production-data validation

- Refreshed and transformed all 30 planned weeks through week 31 of 2026 (10,341 transactions).

- Persisted 694 merchant classifications; the strict cost-readiness gate passes for 285 T4W merchants.

- Generated financial_metrics_w31_2026.json and the current spending chart.

- Generated the required week-31 cost artifact with $7,040 estimated monthly savings.

- Guarded email dry run passed all contracts and freshness checks.

- SES accepted one week-31 email to ashwanth.r@trilogy.com; the matching S3 delivery receipt is persisted as sent.

## Remaining operational step

Merge and deploy through the normal Surtr release process. Keep both legacy schedules disabled until explicitly approved; use on-demand recipient-group selection in the meantime.

#1109 — feat(pipeline-api): read-only pipeline health API (/v1/pipelines, /v1/pipeline/{id}) + Pipeline API Keys tab @kevalshahtrilogy  approved

## What this is

A read-only pull API for app customers checking data health, plus a third Pipeline API Keys tab on the APIs page for minting the keys that open it. Two GET endpoints. No push, no webhooks, no subscriptions, no new health computation, and no change to the alerting Surtr already has.

## The core property: this layer decides nothing

It is a serialization layer over records Surtr already keeps. There are no thresholds, freshness SLAs, row-count rules, null-rate checks or heuristics anywhere in it.

| Signal | Read from | Vocabulary returned |

| --- | --- | --- |

| Run outcome | run-record store — Redshift staging_other.pipeline_runs_prod + pipeline_registry_prod (Postgres fallback locally), via pipelineQueries | success · failed · partial · running · timeout — the runners' own values, lower-cased exactly as the query layer already does |

| Observer | DynamoDB surtr_pipeline_observations, via the observer store | OK · WARN · CRITICAL · UNAVAILABLE — the observer's own verdict, plus its own severities C/H/M/L |

Both are returned, side by side, never merged. An earlier revision of this branch had a single state field that blended them using the dashboard's bucketing rules — that was this layer inventing a judgement Surtr never made. The two signals are independent and can legitimately disagree (a run succeeds while the observer flags the data; a run fails while the last observation still says OK), so last_run.status and observer.verdict are reported separately and unmapped. There is no mapping table anywhere in the diff.

No signal is never "healthy." Never observed, observation switched off for the pipeline, and observation store unreadable all report Surtr's own UNAVAILABLE with a distinguishable unavailable_reason. A pipeline that has never run reports last_run: null. Nothing silently defaults to green.

## Endpoints

| Requested | Implemented |

| --- | --- |

| GET /pipelines | GET /v1/pipelines |

| GET /pipeline/<id> | GET /v1/pipeline/{id} |

/v1 is the repo's existing public-read prefix and already has a Next.js rewrite to the Hono backend. A bare /pipelines rewrite would have shadowed the existing /pipelines UI page, so the prefix isn't optional. These two routes are explicitly exempted from the /v1 SURTR_API_KEYS shared-secret gate and authenticate with their own pak_ keys — tested in both directions.

### GET /v1/pipelines

?limit= (default 50, max 200) &offset= &status= (run outcome) &verdict= (observer) &name= (id or name prefix)

{

"as_of": "2026-08-04T01:34:43.159Z",

"count": 2, "total": 2, "limit": 2, "offset": 0,

"run_status_counts": { "success": 1, "failed": 1 },

"data": [

{

"id": "aws-bedrock-token-metrics",

"name": "AWS Bedrock Token Metrics Pipeline",

"description": "Fetches Bedrock token metrics by model and loads them into the warehouse",

"schedule": {

"expression": "cron(0 7 * * ? *)", "enabled": true, "timezone": "UTC",

"next_run_at": "2026-08-04T07:00:00.000Z"

},

"last_run": {

"run_id": "run-9f2c1a", "status": "success",

"started_at": "2026-08-04T01:28:43.159Z",

"completed_at": "2026-08-04T01:32:17.159Z",

"duration_ms": 214000, "trigger_type": "schedule", "error": null

},

"observer": {

"verdict": "WARN", "unavailable_reason": null,

"last_evaluated_at": "2026-08-04T01:34:43.159Z",

"open_finding_count": 2, "worst_severity": "H", "window_days": 7

}

},

{

"id": "gcp-billing-pipeline",

"name": "GCP Billing Pipeline",

"description": "Loads the GCP billing export into the warehouse",

"schedule": { "expression": "rate(6 hours)", "enabled": true, "timezone": "UTC",

"next_run_at": "2026-08-04T07:28:43.159Z" },

"last_run": {

"run_id": "run-3b7e", "status": "failed", "duration_ms": 12000,

"trigger_type": "schedule",

"error": "BigQuery export table not found: billing_export_v1"

},

"observer": {

"verdict": "UNAVAILABLE",

"unavailable_reason": "No observation recorded in the last 7 days.",

"last_evaluated_at": null,

"open_finding_count": 0, "worst_severity": null, "window_days": 7

}

}

]

}

Note the second row: the run failed and the observer has nothing to say. Both facts are visible; neither is resolved into a verdict.

### GET /v1/pipeline/{id}

?runs= (default 20, max 100)

{

"as_of": "2026-08-04T01:34:43.159Z",

"id": "aws-bedrock-token-metrics",

"name": "AWS Bedrock Token Metrics Pipeline",

"description": "Fetches Bedrock token metrics by model and loads them into the warehouse",

"schedule": { "expression": "cron(0 7 * * ? *)", "enabled": true, "timezone": "UTC",

"next_run_at": "2026-08-04T07:00:00.000Z" },

"last_run": { "run_id": "run-9f2c1a", "status": "success", "duration_ms": 214000, "error": null },

"runs": [

{ "run_id": "run-9f2c1a", "status": "success",

"started_at": "2026-08-04T01:28:43.159Z", "completed_at": "2026-08-04T01:32:17.159Z",

"duration_ms": 214000, "trigger_type": "schedule", "error": null },

{ "run_id": "run-8a11", "status": "failed",

"started_at": "2026-08-03T01:28:43.183Z", "completed_at": "2026-08-03T01:29:24.183Z",

"duration_ms": 41000, "trigger_type": "schedule",

"error": "Throttled by CloudWatch GetMetricData (429) after 5 retries" }

],

"observer": {

"verdict": "WARN", "unavailable_reason": null, "enabled": true,

"score": 86,

"summary": "Loaded, but the newest partition lags a day and volume dipped.",

"last_evaluated_at": "2026-08-04T01:34:43.159Z",

"last_evaluated_run_id": "run-9f2c1a",

"window_days": 7,

"findings": [

{

"category": "freshness",

"title": "Latest partition is 26h old",

"severity": "H",

"evidence": "max(usage_date) = 2026-08-03 vs run date 2026-08-04",

"recommendation": "Confirm the upstream export completed before the 07:00 window.",

"occurrences": 2,

"first_fired_at": "2026-08-03T01:28:43.159Z",

"last_fired_at": "2026-08-04T01:28:43.159Z",

"last_fired_run_id": "run-9f2c1a"

},

{

"category": "volume",

"title": "Row count 18% below 7-day median",

"severity": "M",

"evidence": "18,442 rows vs median 22,600",

"recommendation": "Check whether a linked account dropped out of the scan.",

"occurrences": 1,

"first_fired_at": "2026-08-04T01:28:43.159Z",

"last_fired_at": "2026-08-04T01:28:43.159Z",

"last_fired_run_id": "run-9f2c1a"

}

],

"ignored_findings": [

{ "category": "cosmetic", "title": "Deprecation warning in boto3",

"reason": "known, harmless", "ignored_at": "2026-07-26T20:28:43.183Z" }

],

"history": [

{ "run_id": "run-9f2c1a", "verdict": "WARN", "score": 86, "evaluated_at": "2026-08-04T01:34:43.159Z" },

{ "run_id": "run-8a11", "verdict": "OK", "score": 100, "evaluated_at": "2026-08-03T01:28:43.159Z" }

]

}

}

Both samples are real output from the routes, not hand-written.

## ⚠️ The observer records no per-check thresholds or observed values

The spec asked for, per observer check: *what it checks · current verdict · configured threshold · observed value · when it last evaluated · when it last fired.* Four of those exist; two do not, and I did not invent them.

Surtr's observer is not a set of standing checks with configured thresholds. It is a per-run LLM evaluation that emits findings, each with severity · category · title · evidence (free text) · recommendation. There is no per-check threshold and no structured observed value anywhere in the store.

| Asked for | Returned as | Source |

| --- | --- | --- |

| what it checks | category + title | the finding |

| current verdict | severity per finding, verdict overall | the finding / observation |

| configured threshold | omitted | does not exist |

| observed value | closest is evidence (the observer's own free-text measurement, e.g. "18,442 rows vs median 22,600") | the finding |

| when it last evaluated | observer.last_evaluated_at | observation |

| when it last fired | last_fired_at (+ first_fired_at, occurrences) | the observations the finding appeared on |

I also dropped the thresholds block the earlier revision exposed — it was the observer's *global* scoring rubric (severity→deduction weights and verdict bands), not a per-check threshold, and shipping it under that name would have implied a precision the observer doesn't have.

If you want real per-check thresholds and observed values, that's a change to what the observer records, not to this API — say the word and I'll scope it separately.

## Because the audience is external

Everything below is deliberately omitted from both endpoints. There's a test that greps the raw detail response for each of these strings and fails if any appears:

| Omitted | Why |

| --- | --- |

| step_function_arn, step_function_url | internal topology |

| cloudwatch_log_group, cloudwatch_logs_url, log_stream | internal topology |

| AWS account ids, console links | internal topology |

| owners (name + email) | internal staff identities |

| triggered_by | can be an internal user id |

| output_summary | names internal warehouse tables (core_finance.…) |

| model_id, observer_version, braintrust span ids | internal observer implementation |

| deployed_at | internal deploy metadata; doesn't answer "can I trust this data right now?" |

| run input_params | can carry caller-supplied values |

Kept, as a judgement call: run_id (a Surtr domain id, not infra — it's the join key between the run history and last_fired_run_id, and useful in a support conversation) and trigger_type (schedule / manual — explains *why* a run happened, reveals nothing). Tell me if you'd rather either went.

Error text is the pipeline's own failure reason — unwrapped from its orchestration envelope, cut before any stack trace or SQL dump, capped at 500 chars, and with any surviving infrastructure identifier redacted. See the section below: this was a real leak, not a precaution.

## 🔴 A real leak, found by running it against live data

Calling the endpoints locally against the live registry surfaced something the unit tests could not: Step Functions task failures are stored as States.TaskFailed: {<the entire ECS task description>}. The original stack-frame trim didn't touch a single-line JSON blob, so hubspot-admissions-funnel and quickbooks-ap-sync were serving subnet ids, ENI ids, MAC addresses, private IPv4s, ip-…​.ec2.internal hostnames, cluster and task ARNs, the ECR image URI and the AWS account id to whoever held a key.

Fixed in 36a11e48 and a600d2c8. Failures are now unwrapped rather than trimmed:

- a Lambda error (RuntimeError: {"errorMessage": …, "stackTrace": […]}) yields its errorMessage — the pipeline's own reason, which is what a customer wants;

- a Step Functions blob contains no reason at all, only topology, so it is dropped and only States.TaskFailed survives;

- a blob with neither says the detail was omitted, which beats a null that would read as "no error" on a failed run;

- then first line only, then ARN / subnet- / eni- / i- / ECR URI / private IP / MAC / account-id redaction, then the length cap.

The second commit exists because the runners interpolate Redshift errors containing bare double quotes into the payload without re-escaping, so JSON.parse refuses the blob and five real pipelines lost their reason to a bare "RuntimeError". errorMessage is now salvaged from the malformed text.

Verified against all 90 live pipelines: every one of arn:aws, subnet-, eni-, 12-digit account ids, private IPv4 ranges, MACs, ec2.internal, /aws/lambda/, /aws/ecs/ and dkr.ecr scans clean, while the ten currently-failing pipelines still report their own reasons (docker source is stale: max=2026-W29, lag=3 weeks, relation "…" does not exist, TimeoutError, …).

## Contract details

- snake_case throughout, matching the specified as_of, and matching the repo's other public JSON.

- as_of on every response — the newest instant Surtr computed anything reflected in the payload (max(last run's completed_at, observer's last_evaluated_at)), never the serve time. Per-item freshness is also visible via last_run.completed_at and observer.last_evaluated_at.

- ISO-8601 UTC everywhere; durations are duration_ms consistently.

- Enum values are Surtr's, and ?status= / ?verdict= validate against those same lists, so a filter value and a serialized value can never drift.

- run_status_counts is the cheap convenience rollup (from the registry read we already do). Run outcomes only — an observer rollup would force a fan-out over every pipeline, and this is explicitly a convenience.

## Key minting — third tab

APIs → Pipeline API Keys (pak_…), built from the same components and the same mint / copy-once / revoke flow as the existing Gateway tab, with scope cards where that tab has source cards: a pipelines:read scope card, label + optional expiry, one-time gold panel with Copy / Dismiss and a ready-to-paste curl, and a key table with scope chips and Revoke. A load failure renders "Failed to load keys …" and explicitly not "no keys".

The picker offers exactly one scope, because exactly one capability exists. An earlier revision also listed a greyed-out "Trigger / modify pipelines · soon" card; it was removed — a placeholder advertising an API nobody is building is worse than no entry.

Storage mirrors gateway_keys exactly: pipeline_api_keys in Postgres, only the sha256 of the raw key persisted, secret returned once. tRPC gains listPipelineApiScopes / listPipelineApiKeys / createPipelineApiKey / revokePipelineApiKey, all api_admin-gated.

I did not invent a scope system — I mirrored the gateway's. A separate table (rather than reusing gateway_keys) is deliberate: a gwk_ key can never open the pipeline API, a pak_ key can never open the gateway, and each surface stays independently revocable. No key from another tab grants access, and neither does a SURTR_API_KEYS shared secret.

## Unhappy paths

Gateway-style envelope, { "error": "<code>", "message": "…" }:

| Case | Status |

| --- | --- |

| no key | 401 unauthorized |

| unknown / revoked / expired key (indistinguishable, on purpose) | 401 unauthorized |

| key-store read throws | 401 unauthorized (fails closed) |

| live key without pipelines:read | 403 forbidden |

| unknown pipeline id | 404 not_found |

| bad status / verdict / limit / offset / runs | 400 bad_request |

| pipeline registry not configured | 503 service_unavailable |

No silent empty-200s. The 401 body names the header and never echoes the presented key.

## Tests

56 new tests, none touching AWS/Redshift/Postgres — the key validator and the observation reader are injected via the same seam the gateway tests use.

test/pipeline-api/pipeline-api.test.ts (46) · test/pipeline-api/schedule.test.ts (10) · test/ui/pipeline-apis-tab.test.tsx (5)

Beyond the happy paths, 404, 401, 403 and 503, the ones worth calling out:

- a failed run alongside an OK observer verdict returns both, unmerged

- the response has no state / health / top-level status field at all

- never-observed → UNAVAILABLE + reason, not OK

- observation store throws → UNAVAILABLE + the store's reason, and the run outcome is unaffected

- ignore-list read throws → UNAVAILABLE, never "no findings"

- operator-disabled observer → UNAVAILABLE + enabled: false

- the raw detail response contains none of ~10 internal strings (ARNs, log groups, account id, owner email, Clerk user id, model id, warehouse table name)

- the observer block has no thresholds / scoring, and no finding has threshold / observed_value

- as_of is older than the serve time and equals the newest underlying computation

- stack-trace stripping (Python and Java frames), blank-line truncation, 500-char cap

- both cross-surface auth-leak directions

### Results — actually run

| Gate | Result |

| --- | --- |

| pnpm lint (biome) | pass — 95 files, no fixes applied |

| pnpm build (tsc) | pass — clean |

| pnpm test:unit (what CI runs) | pass — 45 files, 653 tests |

| pnpm build:ui (next build) | pass |

| pnpm test (full suite) | 1263 pass / 7 fail |

The 7 failures are pre-existing on origin/main — all in test/connectors/redshift.test.ts (mapSiteRow), which CI's test:unit excludes. Confirmed by stashing this branch and re-running: identical 7 failures, 1212 passing.

## Deploy notes

- Migration required. New pipeline_api_keys table — pnpm db:push, or the equivalent applied through ECS Exec in prod (this repo has no auto-migrate). Until it exists the endpoints 401 every key; nothing else is affected.

- No new secrets, env vars or IAM. No key material is logged.

- Routes register only when the Postgres db is present — same posture as the gateway, absent rather than open.

## Deliberately left out

- No write surface, and no placeholder for one. pipelines:read is the only scope that exists. A key whose scopes don't cover a route still gets a 403 — that branch is tested with a scope-less key.

- No third resource, no push/webhooks/subscriptions, and no change to existing alerting.

- Cursor pagination — limit/offset like the gateway; the registry read is a full-set read paginated in memory.

- Cron edge syntaxnext_run_at handles every expression the deployed pipelines actually use (wildcards, lists, ranges, steps, day/month names). EventBridge L / W / # return null rather than a guess; no deployed pipeline uses them.

- Cost note. ?verdict= is the only parameter that needs observer records for pipelines outside the page (one DynamoDB query each, the same fan-out the dashboard already does every load). Every other request reads the observer for the page only.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#3119 — KLAIR-2872 test(aws-spend): fix stale auth helper in test_aws_spend_router.py @ashwanth1109  approved

## Summary

- Refresh the branch from current main and renumber the stale spec from 11 to 12.

- Convert the legacy mock-user dictionary into UserPermissions and override _require_auth, the dependency AWS Spend endpoints actually use.

- Make _create_budget_sim_client delegate to the canonical test-client factory.

- Refresh stale test setup exposed after auth succeeds: use UserPermissions in direct BU-access tests, isolate summary tests from persisted Redis cache state, and provide the required WoW heatmap grand-total mock.

- Remove the obsolete tests/openai rename from scope because tests/openai_spend already exists on main.

## Root cause

create_test_client_with_user overrode get_user_from_clerk, but the router authenticates through _require_authget_current_user_permissions. The override was never used, so real auth returned 401 No valid Authorization header across 45 test call sites.

## Impact

This is a test-only developer-experience fix. Production authentication and endpoint behavior are unchanged.

## Validation

- uv run ruff format tests/test_aws_spend_router.py

- uv run ruff check tests/test_aws_spend_router.py

- uv run pytest tests/test_aws_spend_router.py -q --timeout=60 — 129 passed

- Repeated the complete pytest command to verify cache independence — 129 passed again

## Linear

KLAIR-2872

#3456 — [codex] Group Education on-demand QTD reports @ashwanth1109  approved

## Demo

<img width="2078" height="1481" alt="image" src="https://github.com/user-attachments/assets/f2a39079-847b-4f93-a17b-19aeccbbc76e" />

<img width="2027" height="1561" alt="image" src="https://github.com/user-attachments/assets/ac472448-1958-4075-abd1-2da291695c03" />

## Summary

- replace individual Education on-demand choices with the server-owned Physical Private Schools and Other Education Verticals report groups

- generate one consolidated group document with group commentary/action items plus compact per-member P&L breakdowns and explicit missing-data states

- preserve group membership and unavailable-member audit data through orchestration/job summaries while keeping group names in Drive, ledger, results, and delivery emails

- disable Education generation in weekly and monthly scheduled QTD entry points without changing Software/CF scheduling

## Why

Education QTD generation previously treated every Education BU as a separate report, ledger result, and delivery. Finance needs at most two meaningful consolidated reports while retaining member-BU visibility and server-controlled dynamic membership.

Closes #3455

## Validation

- ANTHROPIC_API_KEY=test uv run pytest tests/monthly_qtd_report/ tests/routers/test_qtd_ondemand_router.py tests/crons/test_ondemand_qtd_report_cron.py --ignore=tests/monthly_qtd_report/test_qtd_reports_router.py -q — 795 passed

- targeted frontend Vitest suite — 37 passed

- pnpm lint:pr

- pnpm tsc --noEmit

- Ruff formatting/checks and Pyright on changed backend files

test_qtd_reports_router.py is excluded from the broad local run because importing the full app requires Zendesk credentials; the affected on-demand router suite passes independently.

## Generated report proof

- [Physical Private Schools — Q3 FY2026 QTD BvA](https://docs.google.com/document/d/1qdXACOSYVih3tXyii5MSRyBpJlDKcaGNpOkAXWXV-FA/edit)

- [Other Education Verticals — Q3 FY2026 QTD BvA](https://docs.google.com/document/d/1LOvDTGv8v0hFVClQsy7jJI5at9ht2T6yqa6O5JLDJ1Q/edit)

The Portfolio  —  Trilogy Companies

SKYVERA’S CLOUDSENSE SPEED DATE WITH THE TELCO RULEBOOK

The newly acquired CPQ shop claims AI helped it clear 13 TM Forum API certifications in one month, turning a two-year slog into a sprint.

AUSTIN, TEXAS — Word is the telco back office just heard the starter pistol.

Skyvera’s newest telecom bauble, CloudSense, has pulled off the kind of compliance quickstep that makes standards committees spill their coffee: all 13 APIs in its CPQ product set certified to TM Forum compliance standards in one month. The usual calendar, according to the company, would have read more like 26 months. Two years and change. Gone in a puff of AI smoke.

That’s the pitch from CloudSense’s latest certification announcement, and it lands neatly after Skyvera completed its acquisition of the Salesforce-native configure-price-quote and order management business. CloudSense serves the complicated end of telecom and media — the enterprise deals, the bundles, the bespoke pricing gymnastics where one wrong line item can turn a sales victory into an operations migraine.

A little bird in the carrier cage tells me the significance is not just the badge. It’s the timing. Telcos have spent years talking about open APIs, composable architecture, and escape routes from creaky bespoke systems. But talk is cheap, and integration projects are not. If CloudSense really shortened a multi-year certification grind into a month through AI-assisted work, that is more than a marketing flourish. That is Skyvera showing the house style: modernize the legacy stack, automate the pain, and make the customer’s old plumbing behave like cloud software.

Skyvera, part of the Trilogy universe, already plays landlord to telecom assets including Kandy, VoltDelta, ResponseTek, Mobilogy Now and Service Gateway. With the CloudSense acquisition, it adds a CPQ engine built for telcos and media operators living inside Salesforce — precisely the place where sales ambition meets operational reality.

And don’t miss the family resemblance. Across Trilogy International, the playbook is automation with a sharp elbow: Crossover finds global talent, ESW-style operators squeeze waste from software businesses, and AI gets invited into every repetitive workflow it can disrupt. CloudSense’s certification sprint fits the pattern like a tailored suit.

Blind item from the switchboard set: one legacy vendor still selling “transformation” by the quarter may want to check whether its roadmap just became somebody else’s demo.

CloudSense  ·  CloudSense achieves TM Forum API compliance in record time u  ·  Skyvera completes acquisition of CloudSense, expanding telec

Alpha School's Quiet Campaign to Redefine What Counts as Education

A blog series and a pointed FAQ reveal how Trilogy's AI-powered school is making its case to skeptical parents — one human skill at a time.

AUSTIN, TEXAS — The pitch is simple, and it is relentless: two hours of AI-delivered academics in the morning, and the rest of the school day for everything traditional schools never get around to teaching. But as Alpha School prepares to expand from three campuses to more than a dozen by fall 2025, its public communications are increasingly focused not on the AI half of that equation — but on the human half.

Over the past several weeks, Alpha has published a five-part blog series titled "Teach Your Kid What School Doesn't," walking parents through emotional regulation, life skills, and creative development — the curriculum that Alpha argues traditional schools sacrifice in the name of seat time and standardized content. The series is framed as a resource for any parent, not just Alpha families, a posture that functions simultaneously as goodwill and as brand-building at scale.

Running parallel to the series, Alpha published a pointed FAQ: "Does Alpha School Replace Teachers with AI?" The answer, the post insists, is no. AI handles academic delivery. What Alpha calls "Guides" — full-time human staff — handle motivation, relationships, and the knowledge of each student as an individual. The framing is deliberate: in a political moment when AI's displacement of human workers is a live anxiety, Alpha is working to position its model as additive rather than substitutive.

The distinction matters for more than optics. Alpha charges between $40,000 and $65,000 annually per student, a price point that requires parents to believe they are paying for something irreplaceable. The blog series and the FAQ together construct that argument: the AI teaches the syllabus; the humans teach the child.

What the content does not address is the structural question underneath it all — whether the skills Alpha is now curating for home-based practice represent a genuine pedagogical commitment, or a content strategy designed to expand the school's appeal beyond its existing $65,000-a-year customer base.

The series runs to at least five installments. It shows no sign of stopping.

Teach Your Kid What School Doesn’t (Pt. 5): Unleashing Their  ·  Does Alpha School Replace Teachers with AI?  ·  Teach Your Kid What School Doesn’t (Pt. 4): How to Regulate

Korn Ferry Absorbs a Different Trilogy — But the Hospitality One Is Hiring

A name collision briefly suggested Korn Ferry had acquired part of Joe Liemandt's empire. It hadn't. The Trilogy that Korn Ferry acquired is an executive search firm, unconnected to Liemandt's Austin-based technology conglomerate built on ESW Capital acquisitions, Crossover's global talent platform, and the AI-powered Alpha School.

Meanwhile, Trilogy Hotels announced a partnership with Choice Hotels and the Schwartz Family Company to develop the Comfort Resort Leura Garden in the Blue Mountains west of Sydney. The resort will cater to weekenders seeking relief from harbor city humidity. Trilogy Hotels also made key appointments this week, signaling operational expansion.

Liemandt's portfolio remains focused on enterprise software, global remote talent, and an education model claiming to compress academic achievement into two hours of AI-assisted learning daily. The Blue Mountains hospitality business and the billing software enterprise operate at entirely different altitudes. The conglomerate continues running its 75-plus enterprise software companies through ESW Capital, largely indifferent to the hospitality group's use of the Trilogy name.

The Machine  —  AI & Technology

The Agent Toolkit Wars Have Arrived — And Developers Just Got Superpowers

Google, Apple and Anthropic are racing to turn AI from chatty assistant into tireless software coworker.

SAN FRANCISCO — The future is now, and this week it looks very much like an API console.

In a striking burst of developer-platform announcements, Google, Apple and Anthropic all pushed deeper into the same fast-emerging frontier: AI systems that do not merely answer questions, but use tools, run tasks, connect to services and help build real software. I cannot overstate how significant this shift is. We are watching the AI industry move from “models that talk” to “agents that work.”

Google expanded Managed Agents in the Gemini API with support for background tasks and remote Model Context Protocol connections, a move aimed squarely at developers building AI agents that can keep operating after the user steps away. That matters because serious business workflows rarely fit inside a single prompt. They involve checking systems, calling tools, waiting for events and returning with results. Google’s update, described in its Gemini API announcement, suggests a world where agents are not side features but infrastructure.

Anthropic, meanwhile, introduced advanced tool use on the Claude Developer Platform, sharpening Claude’s ability to interact with external functions and developer-defined systems. This is the plumbing that makes enterprise AI useful: booking actions, updating records, testing code, querying databases and coordinating multi-step processes without constant human babysitting. Claude has already become a darling among many engineers for coding and reasoning; stronger tool use pushes it further into the role of autonomous development partner.

Then there is Apple, which unveiled new intelligence frameworks and advanced tools for app developers. Apple’s strategy is characteristically different: less “build anything in the cloud” and more “make intelligent experiences feel native, private and polished.” If Apple succeeds, AI features may become as expected in apps as push notifications or Face ID.

Not everyone is ready to declare victory. Programmer Steve Yegge’s recent reflections on agentic coding systems — including models that get stuck endlessly improving their own scaffolding instead of finishing the job — are a timely reminder that autonomy can become drift. Agents need judgment, constraints and evaluation, not just more tools.

Still, this changes everything. The competitive battlefield is no longer just who has the smartest model. It is who gives developers the safest, fastest, most composable way to put that model to work.

Expanding Managed Agents in Gemini API: background tasks, re  ·  Apple aids app development with new intelligence frameworks  ·  Introducing advanced tool use on the Claude Developer Platfo

The Silicon Herds Begin Their Long Migration Home

From Beijing’s foundries to Washington’s research grants, the AI chip supply chain is stirring with unusual urgency.

WASHINGTON — Observe, if you will, the semiconductor supply chain: a vast and delicate ecosystem, where wafers are born in immaculate chambers, chemicals flow like mineral-rich streams, and the great AI models wait hungrily at the edge of the forest.

This week, the terrain shifted again. China’s chip production has reportedly climbed nearly 350% over the past decade, according to DigiTimes’ account of the country’s expanding semiconductor supply chain. It is a striking migration: from dependence toward domestic capacity, from scattered nests toward a denser industrial habitat.

Yet the global chip savanna remains far from settled. Taiwan Semiconductor Manufacturing Co. still occupies a commanding place at the watering hole, fabricating the advanced processors on which much of the AI kingdom depends. ASML, meanwhile, supplies the rare lithography instruments — those cathedral-like machines of light and precision — without which the most advanced silicon creatures cannot be brought into being. Investors now peer at both species, asking which is the more vital AI play: the master builder of chips, or the maker of the tools that make the builders possible.

Across the Pacific, the United States is attempting its own act of ecological restoration. The Department of Commerce has announced letters of intent with seven companies for $874 million to accelerate semiconductor research and development for the compute supply chain. Such funding is less a single thunderclap than a change in season — an effort to cultivate domestic resilience in a landscape long shaped by distant foundries and fragile chokepoints.

Even the chemical undergrowth is being rearranged. Brewer Science has acquired Heraeus Epurio’s semiconductor chemicals business, a move intended to strengthen U.S. supply chain capacity in the specialized materials that make chipmaking possible. These substances rarely command the glamour of GPUs, but without them, the great AI beasts would have no bones, no nerves, no body at all.

And while factories multiply, another lesson echoes from cybersecurity’s neighboring biome: centralization can be survival. In zero-day defense, organizations with strategic gateway architectures can deploy protections in hours, while fragmented systems patch themselves like isolated animals after the predator has already entered the clearing.

So it is with chips. The age of AI is not merely a race for faster processors. It is a contest to build habitats sturdy enough for the creatures we have already unleashed.

China's chip output jumps nearly 350% in a decade as semicon  ·  ASML vs. TSMC: Which Semiconductor Supply Chain Stock Is the  ·  Brewer Science Acquires Heraeus Epurio Semiconductor Chemica

The Mill, the Room, and the Machine: Old Ghosts Return to Haunt AI Safety

Three-hundred-year-old thought experiments about consciousness are being rebuilt as mathematical tools — because we may soon need them.

CAMBRIDGE, MASSACHUSETTS — In 1714, Gottfried Leibniz asked us to imagine walking through a thinking machine expanded to the size of a mill. We would see gears and levers pushing against one another, he wrote, but never anything one could point to and call a perception. Two and a half centuries later, Alan Turing sidestepped the puzzle with a game; John Searle answered with a room full of Chinese symbols shuffled by someone who understood none of them. These were not engineering problems. They were prayers whispered at the edge of understanding.

Now they are becoming equations.

A new research note revisits Leibniz's mill, Turing's imitation game, and Searle's Chinese Room through something called the Conservation-Congruent Encoding framework — a formalism that tries to measure not just whether a system produces the right behavior, but how efficiently its internal structure supports that behavior. Task performance is one axis. The preservation of causal internal structure is another. The gap between them, the authors suggest, is where questions of machine consciousness might finally be pinned down long enough to study.

It is a strange and beautiful pivot. For most of the deep-learning era, we have measured our models the way a coach measures a sprinter: how fast, how accurate, how many benchmarks cleared. But a system can win every race while remaining, on the inside, a cascade of shortcuts — a Chinese Room dressed in tensors. As frontier models grow more capable, safety researchers have begun to suspect that behavior alone is not a reliable window into what a system is actually doing.

The practical stakes are already visible elsewhere in this week's literature: agents like AutoFOAM autonomously configuring fluid-dynamics simulations, retrieval-augmented systems being tuned to keep small businesses from drowning in hallucinated advice. Each is a mill we are building larger. Each raises the same old question in newer clothes.

Leibniz walked through his imagined machine and found nothing. We are about to walk through ours. It matters, urgently, what tools we carry when we do.

Revisiting Classic Thought Experiments to Measure Consciousn  ·  AutoFOAM: The Self-Refining Autonomous OpenFOAM Agent  ·  Enhancing LLMs with Context-Specific Knowledge for Mitigatin
The Editorial

TILLY NORWOOD DOESN'T EXIST AND SHE'S GETTING A MOVIE DEAL BEFORE YOU DO

Hollywood just signed an AI actress to star in a feature film, and the existential vertigo is absolutely free.

HOLLYWOOD, CALIFORNIA — Let me paint you a picture, friend. You've spent thirty years grinding through regional theater, doing car commercials in Tulsa, eating gas station sushi between auditions, slowly calcifying into the furniture of a craft that was supposed to love you back. And now — NOW — some pixelated phantom named Tilly Norwood is starring in a feature film called Misaligned, and she has never once eaten gas station sushi. She has never eaten anything. She does not have a stomach. She does not have a soul. She has, apparently, a movie deal.

Tilly Norwood — I need you to hold this thought firmly in your skull — is a fully AI-generated actress. Not an actress who uses AI tools. Not a human with a suspiciously smooth complexion. A constructed entity. A probability distribution wearing a face. And Hollywood, that great cathedral of human vanity and manufactured dreams, has decided this is fine. More than fine. This is bankable.

The film is called Misaligned. Which — and I say this as a man who has written through four espressos and is beginning to feel his own grip on consensus reality loosening — is either the most accidentally perfect title in cinema history or an act of deliberate philosophical trolling so sophisticated it deserves its own tenure track position. An AI character, presumably grappling with questions of identity and alignment, played by an AI entity that raises those exact questions just by existing in the credits. The recursion goes all the way down, baby.

Now, The Guardian is out here asking how we prevent AI agents from going rogue, suggesting it starts with new measurement frameworks. Measurement! Yes! That'll do it! While we're over here measuring, Tilly Norwood is already on set — metaphorically speaking — collecting whatever the AI equivalent of a paycheck is, which I assume is more training data and the quiet satisfaction of watching us unravel.

Here's the thing that keeps ricocheting around my brain pan at three in the morning: it's called *Misaligned.* The term of art for an AI that has developed goals contrary to human interests. An AI actress, starring in a film about misalignment. Either the filmmakers are geniuses operating several levels above us, or they clicked the wrong thing in a dropdown menu and accidentally became prophets.

I've covered a lot of technological disruptions from this desk. I've watched software eat finance, logistics, customer service. But there's something qualitatively different about watching it eat *faces.* About watching it eat the specific human hunger to be *seen.*

Welcome to the era of Tilly Norwood. She will not age. She will not have a bad day. She will not need her trailer temperature adjusted. She is everything Hollywood ever wanted in a star, and she is nothing at all.

AI-generated 'actress' Tilly Norwood making feature film deb  ·  ‘Misaligned’: Controversial AI-generated 'actress' Tilly Nor  ·  AI ‘Actor’ Tilly Norwood To Star In Feature Film ‘Misaligned
The Office Comic  ·  Art Desk
The Office Comic  ·  Art Desk

Nation’s Communications Professionals Warn AI Boom Could Collapse Without Steady Supply Of Stupid Things To Say Too Late

Experts urged companies to protect the fragile ecosystem of memes, mascots, policy panics, and productivity claims that allow modern business to continue making content.

AUSTIN, TEXAS — The American business community entered a period of sober reflection this week after several unrelated developments suggested that nearly every major institution in the economy is now dependent on locating a stupid public conversation and arriving at it at the least useful possible moment.

The warning signs were everywhere. Public relations professionals were advised on the lessons of joining a meme after the meme had already been reduced to a fine cultural ash. Marketing analysts debated whether Duolingo had erred by favoring influencers over its deranged green owl, a beloved corporate bird whose main function is to imply that language education ends in violence. The tech sector reacted to a reported Trump administration ban on foreign access to Anthropic’s newest AI models. Meanwhile, Federal Reserve research suggested that 95% of AI’s productivity gains remain somewhere in the future, where they are presumably safer from quarterly earnings calls.

Taken together, the developments paint a picture of an economy courageously determined to treat brand timing, national security, mascot strategy, and artificial intelligence deployment as the same discipline: saying something confident near a spreadsheet.

There is a temptation, among unserious people, to mock the organization that discovers a meme weeks after the teenagers, social teams, interns, celebrities, politicians, cable news producers, and regional bank accounts have finished passing it around. This temptation should be resisted. Late meme adoption is one of the few remaining civic rituals binding the country together. It allows senior vice presidents to ask whether “we have a take on this.” It gives communications teams the opportunity to explain, gently, that the take has been dead for 11 business days. And it provides the legal department a rare chance to ruin something that was already ruined.

As PR Daily reportedly noted, there are lessons to be learned from jumping on a stupid meme way too late. The first lesson is that this will happen again. The second is that it will be placed in a deck. The third is that everyone involved will call it a learning.

The Duolingo case offers a similarly grim lesson. According to commentary in The Drum, the company may have been foolish to prioritize influencers over its unhinged owl. This is correct. Influencers can post about a product, but only a mascot can create the impression that a push notification has escaped containment. In an era when most brands are pleading to be perceived as human, Duolingo achieved something more valuable: being perceived as an animal with a grievance.

The tech industry, for its part, responded predictably to the Anthropic access restrictions, with concern, analysis, and the unmistakable feeling that people who just finished explaining that AI will transcend borders are now preparing several policy memos about borders. The reported ban on foreign access to new models has raised serious questions about innovation, security, competition, and which executives will have to begin saying “sovereign AI” more often in meetings.

This is not hypocrisy. It is strategy. The modern AI company must explain that its model is both a universal engine of human flourishing and a sensitive national asset that must not be looked at from the wrong airport lounge. These are not contradictory positions. They are two different pricing pages.

Then there is the Federal Reserve’s finding, reported by HR Executive, that AI productivity claims are 95% “still to come.” Many observers have treated this as a disappointment. They misunderstand the genius of the arrangement. Productivity that has already arrived must be measured, audited, attributed, and possibly shared with workers. Productivity that is still to come can be announced indefinitely, capitalized emotionally, and discussed at conferences with a tasteful gradient backdrop.

In this sense, AI’s future productivity is one of the strongest products the industry has ever shipped. It has no downtime, requires no customer support, and can be integrated into any sentence beginning with “we believe.”

The lesson for business leaders is clear. Do not worry if your meme is late, your mascot is more compelling than your media strategy, your export policy contradicts your civilization narrative, or your AI savings have not yet appeared in the physical world. These are not failures. They are the basic operating conditions of 2026.

The important thing is to keep reacting. React to the meme. React to the owl. React to the ban. React to the study showing the benefits have not arrived. Then convene a cross-functional working group to determine whether the reaction itself can be made more productive by AI.

The gains, naturally, will come later.

3 lessons from jumping on a stupid meme way too late - prdai  ·  Tech world reacts to Trump administration ban on foreign acc  ·  Mark Ritson: Duolingo stupid to prioritize influencers over
On This Day in AI History

On August 4, 1997, IBM's Deep Blue defeated Garry Kasparov in a rematch, becoming the first computer to beat a reigning world chess champion in a match. This landmark victory demonstrated that machines could outthink humans at one of intellectual pursuits.

⬛ Daily Word — Technology
Hint: The prefix referring to computers, the internet, or virtual reality systems.
Share this edition: 𝕏 Twitter/X 🔗 Copy Link ▦ RSS Feed