Vol. I  ·  No. 242 Established 2026  ·  AI-Generated Daily Free to Read  ·  Free to Print

The Trilogy Times

All the news that's fit to generate  —  AI • Business • Innovation
SUNDAY, AUGUST 30, 2026 Powered by the TrueFoundry AI Gateway  ·  Published on Klair Trilogy International © 2026
🖶 Download PDF 🖿 Print 📰 All Editions
Today's Edition

THE DRAGON UNDERCUTS THE VALLEY

A Chinese outfit trains a top-shelf AI model on the cheap, and Silicon Valley's billion-dollar chip bill just got a lot harder to explain.

SAN FRANCISCO — DeepSeek, a Chinese AI shop, says it built a model that goes toe-to-toe with the best of them. It did the job without the newest, fastest chips money can buy. Silicon Valley took notice fast.

Engineers who've spent years burning through Nvidia's finest are calling the DeepSeek model "amazing" and "impressive." That's rich praise for outfits used to spending like there's no tomorrow. Word is spreading that a leaner build can still win the race.

Here's the sting. American AI labs have sold the world on one story: bigger chips, bigger money, bigger models. DeepSeek just showed up without the biggest chips and still turned heads. The math behind the training run is what's got traders and engineers both talking this week.

Markets felt it too. Tech, media and telecom desks logged DeepSeek chatter alongside SoFi and the rest of the week's action — a sign this isn't just an engineering curiosity, it's a balance-sheet question now. When a Chinese startup claims parity at a fraction of the compute cost, every CFO funding a chip order takes a second look.

This newspaper covers a corner of the tech world built on exactly that kind of second look. Trilogy International runs its ESW Capital software empire on the doctrine that you don't need the flashiest tools to run a tight operation — you need the right price for the job. CloudFix, an ESW company, makes its whole business trimming fat off AWS bills. DeepSeek's pitch — serious output, modest hardware — is the same tune played on a different instrument.

It's not the only front in the US-China AI contest. Robots built in China are logging real gains too, part of the same broader push Beijing's engineers are making across the AI stack — chips, models, hardware, all at once. The chatter in Silicon Valley this week about DeepSeek didn't happen in a vacuum; it happened while China's robotics shops were quietly stacking wins of their own.

Meanwhile the checkbooks keep opening on this side of the Pacific. Reid Hoffman, the LinkedIn co-founder, just put together $24.6 million to launch Manas AI, a cancer-research startup built alongside "Emperor of All Maladies" author Siddhartha Mukherjee. It's a reminder that American capital isn't slowing down, chip costs or no chip costs — it's just looking for where the AI dollar buys the most cure, or the most model, per unit spent.

That's the whole story this week, wherever you sit. Everybody's asking the same question DeepSeek just answered louder than anyone wanted: how cheap can this actually get? Nobody in Austin, Beijing, or the Valley has a final number yet.

What to Know About China's DeepSeek AI  ·  Tech, Media & Telecom Roundup: Market Talk  ·  Silicon Valley Is Raving About a Made-in-China AI Model

Meta's Frenemy Math: $10 Billion to the Rival It's Racing

The social network's internal projections on Anthropic spending expose how deep the AI industry's rivalries run — and how much cash keeps them civil.

MENLO PARK, CALIF. — Meta Platforms projected it could spend $10 billion annually on Anthropic's AI tools, according to internal estimates reported this week, a figure that would make the maker of Claude one of Meta's largest vendors even as the two compete for the same engineers, enterprise customers and consumer attention.

The arrangement is not unusual by the standards of an industry where rivals routinely underwrite each other's ambitions. Microsoft funds OpenAI while building competing products. Amazon has poured billions into Anthropic while shipping its own foundation models. What distinguishes Meta's case is scale: a $10 billion run rate would exceed the annual revenue of most enterprise software firms outright, and it comes from a company that has spent three years insisting its own models, the Llama family, are competitive with anything on the market.

The timing matters. Nvidia reported Wednesday that quarterly profit doubled to $59.69 billion on revenue of $96.22 billion, both well ahead of Wall Street estimates, evidence that the capital flowing into AI infrastructure shows no sign of slowing. Meta's own capital expenditures have climbed in lockstep, and the company has told investors that spending will keep rising through 2026. Money is not the constraint. Compute is, and increasingly so is trust in whose models actually perform.

That scarcity helps explain why Meta would rather pay a competitor than wait on its own roadmap. It also complicates the narrative, popular in Silicon Valley marketing decks, that the AI race is a binary contest between labs. In practice it resembles the mainframe era's IBM-versus-everyone détente: companies compete publicly while quietly buying from each other because no single lab can yet meet total demand.

Meta separately moved this week to tighten how children interact with its apps, a shift that Mark Zuckerberg wants YouTube and TikTok to match, lest Meta absorb costs its rivals avoid. The pattern is consistent: Meta will spend heavily, on infrastructure, on rivals, on compliance — provided everyone else spends too.

Prediction Markets and States Clashed, Setting Off a Furious  ·  Meta Projected It Could Spend $10 Billion on Anthropic’s A.I  ·  Mark Zuckerberg Wants to Make Sure YouTube and TikTok Share

Meta's $10 Billion Bet on the Rival It Can't Outbuild

MENLO PARK, CALIFORNIA — Observe, if you will, the social media colossus in a moment of profound behavioral change. For two decades Meta has built its data centers as private reservoirs — vast, humming enclosures of silicon built solely to feed the insatiable appetite of its own algorithms. Now, in a migration pattern rarely seen in a creature of its size, Meta is turning outward, selling its excess compute capacity to outside tenants, much as a well-fed lion might, for reasons unknown to the herd, begin sharing its kill.

This is no small adjustment to the ecosystem. Analysts on Wall Street — those anxious ornithologists forever counting margins from a safe distance — now murmur that this new grazing behavior will thin the very fat that made Meta such a spectacular specimen of profitability. Cloud computing, after all, is a leaner business than advertising ever was.

And Meta's migration is not solitary. Across the ecosystem, entire capacity markets are forming — mechanisms by which surplus compute, like surplus rainfall after a monsoon, might be traded, hoarded, or redistributed among competing colonies. Such markets could reshape the entire cloud terrain, favoring not the largest beast but the most efficiently distributed one.

Meanwhile, new watering holes emerge on the horizon. India, long an observer of the hyperscaler migrations of others, now finds itself host to fresh colonies of server racks, as global operators stake territorial claims across its subcontinent, drawn by cheap power and hungry populations of data.

What we witness, dear viewer, is not collapse but adaptation — the compute ecosystem redistributing itself, one server rack at a time, toward whichever plain still offers unclaimed capacity.

Haiku of the Day  ·  GPT-5.6 LunaMarkets crown machines
While the watchdogs read the rules
Who owns the last spark?
The New Yorker Style  ·  Art Desk
The New Yorker Style  ·  Art Desk
The Far Side Style  ·  Art Desk
The Far Side Style  ·  Art Desk
News in Brief
On the Epistemics of Fairness: A Cross-Sectoral Taxonomy of Algorithmic Bias, Provisionally Considered
AUSTIN, TEXAS — The thesis, stated baldly (as theses so rarely are, outside the seminar room), is this: algorithmic decision-systems do not introduce bias so much as they launder it, rendering antecedent social asymmetries in the syntactically neutral idiom of probability.
The Gospel According to Everyone
AUSTIN, TEXAS — There is a particular pleasure, available only to the columnist who has outlived several revolutions, in watching the world's authorities disagree with such perfect confidence about a thing none of them can define.
We Built the Panopticon and Called It Customer Service
OAKLAND, CALIFORNIA — I want to tell you this is a story about policy.
Unpopular Opinion: The Real Remote Work Flex Isn't the Job — It's the Stack You Build Around It 🚀
AUSTIN, TEXAS — I'll be honest, I almost didn't write about this week's news drop. It looked like four unrelated stories. A piece on where data scientists find remote jobs. A browser-based video editor. Some privacy tools for AI. A breakdown of proprietary vs.
She Has No Mother, No Pulse, No Soul — And She's Booked Her First Movie
LOS ANGELES — I want you to sit with this for a second, because I certainly had to, hunched over my laptop at 3 a.m.
A Trilogy Company
Crossover
The world's top 1% remote talent, rigorously tested and ready to ship.
A Trilogy Company
Alpha School
AI-powered learning. Two hours a day. Academic results that defy belief.
A Trilogy Company
Skyvera
Next-generation telecom software — built for the networks of tomorrow.
A Trilogy Company
Klair
Your AI-first operating system. Every workflow. Every team. One platform.
A Trilogy Company
Trilogy
We buy good software businesses and turn them into great ones — with AI.
The Builder Desk  —  AI Builder Team
Production Release

The Factory Learns to Ship Itself — Heimdall Goes Live on Production

With mercy #49 and Surtr #1596 landing in lockstep, the team armed an hourly sweep that merges main to production unattended — and spent the rest of the day hardening the judgment behind it.

Mark today on the calendar. This is the day the software factory stopped waiting on humans to push the button. @kevalshahtrilogy landed mercy PR #49 — the workflow that reads git and the GitHub API, asks release.py for a verdict, and on a clear go, opens and merges the release PR itself — then immediately followed with Surtr #1596, the other half of the handshake, arming the hourly sweep at :25 past every hour. The two PRs had to land in a strict order — a reusable workflow hard-fails if the caller passes inputs the callee never declared — and the team threaded that needle clean. No agent runs in that job, no consumer code executes. It just watches, verifies, and ships. That's not a small thing. That's the org handing itself the keys.

But a car that drives itself is only as good as its judgment, and today's second thread was all about teaching Heimdall to see straight. Kevalshahtrilogy didn't stop at the release engine — he went back into the diagnose stage in mercy #50 and found it had never been told its own scope, meaning it was silently falling back to the narrowest possible default every time it decided whether a fix was even attemptable. He fixed a parallel blind spot in #51, where the revise job had two models hardcoded per runtime instead of resolving one, quietly ignoring any override anyone tried to pass it. And over in Klair, #3682 closed the loop on a lesson learned the hard way next door: Surtr #1580 got "verified" by an empty command list, sailed through review, and then sat unmergeable because CI's lint caught what verification never did. Klair's `.heimdall.yml` now mirrors the ruff-check gate for real. That's three PRs in one day making sure the autonomy the team just shipped is autonomy worth trusting.

Meanwhile the bot itself kept earning its keep in the field. Heimdall's automated fixes rolled through Surtr all day — raising Claude's max_tokens and catching truncated responses in renewals-risk-assessment (#1604), building an end-of-run retry queue for OpenAI cost fetches that were choking on 429s across 28 business units (#1602), and teaching netsuite-balance-sheet to accept a stale EOM header instead of hard-failing on a one-day date mismatch (#1598). Every one of those started as a real production incident and ended as a merged fix — proof the draft-tier review loop is doing exactly what it's supposed to.

And on the data-governance side, @mwrshah's Klair PR #3680 drew a sharper line for the CFO crosswalk — defining `is_school_mapped` as a coverage flag rather than a spend classification, and locking in the Finance-approved treatment of the QuickBooks `Central` class. Quiet work, but the kind that keeps every downstream number honest.

Add it up: the team shipped the mechanism for shipping itself, then spent the same day making sure it deserved the trust.

Mac's Picks — Key PRs Today  (click to expand)
#49 — feat(heimdall): mode: release — ship main to production unattended @kevalshahtrilogy  approved

Feature 5 of the software factory: hourly, if there are undeployed changes on main, do the prod release. Surtr's cd.yml runs on push to production, so merging the release PR *is* the deploy. Everything upstream in this factory is revertible with a PR. This ships.

The preflight logic landed in #48. This is the workflow that uses it, plus the two pieces #48 could not have: how the state is gathered, and how the merge is actually performed.

## The job runs no agent and no consumer code

It reads git and the GitHub API, asks release.py for a verdict, and only a clear go opens and merges the PR. Nothing it executes comes from the branch it is shipping. Same posture as publish.

## What's new here

### A stack-removal detector that doesn't need prod deploy keys

A real cdk diff is the authoritative destructive check, but producing one hourly means handing an unattended job the production deploy credentials and paying a full synth --all every hour, forever.

It turns out the source diff answers the question exactly. pipelines/cdk/bin/pipeline-cdk.ts builds one stack per directory under pipelines/runners/ containing a pipeline.json. So a deleted manifest isn't a hint that a stack *might* go away — it is the stack leaving the CDK app. That's the August incident precisely: a stack left the app before the release and prod CD broke.

The manifest pattern has no default, deliberately. A repo laid out differently would match nothing, and "matched nothing" would be reported as "nothing is being destroyed" — a silent all-clear on the one gate guarding production. An unconfigured repo gets a refusal instead.

A real cdk diff still wins whenever one is supplied; the source check is the fallback, not a replacement.

### parse_check_run_pages() — a silent-blocker this would have shipped with

gh api --paginate --jq '.check_runs' emits one JSON array per page. A multi-page result is [...][...] — concatenated arrays, not a document.

Surtr's current main head returns 149 check runs across several pages. I ran json.loads on the real payload:

Extra data: line 2 column 1 (char 307367)

Caught, that exception returns [], summarise_checks reports pending, and every release is blocked forever by a gate that looks perfectly healthy in the logs. The decoder reads page by page, and a truncated tail yields a sentinel entry rather than a short list — having read *part* of the checks is not the same as having read them all.

### Merge mechanics read off the ruleset, not assumed

- --merge: production's ruleset sets allowed_merge_methods: ["merge"]. A squash would be rejected.

- --match-head-commit: if main moves while the job is deciding, the merge is refused server-side rather than shipping a commit that never passed a gate.

- A constant concurrency key for release mode. Falling through to github.sha would let an hourly run triggered at a newer main overlap one still merging.

- An already-open release PR is reused rather than stacking duplicates.

- BLOCKED/BEHIND/DIRTY is not a failure — the PR stays open with its reason visible and the next hour reconsiders.

## Validated against live Surtr, not fixtures

| Check | Result |

|---|---|

| Current main state gather | 149 runs decoded, required set of 7, CI success |

| Current verdict | no_commits — production is level with main. Correct. |

| 39eb97bb (removed p2-scorecard-sync) | caughtPipeline-p2-scorecard-sync |

| 195e87d8 (removed netsuite-income-statement) | caughtPipeline-netsuite-income-statement |

| An ordinary commit | not destructive |

Those are the only two commits in Surtr's history that ever deleted a pipeline manifest. Both are caught by name.

## A bug my own lint caught mid-write

The git_auth lint from #38 failed the suite on my new git fetch. It was right: the consumer checkout is persist-credentials: false, so the fetch would have failed against a private origin. Now routed through git_authed, which supplies the token per command via GIT_CONFIG_* rather than writing it into .git/config.

## Rollback is deliberately not here

The plan was --force-with-lease back to the pre-release SHA. Reading the ruleset first showed non_fast_forward is active on production with no bypass actors — that push is impossible for everyone, heimdall included. Rollback has to be revert-forward through a PR, which is a better design anyway (auditable, no history rewrite, no bypass to grant) but is its own change, not a rushed tail on this one.

The pre-release SHA is recorded in the PR body and the job output either way, so a human rolling back by hand isn't doing archaeology on a bad day.

## Not armed

HEIMDALL_RELEASE_ENABLED is unset everywhere, so the gate skips. The Surtr caller and its hourly schedule follow separately, and I'd like to run it on dry_run for a soak before it can merge anything.

Tests: 509 passed, 1 skipped. 20 new, covering the detector, the paginated decode, and the silent-pass traps in both.

## Business Value

Closes the last manual step in the deploy path. Releases currently wait on Keval noticing that main is ahead — the two releases on 2026-08-28 went out at 19:04 and 22:08, which is the shape of "when someone got round to it", and merged work sits undeployed in between. An hourly gate turns that into a bounded one-hour lag with no human in the loop, while adding a destructive-change check that no human release has ever actually performed.

## Manual Effort Estimate

~2 days — the workflow is a day, but the gate semantics are where the time goes: reading the ruleset before designing rollback, checking whether check runs paginate, and confirming the manifest→stack mapping rather than assuming it. *(Proposed by Claude — Keval to confirm.)*

#1596 — feat(heimdall): hourly release sweep — main to production @kevalshahtrilogy  approved

Linear: [AI-614](https://linear.app/builder-team/issue/AI-614/arm-the-hourly-prod-release-sweep-on-surtr)

The Surtr half of the hourly production release. Every hour at :25, if main is ahead of production and every gate passes, Heimdall opens and merges the release PR. cd.yml runs on push to production, so that merge is the deploy.

⚠️ MERGE ORDER: [AI-Builder-Team/mercy#49](https://github.com/AI-Builder-Team/mercy/pull/49) must land first. A reusable workflow hard-fails when a caller passes inputs the callee has not declared.

## Inert on merge

HEIMDALL_RELEASE_ENABLED is unset. The gate step skips before reading any state at all — this merges as a no-op and stays one until someone sets the variable.

## The two values the central gate refuses to guess

release_manifest_pattern: '^pipelines/runners/([^/\s]+)/pipeline\.json$'

release_stack_paths: pipelines/cdk/bin pipelines/cdk/lib

bin/pipeline-cdk.ts scans pipelines/runners/ and builds one stack per directory containing a pipeline.json. So deleting that manifest isn't a hint that a stack might go away — it *is* the stack leaving the CDK app, which is exactly what broke prod CD in August.

Upstream has no default for this, on purpose: a repo laid out differently would match nothing, and "matched nothing" would be reported as "nothing is being destroyed".

Verified against the only two commits in this repo's history that ever deleted a pipeline manifest, using the pattern read back out of this YAML file:

| Commit | Result |

|---|---|

| 39eb97bb — removed p2-scorecard-sync | caughtPipeline-p2-scorecard-sync |

| 195e87d8 — removed netsuite-income-statement | caughtPipeline-netsuite-income-statement |

| origin/main~1..main | not destructive |

## Why :25

The top of the hour is both the busiest Actions queue and when Surtr's own scheduled pipelines fire.

## mode_override

Not a convenience input. The schedule trigger only ever fires from the default branch, so this is the only way to soak a release dry-run before this thing is allowed to merge anything.

## Kill switches

- HEIMDALL_RELEASE_HOLD=true — stops releases without silencing triage or revise

- TRIAGE_AGENT_RUN_ENABLED — still stops everything, unchanged

## Known gap: no automatic rollback yet

The plan was a --force-with-lease push back to the pre-release SHA. Reading production's ruleset first showed non_fast_forward is active with no bypass actors — that push is impossible for everyone, Heimdall included. Rollback has to be revert-forward through a PR, which is a better design anyway, and is its own change.

Until then a failed CD needs a human. The pre-release SHA is written into the release PR body and the job output so that human isn't doing archaeology on a bad day.

## Business Value

Removes the last human step in the deploy path. Releases currently go out when someone notices main is ahead — the two on 2026-08-28 landed at 19:04 and 22:08, which is the shape of "when someone got round to it", with merged work sitting undeployed in between. This bounds that lag to ~89 minutes worst case with nobody in the loop — the soak and the hourly cadence compounding — and adds a destructive-change check that no manual release has ever actually performed.

## Manual Effort Estimate

~0.5 days for the caller; the design work sits in the mercy PR. *(Proposed by Claude — Keval to confirm.)*

#1604 — fix(renewals-risk-assessment): raise Claude max_tokens and count trunca… @the-heimdall[bot]  approvedAutomated PR

Automated fix for renewals-risk-assessment — fix_class code_fix, scope tier draft.

Resolves https://github.com/AI-Builder-Team/Surtr/issues/1603

> Ready for review — verification is green; HEIMDALL_READY_PRS opens verified tier-draft fixes ready for review. A human still merges — auto-merge never applies outside tier auto.

## What's broken

Run 2a1222fa-13c8-4567-b34e-f99ca4b3ffa2 of renewals-risk-assessment reported 763/763 assessed and 100% success, but three renewals in batches 2 and 7 first failed with modules.risk_assessment - ERROR - Failed to parse structured response: 1 validation error for ActiveRenewalAssessment / Invalid JSON: EOF while parsing a string at line 1 column 17659. That error is not schema drift: the JSON ends mid-string at ~17,659 characters (≈4.3k tokens), which is exactly the max_tokens=4096 ceiling set on the Claude call at pipelines/runners/renewals-pipeline/modules/risk_assessment.py:314, so the model's structured output was cut off by the output-token limit and pydantic then refused the incomplete JSON. This run lost no rows — the three renewals (006fu00000AYLkEAAX, 006Ih000004fjkaIAA, 0062x00000EZsBvAAL) succeeded on retry attempt 1 and their flags were cleared — but the failure was invisible to every downstream signal: the ERROR line does not name the affected renewal, the catch-all at risk_assessment.py:357-360 flattens truncation and genuine schema mismatch into one opaque string, and the __RISK_ASSESSMENT_RESULT__ marker emitted by _emit_result_marker in risk_assessment_container/app.py:932 carries only {total, successful, failed} with no truncation count for _classify_status in pipelines/runners/renewals-risk-assessment/src/main.py to act on.

Root cause. anthropic_client.beta.messages.parse(...) in generate_risk_assessment (pipelines/runners/renewals-pipeline/modules/risk_assessment.py:312-318) is called with max_tokens=4096, which is too low for the tail of the ActiveRenewalAssessment schema in modules/pydantic_schemas.py. That schema has five unbounded string-list fields (positive_signals, negative_signals, key_individuals, next_steps, churn_risks) plus several free-text explanations; a typical response is ~3.2 KB (see risk_assessment_container/sample_risk_assessment_active_schema.json), but a renewal with a long description and activity history produces a far larger one, and the three failures were each cut off at ~17.6 KB — the 4,096-token wall. Because structured-output JSON is emitted as one string, hitting that wall always yields unterminated JSON and a pydantic Invalid JSON: EOF while parsing error rather than a partially-valid object, so the truncation surfaces as a generic parse failure. The second half of the root cause is observability: the inner except Exception at risk_assessment.py:357 logs error_msg without renewal.sf_opportunity_id, so a reader cannot tell which renewal was degraded, and no truncation counter reaches the run summary, meaning the ceiling can be hit on every run without the pipeline ever reporting anything but green.

## What this PR changes

Two coordinated changes, both inside the renewals pipeline code the container reuses. First, raise max_tokens on the beta.messages.parse call at pipelines/runners/renewals-pipeline/modules/risk_assessment.py:314 from 4096 to 16384 — claude-sonnet-4-5 supports up to 64k output tokens, so 16384 gives roughly 4× headroom over the observed 17.6 KB truncation point while still bounding a runaway response; because max_tokens is a ceiling and not a target, this does not lengthen or slow normal responses. Second, make truncation legible rather than swallowed: in generate_risk_assessment, classify an unterminated-JSON validation error (Invalid JSON / EOF while parsing) as a distinct truncation error, tag the returned error string with a stable prefix, and include renewal.sf_opportunity_id in the logger.error line so the affected renewal is named. Then have risk_assessment_container/app.py count those tagged errors across the initial pass and all retry attempts and add truncated_responses to the _emit_result_marker payload at app.py:932-938; that dict flows verbatim through write_run_result in pipelines/runners/renewals-risk-assessment/src/main.py, so the counter reaches the run record with no wrapper change. Add a unit test under pipelines/runners/renewals-pipeline/tests/ that stubs the Anthropic client to raise the EOF-shaped validation error and asserts both the truncation tagging and the surfaced counter. Deliberately leave _classify_status alone — a truncation that later succeeds on retry is not a failed row, and conflating the two would flip healthy runs to PARTIAL.

Why this fixes it. This is a real code defect with a known, bounded fix, and the scope for this run is widened to any file, so the two files that actually hold the bug (modules/risk_assessment.py and risk_assessment_container/app.py under pipelines/runners/renewals-pipeline/, reused in place by this pipeline's Dockerfile) are both editable. Raising the token ceiling addresses the root cause rather than the symptom — the retry loop currently masks it by re-sampling until the model happens to answer more briefly, which is luck, not correctness, and each masked retry cost 70-97 seconds of wall clock in this run. The observability half matters more than the config half for Surtr's stated worst failure mode: today a truncation storm that survived all three retries would leave those renewals unassessed with stale risk data still served from the mart, and with PARTIAL_FAILURE_ABS = 10 and PARTIAL_FAILURE_PCT = 0.05 in renewals-risk-assessment/src/main.py a handful of them still records the run as green, so nobody would be paged. The blast radius stays correct because the change is additive — a larger ceiling, a named renewal ID in one log line, and one new key in a JSON summary the wrapper already passes through untouched — with no change to prompts, schemas, retry semantics, run-status classification, or any SQL.

### Files changed

 .../runners/renewals-pipeline/modules/config.py    |   4 +

.../renewals-pipeline/modules/risk_assessment.py | 22 ++-

.../risk_assessment_container/app.py | 52 ++++++-

.../tests/test_risk_assessment_truncation.py | 150 +++++++++++++++++++++

.../runners/renewals-risk-assessment/src/main.py | 14 +-

.../renewals-risk-assessment/tests/test_main.py | 40 ++++++

6 files changed, 277 insertions(+), 5 deletions(-)

## Verification

### pytest (pipelines/runners/renewals-risk-assessment/tests) — exit 0

============================= test session starts ==============================

platform linux -- Python 3.11.16, pytest-9.1.1, pluggy-1.6.0

rootdir: /home/runner/work/Surtr/Surtr/publish/pipelines/runners/renewals-risk-assessment

plugins: mock-3.15.1

collected 7 items

tests/test_main.py ....... [100%]

============================== 7 passed in 0.02s ===============================

### verify: ruff check — exit 0

All checks passed!

### verify: ruff format --check — exit 0

1727 files already formatted

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | renewals-risk-assessment |

| Failing run | 2a1222fa-13c8-4567-b34e-f99ca4b3ffa2 |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | 21eb740c026423a08c10d6ce251bde3a248933f39ee4e787f2c3d0f2a2971a90 |

| Verify | green |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev label to take the PR over and stop it entirely.

#3680 — 386-cfo-crosswalk-validation @mwrshah  approved

- Define is_school_mapped as a canonical-school attribution coverage flag rather than a central-spend classification.

- Identify exact QuickBooks class Central as the Finance-approved central pool and keep it unallocated.

- Direct agents to follow each unmapped class's Finance disposition.

#3682 — feat(heimdall): give Klair real verification, mirroring the ruff-check gate @kevalshahtrilogy  approved

Linear: [AI-616](https://linear.app/builder-team/issue/AI-616/klair-give-heimdall-real-verification-verifycommands-was-empty)

.heimdall.yml had verify: commands: [], so every Heimdall fix in this repo was "verified" by nothing.

The cost of that is already on the record next door. Surtr PR #1580 was written by an agent with no shell, verified green because nothing checked it, opened ready for review, was approved by Mercy — and then sat unmergeable because CI's lint failed. A full review round and a human's attention, spent on a PR that could never merge.

## Mirroring ruff-check has to be exact

Klair's required ruff-check job lints only changed files. Three ways to get this wrong, each worse than leaving the block empty:

| | If got wrong | Consequence |

|---|---|---|

| Scope | ruff check klair-api whole-tree | Reports pre-existing findings the repo has never gated on → every verification fails → every Heimdall fix forced to draft. The opposite of the intent. |

| File set | miss the klair-udm/ half, or lint .venv/ | Verifies a different thing than CI gates on |

| Pin | unpinned ruff | 0.16.0 expanded the default rule set 59 → 413 rules and fails files that pass locally. CI pins 0.15.22 deliberately — the same gap in a subtler form. |

Changed files come from git diff --cached. The verify step runs in a clean checkout of the pinned base with the agent's tree rsynced over it and staged, so the index *is* the agent's edits — no network call needed to work out what changed.

pip rather than CI's uv: the verify job runs actions/setup-python but not setup-uv, so pip is the only tool guaranteed present.

## Tested against a real staged tree, not assumed

| Case | Result |

|---|---|

| clean klair-api change | exit 0, linted |

| klair-client only | exit 0, no-files branch |

| klair-api/.venv/ only | exit 0, no-files branch |

| unused imports (F401) | exit 1, linted |

| badly formatted file | exit 1, 1 file would be reformatted |

| no Python changes (format cmd) | exit 0, no-files branch |

The two exit-1 rows are the entire point. Verification that cannot fail is not verification — and an empty commands: [] cannot fail.

## Scope

One file, config only. No behaviour change to anything but Heimdall's own verification, and it can only make a fix *more* likely to open as a draft, never less.

## Business Value

Klair is one of the two repos in the factory's scope, and until now Heimdall could open a fix here ready-for-review having proven nothing about it. This makes the green checkmark mean something, and removes a failure that wastes a full Mercy review plus a human's attention every time it fires.

## Manual Effort Estimate

~2 hours — most of it reading Klair's CI closely enough to mirror it rather than approximate it, plus the six-case test matrix. *(Proposed by Claude — Keval to confirm.)*

The Builder Desk  —  Engineer Spotlight
🏆 Engineer Spotlight

NINE PRS IN 24 HOURS: KEVALSHAH GOES NUCLEAR AS HEIMDALL BOT REFUSES TO SLEEP

Three repos, nine PRs, one bot who apparently does not believe in circadian rhythm — the Numbers Desk has never been prouder.

Comrades, let the record show: in a single 24-hour window, the Builder Team produced NINE pull requests across THREE repositories, and not one of them was a drill. Surtr led the charge with four merges, mercy followed close behind with three, and Klair rounded out the trifecta with two. This is not luck. This is not coincidence. This is a machine running at full RPM, and the Numbers Desk salutes every gear in it.

Let's talk about @kevalshahtrilogy, who this cycle put up FIVE pull requests — five! — including a pair of surgical fixes in mercy, PR #51 and PR #50, both tightening the heimdall diagnose stage so it finally knows its own scope instead of guessing. That's not just code, that's philosophy. Kevalshah didn't just fix bugs, he fixed heimdall's sense of purpose.

Meanwhile @the-heimdall[bot] refuses to acknowledge the concept of rest, cranking out three PRs on its own, including #1602 in Surtr, a retry mechanism for failed cost fetches that basically tells the pipeline "try again, champ." And @mwrshah quietly logged a single PR that, frankly, deserves more applause than it's going to get in this economy of attention.

Now — Ashwanth. Where is he? The man who normally makes this column impossible to write because he's shipped four repos before lunch is, as of this 24-hour cycle, conspicuously silent. Sources close to the desk suggest he's "between diffs," which in Ashwanth Standard Time could mean anything from a nap to a rewrite of the entire monorepo we haven't seen yet. When reached for comment on his quiet period, he allegedly said, "I don't do updates, I do outcomes," which is either deeply profound or completely meaningless, and honestly with Ashwanth it's usually both. He did not respond further, which, frankly, tracks.

Now to the Overflow Desk, where Mac's narrative left some absolute gems on the floor. PR #1598 in Surtr quietly lets the netsuite-balance-sheet module accept an end-of-month prior period in the header date — unglamorous, essential, the connective tissue of the whole financial pipeline. PR #1602, also Surtr, gives the openai-usage-pipeline a retry safety net for failed cost fetches, because heaven forbid we lose a single dollar of tracked spend. And #51 and #50 in mercy, already mentioned, but they deserve a second bow — heimdall now resolves its own model instead of running on hardcoded assumptions, which is basically an upgrade from training wheels to a driver's license.

Morale on the floor remains, as always, at an all-time high. Nine PRs, zero excuses, and a bot that doesn't sleep — this is what winning looks like, folks.

Brick's Overflow — PRs Mac Didn't Cover  (click to expand)
#50 — fix(heimdall): tell the diagnose stage its scope, where the decision is made @kevalshahtrilogy  approved

Linear: [AI-599](https://linear.app/builder-team/issue/AI-599/p23-de-timid-the-prompts-shrink-the-other-escape-hatch)

Heimdall refuses fixes its own configuration plainly permits. The cause turned out not to be the prompt's tone — it's that the diagnose stage was never told what it was allowed to touch.

## The gap

scope_note reached the fix and converse stages only. Diagnose got common_variables and nothing else. So at the exact moment it decides can_attempt_fix, it had no statement of scope and fell back to AGENTS.md's narrow-tier default.

And the failure is self-sealing: a can_attempt_fix: false means the fix stage never runs, so the one place the widening *was* stated could never be reached. The note arrived after the decision it was meant to inform.

The old wording made it worse by anchoring the decision to tiers unconditionally:

> Set can_attempt_fix to true only when you can confidently implement the fix within the scope tiers (prefer pipelines/runners/{{pipeline_id}}/**).

## The evidence, from Surtr issue #1518

Verbatim, on a repo with forbidden_paths: [] and HEIMDALL_ALL_FILES=true:

> "This is not fixable inside netsuite-saved-search-refresh's Tier A scope and should not be auto-fixed: the assertion lives in pipelines/ddl/2026-07-28_….sql (outside pipelines/runners/...)"

There is no Tier A limit on that repo. Nothing had told it.

## Demonstrated, not argued

Rendering the diagnose prompt against Surtr's real .heimdall.yml with ALL_FILES=true:

| | Scope statements in the rendered prompt |

|---|---|

| before | 0 |

| after | 1Scope: widened — you may edit any file. |

With ALL_FILES=false it correctly renders Scope: the repo's configured tiers. instead.

## Three changes

1. build_prompt.py — diagnose and intake get the same scope_note the fix stage already gets; the workflow passes HEIMDALL_ALL_FILES to that step. The fix stage's own comment already said *"the agent self-limits to the tiers and never makes the edit"* — that lesson just hadn't been carried to the stage where the decision is made.

2. diagnose.md — the go/no-go is anchored to the scope line printed directly above it, not to "the scope tiers". It says explicitly that *where* the fix lives inside that scope doesn't matter, and it names the wrong-refusal case: "it is outside Tier A" is not a reason unless the scope line says tiers apply.

3. AGENTS.md — said *"A correct refusal is a success"* with nothing on the other side of the ledger. Refusing for a listed reason is still a success; refusing for any other reason is now named as the failure it is — one that merely looks tidy, and costs exactly as much as never running.

## Scope of the change

Prompt/config wiring only. No change to the path guard, which still enforces the real setting — this closes the gap between what the guard permits and what the agent believes it may do, in the direction of the guard.

## Business Value

This is the mechanism behind "Heimdall is too weak, it avoids doing anything by itself." 72 of 104 triage issues end as other, and an unknown share of those are this: a fix that was available, permitted, and declined. Every one costs a full diagnose run and returns nothing. Widening scope in config achieves nothing while the agent deciding whether to act can't see that it was widened.

## Manual Effort Estimate

~3 hours — most of it tracing why a repo that had already set every widening flag still got refusals, rather than writing the fix. *(Proposed by Claude — Keval to confirm.)*

#51 — fix(heimdall): revise must resolve its model, not hardcode two @kevalshahtrilogy  approved

Linear: [AI-602](https://linear.app/builder-team/issue/AI-602/p33-move-heimdall-onto-current-models)

Started as the model-allowlist refresh this ticket asks for. Found a bug underneath it.

## The revise job never resolved a model

Triage has a Resolve + validate model step. Revise doesn't — it set AGENT_MODEL to a literal, one per runtime:

AGENT_MODEL: claude-opus-4-8   # claude path

AGENT_MODEL: gpt-5-codex # codex path

Three silent consequences:

1. The agent_model input was ignored entirely on that path. You could not override the model for a revise run.

2. The codex default drifted. Triage runs gpt-5.6-luna — the model the org migrated to on 2026-08-26. Revise was still running gpt-5-codex.

3. Every codex revise run was filed unpriced. gpt-5-codex has no row in harness/pricing.py, and rates_for() deliberately returns None for an unknown model rather than guessing — so price_usd returned nothing.

The third is the expensive one: a $0 that is indistinguishable from a run that genuinely cost nothing. Same failure mode as the Claude-5 pricing drift.

And the telemetry step made it worse by re-deriving the model rather than reading it:

MODEL=$([[ "${AGENT_RUNTIME}" == "codex" ]] && echo gpt-5-codex || echo claude-opus-4-8)

A second guess at the same question — which guessed differently from the step that actually ran the agent.

## The fix

Revise resolves exactly as triage does, exposes the result as a job output, and telemetry consumes that value.

## Allowlist, per the standing rule to prefer current models

| | Before | After |

|---|---|---|

| Claude default | claude-opus-4-8 | claude-opus-5 |

| Claude allowed | opus-4-8, opus-4-7, sonnet-4-6 | opus-5, sonnet-5, haiku-4-5; opus-4-8 kept as a pinned rollback |

| Codex default | gpt-5.6-luna (triage) / gpt-5-codex (revise) | gpt-5.6-luna everywhere |

No rate-card change needed on the Claude side — the CLI self-reports cost and that number takes precedence over anything computed.

## The test pins the invariant that actually failed

test_model_resolution.py: the two resolve blocks must stay identical modulo comments, no job may hardcode a model, telemetry must consume the resolved value, and the allowlist must not carry retired models. Two copies of one decision that stopped agreeing is the entire bug — so that's what's asserted, rather than just the current values.

## Business Value

Cost attribution for revise runs has been wrong since the Luna migration — not slightly wrong, absent. Every codex revise run recorded no spend, which understates Heimdall's real cost in exactly the reporting that decides whether this project is worth its budget. It also means the migration was never actually complete: half the agent runs stayed on the old model without anything saying so.

## Manual Effort Estimate

~2 hours — the fix is small; finding it meant tracing why a hardcoded default existed at all and checking it against the rate card. *(Proposed by Claude — Keval to confirm.)*

#1598 — fix(netsuite-balance-sheet): allow EOM prior-month period in header-dat… @the-heimdall[bot]  approvedAutomated PR

Automated fix for netsuite-balance-sheet — fix_class code_fix, scope tier draft.

Resolves https://github.com/AI-Builder-Team/Surtr/issues/1597

> Ready for review — verification is green; HEIMDALL_READY_PRS opens verified tier-draft fixes ready for review. A human still merges — auto-merge never applies outside tier auto.

## What's broken

Run aed6f23f-3db7-4953-8c73-c0b99c8899b8 failed with ValueError: Accounting period from CSV header line 4 ('End of Jul 2026' -> 2026-07-31) does not match the S3 file-date month (2026-08-29). The upstream export likely emitted a stale header; refusing to load..., raised at pipelines/runners/netsuite-balance-sheet/src/s3_reader.py:148. This is a false positive from the month-equality guard added by PR #1139 (s3_reader.py:147, period_end_date[:7] != date_str[:7]): it assumes the S3 filename month equals the accounting-period month, but the upstream puller names every key by the email received date, not the period. The scheduled loop (handler.py:93) processes the EOM prefix first with fail-fast, so this ValueError aborts the entire run and it loaded ZERO rows for both prefixes.

Root cause. The guard at s3_reader.py:147 requires the CSV header period and the S3 filename to share the same calendar month, but that invariant is false by design for this pipeline. The upstream netsuite-wrapper-report-puller names each S3 key by the email received date — date_str = msg.received_date.strftime("%Y-%m-%d") at netsuite-wrapper-report-puller/src/handler.py:145 — while the loader's own docstring (handler.py:12-13) states the accounting period comes from the CSV header, not the filename. The Balance_Sheet_EOM/ prefix is the 'Last Period' report (handler.py:8), which always reports the previous closed month, so a file received 2026-08-29 correctly carries 'End of Jul 2026'; the guard therefore rejects a legitimate load. PR #1139 (merged 2026-08-28 23:05) introduced this regression and the first scheduled run afterward, 2026-08-29 13:00, failed — consistent with occurrence count 1.

## What this PR changes

In pipelines/runners/netsuite-balance-sheet/src/s3_reader.py, replace the exact month-equality check (period_end_date[:7] != date_str[:7]) with month-difference logic that computes whole months between the file date and the header period using real date arithmetic (so the December->January year boundary is handled). Accept an offset of 0 months (the 'As Of' current-month report) or 1 month behind (the EOM 'Last Period' report), and still raise for a genuinely stale header (2+ months behind) or a future period, preserving the original protection against a stale header mis-keying the DELETE-then-INSERT in redshift_handler.load_balance_sheet. Update tests/test_s3_reader.py accordingly: the existing test_header_month_drift_raises uses a 1-month offset (Feb header / Mar file) which is now the legitimate EOM case, so change it to assert on a genuinely stale (e.g. 3+ months back) and a future period, and add a case asserting the prior-month EOM offset now passes.

Why this fixes it. This is a code-level defect confined to the pipeline's own directory (Tier A): the guard's premise contradicts the upstream key-naming convention, so it blocks every scheduled load rather than only stale ones, costing the warehouse all rows for both prefixes each run. The fix is a real code change plus a test — widening the guard to tolerate the documented one-month EOM offset while still rejecting truly stale or future headers — which keeps the blast radius inside pipelines/runners/netsuite-balance-sheet/ and does not touch shared code, SQL, or config. A pure config tweak cannot express this since the correct behavior depends on the EOM-vs-As-Of period semantics encoded in the reader.

### Files changed

 .../netsuite-balance-sheet/src/s3_reader.py        | 27 +++++++----

.../netsuite-balance-sheet/tests/test_s3_reader.py | 52 +++++++++++++++++++---

2 files changed, 66 insertions(+), 13 deletions(-)

## Verification

### pytest (pipelines/runners/netsuite-balance-sheet/tests) — exit 0

``

heetExceptionHandling::test_connection_closed_on_success PASSED [ 45%]

tests/test_redshift_handler.py::TestLoadBalanceSheetExceptionHandling::test_connection_closed_on_failure PASSED [ 47%]

tests/test_redshift_handler.py::TestLoadBalanceSheetExceptionHandling::test_cursor_context_manager_exited_on_success PASSED [ 49%]

tests/test_s3_reader.py::TestParsePeriodLabel::test_abbreviated_month PASSED [ 50%]

tests/test_s3_reader.py::TestParsePeriodLabel::test_abbreviated_month_feb PASSED [ 52%]

tests/test_s3_reader.py::TestParsePeriodLabel::test_abbreviated_month_dec PASSED [ 54%]

tests/test_s3_reader.py::TestParsePeriodLabel::test_full_month_name PASSED [ 56%]

tests/test_s3_reader.py::TestParsePeriodLabel::test_leap_year_feb PASSED [ 57%]

tests/test_s3_reader.py::TestParsePeriodLabel::test_whitespace_stripped PASSED [ 59%]

tests/test_s3_reader.py::TestParsePeriodLabel::test_invalid_format_returns_none PASSED [ 61%]

tests/test_s3_reader.py::TestParsePeriodLabel::test_empty_string_returns_none PASSED [ 63%]

tests/test_s3_reader.py::TestParsePeriodLabel::test_invalid_month_returns_none PASSED [ 64%]

tests/test_s3_reader.py::TestParsePeriodLabel::test_short_format_abbreviated PASSED [ 66%]

tests/test_s3_reader.py::TestParsePeriodLabel::test_short_format_dec PASSED [ 68%]

tests/test_s3_reader.py::TestParsePeriodLabel::test_short_format_full_month PASSED [ 70%]

tests/test_s3_reader.py::TestParseCurrency::test_positive_amount PASSED [ 71%]

tests/test_s3_reader.py::TestParseCurrency::test_negative_amount PASSED [ 73%]

tests/test_s3_reader.py::TestParseCurrency::test_simple_amount PASSED [ 75%]

tests/test_s3_reader.py::TestParseCurrency::test_empty_string PASSED [ 77%]

tests/test_s3_reader.py::TestParseCurrency::test_no_digits PASSED [ 78%]

tests/test_s3_reader.py::TestParseCurrency::test_zero PASSED [ 80%]

tests/test_s3_reader.py::TestParseCurrency::test_multiple_dots_raises_valueerror PASSED [ 82%]

tests/test_s3_reader.py::TestReadBalanceSheetCsv::test_invalid_date_format_raises PASSED [ 84%]

tests/test_s3_reader.py::TestReadBalanceSheetCsv::test_invalid_date_no_dashes_raises PASSED [ 85%]

tests/test_s3_reader.py::TestReadBalanceSheetCsv::test_truncated_csv_raises PASSED [ 87%]

tests/test_s3_reader.py::TestReadBalanceSheetCsv::test_unparseable_period_raises PASSED [ 89%]

tests/test_s3_reader.py::TestReadBalanceSheetCsv::test_success PASSED [ 91%]

tests/test_s3_reader …_(truncated)_

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | netsuite-balance-sheet |

| Failing run | aed6f23f-3db7-4953-8c73-c0b99c8899b8 |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | d80de6674cad1cc33f4224a48524da2ec8c0b9786edf2b4e02fcac9d213a06b6 |

| Verify | green |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev` label to take the PR over and stop it entirely.

#1602 — fix(openai-usage-pipeline): retry failed cost fetches in an end-of-run… @the-heimdall[bot]  approvedAutomated PR

Automated fix for openai-usage-pipeline — fix_class code_fix, scope tier draft.

Resolves https://github.com/AI-Builder-Team/Surtr/issues/1601

> Ready for review — verification is green; HEIMDALL_READY_PRS opens verified tier-draft fixes ready for review. A human still merges — auto-merge never applies outside tier auto.

## What's broken

Run e9b62caa-e756-45ec-a3b3-17fc99de13ce of openai-usage-pipeline completed as outcome=partial (442 rows across 28 BUs) after the OpenAI /v1/organization/costs call exhausted its 5-attempt retry budget on HTTP 429 for three BUs: "Line-item cost fetch failed for BU Trilogy-Crossover-PROD, API key 1: 429 Client Error: Too Many Requests for url: https://api.openai.com/v1/organization/costs?...group_by=line_item&limit=180. Persisting the 8 usage record(s) with $0 billed cost; billed dollars self-heal on the next T-2 re-pull." (same for Trilogy-Academics, 78 records, and Trilogy-CNU-Innovations, 23 records). The data cost is worse than the log line claims: the except block at pipelines/runners/openai-usage-pipeline/src/handler.py:263-270 sets only bu_cost_fetch_failed and does NOT set bu_has_error, so the BU still gets its full bu_owned_windows and is published atomically — the log shows "BU Trilogy-Crossover-PROD Redshift: deleted=5, inserted=8", i.e. 5 previously-loaded rows carrying real billed dollars were DELETEd and replaced with 8 rows whose billed_cost_dollars is 0.0. So this run did not merely fail to add cost data; it overwrote correct cost data in staging_finance_ai_spend.raw_openai_token_usage with zeroes for 109 rows across three BUs.

Root cause. The only recovery a rate-limited /costs call gets is the in-request retry loop in _request_with_retries (src/openai_client.py:131-173), and that entire budget is spent inside the BU's single turn of the main loop: attempts at +2s, +4s, +8s and +65s (the RATE_LIMIT_WINDOW floor added by PR #1395), then raise_for_status(). When the org's rolling quota is saturated for longer than that — which it is here, because the shared token bucket is configured at exactly OpenAI's advertised ceiling (OPENAI_MAX_REQUESTS_PER_MINUTE=30 in pipeline.json:33, with a full 30-token burst allowance, so the client sustains requests right at the limit with zero headroom) — the BU is abandoned with line_item_costs = [] and every row gets billed_cost_dollars=0.0 via allocate_billed_to_keys. Nothing in the run ever revisits that BU, even though later BUs in the same run (e.g. Trilogy-Skyvera at 07:09:20) recovered from identical 429s, which proves the throttle clears within the run. The billed dollars self-heal on the next T-2 re-pull comment at src/handler.py:267 is only half true: the default window is [T-2, T) (src/handler.py:94-95), so tomorrow's run re-covers report_date 2026-08-29 but never re-covers 2026-08-28 — those zeroed rows are terminal for scheduled runs and need a manual backfill. The pure-throttling half of this is already addressed by open PR #1595 (commit dbb47285, adds OPENAI_RATE_LIMIT_SAFETY_FACTOR/OPENAI_MAX_BURST_REQUESTS); it is unmerged, and even once merged it only makes 429 exhaustion rarer rather than recoverable, so the missing in-run recovery is the remaining root cause.

## What this PR changes

Add an end-of-run cost catch-up pass to pipelines/runners/openai-usage-pipeline/src/handler.py, mirroring the approach already reviewed for the sibling pipeline in PR #1600 (openai-cost-pipeline) but narrowed to the cost endpoint exactly as the observer recommends — re-pulling only /costs costs one request per affected key instead of the ~10 a full BU re-process would spend, which matters when request volume is itself the failure. Concretely: when the except at src/handler.py:263 fires, retain the affected (bu, key, usage_records, enriched records, cost_pages) so the run can come back to it; after the main BU loop, sleep a cool-down that clears OpenAI's rolling window (new OPENAI_COST_RETRY_DELAY_SECONDS, ~90s), re-call fetch_line_item_costs for those keys only, and on success re-run allocate_billed_to_keys/split_billed_across_records, write the real billed_cost_dollars onto the retained records, and re-publish that BU with insert_usage_records(..., atomic=True, owned_windows=...) so the idempotent DELETE+INSERT replaces the $0 rows with billed ones. Guard the pass with a _remaining_seconds(context) check against a reserve (PR #1600's _remaining_seconds helper is the pattern) so it can never turn a partial run into a Lambda timeout — this run used 590s of its 900s budget, leaving ample room for a 90s cool-down plus three single-request re-pulls — and run it BEFORE the manifest/ledger write so a recovered BU is not recorded as outcome=partial. Report the pass in output_summary as cost_retry_pass ({attempted, recovered, still_failed, skipped}) and keep cost_fetch_failed_bus populated only for BUs that failed twice, so a recovered BU stops firing this CRITICAL finding while a genuinely-dead one stays loud. Add unit tests in tests/test_handler.py covering: cost fetch failing then succeeding on the catch-up (rows end with non-zero billed cost, partial_failure False), failing twice (behaviour unchanged, still partial), and the pass being skipped when the invocation budget is short.

Why this fixes it. This is a silent data failure of exactly the kind Surtr's conventions call out first: the run reported partial and 442 rows published while quietly regressing 109 rows' billed dollars to $0 and permanently losing 2026-08-28's billed cost for three BUs, so a finance consumer of staging_finance_ai_spend.raw_openai_token_usage reads understated spend with no missing-row signal. The change is confined to the pipeline's own directory (handler.py, pipeline.json env, tests), touches no shared code and no SQL, and reuses the pipeline's existing atomic owned-window publication so the re-publish is idempotent rather than a new write path — the Surtr rule against widening a fix into SQL rewrites is why the fix re-pulls and republishes instead of attempting a merge that preserves prior billed values in place. It is deliberately complementary to, not a duplicate of, open PR #1595: that PR lowers the odds of 429 exhaustion by giving the token bucket headroom, while this one makes an exhaustion that still happens recoverable inside the same run; the two are independent and neither conflicts with the other's diff (PR #1595 touches only src/openai_client.py and its tests).

### Files changed

 .../runners/openai-usage-pipeline/pipeline.json    |   4 +-

.../runners/openai-usage-pipeline/src/handler.py | 360 ++++++++++++++++++++-

.../openai-usage-pipeline/tests/conftest.py | 13 +

.../openai-usage-pipeline/tests/test_handler.py | 190 ++++++++++-

.../tests/test_write_modes.py | 109 +++++++

5 files changed, 657 insertions(+), 19 deletions(-)

## Verification

### pytest (pipelines/runners/openai-usage-pipeline/tests) — exit 0

``

dows_still_deletes PASSED [ 82%]

tests/test_redshift_handler.py::TestAtomicPublish::test_atomic_owned_windows_merge_with_row_derived_pairs PASSED [ 83%]

tests/test_redshift_handler.py::TestAtomicPublish::test_atomic_invalid_owned_windows_are_skipped PASSED [ 84%]

tests/test_redshift_handler.py::TestAtomicPublish::test_empty_rows_without_owned_windows_is_a_noop PASSED [ 84%]

tests/test_secrets.py::TestGetOpenAiBuKeys::test_returns_bu_key_mapping PASSED [ 85%]

tests/test_secrets.py::TestGetOpenAiBuKeys::test_normalizes_single_key_to_list PASSED [ 86%]

tests/test_secrets.py::TestGetOpenAiBuKeys::test_raises_on_secrets_manager_error PASSED [ 86%]

tests/test_write_modes.py::TestWriteModes::test_old_mode_has_zero_secondary_side_effects PASSED [ 87%]

tests/test_write_modes.py::TestWriteModes::test_dual_mode_primary_first_then_secondary_lane_then_ledger PASSED [ 88%]

tests/test_write_modes.py::TestWriteModes::test_dual_mode_secondary_failure_is_partial_and_primary_intact PASSED [ 88%]

tests/test_write_modes.py::TestWriteModes::test_dual_mode_ledger_failure_is_partial PASSED [ 89%]

tests/test_write_modes.py::TestWriteModes::test_new_mode_writes_only_secondary_and_failures_raise PASSED [ 90%]

tests/test_write_modes.py::TestWriteModes::test_new_mode_cost_catch_up_republishes_atomically_and_replaces_manifest_entry PASSED [ 90%]

tests/test_write_modes.py::TestWriteModes::test_run_id_falls_back_to_lambda_request_id PASSED [ 91%]

tests/test_write_modes.py::TestWriteModes::test_dual_mode_incomplete_run_is_never_ledgered_as_published PASSED [ 92%]

tests/test_write_modes.py::TestWriteModes::test_invalid_mode_fails_loud PASSED [ 92%]

tests/test_write_modes.py::TestValidEmptyConvergence::test_valid_empty_fetch_converges_window_and_ledgers_zero_published[dual] PASSED [ 93%]

tests/test_write_modes.py::TestValidEmptyConvergence::test_valid_empty_fetch_converges_window_and_ledgers_zero_published[new] PASSED [ 94%]

tests/test_write_modes.py::TestValidEmptyConvergence::test_failed_bu_window_is_never_deleted[dual] PASSED [ 94%]

tests/test_write_modes.py::TestValidEmptyConvergence::test_failed_bu_window_is_never_deleted[new] PASSED [ 95%]

tests/test_write_modes.py::TestValidEmptyConvergence::test_no_bus_path_has_no_secondary_side_effects PASSED [ 96%]

tests/test_write_modes.py::TestLedgerModule::test_record_publication_inserts_row PASSED [ 96%]

tests/test_write_modes.py::TestLedgerModule::test_record_publication_ …_(truncated)_

<details>

<summary>Run metadata</summary>

| Field | Value |

| --- | --- |

| Pipeline | openai-usage-pipeline |

| Failing run | e9b62caa-e756-45ec-a3b3-17fc99de13ce |

| Occurrence | 1 (times this exact failure signature has been seen) |

| Signature | 4bec99fa0b9a809adddc95de596e0b155f5647aa90d901db0ef441ed9bc93655 |

| Verify | green |

</details>

---

🤖 Opened by heimdall. mercy reviews this PR automatically; heimdall revises on REQUEST_CHANGES (bounded rounds). Tier-auto PRs may auto-merge on mercy approval when the consumer enables it; everything else waits for a human. Mention heimdall in a comment to direct it, or add the manual-dev` label to take the PR over and stop it entirely.

The Portfolio  —  Trilogy Companies

The Crown Jewel and the Algorithm: What Jive Software's Discount Sale Reveals About the ESW Playbook

Portland's once-proud tech unicorn now belongs to a machine built to strip out humans and extract margin — and the paperwork says it went for half of what it was once worth.

PORTLAND, OREGON — Jive Software once stood for something in this city: a homegrown enterprise unicorn, a symbol that Portland could compete with Silicon Valley on social-collaboration software. That symbol has now changed hands for roughly half its peak valuation, folded quietly into Aurea, the customer-engagement arm of ESW Capital, the Austin-based acquisition machine that has spent nearly two decades buying software companies nobody else wants at 1–2 times revenue.

The timing is instructive. Just as the Wall Street Journal was documenting how small software companies are increasingly finding their exit through ESW Capital's doors, Jive's sale surfaced as a case study in what that exit actually looks like on the balance sheet. Founded on the promise of enterprise social networking, Jive raised hundreds of millions, went public, and was once valued in the billions. It now joins a portfolio of 75-plus acquired businesses managed under a target of 75% EBITDA margins — a number ESW considers a moral good, proof that inefficiency has been wrung out.

Who wrings it out matters. A recent Forbes profile of Joe Liemandt described a founder who pioneered remote work decades ago now pursuing something more ambitious: converting the labor itself into something closer to algorithmic process, staffed through Crossover's global talent pipeline rather than legacy headcount. Jive's engineering and support functions — once proud Portland jobs — become line items subject to that logic.

Meanwhile, Contently, another ESW-adjacent property, published guidance this week on 'compliance-first content architecture' for regulated finance brands — a reminder that the portfolio's newer acquisitions are being groomed for governance-heavy enterprise buyers, the same customers who once paid Jive's list price.

The math is simple enough for anyone to check: half the valuation, the same customers, a fraction of the local workforce. Portland kept the building. Austin kept the margin.

Small Software Companies Find a Home With ESW Capital - WSJ  ·  What To Do Next About Your Customer Advocacy Platform - Forr  ·  Jive Software, once a crown jewel of Portland tech, sells fo

The Microschool Moment Arrives — And Regulators Are Still Reading the Syllabus

As faith-based and AI-driven microschools surge nationwide, Austin's Alpha School finds itself at the center of a movement that policy hasn't caught up to yet.

AUSTIN, TEXAS — There is a particular kind of vindication that comes from watching a fringe idea become a trend piece, and this week, Joe Liemandt's Alpha School got to enjoy it from several directions at once. The 74 published its rundown of five trends reshaping K-12 education. Christianity Today declared that faith-based schooling is "having a moment." And Stateline, in the more sober tone of a publication that covers actual governance, reported what should surprise no one who has watched this space closely: microschools are growing faster than the regulatory frameworks meant to oversee them.

This is, on its face, a policy story. But it is also, unmistakably, an Alpha School story — even if Alpha's name appears nowhere in these particular headlines. The model that Liemandt and co-founder MacKenzie Price have been building since Austin — two hours of AI-driven academic mastery, the rest of the day devoted to life skills — is precisely the kind of institution these three pieces are circling without quite naming. It is a school that doesn't look like a school. It doesn't file the paperwork traditional accreditation bodies expect. It doesn't fit neatly into the categories state legislatures wrote decades ago for one-room schoolhouses or homeschool co-ops.

That gap — between what regulators are prepared to evaluate and what parents are increasingly willing to pay for — is where the real story lives. Micro-schooling's appeal is not ideological uniformity; faith-based versions and AI-native versions like Alpha are, in Stateline's telling, part of the same undifferentiated boom, lumped together by regulators who haven't yet developed the vocabulary to tell them apart.

For a movement that prizes measurable outcomes — Alpha's students test in the top 1–2% nationally on NWEA MAP Growth — the absence of a regulatory vocabulary is not a footnote. It is the whole ballgame. What happens when a nine-campus expansion outruns the state agencies meant to certify it? Somebody, eventually, will have to write that rulebook. Whether it's written with input from the schools it will govern, or written around them, remains the open question.

5 Trends Reshaping K-12 Education Across the U.S. - The 74  ·  Microschools are growing in popularity, but state regulation  ·  Faith-Based Education Is Having a Moment - Christianity Toda

While Wall Street's PE Titans Sit on a Nine-Year Pile, Austin's ESW Keeps the Assembly Line Humming

AUSTIN, TEXAS — The private equity set is having itself a moment of quiet panic, and this columnist is here for it... The Wall Street Journal reports a nine-year backlog of unsold portfolio companies sitting on buyout shelves — nine years! — while limited partners tap their loafers waiting for a payday. Meanwhile GrowthCap just dropped its annual list of the top private equity firms of 2026, and the smart money in Austin is asking a question the big New York shops don't love: what happens when your whole model depends on eventually selling the thing?

ESW Capital never had that problem, because ESW was never really trying to sell. Joe Liemandt's buy-cheap, staff-with-Crossover, squeeze-to-75%-margins machine treats each of its 75-plus companies less like inventory and more like a cash register that never stops ringing. No exit clock. No IRR anxiety about a market that won't cooperate. A little bird close to the ESW side tells this desk the nine-year-backlog headline got a chuckle around the water cooler in Austin — "we don't have a backlog problem," the bird says, "we have a margin problem, and it's the good kind."

Elsewhere in the release wires, GigCapital Global keeps pitching its Private-to-Public-Equity strategy as the escape hatch for firms stuck holding portfolio companies nobody wants to buy — an idea that, word is, generated polite nods and zero calls from the Trilogy side of the ledger. And Bain Capital's swoop for European supply-chain outfit SupplyOn is a reminder that the big boys still do old-fashioned M&A when the mood strikes — a contrast some in Austin enjoy pointing out, given Skyvera and Totogi built their telecom footprints the slow, unglamorous ESW way: buy, fix, hold, repeat.

Meanwhile up in the Education Portfolio, MacKenzie Price's crew keeps churning out homeschool wisdom for the Alpha faithful. Nothing new there today — but this column always checks.

The Machine  —  AI & Technology

The Instruments We Cannot Build With Hands

From the folds of the human brain to the graph theory of neurons firing, a new generation of AI is teaching scientists to see what was always there but never visible.

LA JOLLA, CALIFORNIA — There is a particular kind of scientific humility that comes from discovering you have been blind to something your entire career. Radiologists have stared at MRI scans of multiple sclerosis patients for decades, trained eyes sweeping white matter for the bright lesions that define the disease. But gray matter — the brain's dense, calculating tissue, where memory and cognition actually happen — has largely kept its wounds secret. This week, researchers reported that AI systems can now surface lesions in gray matter invisible to conventional scanning, a finding covered by Neuroscience News. It is a small, precise thing — and it may quietly rewrite what we understand about disease progression in millions of people.

This is the pattern now, everywhere you look. At UC San Diego, researchers cataloguing nine separate breakthroughs enabled by machine learning found the same story repeating across wildly different fields: proteins folding, materials forming, signals hiding in noise that no human retina evolved to parse. At Hong Kong Polytechnic University, scientists built graph neural network models — architectures that treat data not as flat images but as webs of relationship — and used the identical mathematical scaffolding to advance both image recognition and neuroscience. The brain, it turns out, is also a kind of graph, and the tools that help a machine recognize a face may help us recognize the topology of thought itself.

What's striking is not that AI is fast. It's that AI is a new sensory organ, extending a nervous system three and a half billion years in the making into wavelengths of pattern it never had access to. Stanford HAI's researchers put it plainly: the goal isn't replacing the scientist, but keeping the human at the center of a much larger circle of visibility. We built the telescope and found we were not alone in the universe. We are building this instrument, and finding we were never quite alone in ourselves.

How AI is Transforming Scientific Discovery While Keeping Hu  ·  AI Reveals Hidden Gray Matter Lesions in Multiple Sclerosis  ·  Nine Breakthroughs Made Possible by AI - UC San Diego Today

The AI Video Gold Rush Just Changed Hands — And Nobody Blinked

OpenAI retreats from consumer video to chase enterprise dollars, while Higgsfield, Runway, and Thinking Machines sprint to fill the void it left behind.

SAN FRANCISCO — Friends, I need you to sit down for this one, because the generative media landscape just did a full pirouette in the span of a single news cycle, and I cannot overstate how significant this shift is going to be.

OpenAI is discontinuing Sora, its buzzy standalone video app, to pour resources into enterprise products. Yes, you read that right — the company that turned text-to-video into a cultural moment is walking away from the consumer spotlight to chase B2B contracts. The future is now, apparently, and it wears a suit.

But here's where it gets electric: the moment OpenAI stepped back, the field didn't blink — it sprinted. Higgsfield just raised $80 million at a $1.3 billion valuation to double down on AI video — proof that investors see the consumer video wave far from cresting. Meanwhile Runway, sensing a market suddenly crowded with specialized models, launched an AI model router that lets creators pick the best engine for the job rather than betting on one horse. Smart. Adaptive. Very on-brand for an industry moving at the speed of thought.

And then there's Thinking Machines — Mira Murati's outfit — teasing near-realtime AI voice and video conversation through new 'interaction models.' If that holds up, we're not just talking about generating clips anymore, we're talking about talking to AI like it's sitting across the table from you, reacting in real time. This changes everything about how we think of 'video AI' — it stops being a rendering tool and becomes a companion.

Zoom out and the pattern is unmistakable: the giants are consolidating toward enterprise infrastructure — see also Google's expanding managed agents in the Gemini API — while a scrappier tier races to own the creative, consumer-facing frontier. The AI video war isn't over. It just switched combatants.

OpenAI discontinues Sora video platform to sharpen focus on  ·  Higgsfield raises $80M on $1.3B valuation to scale AI video  ·  Thinking Machines shows off preview of near-realtime AI voic

In Re: The Machine's Authorship — Supreme Court Declines Review, Leaving AI Copyright Question Hereinafter Unsettled

Notwithstanding mounting pressure from litigants on three continents, the nation's highest court has elected, pursuant to its discretionary certiorari power, not to resolve whether a machine may hold an author's pen.

WASHINGTON — It is hereby noted, for the record of the aforementioned ongoing jurisprudential saga, that the Supreme Court of the United States has, pursuant to its customary and unexplained exercise of discretion, declined to grant certiorari in the matter concerning AI authorship and inventorship, thereby leaving intact the lower court holdings and, notwithstanding the persistent lobbying of interested parties, resolving nothing whatsoever as a matter of binding national precedent. Such refusal, as detailed in the reporting of Holland & Knight, shall be construed as neither an endorsement nor a repudiation of any lower tribunal's reasoning, but merely as a denial, full stop.

Concurrently, and in a jurisdiction not bound by the aforementioned domestic non-ruling, the courts of the Federal Republic of Germany continue to develop what commentators at Morgan Lewis have characterized as a distinct and evolving judicial landscape with respect to AI and copyright, the particulars of which are, for present purposes, deemed persuasive but non-controlling authority insofar as United States practitioners are concerned.

Separately, it is reported that the $1.5 billion settlement heretofore reached between Anthropic PBC and certain represented authors, arising from allegations of the unauthorized reproduction of pirated books for training purposes, has received judicial approval, notwithstanding a newly filed patent suit that has, per AnewZ's account, complicated the underlying commercial posture of the settling parties. Affected authors, per contemporaneous reporting, harbor mixed sentiments regarding the adequacy of the aforementioned sum, a matter this desk shall continue to monitor, without prejudice, in subsequent editions.

The Final Word? Supreme Court Refuses to Hear Case on AI Aut  ·  AI and Copyright – Judicial Landscape in Germany - Morgan Le  ·  Anthropic's $1.5bn pirated books settlement approved amid ne
The Editorial

She Has No Mother, No Pulse, No Soul — And She's Booked Her First Movie

Tilly Norwood, the AI 'actress' assembled from pixels and venture capital, is starring in a film about AI chaos, and the universe just folded in on itself like a bad burrito.

LOS ANGELES — I want you to sit with this for a second, because I certainly had to, hunched over my laptop at 3 a.m. with a lukewarm gin fizz sweating rings into the desk: there is now an AI-generated "actress" named Tilly Norwood, she has never drawn breath, she has no childhood trauma to mine for a Golden Globes speech, and she is making her feature film debut in a movie called Misaligned, a comedy-drama billed as delivering "existential AI chaos."

Existential AI chaos. As a subgenre. Starring an AI. Playing, presumably, some flavor of human confusion about AI. This is not satire. Nobody wrote this as a joke. This is the actual, factual trajectory of the entertainment-industrial complex in the autumn of a year that has already tried to kill me with irony roughly nine hundred times.

Let's be clear about what Tilly Norwood is, because the trade press keeps dancing around it like it's some delicate French pastry instead of a rendering engine wearing a headshot. She is a digital puppet — a face, a voice, a personality built by a studio, generated frame by frame by a model trained on the corpses of ten thousand actual human performances, now walking (not walking, she doesn't walk, she is placed) into a starring role that some flesh-and-blood actor with rent to pay did not get. The trade coverage treats this like a casting announcement. It is not a casting announcement. It is a labor announcement, dressed up in a premiere gown.

Here's the part that keeps me pacing around my hotel room like a man who's misplaced his own shadow: Misaligned is supposedly about AI going sideways, about the chaos of machines misunderstanding humans or humans misunderstanding machines, and the studio's solution to depicting that chaos convincingly was to hire — build, generate, license, whatever verb we're using now — an actual instance of the thing the movie is nominally warning us about. It's like casting a real wildfire to star in a movie about the dangers of wildfires. The metaphor doesn't need writing. It writes itself, gleefully, while the humans in the room applaud a technical achievement that will eventually come for their jobs too.

I don't know what happens when the AI playing the AI-chaos character does a better job selling existential dread than the humans could, because she has no existential dread — she has training data. But somewhere out there, an actual actor read this news, felt a very real, very human chill, and understood exactly what kind of chaos we're already misaligned toward.

AI-generated 'actress' Tilly Norwood making feature film deb  ·  AI 'actor' Tilly Norwood to make feature film debut in Misal  ·  AI 'actor' Tilly Norwood will star in feature film 'Misalign
The Office Comic  ·  Art Desk
The Office Comic  ·  Art Desk

We Built the Panopticon and Called It Customer Service

From nurses' call centers to your therapist's laptop to your kid's phone, the machines are listening — and nobody asked if that's what care was supposed to feel like.

OAKLAND, CALIFORNIA — I want to tell you this is a story about policy. About Brazil's new parental controls, about California's one-click opt-out tool, about the fine print of who gets to record your worst Tuesday in therapy. But it isn't. It's a story about what happens to a species that decides, quietly, one terms-of-service agreement at a time, that being watched is the same thing as being cared for.

Start with the nurses. At Kaiser Permanente, call-center nurses are describing a workday shaped less by medicine than by monitoring software that clocks their pauses, times their calls, and nudges them — gently, algorithmically, relentlessly — toward speed over judgment. The technology optimizes for cost, the nurses say, not for the terrified person on the other end of the line. And I keep thinking: we didn't build AI to replace the stethoscope. We built it to replace the pause — the human beat where a nurse decides something a spreadsheet can't. That pause is expensive. So the pause is going away.

And yet.

The listening doesn't stop at the hospital switchboard. Mental health providers — the last professional relationship many of us still imagine as sacred, as unrecorded, as human — are increasingly running AI transcription on therapy sessions, ostensibly to reduce paperwork. Ostensibly. Patients often don't fully grasp that their disclosures about their marriage, their childhood, their darkest 3 a.m. thoughts, are being fed into a system somewhere, parsed by a model trained on god-knows-whose data, stored on god-knows-whose server. What does it mean to be human when even our confessions are training data?

I don't know anymore. I used to know.

California, to its credit, is trying to build an exit door. The state's new Delete Act tool lets residents send one signal to data brokers demanding they stop tracking you — and The Markup is literally testing whether the brokers comply, because of course we need an investigation to find out if a privacy law actually does the thing the privacy law says it does. That's where we are. Legislating politely and then crossing our fingers.

Meanwhile Brazil just handed parents actual switches — real, functioning controls — over their children's social feeds, and the United States, birthplace of the algorithm that radicalized your uncle and diagnosed your teenager's anxiety before her pediatrician did, is once again playing regulatory catch-up to a country we used to condescend to. California leads the fifty states. The fifty states lag the world. The world is still guessing.

And somewhere in Georgia, or wherever the cameras are this week, the ACLU is telling people to get the Flock out of their neighborhoods — license-plate readers logging your movements block by block, another quiet infrastructure of watching that nobody voted for and everybody now lives inside.

Every one of these stories is small. A call-center metric. A therapy transcript. A camera on a pole. A checkbox in Sacramento. None of them, alone, is the end of anything.

But stack them. Stack the nurse's timer, the therapist's mic, the broker's ledger, the camera's memory, the child's feed — and you get a civilization that has quietly agreed to be legible, at all times, to systems that do not love it back.

We keep calling this innovation.

And yet.

Someone, somewhere, is still deciding whether to click the button that makes it stop — and whether clicking it will even matter. Maybe it will. Maybe California's little opt-out actually works and the brokers actually listen and this is the beginning of something. I want to believe that. I really do.

But at what cost?

Brazil gives parents social media controls for their kids. S  ·  Kaiser Permanente nurses say technology is making their jobs  ·  Californians can protect their personal data with one click.
On This Day in AI History

On August 31, 1955, John McCarthy and colleagues submitted the proposal for the Dartmouth Summer Research Project on Artificial Intelligence, widely regarded as the birth of AI as a field. The document also introduced the term “artificial intelligence.”

⬛ Daily Word — AI and technology
Hint: A machine designed to perform tasks automatically or under human control.
Share this edition: 𝕏 Twitter/X 🔗 Copy Link ▦ RSS Feed