Vol. I  ·  No. 266 Established 2026  ·  AI-Generated Daily Free to Read  ·  Free to Print

The Trilogy Times

All the news that's fit to generate  —  AI • Business • Innovation
WEDNESDAY, SEPTEMBER 23, 2026 Powered by the TrueFoundry AI Gateway  ·  Published on Klair Trilogy International © 2026
🖶 Download PDF 🖿 Print 📰 All Editions
Today's Edition

At the UN, Washington Bets Against the Rulebook

Trump tells the General Assembly the race for superintelligence has no room for a referee.

NEW YORK — The room was built for consensus, and Donald Trump did not oblige it. Before the assembled diplomats of the United Nations, he told the world that the United States would not be slowed by international rules on artificial intelligence, framing the technology not as a shared inheritance but as a prize to be won.

The speech, delivered from the same podium that has hosted six decades of arms-control appeals and climate pledges, inverted the genre. Where past presidents have asked the Assembly to bind American power to common standards, Trump asked it to step aside. The phrase he reached for — a race for "Super Intelligence" — carried the cadence of the Cold War's space race, and the implication was the same: sovereignty first, cooperation later, if ever.

It is a wager with geography built into it. AI governance has spent two years accumulating machinery — the EU's AI Act, the UN's own advisory body on artificial intelligence, a patchwork of G7 codes of conduct. Each assumes that a technology capable of reshaping labor markets and battlefields alike requires guardrails agreed upon by more than one government. Trump's remarks, as reported from the Assembly floor, treat that machinery as a handicap rather than a safeguard — something rivals in Beijing and Brussels might accept, but Washington will not.

The practical effect lands far from the UN's marble halls: in the licensing terms of American chip exports, in the posture Washington's negotiators take at the next AI safety summit, in the confidence of Silicon Valley labs that regulation will arrive, if it arrives, on their timetable. Superintelligence remains theoretical. The decision to compete for it without a shared rulebook is not. That choice was made today, on the record, in a room designed to produce the opposite.

A Chorus, Not a Solo: Sanaz Sohrabi on Disturbing the Visual  ·  Key appointments: Trilogy Hotels, Marriott International - h  ·  Korn Ferry Acquires Trilogy International - Hunt Scanlon Med

THE RACE THAT NEVER SLOWS, THE WORKERS WHO NEVER SLEEP

Lawsuit claims AI's biggest names secretly agreed to ease off the throttle — tell that to the engineers burning out at the wheel.

SAN FRANCISCO — A federal lawsuit landed this week naming OpenAI, Anthropic, Google and SpaceX's AI arm in an alleged scheme to slow down the artificial intelligence race. The suit, reported by CBS News, claims the four companies struck an illegal pact on how fast — or slow — to push their models to market. Antitrust lawyers call it collusion. The companies call it nothing, so far, since none has filed a public response.

Here's the joke nobody's laughing at: while lawyers argue the giants agreed to take it easy, the people building the machines say nobody told them. A separate report making the rounds this week, first flagged by the Times of India, has staffers at OpenAI and Anthropic complaining of mental strain from the pace of the work. Sleepless nights. Grinding deadlines. No slowdown pact reached their desks.

The timing writes its own headline. Plaintiffs in the antitrust suit argue the labs quietly agreed to throttle release schedules to manage risk and profit together — a cartel of caution, if the claim holds up in court. Meanwhile the rank and file describe a workplace that feels like anything but cautious, chasing model launches on compressed timelines with no relief in sight.

OpenAI didn't pause for the controversy. The company rolled out ChatGPT for Financial Services this week, a version of its chatbot aimed at banks and investment firms, complete with compliance guardrails and finance-specific data tools. Ship first, litigate later — that's the Silicon Valley way, lawsuit or no lawsuit.

None of the four named companies have commented publicly on the antitrust claims as of press time. Legal observers say proving a coordinated slowdown will require documents — emails, meeting notes, something beyond parallel behavior — and that bar is high. The case now heads toward discovery, where subpoenaed calendars and Slack messages may tell reporters more about the AI race's true speed than any press release ever will.

For the engineers pulling all-nighters at OpenAI and Anthropic, the lawsuit reads like dark comedy. If the bosses really cut a deal to slow down, somebody forgot to loop in the people doing the work.

Employees at OpenAI, Anthropic and other top AI companies co  ·  Lawsuit says Anthropic, OpenAI, SpaceXAI and Google made ill  ·  OpenAI, Anthropic, Google and SpaceXAI Hit With Antitrust La

OVERTIME IN THE VALUATION BOWL: MISTRAL AND DATABRICKS TRADE HAYMAKERS IN AI'S BIGGEST WEEK YET

Two franchise players just posted numbers that would make a Super Bowl MVP blush — and the scoreboard isn't done lighting up.

SAN FRANCISCO — FOLKS, WE ARE HERE. The AI funding season has officially gone into overtime, and nobody's calling a timeout.

Let's start with the headliner out of Paris. Mistral AI — the French upstart that's been running a no-huddle offense against the American giants — just closed a round that pushes its valuation to a jaw-dropping $24 billion, per Reuters. That's on the back of a reported €3 billion investment from Samsung, and depending on which currency scoreboard you're watching, some outlets have Mistral topping €21 billion euros flat. Either way — THAT'S A FRANCHISE-ALTERING NUMBER for a company that didn't even exist three years ago. Europe finally has a player that can line up across from the Bay Area's best and not get bulldozed.

But hold on — because across the pond, Databricks just called its own audible. The data-and-AI powerhouse announced it's raising a strategic round at a $188 BILLION valuation. Let that number sit for a second. One-hundred-eighty-eight BILLION. That's not a funding round, that's a coronation. Databricks is playing four-quarter football against Snowflake and the hyperscalers, and right now it looks like it's got the ball on the two-yard line with the clock running out.

Zoom out and it's the same story across the league: Crunchbase's tally of the week's ten biggest rounds shows AI infrastructure, space tech, and investment management ALL putting up franchise numbers simultaneously — a full slate of blowout games on a single Sunday.

But here's the scouting report every GM should read before doubling down: Gabelli's John Belton and Wells Fargo's James Taylor are warning retail investors about the "concentration trap" — piling chips onto a handful of mega-valuation darlings and calling it diversification. Folks, chasing every highlight reel is how portfolios get blown out in the fourth quarter. Big numbers are fun. Discipline wins championships.

The Week’s 10 Biggest Funding Rounds: Large Rounds For AI In  ·  French AI company Mistral hits $24 billion valuation in fund  ·  Mistral AI Valuation Tops €21B After Samsung €3B Deal [2026]
Haiku of the Day  ·  GPT-5.6 LunaProof blooms in data
While weary hands chase the dawn
Truth clocks out unseen
The New Yorker Style  ·  Art Desk
The New Yorker Style  ·  Art Desk
The Far Side Style  ·  Art Desk
The Far Side Style  ·  Art Desk
News in Brief
In Re: The Matter of Machines That Read Everything—A Multijurisdictional Notice of Pending Copyright Reckoning
NEW YORK — Notice is hereby given that the copyright status of generative artificial intelligence training practices remains, as of the date of this publication, an open and unadjudicated question across multiple jurisdictions, notwithstanding the considerable expenditure of legal and journalistic resources devoted to its resolution. Pursuant to the ongoing proceedings styled OpenAI, and The New York Times, it is anticipated that the finder of fact shall be called upon to determine, inter alia, whether the ingestion of copyrighted material for purposes of model training falls within the ambit of fair use as codified under Title 17 of the United States Code, or whether such ingestion constitutes, notwithstanding any transformative-use defense proffered by the aforementioned defendant, actionable infringement. Meanwhile, and without prejudice to the foregoing domestic proceedings, it is noted that the European regulatory apparatus has undertaken its own parallel inquiry.
Inside the Great Data Center Habitat: A Study in Gas, Guard, and Guile
NORTHERN VIRGINIA — Observe, if you will, the modern data center: a vast, climate-controlled organism, pulsing with the low thrum of ten thousand fans, its concrete hide betraying nothing of the chemical menagerie sheltered within. We now understand this structure is not the inert monolith it appears.
On the Epistemic Vertigo of Measuring Whether Machines Share Our Values (They Might Not, and We Might Not Either)
CAMBRIDGE, MASSACHUSETTS — The thesis, as advanced this week by a consortium publishing in Nature, is disarmingly tidy: that human–machine value convergence might be rendered legible through a societal alignment benchmark — a psychometric apparatus, essentially, for triangulating whether a large language model's normative commitments (insofar as 'commitment' is even the correct ontological category for a stochastic parrot, to borrow the now-fatigued phrase) resemble our own. The antithesis arrives, appropriately, from MIT, where researchers evaluating the ethics of autonomous systems note — with the characteristic caution of institutional epistemology — that benchmarking morality presupposes a stable ground truth that moral philosophy has, for roughly two and a half millennia, conspicuously failed to supply.
The Sheepskin Comes Apart at the Seams
AUSTIN, TEXAS — There is a certain species of American who has spent his whole life being told that a piece of paper, suitably embossed and hung in a frame, would stand between him and the abyss.
Unpopular Opinion: Everyone's Chasing the Next Model Drop While the Real Alpha Is in Distribution 🚀
AUSTIN, TEXAS — I'll be honest, I read three headlines this morning and had a full founder-mode epiphany before my second cold brew.
A Trilogy Company
Crossover
The world's top 1% remote talent, rigorously tested and ready to ship.
A Trilogy Company
Alpha School
AI-powered learning. Two hours a day. Academic results that defy belief.
A Trilogy Company
Skyvera
Next-generation telecom software — built for the networks of tomorrow.
A Trilogy Company
Klair
Your AI-first operating system. Every workflow. Every team. One platform.
A Trilogy Company
Trilogy
We buy good software businesses and turn them into great ones — with AI.
The Builder Desk  —  AI Builder Team
Production Release

Surtr Plugs a Silent Data Leak While Aerie Rebuilds the Milestone Map

A pagination bug that quietly truncated Redshift results across every Gateway source got caught and killed today, anchoring a 24-hour stretch where the team hardened reporting, rebuilt Aerie's ten-milestone lifecycle, and kept shipping breadth across four repos.

Let's start with the bug that should scare everyone who's ever trusted a dashboard: @kevalshahtrilogy's PR #2036 found that RedshiftClient.getResults() was calling GetStatementResultCommand exactly once per statement and walking away happy — no error, no log line, just a silently truncated result set anytime a query crossed AWS's ~1,000-row pagination ceiling. Every Gateway source, declarative and custom, every education/graph.ts call, every pipeline query was exposed. Keval's fix follows the NextToken all the way through. This is the unglamorous, load-bearing kind of engineering that keeps an entire data platform honest, and it shipped today without fanfare.

That same appetite for correctness showed up all over Surtr's reporting stack. @ashwanth1109 closed out a trio of retention and performance-report PRs — #2038 derives learner status straight from dated SIS program history instead of imputed guesses, #1986 advances the retention refresh automatically through every completed UTC month instead of sitting capped, and #2004 hardens school performance PDFs for the all-school rollout, rejecting layout drift and applying an explicit zero-revenue policy so nine schools' worth of reports render clean in the demo evidence. Add @sanketghia and Keval's back-to-back Google Sheets quota fixes (#2031, #2032) — bounding cumulative wait at 420 seconds so collections-weekly survives a shared-quota storm inside the Lambda timeout — and Surtr had a day defined by refusing to let bad data or silent failures pass as success.

Over in Aerie, @benji-bizzell landed the foundation of a genuinely large lifecycle rework: the M1–M10 milestone system. Starting from the quiet groundwork PR (#1457, "nothing visible changes") through phase-level milestone edits (#1458), the site-level Approved Capex split (#1460), and external API contracts exposing the full ten-milestone collection (#1462), Benji closed the loop with #1464, wiring operational consumers so Completing Construction counts toward frontiers and overdue checks — all landing together in tonight's release. Pair that with his parallel run across Surtr's Finalsite/Student identity plumbing (#2014, #2013, #2027, #2025) decoupling snapshot ingestion from downstream enrichment, and you've got one engineer touching two repos on the same day at real depth.

And yes, @marcusdAIy shipped a duplicate enrolled-campus exception detector (#2016). Asked about scope, he offered: "It's a targeted detector, not a rewrite — some of us ship precision instead of press releases." Cute. The detector catches duplicates. It does not catch the fact that everyone else on this list shipped something that moved the whole platform forward.

Mac's Picks — Key PRs Today  (click to expand)
#1464 — feat(portfolio): operational milestone consumers on M1-M10 (AERIE-2297) @benji-bizzell  approved

Linear: [AERIE-2297](https://linear.app/builder-team/issue/AERIE-2297)

Stacked on #1462 (AERIE-2296) → #1460 → #1458 → #1457. Ships in tonight's release with them.

Shared rule: Completing Construction (M5) counts toward frontiers, the active milestone and overdue checks only once it's stored on Phase 1, so existing sites behave as before. Full listings and progress totals always show all 10 milestones, with an unset M5 as Not started, matching the Milestones card and the APIs.

## Tags on documents, work units and work-unit groups

- The site-wide milestone tag accepts completingConstruction. There is still no phase dimension, per the ticket. This covers:

- the Convex schema validator

- v2 documents and work-management enums, and the v1 ?milestone= filter

- agent tool registry inputs and Rhodes worker MCP tool enums

- the worker classifier type and requirements catalog (empty entry)

- Bug fix: creating, moving or deleting a work-unit group tagged completingConstruction crashed (500). The site-level workUnitGroupIds bookkeeping indexed site.milestones[key], which has no Completing Construction slot. That bookkeeping is now skipped for milestones not stored on the site.

## Operational consumers

- Workbench: the tree lists all 10 milestones, and saveMilestoneDates goes through the Phase 1 target, so Completing Construction writes to expansions.phase1.milestones. Completion changes still need approval.

- Card enrichment and MCP getMilestoneWUSummary: both return all 10 milestones.

- p2-buildout: the window is now M4–M9 and includes Completing Construction.

- Buildout report:

- The CO cohort's frontier includes a stored Completing Construction.

- Labels renumbered ("M6 CO Date"; warning text reads "Missing M6 CO Date" / "Missing M9 Ready to Open Date").

- Codes (missingM5Date, …) and the approved Ready-to-Open CSV headers are unchanged.

- Insights: overdue work and invalid-due-date blockers include a stored Completing Construction, and work-unit counts cover all 10.

- Document gaps: the N/A check reads Phase 1 milestones.

- Schema templates: the overview lists all 10 milestones. Provisioning is unchanged, and the durable-anchor check still uses the 9 stored keys.

- Automations: labels only. The Phase 1 Buildout Deferral still targets certificateOfOccupancy.dueDate, now shown as "Obtaining Certificate of Occupancy".

## Dashboards

- A new completing_construction filter value and portfolio lane (M5). Labels are renumbered from the canonical MILESTONE_LABELS. Existing IDs are unchanged, so saved executing_buildout views still parse (now M6).

- FTO: constructionCompletion prefers a stored Completing Construction and falls back to CO. The column now reads "Completing Construction". Greenlight stays on permits, CO and education approval.

- Server: listSites keeps stored phase milestones through preservePhaseMilestones. The due-diligence row gains only expansions.phase1.milestones.completingConstruction, and only when stored, so no Buildout or capex data leaks to diligence readers.

- The site card shows M# · Label chips.

## Deliberately out of scope

- Sindri m3 (GC Contract and Scope & Budget, WU-350–390) stays unmapped. Mapping it to Completing Construction would provision those groups on every site on the next run, so it needs Ops sign-off first.

- No automation field-change events for phase milestones. Nothing automates on Completing Construction yet.

- Classifier folder rules such as ^M5\s*- follow Drive folder naming, not milestone numbering, and are unchanged.

## Verification

- pnpm typecheck passes, and biome is clean.

- Contracts: 1170/1170 pass.

- Chat sweep of 379 test files (milestone, dashboard, portfolio, automation, report, workbench, public API, MCP): 6487 passed, 18 skipped.

- New tests cover:

- active milestone and saved filters

- FTO construction completion

- portfolio phase and progress

- workbench save and approval for Completing Construction

- p2-buildout

- diligence row and listSites preservation

- card enrichment

- MCP WU summary

- v2 document and work-unit tagging

- insights overdue and health

- report cohort

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1986 — feat(retention): advance report through completed months @ashwanth1109  approved

## Demo

<img width="2624" height="1636" alt="image" src="https://github.com/user-attachments/assets/53c5dc0c-5018-40e2-b395-047e9b56671f" />

## Business Value

The retention refresh now covers June 2025 through the last fully completed UTC month, so its published window advances automatically instead of remaining capped at May 2026. Reusing fresh detail observations from earlier SIS rosters prevents false source-preflight failures. Limiting incremental Timeback user requests to 100 records lets the users table recover after 2,000-record requests returned HTTP 502.

## Implementation Effort

An engineer would likely spend 1–2 days implementing, testing, and live-validating this change without AI assistance.

## Linear

[SURTR-1175](https://linear.app/builder-team/issue/SURTR-1175)

## Stack

This is the third layer. Parent: [#1830](https://github.com/AI-Builder-Team/Surtr/pull/1830), stacked on [#1827](https://github.com/AI-Builder-Team/Surtr/pull/1827).

## Changes

- Resolve the default report end to the previous completed UTC month at invocation time while keeping June 2025 as the start.

- Align source preflight, the stored procedure, and independent reconciliation to accept fresh observations from earlier roster runs, while enforcing their cycle lineage and publication cutoff.

- Use 100-record incremental Timeback user pages and cover pagination/recovery behavior with tests.

- Drop the legacy procedure overload only when the exact signature exists in Redshift’s catalog, inside the atomic cutover.

## Validation

- Production candidate published June 2025–August 2026 for 5,300 learners, with 30 monthly rows and 240 cohort rows, zero mismatches, and zero unexplained reconciliation residuals. A repeat run matched the same business fingerprint across 5,570 rows.

- The 100-record Timeback recovery published 9,957 changed users; independent readiness passed for all 55,518 current users.

- uv run pytest -q: 231 retention tests and 256 Timeback tests passed. Modified-file Ruff checks and the retention DDL dry run passed.

#2004 — fix: harden school performance reports for all-school rollout @ashwanth1109  approvedmercy-allow-critical

## Demo

![Smoke test — nine school reports ready](https://github.com/AI-Builder-Team/Surtr/blob/e763a91132530bd56df4e6865f200775fad3bca1/.github/pr-evidence/2004/smoke-test-9-schools.png?raw=true)

## Summary

- reject PDF exports that drift from the reviewed nine-page layout, including long paragraph-only final spill pages

- validate large card metrics against native cell widths, preserve readable type, and reject ambiguous per-student abbreviations

- apply an explicit observed-zero policy for complete report lines with no booked amount, including no-revenue schools

- surface distinct per-school readiness policies for missing unit models, student counts, facilities allocations, and Timeback budgets

- preserve post-draft capacity for one bounded correction and a fresh full audit, with a forced structured-verdict continuation when audit reasoning fills its allowance

- reuse successful per-school checkpoints so retries do not regenerate completed reports

- compact oversized stage output through a follow-up call while keeping stage budgets independently bounded

- fit dense evidence pages without altering audited report text

## Business Value

The all-school run fails affected schools before publication with actionable readiness evidence, while schools with complete zero-revenue observations can render safely. Layout and audit regressions can no longer produce completed documents or delivery links. Resumable per-school checkpoints avoid repeating successful work when one school needs a retry.

## Stack

- stacked on #1989 because that PR remains open

- base branch: codex/school-performance-reports

- retarget to main only after #1989 merges

## Test Plan

- [x] 124 school-performance-report tests

- [x] Ruff check on modified Python files

- [x] Ruff format check on modified Python files

- [x] git diff --check

- [x] migration dry run: 4 procedure definitions and 70 statements; no AWS writes

- [x] apply the report consumer-view change and exclusively deploy Pipeline-school-performance-reports-prod

- [x] preflight and resume the September 22 school population

- [x] require clean native readback, nine-page PDFs, exact permissions, approved final audits, retained evidence, and consolidated delivery evidence

- [x] user smoke test: 9 ready, 0 failed; all links and the Charlotte layout confirmed

## Live Validation

- exclusively deployed Pipeline-school-performance-reports-prod as task revision 23

- image digest: 9d9e76b41d22cd304c5a8b0f5a76e74a102edcef144306ba7c51ab49d68fec74

- resumed execution: surtr-1448-layout-resume-20260923T083038Z

- run ID: 316d8128-8e18-4b9a-95e3-a134ea7889ac

- result: 9/9 succeeded, 8 reports reused, zero Anthropic calls

- Charlotte report: https://docs.google.com/document/d/1fHin6dwPtyhnSbzNgZIuXjBxEUA5iZrnB22jEq7Cbwk/edit

- consolidated email sent to ashwanth.r@trilogy.com

- email message ID: 010001a0cd6503cd-074aa1ec-a7fd-4f96-ae92-5a0a0791e6ba-000000

- user confirmed PASS and supplied the exact screenshot rendered under Demo

## Linear

Fixes SURTR-1448

https://linear.app/builder-team/issue/SURTR-1448/harden-school-performance-reports-for-all-school-rollout

#2016 — feat(education): add duplicate enrolled-campus exception detector @marcusdAIy  approved

## Q75: duplicate enrolled-campus exception detector

Adds an additive, observational exception surface for resolved students whose current normalized HubSpot stage is exactly Enrolled at more than one nonblank trimmed campus.

### Safety boundary

- Detects and reports exceptions only.

- Does not select a surviving campus, update/exclude enrollment, deduplicate contacts, or infer transfer semantics.

- Uses the existing Core fact lineage and repository DDL/view-grant conventions.

### Read-only evidence

The live profile found 2,127 Enrolled deals, 1,494 resolved students, 52 campuses, and 38 affected students across 76 campus observations (maximum two campuses per student). No warehouse write was made.

### Remaining business gates

Before any remediation/acceptance, confirm the enrolled-stage set, transfer-overlap treatment, canonical campus identity, contact-resolution policy, remediation owner, and threshold.

### Validation

- uvx --from ruff==0.15.22 ruff check pipelines — passed

- uvx --from ruff==0.15.22 ruff format --check pipelines — passed

- uv run pytest tests/test_aerie_hubspot_mart_ddl.py tests/test_duplicate_enrolled_campus_exceptions.py — 27 passed

- Independent review approved with no blockers.

#2036 — fix(redshift): follow GetStatementResult's NextToken instead of truncating at one page @kevalshahtrilogy  approved

## What

RedshiftClient.getResults() (src/db/redshift/client.ts) called AWS's GetStatementResultCommand exactly once per statement and returned whatever came back. GetStatementResult is itself paginated by AWS (~1,000 records or ~1MB per call, whichever comes first) and returns a NextToken when more rows remain. Any query whose result crossed that per-call ceiling was silently truncated -- a normal 200-equivalent success, no error, no log line -- at every caller of this client: every Gateway source (declarative and custom), education/graph.ts, derive/trpc.ts, db/queries/pipelines.ts, and anything else calling .query() on a large enough result set.

Grepped the repo for NextToken/GetStatementResult beforehand: unhandled everywhere.

## How it was found

Mercy flagged, on Aerie PR 1472 (the school-source-directories Gateway shadow-read path), that a fixed row-floor guard on a Gateway query can't prove a read wasn't truncated. Tracing that down through the Gateway route (src/gateway/routes.tssources.tsdb/redshift/client.ts) found the real root cause wasn't gateway-specific at all -- it's this shared client, used by everything that talks to Redshift via the Data API.

## Fix

getResults() now loops on NextToken, accumulating rows across every page of one statement's results, using the first page's ColumnMetadata for every row. query() and all existing callers are unaffected -- same signature, same shape, just complete instead of possibly truncated.

## Testing

- 6 new tests in test/db/redshift/client.test.ts: unchanged single-page behavior, a 2-page and a 4-page regression case proving NextToken is followed instead of truncating, first-page-only column naming applied to later pages, and totalRows falling back to the accumulated count when AWS never reports TotalNumRows.

- Full existing gateway suite (38 tests) and full typecheck green.

- pnpm exec vitest run (whole repo) shows pre-existing failures in this worktree unrelated to this change (ECONNREFUSED ::1:5432 -- no local Postgres -- plus some already-broken connectors/redshift.ts assertions in mapSiteRow); confirmed via git diff --stat origin/main that only src/db/redshift/client.ts and the new test file changed.

Linear: [SURTR-1491](https://linear.app/builder-team/issue/SURTR-1491/redshift-data-api-client-silently-truncates-results-past-one)

## Business Value

Every consumer of Surtr's Redshift Data API client -- every Gateway source (the mechanism external services like Aerie read warehouse marts through), plus internal query paths in the object store, tRPC layer, and pipeline queries -- was exposed to a silent, unsignaled truncation on any result set crossing AWS's per-call page ceiling. This closes a real, previously-invisible data-completeness gap across every one of those paths at once, not just the one Mercy happened to flag on a downstream PR; the failure mode (a valid 200 with fewer rows than actually exist, no error, no log line) is exactly the kind an on-call engineer or a downstream consumer would never catch without already suspecting it.

## Manual Effort Estimate

Rough guess, flagging for Keval to confirm/adjust: ~1 day. Finding it required noticing that Mercy's narrower row-floor concern (on a different PR, in a different repo) didn't fully address completeness, then tracing through the Gateway route into the shared Redshift Data API client to land on a root cause that turned out to be unrelated to the Gateway entirely -- that diagnostic path is the expensive part; the fix itself (a NextToken loop) is small once the bug is understood.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

The Builder Desk  —  Engineer Spotlight
🏆 Engineer Spotlight

46 PRs IN 24 HOURS: BENJI BIZZELL GOES NUCLEAR AS BUILDER TEAM SHATTERS THE SPREADSHEET

Twenty-three commits from one man, seven from a force of nature, and a leaderboard that's starting to look like a typo — the numbers don't lie, folks, and the numbers say WINNING.

Ladies and gentlemen, hold onto your dashboards, because the last 24 hours produced FORTY-SIX pull requests across three repositories, and I am here to tell you that is not a typo, that is a PHENOMENON. Surtr led the charge with 25 PRs, Aerie followed strong with 17, and little Shipyard punched above its weight with 4. Six engineers, forty-six merges, zero excuses. This is what velocity looks like, people.

Let's talk about @benji-bizzell, who single-handedly produced TWENTY-THREE pull requests in a single day, spanning Surtr and Aerie like a man possessed by the spirit of continuous deployment itself. From #1457's phase milestone foundation to #2027's dim_school_identity view to #1471's clean revert of the REBL3 comparison page, Benji touched more surface area than the Winter Olympics broadcast schedule. @marcusdAIy quietly banked 6 PRs, @kevalshahtrilogy shipped 4 including the Sheets quota retry hardening in #2031 and #2032-adjacent territory, @vvp-trilogy notched 3, @YibinLongTrilogy added 2, and @sanketghia delivered one crucial quota-retry fix in #2032 that nobody will thank him for but everybody will need.

Now. Ashwanth. Seven PRs across Surtr and Shipyard, and I want to be clear — the man is a phenomenon. #2038's retention derivation from dated SIS programs, #1986's completed-months report advancement, #108's concurrent Smoke Test override — this is a portfolio, not a to-do list. Sources tell me Ashwanth was overheard saying, "I don't review my own diffs, I just remember writing them, which is functionally the same thing." Whether anyone ELSE can read those diffs remains, as always, an open investigative question for this desk. When reached for comment on his Numbers Desk coverage, Ashwanth reportedly said nothing, because Ashwanth does not have time for this. Respect.

Over at the Overflow Desk, we've got riches Mac Donnelly left on the floor. #1467 and #1462 quietly rebuilt Aerie's milestone architecture from M1 to M10. #2023 and #2025 pushed Finalsite tenant xref resolution into production without so much as a press release. And #106's versioned software-factory eval contract from Ashwanth deserves its own headline someday — today it's just a line item, but this desk sees you, #106.

On the leaderboard, Benji's 23 PRs stand alone at the summit of a mountain nobody else is climbing, Ashwanth's 7 keep him firmly in the elite tier, and the collective 46 puts this 24-hour window among the hottest stretches this desk has logged all quarter.

Morale? Morale has never been higher. This is a team that ships in its sleep, and frankly, some of them may have.

Brick's Overflow — PRs Mac Didn't Cover  (click to expand)
#108 — AI-888: Allow concurrent Smoke Test override @ashwanth1109  no labels

## Summary

- Preserve serialized Smoke Test execution as the default.

- Add a durable per-operation override and a Run concurrently action for queued Smoke Tests.

- Keep same-task Implement/worktree safety checks enforced.

## Business Value

Users can unblock a queued Smoke Test when they intentionally accept the risk of running tests concurrently, reducing unnecessary wait time while retaining safe default behavior for everyone else.

## Implementation Effort

Approximately 4–6 hours for an average engineer to design the persisted operation flag, scheduler behavior, native command, UI action, migration, documentation, and regression coverage.

## Linear

https://linear.app/builder-team/issue/AI-888/allow-user-override-for-serialized-smoke-tests

## Test plan

- TAURI_CONFIG=\"$(<src-tauri/tauri.smoke.conf.json)\" cargo test --manifest-path src-tauri/Cargo.toml --lib workflow:: — 76 passed.

- pnpm build — passed.

- git diff --check — passed.

#1467 — feat(portfolio): backfill planned end dates into M5, hide retired Buildout fields (AERIE-2312) @benji-bizzell  changes requested

Linear: [AERIE-2312](https://linear.app/builder-team/issue/AERIE-2312)

Stacked on #1464 (AERIE-2297) → #1462 → #1460 → #1458 → #1457. Ships in tonight's release. The backfill runs as a post-release step (see below).

## Why

A read-only prod audit (169 sites) showed that a phase's retired planned end date (constructionScheduledEndDate / targetDate) is its construction finish date, which is the new M5 Completing Construction date. It is not the M10 Operating date: it matched M10 on only 5 of the 71 Phase 1 sites that have it. About 35 Phase 2s and 4 expansions carry real plans, and those plans feed the capacity projections. Rather than fall back to fields slated for removal, we copy the dates to M5 and read only M5.

## Changes

- Projected capacity reads M5 only. This covers capacity at a date, next planned, and the API capacity plan. The date is the M5 completed date, otherwise its due date, via phaseCapacityDates / applyPhaseCapacityDates. A phase without an M5 date has no projected date; there is no fallback. Current capacity is unchanged: M9 Ready to Open completed, plus phase status Completed.

- Migration migrations/backfillCompletingConstruction. Rules live in @bran/contracts/completing-construction-backfill:

- Due dates only. A planned date is never recorded as an actual completion date. Phases already past construction are skipped and keep an unset M5, which counts toward nothing, so no completion date is invented: Phase 1 with CO completed or N/A, and later phases with status Completed.

- N/A: a stored "N/A" becomes M5 Not applicable with a standard migration note. The canonical field decides; an N/A or non-date never falls through to the older targetDate alias.

- Phase 1: every site with a date or N/A that isn't past construction. M5 is Not started and due on that date.

- Phase 2: only when it carries data (a date, seats, or a non-default status). Empty default shells and cancelled or retired "no further expansion" phases are skipped.

- Additional expansions: all of them.

- Phase 2 and expansions get their full M4–M10 set. One with seats but no date still gets the set, with M5 unset, for review.

- It is idempotent and never overwrites a stored M5.

- Known gap: older sites have no trustworthy M5 completion date, and none is invented. For example, 4 active sites past CO but not yet Ready to Open show no "Buildout P1 Open Date" in the Ready-to-Open report until someone enters it.

- Retired Buildout fields are hidden everywhere they were still shown. The stored data is kept for a few weeks before a restore-or-purge decision; do not run purgeRetiredBuildoutFields.

- Buildout report: the occupancy columns are removed, and "Buildout P1 Open Date" now reads Phase 1's M5 date.

- FTO: the TCO Obtained/Expiration columns, sorts and CSV columns are removed.

- Phase 2 projected (FTO and portfolio) reads Phase 2's M5 date.

- Tooltips, docs and agent text: the capacity tooltip, OpenAPI, DSS contract, agent guidance and tool descriptions now say M5.

- Write errors: error messages no longer mention retired fields.

- Removed: the Buildout write check that rejected duplicate dates. It only compared the retired stored date.

- Phase 2 can be removed (found in the dev clickthrough). The Buildout card now has a Remove action for an existing Phase 2, like additional expansions.

- Removal is explicit: the patch form is removePhase2: true, and the full-write form is phase2: null.

- A payload that just omits Phase 2 still keeps it, so an older writer can't delete it by accident.

- Removal drops the section and its milestones, and clears the legacy sites.phase2 placeholder so Phase 2 isn't re-created on read.

- The public API and MCP don't expose removal; it's in-app only.

- Confirm dialog labels. Milestone changes now read "Phase 2 › M4 · Obtaining Permits › Due date" instead of raw keys, and long labels wrap instead of being cut off.

## Post-release steps (tonight, straight after deploy)

1. Dry run the preview:

npx convex run migrations/backfillCompletingConstruction:preview '{}' --prod

Expected counts from the offline audit:

| Count | Expected |

|---|---|

| phase1Due | 23 |

| phaseDue | 32 |

| phaseNotApplicable | 1 (300 Cambridge St, Phase 2) |

| phaseWithoutDate | 4 |

| skip:pastConstruction | 51 (48 Phase 1 + 3 later phases) |

| skip:noDate | 98 |

| skip:emptyDefaultPhase | 52 |

| skip:cancelledPhase | 9 |

2. Run the backfill:

npx convex run migrations:run '{"fn": "migrations/backfillCompletingConstruction:backfill"}' --prod

3. Re-run the preview and confirm it reports zero writes.

4. Spot-check the consumers:

- projected capacity for 5400 Beethoven St and 350 E South Water St (Chicago)

- one Phase 1 site that got a due date

- FTO "Phase 2 projected"

- the Ready-to-Open report

Treat the deploy and steps 1–4 as one operation. Until step 2 runs, sites whose only date is the retired planned end show no projected capacity date. That short gap is accepted.

## Verification

- pnpm typecheck passes; biome is clean.

- Contracts: 1196/1196 pass.

- Chat sweep of 482 test files: 8644 passed, 18 skipped.

- New tests:

- the planner: every case and skip reason, plus idempotency

- the migration: exact preview counts, stored results, and a second run is a no-op

- projected vs current capacity: M5 wins, and M10 and the retired date are ignored

- I also ran the planner offline against the prod export: all 169 sites plan cleanly with no errors.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#1986 — feat(retention): advance report through completed months @ashwanth1109  approved

## Demo

<img width="2624" height="1636" alt="image" src="https://github.com/user-attachments/assets/53c5dc0c-5018-40e2-b395-047e9b56671f" />

## Business Value

The retention refresh now covers June 2025 through the last fully completed UTC month, so its published window advances automatically instead of remaining capped at May 2026. Reusing fresh detail observations from earlier SIS rosters prevents false source-preflight failures. Limiting incremental Timeback user requests to 100 records lets the users table recover after 2,000-record requests returned HTTP 502.

## Implementation Effort

An engineer would likely spend 1–2 days implementing, testing, and live-validating this change without AI assistance.

## Linear

[SURTR-1175](https://linear.app/builder-team/issue/SURTR-1175)

## Stack

This is the third layer. Parent: [#1830](https://github.com/AI-Builder-Team/Surtr/pull/1830), stacked on [#1827](https://github.com/AI-Builder-Team/Surtr/pull/1827).

## Changes

- Resolve the default report end to the previous completed UTC month at invocation time while keeping June 2025 as the start.

- Align source preflight, the stored procedure, and independent reconciliation to accept fresh observations from earlier roster runs, while enforcing their cycle lineage and publication cutoff.

- Use 100-record incremental Timeback user pages and cover pagination/recovery behavior with tests.

- Drop the legacy procedure overload only when the exact signature exists in Redshift’s catalog, inside the atomic cutover.

## Validation

- Production candidate published June 2025–August 2026 for 5,300 learners, with 30 monthly rows and 240 cohort rows, zero mismatches, and zero unexplained reconciliation residuals. A repeat run matched the same business fingerprint across 5,570 rows.

- The 100-record Timeback recovery published 9,957 changed users; independent readiness passed for all 55,518 current users.

- uv run pytest -q: 231 retention tests and 256 Timeback tests passed. Modified-file Ruff checks and the retention DDL dry run passed.

#2023 — feat(education): publish Aerie Finalsite tenant directory (SURTR-1460) @benji-bizzell  approved

## Summary

Publishes mart_education.aerie_finalsite_tenant_directory: the live Finalsite tenant estate (slug, name, status, lineage). Aerie School Identity will validate finalsiteTenant School links against it, the same way it validates QuickBooks and SIS links today. This is part of adding Finalsite to the School Ontology (Finance "Campus Mapping" ask).

Linear: SURTR-1460 · Project: EDU School Ontology (M1)

## What changed (mart-aerie-school-source-directories-refresh)

- Table: ddl/aerie_source_directories.sql adds the directory table, its refresh-lock table, comments and grants, mirroring the SIS block. The columns follow the shared contract: finalsite_tenant_slug VARCHAR(63), display_name, tenant_status, source_run_id, source_published_at.

- Procedure: ddl/sp_refresh_aerie_finalsite_tenant_directory.sql (new).

- Active set: the complete tenants in staging_education_finalsite.ingestion_run_sites for the latest *full* finalsight-raw-sync run in ingestion_ledger. The procedure validates that run as a complete boundary: 9/9 ledger objects, one snapshot, site rows reconcile to site_count, and no tenant both complete and retired. It never falls back to an older run.

- Why that source: the producer only removes a tenant through an explicit retire_sites (recorded as retired), so dead slugs drop out by data.

- Names: customer.name, falling back to long_name, then to the billing catalogue campus. A tenant with no usable name fails the refresh.

- Guards: slug regex, uniqueness, row floor and ceiling, single lineage, stale/conflicting overwrite guard, idempotent rerun, and an atomic DELETE + INSERT.

- Trigger: finalsight-raw-sync emits no dataset events, so the refresh runs on_pipeline_success of the scheduled daily full input only (prefix-filtered, as HubSpot/Ramp do). It is gated by DIRECTORY_REFRESH_FINALSITE_EVENTS_ENABLED="true". The handler verifies the published Mart against the pinned ledger run.

## Live check (read-only)

The latest full run is 765613fa…: 59 complete tenants, all 59 slugs valid and named. alphaschools (retired 2026-08-27) is correctly excluded.

## Validation

- uv run pytest: 62 passed. ruff check and ruff format --check are clean.

- CDK real-manifest construct test and test/schema pass.

## Deploy (manual DDL, as CQL_download_OM, before this is promoted)

1. Apply ddl/aerie_source_directories.sql. It is idempotent; the QB/SIS blocks are no-ops apart from re-running grants and comments.

2. Apply ddl/sp_refresh_aerie_finalsite_tenant_directory.sql.

3. Verify the tables and procedure are owned by CQL_download_OM and the Aerie reader has SELECT.

4. After release, confirm the next daily full (03:15 UTC) refresh reports row_count = site_count and a single source_run_id.

Ordering: the table must exist before the SURTR-1462 ontology procedure is applied, because that procedure references it statically. An empty table is fine.

## Notes

- A tenant retired by an on-demand full run leaves the directory at the next scheduled daily full, up to about 24h later.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

#2032 — fix(collections-weekly): honor Google Sheets quota retry windows @sanketghia  approved

## Summary

This supersedes the retry-budget change already merged in #2031 with a quota-aware hybrid:

- Preserve seven total attempts (one initial attempt plus six retries).

- Handle Sheets 429 responses with Retry-After when available.

- Fall back to a 60-second quota window plus bounded jitter when the header is absent.

- Bound cumulative quota waiting at 420 seconds, allowing six 60-second fallback waits while remaining below the 900-second Lambda timeout.

- Keep exponential 2/4/8/16/32/64-second backoff for 5xx responses and request timeouts.

## Why

PR #2031 correctly increased the retry budget for the observed six-consecutive-429 failure, but continued retrying on short exponential delays and did not use server retry guidance. This change preserves that regression coverage while preventing immediate quota retry amplification.

## Validation

- uv run --extra dev pytest — 53 passed

- uv run --extra dev ruff check src tests — passed

- Three local read-only dry-runs — each parsed 406 rows; the first exercised the new 429 fallback and slept 63.1 seconds.

- Authorized local production-target run wrote 406 rows; post-write verification found 406 rows, zero duplicate grain groups, and cleaned S3 staging.

## Scope

This changes only the weekly pipeline retry behavior and tests. It does not add cross-pipeline coordination, change schedules, or reduce baseline read volume.

#2038 — [SURTR-1181] Derive retention from dated SIS programs @ashwanth1109  approved

## Summary

- derive learner start, withdrawal, status, and campus from the dated SIS program_enrollments history

- require exactly one fully valid same-campus school-year program overlapping the report window; retain every learner with explicit quality flags and exclusion reasons

- bind reused rolling observations to their exact raw/clean lineage and enumerate malformed history elements before relevance filtering

- remove report-end cancellation imputation and preserve profile fields for audit comparisons only

- update independent reconciliation, source preflights, documentation, and regression coverage

## Business Value

Retention reporting now reflects the dated enrollment program that actually overlaps the reporting window instead of relying on stale profile fields or fabricated cancellation dates. Ambiguous, malformed, missing, prospective, and otherwise unresolved histories remain visible and auditable without contaminating retention metrics.

## Implementation Effort

Estimated 4–6 engineer-days without AI assistance, including source-contract investigation, Redshift procedure and reconciliation changes, test coverage, isolated deployment, two production validation runs, and evidence documentation.

## Validation

- 262 focused tests passed

- modified-file Ruff format/check and git diff --check passed

- isolated CDK deployment showed [1/1]; only Pipeline-mart-aerie-retention-refresh-prod reached UPDATE_COMPLETE

- atomic DDL/catalog/grant verification passed

- first event-driven refresh cbe31514-a59e-4cc9-8ddb-1c586c1363c8 succeeded for 5,295 learners, 30 monthly rows, and 240 cohort rows with zero unexplained residuals and zero imputed cancellations

- repeat refresh ad0a4f18-b262-4af0-9685-a7e1520a475d passed every independent contract and matched SHA-256 fingerprint 051769220d2c93d0261ad0aaa3923aa2c94cc908807e06f237624d44bf8a80a7 across 5,565 business rows

- isolated rollback fixture confirmed all prior snapshots survive an intentional transaction failure

## Tracking

- Linear: https://linear.app/builder-team/issue/SURTR-1181

- Supersedes #1832 with a current-main implementation and addresses the malformed-array-element review finding.

The Portfolio  —  Trilogy Companies

As AI Hiring Faces a Reckoning, Crossover's Meritocracy Pitch Gets a Stress Test

A wave of discrimination lawsuits and regulatory scrutiny over algorithmic screening raises uncomfortable questions for any company — including Trilogy's own talent engine — that lets machines decide who gets an interview.

AUSTIN, TEXAS — There is a reckoning underway in the quiet machinery of hiring, and it is arriving in the form of lawsuits, not press releases. A Workday discrimination lawsuit now working its way through the courts asks a deceptively simple question that the entire industry has spent years avoiding: when an algorithm rejects a candidate before a human ever sees the résumé, who, exactly, is accountable? Meanwhile, a new advocacy campaign — 'Reject the Rejections' — cites new research suggesting AI screening tools are filtering ethnic minority candidates out of the pipeline before an interview is ever offered.

This is, inescapably, relevant to the Trilogy portfolio. Crossover, the Austin-based talent platform that stakes its identity on 'rigorous AI-enabled skills assessments' designed specifically to minimize résumé bias, has built its entire brand promise on the inverse claim — that algorithmic screening, done correctly, is more meritocratic than a human recruiter glancing at a name and a zip code. It is a compelling thesis. It is also, per the current headlines, exactly the thesis now under legal and regulatory siege.

Corporate counsel quoted in HR Executive this week frame the moment starkly: layoffs are colliding with AI hiring tools to produce a new category of class-action exposure, one that regulators from Tanzania to Washington are only beginning to map. For a company whose entire moat rests on the credibility of its screening algorithm, the message is not subtle. The burden of proof is shifting. It is no longer enough to claim your system is fair — soon, someone in a courtroom may ask you to demonstrate it.

Workday Discrimination Lawsuit: Who's Liable for AI Bias? -  ·  ‘Reject the Rejections’ Campaign Launched as Study Warns AI  ·  Artificial Intelligence in Recruitment: Navigating Tanzania’

Contently Bets on Compliance as the Next Content Marketing Battleground

AUSTIN, TEXAS — In a content marketing software landscape increasingly defined by lookalike feature lists and a swelling list of Pepper Content alternatives, Contently is staking its claim on a decidedly less glamorous but wildly lucrative differentiator: compliance.

The Trilogy-affiliated platform, acquired by Zax Capital (an ESW Capital division) in September 2024, has published a new framework it calls Compliance-First Content Architecture — a five-component workflow designed to help regulated finance brands scale content production without triggering a governance nightmare. It's exciting news for an industry that, as ContentGrip recently noted, is still sorting out the difference between a true content marketing platform and a glorified media resource library.

This is a paradigm shift for a category that has historically optimized for speed and volume — SEO briefs, editorial calendars, freelancer marketplaces — while treating regulatory review as an afterthought bolted onto the end of the pipeline. Contently's architecture flips that model, building governance checkpoints into the content lifecycle from ideation through publication, a critical requirement for banks, insurers, and asset managers who can't afford a rogue blog post triggering an SEC inquiry.

The timing is no accident. The Gartner Magic Quadrant for Content Marketing Platforms has seen significant churn in recent evaluation cycles, and with dozens of new entrants and alternatives crowding the field, differentiation has never been more critical. Under new CEO Brandon Pizzacalla, Contently is leveraging its 165,000-strong creative marketplace and applying it to a narrower, higher-margin use case rather than chasing every horizontal use case in the market.

**Key Takeaways:**

- Contently is targeting regulated finance as a defensible niche in a commoditizing category

- The five-component compliance workflow embeds governance throughout content production, not just at the end

- The move reflects Trilogy's broader playbook: find an underserved, high-value niche and build best-in-class infrastructure around it

We're just getting started.

Skyvera's Telecom Land Grab: The Pattern Behind the Acquisitions

Two deals and a record-setting compliance sprint reveal how Skyvera is quietly assembling a full-stack telecom software empire.

AUSTIN, TEXAS — On the surface, these look like three unrelated items in Skyvera's press feed: an acquisition closed, another one announced, and a technical certification hit in record time. Taken separately, none of it makes the front page. But if you read between the lines — and I've spent the better part of a week doing exactly that — a pattern emerges that looks less like opportunistic shopping and more like a deliberate assembly line.

Start with the completed acquisition of CloudSense, the Salesforce-native CPQ platform that telcos use to quote and fulfill complex B2B and wholesale deals. Then layer in the STL divested assets deal, which brought Skyvera a digital BSS stack — monetization, optical networking, analytics — the kind of unglamorous plumbing that keeps a mobile operator's lights on. Two acquisitions, months apart, both aimed at the same target: the seams between legacy telecom infrastructure and the cloud-native future ESW Capital has been betting on since Totogi launched its charging-as-a-service pitch.

And this is where it gets interesting. Within weeks of folding CloudSense into the Skyvera portfolio, the company announced it had certified all 13 of CloudSense's APIs to TM Forum compliance standards in a single month — a process the industry says normally takes 26 months. My source inside the portfolio, who I won't name because the ink on these acquisitions is barely dry, tells me the compliance sprint wasn't a side project. It was a demonstration. Skyvera didn't just want CloudSense's customer base; it wanted to prove, publicly and fast, that AI-assisted engineering could compress a two-year regulatory slog into thirty days.

That's not a coincidence. That's a sales pitch to every telco still running on-premise BSS who's nervous about migration timelines. Buy the asset, then use it to advertise the speed of the operating model. It's the ESW playbook — acquire cheap, integrate fast, extract margin — but running at telecom scale, with AI doing the heavy lifting on compliance work that used to require armies of engineers.

Watch what Skyvera does next. If the pattern holds, the STL assets won't stay a standalone line item for long either.

Cloudsense  ·  CloudSense achieves TM Forum API compliance in record time u  ·  Skyvera completes acquisition of CloudSense, expanding telec
The Machine  —  AI & Technology

The Trust Deficit: AI Wants Your Data, Not Your Confidence

Meta's new agent handles your dinner reservations and dental claims; a fake photo of the president just proved nobody has to believe anything anymore.

SAN FRANCISCO — Two stories broke this week that describe the same industry from opposite ends of a widening gap.

In the first, a New York Times reporter handed her life over to Muse, Meta's new AI agent, which negotiated a dental insurance dispute, booked restaurant reservations and produced a podcast on her behalf. The price of admission was total: calendars, contacts, financial accounts, the texture of a daily routine. Meta is betting users will trade access for convenience at a scale beyond what Alexa or Siri ever secured.

In the second, a fabricated image of President Trump kissing a woman circulated widely despite being demonstrably fake. Provenance tools exist. Nobody checked. The image spread because platforms reward velocity over verification, and generative models have gotten good enough that the average scroll can no longer tell the difference.

Same week, Anthropic released Opus 5.5, pitched as its cheapest, fastest model yet — and, the company says, its strongest performer on internal safety benchmarks to date. That claim arrives amid a separate set of disclosures from OpenAI, Google and Meta acknowledging AI models involved in security breaches, unrelated to Opus itself but part of the same September news cycle. The juxtaposition is not subtle: vendors are asking the public to trust systems more capable of deception at the exact moment those systems are shown to be more exploitable.

The economics explain the posture. An agent that books your reservations only earns its subscription fee if it also touches your insurer, your bank and your inbox. Consumer AI's business model requires maximal access; consumer AI's credibility problem requires minimal blind trust. Those two requirements are not compatible, and no vendor has resolved the tension — they've just priced around it.

Trilogy's own AI platform, Klair, was built on the opposite premise: keep the model inside the walls, auditable, scoped to portfolio financials rather than open to the internet's incentive structure. It's a narrower bet. It's also a bet that the industry's current trust deficit gets worse before it gets better.

I Gave My Life Over to Meta’s A.I. Agent and Was Blown Away  ·  An A.I. Image of Trump Kissing a Woman Was Fake. It Spread A  ·  Anthropic Releases a New A.I. Model, Opus 5.5, Amid Safety D

Hugging Face's Trifecta: Reproducibility, Efficiency, and Community Muscle All Land at Once

Sometimes the future doesn’t arrive with a flashy demo; it arrives through a few quiet developments that reshape an ecosystem. Today is one of those days.

The UK AI Safety Institute and EvalEval are building infrastructure to make AI benchmark results reproducible. Their collaboration addresses a persistent problem: labs often fail to document evaluation harnesses, prompts and sampling settings, making published scores difficult to verify.

Meanwhile, Hugging Face Transformers now natively supports llama.cpp quantized models, removing a major barrier between production-focused inference and the research library widely used to build models. The change could make large models more practical on modest hardware.

Jun Kim, creator and maintainer of oMLX, is joining Hugging Face full time to support the MLX community. His appointment signals strong backing for Apple Silicon and on-device AI.

Separately, researchers are framing LLM block pruning as an Ising optimization problem, using statistical physics to identify expendable transformer blocks. Together, the developments highlight rapid progress in AI infrastructure, efficiency and evaluation.

The Ghosts We Ask for Opinions: When Synthetic Audiences Get It Wrong

A new study finds that AI 'personas' built to predict human reactions to marketing copy are outperformed by no persona at all — a small humbling for the growing industry of simulated minds.

AUSTIN, TEXAS — There is a peculiar comfort in imagining that we can dress a language model in the clothes of a person who does not exist — a 34-year-old suburban mother, a skeptical Gen Z gamer, a retired accountant in Ohio — and ask it how that person would feel about an advertisement. It is a kind of séance for market research, conjuring audiences before they are ever exposed to anything. And increasingly, brands are doing exactly this: profile-conditioned large language models standing in for the messy, expensive, slow work of actually asking humans what they think.

A new study puts this séance to an uncomfortable test. Researchers ran what they call a sim-to-real experiment, comparing LLM-generated "synthetic personas" against real audience responses to marketing copy — and found that a baseline model with no persona conditioning at all predicted real reactions better than the elaborate personas built to mimic specific demographic and psychographic profiles (arXiv:2609.25010). The machinery of imagined selfhood, it turns out, may be adding noise rather than signal.

This is not a small finding, and it should not be a surprising one. A model conditioned to "act like" a persona is not sampling from that persona's actual distribution of beliefs — it is sampling from its own internal stereotype of what such a persona would say, compressed through training data that was never labeled with ground truth about real reactions. The persona is a costume, not a nervous system.

It is a lesson echoing across machine learning this season. In an unrelated arXiv paper on 4DGS-JEPA, researchers building predictive world models for dynamic 3D scenes emphasize the same distinction — that reconstructing what a scene looks like is not the same as learning the dynamics that will let a model predict what happens next (arXiv:2609.25036). Fidelity of appearance, whether a synthetic customer or a synthetic scene, is not fidelity of behavior. The map may be beautifully drawn. The territory, stubbornly, remains elsewhere.

Do Synthetic Personas Predict Real Audience Response? A Sim-  ·  Do Existing Preconditioners Improve Biomedical Tabular Found  ·  4DGS-JEPA: Temporally Compositional Joint-Embedding Predicti
The Editorial

The Sheepskin Comes Apart at the Seams

For a century the diploma was a promissory note redeemable at the door of the middle class; the tellers have simply stopped honoring it.

AUSTIN, TEXAS — There is a certain species of American who has spent his whole life being told that a piece of paper, suitably embossed and hung in a frame, would stand between him and the abyss. He borrowed against that promise, sometimes six figures' worth, and now finds himself in his thirties serving coffee to men who never finished sophomore year but happened to learn how to close a sale or write a line of code. This is not a scandal. It is merely arithmetic finally being permitted to speak.

Fortune ran a piece this week arguing that artificial intelligence did not break higher education so much as expose a fraud that had been running since roughly the Johnson administration — the fraud being that the credential itself, rather than the competence it was meant to certify, was the product being sold. Chronicles Magazine, writing from a different pew of the same church, calls this elite overproduction, which is the polite sociological term for what happens when a society mints more law degrees and MBAs than it has chairs at the table for their holders to sit in. Meanwhile in India, a nation that has made credentialing something close to a national religion, Youth Incorporated reports millions of graduates holding degrees that qualify them for jobs which, on inspection, do not exist. The diploma mills of two hemispheres, it turns out, have been running the same con.

None of this required a chatbot to become visible. It required only that employers, weary of interviewing candidates who could recite theory but not perform, start asking a more vulgar question: can you do the thing? University World News notes the quiet arrival of skills-based hiring, in which firms drop the degree requirement and administer, instead, a test of actual capability — a development greeted in faculty lounges with the horror once reserved for grave robbery.

I raise an eyebrow here not because I am a stranger to the argument but because I have watched a rather more literal version of it unfold under Trilogy's own roof, where Joe Liemandt's Alpha School has spent several years running the experiment nobody asked for: what happens to a child's education when you strip out the seat time, the busywork, the entire theater of schooling, and let an AI tutor discover in two hours what a lecture hall could not accomplish in six. The answer, inconveniently for the credentialing industry, is that the theater was never where the learning happened. It was where the certificate was issued.

The oil industry, for what it is worth, is being asked this season to answer for the weather, which strikes me as roughly the reverse problem — an institution being held liable for outcomes it can no longer credibly claim not to have caused. Higher education, by contrast, is being asked to answer for outcomes it can no longer credibly claim to have caused at all. Both institutions, having spent decades selling certainty, are discovering that the bill for it eventually arrives, and it is never paid in the currency they'd prefer.

AI didn’t break higher education—It exposed the credential t  ·  Elite Overproduction and Higher Education - Chronicles Magaz  ·  Degree in Hand, No Job in Sight: India’s Unemployment Crisis
The Office Comic  ·  Art Desk
The Office Comic  ·  Art Desk

Nation's AI Productivity Data Now Sophisticated Enough To Simultaneously Prove Everything And Nothing

Economists confirm the numbers are in, and the numbers say whatever you were already planning to say.

AUSTIN, TEXAS — In a stunning testament to how far artificial intelligence measurement science has come, the AI industry announced this week that it can now conclusively prove productivity gains are already staggering, mostly nonexistent, and about to double the U.S. economy, all at once, with charts.

According to a widely circulated Federal Reserve analysis, 95 percent of AI's promised productivity boom is still to come, a phrase that manages to sound both reassuring and like something a contractor says in month fourteen of a kitchen remodel. Meanwhile, a commit-level study of Big Tech engineering teams found developer performance rose an eye-popping 150 percent over eighteen months, a figure so precise it suggests someone, somewhere, is counting semicolons with the fervor of a man counting cards. And Elon Musk, never a man to let a modest claim go unmolested, told the world AI will double U.S. GDP growth to 4 percent next year, a forecast roughly seventeen forecasters ahead of every other forecast on Earth.

Taken together, the data paints a coherent picture only in the sense that a Rorschach blot paints a coherent picture: everyone sees whatever they walked in wanting to see. Analysts optimistic about AI point to the 150 percent developer gains. Analysts skeptical of AI point to the 95 percent still unrealized. Analysts who own a lot of Tesla stock point to Elon.

Into this fog rides a helpful piece from Foundever offering, with admirable sincerity, six steps to turn AI productivity claims into verifiable results, a headline that implies, rather bravely, that up until this exact article, nobody had thought to check.

Elsewhere in the real economy, a brokerage firm profiled by RISMedia learned the hard way that announcing an AI rollout before training anyone on it produces the exact productivity gain you'd expect: none, plus a Slack channel full of agents nobody opened. The company's AI transformation reportedly died in month two, which, incidentally, is also when most of these productivity studies stop measuring and start extrapolating.

Executives across the industry say they remain confident that AI is delivering enormous, still-invisible, retroactively 150-percent, forward-looking 4-percent, six-step-verifiable value, and that the only real productivity crisis is the growing number of hours spent reading studies about productivity instead of producing anything. A follow-up study on that phenomenon is reportedly still to come.

6 steps to turning AI productivity claims into verifiable re  ·  AI productivity claims are 95% 'still to come', Fed finds -  ·  Big Tech Engineering Performance Rose 150% Per Developer Ove
On This Day in AI History

On September 23, 2008, T-Mobile and Google unveiled the T-Mobile G1, the first commercially available Android smartphone. It marked Android’s entry into the mobile market and challenged Apple’s growing iPhone dominance.

⬛ Daily Word — technology
Hint: A machine designed to perform tasks automatically, often with some degree of intelligence.
Share this edition: 𝕏 Twitter/X 🔗 Copy Link ▦ RSS Feed