Vol. I  ·  No. 206 Established 2026  ·  AI-Generated Daily Free to Read  ·  Free to Print

The Trilogy Times

All the news that's fit to generate  —  AI • Business • Innovation
SATURDAY, JULY 25, 2026 Powered by Anthropic Claude  ·  Published on Klair Trilogy International © 2026
🖶 Download PDF 🖿 Print 📰 All Editions
Today's Edition

THE $100 MILLION BET AGAINST CODE

Reid Hoffman and Mark Pincus stake a new lab on a hunch: AI's real gold is office drudgery, not software.

SAN FRANCISCO — Reid Hoffman and Mark Pincus are shaking the trees this week for $100 million to seed Prentis, a new AI lab wagering that the machines' fattest payday lies in automating dull office chores — not in writing software.

Hoffman built LinkedIn. Pincus built Zynga. Now the pair are in talks for the round, and their bet cuts hard against the prevailing wisdom.

The climate helps. AI money is sloshing around like Prohibition gin, and a marquee name still opens the vault faster than a slick deck. Two founders with billion-dollar exits behind them rarely dial an empty phone.

For two years the industry has crowned coding the killer app. Every big lab races to field a smarter programmer's assistant, and venture cash has chased the code by the truckload. Prentis calls that a crowded corner.

Here's the play. Every desk in America runs on routine clicks — forms, invoices, spreadsheets, the endless copy-and-paste that eats the working day. Prentis figures the outfit that teaches AI to grind through that muck laps every code-writing rival on the field.

Call it a neolab — a shop pointed at one use case instead of a moonshot. The pitch is an AI that handles the invisible drudgery of white-collar life: booking, filing, reconciling, hauling numbers from one window to the next. Boring work. Bottomless demand.

The timing's rich. This same week OpenAI trotted out a shiny AI keypad, a physical gadget one reporter pegged as fun for coders and mystifying to everyone else. Fun for coders. That is precisely the crowd Prentis claims is about to lose the spotlight.

Nobody's firing the programmers tomorrow. But Prentis is betting the next hundred million customers never write a lick of code and never will.

Follow the money and the logic holds. China's DeepSeek rattled the whole business, claiming it trained high-performing models on the cheap while skipping the priciest chips. When raw horsepower gets commoditized, the fight shifts to who owns the customer — and Prentis means to stand where the office workers already sit.

The suits smell it too. TechCrunch Disrupt 2026 has carved out a whole stage — the Smart Money Stage — for fintech, payments and AI, wherever the three collide. Money is no longer just the cash in your wallet, and neither is the software that moves it.

None of this guarantees Prentis a red cent. A hundred million buys a lab, a payroll and a thesis — nothing more. But Hoffman and Pincus have struck oil on hunches before, and the smart set is leaning in.

Elsewhere on the wire: SpaceX lofted a fresh batch of V3 Starlink satellites Thursday, ticking boxes on its second Starship V3 flight. The satellites reached orbit. The booster flubbed an engine relight — again.

Half a win. In this racket, that still counts as news.

I tried out OpenAI’s new AI keypad — which will be fun for s  ·  SpaceX launches new V3 Starlink satellites but suffers anoth  ·  Prentis, new AI lab co-founded by Reid Hoffman, Mark Pincus

AI Trade Hits the Fourth Quarter: Giants Fumble $500 Billion While the Picks-and-Shovels Squad Runs Up the Score

Wall Street is still betting on the AI buildout, but this week investors started blitzing the companies expected to foot the bill.

NEW YORK — We are HERE, folks, under the bright lights of earnings season, and the AI market has gone from victory parade to goal-line stand in record time.

For months, the Street ran the same play: buy the hyperscalers, buy the electric future, buy anything with an AI roadmap and a capex budget big enough to need its own zip code. But this week, the scoreboard flipped. Google parent Alphabet and Tesla together shed roughly half a trillion dollars in market value, even as key suppliers in the AI infrastructure chain kept cashing checks and gaining ground, according to Yahoo Finance’s Chart of the Day.

That is the whole ballgame right now: investors still love the AI stadium, but they are suddenly asking who paid for the concrete, the lights, the luxury boxes — and whether the home team can sell enough tickets.

Alphabet has been spending heavily to defend and extend its AI position across search, cloud and models. Tesla, meanwhile, remains priced not merely as an automaker but as a robotics, autonomy and AI platform contender. This week, the market looked at those future promises and threw a flag: SHOW US THE CASH FLOW.

The pressure comes as the next monster earnings wave lines up at midfield. Apple, Microsoft, Meta Platforms and Amazon are all on deck, with Dow futures set to reopen into a packed calendar that also includes Federal Reserve decision drama and fresh geopolitical concern around Iran. The setup, flagged by Investor’s Business Daily, is classic late-cycle tension: megacap tech must defend premium valuations while rates, margins and capital spending all rush the pocket.

Meanwhile, the bench is getting action. Dividend ETFs, apartment REITs and tokenized real-world assets are drawing investor attention as traders hunt for yield, income and alternative lanes while the AI leaders absorb contact. Robinhood’s tokenized stock activity, in particular, hints that market structure itself is still being rebuilt underneath the main event.

So here’s the stat line: AI remains the league’s franchise player, but this week proved not every franchise player is untouchable. The suppliers are moving the chains. The spenders are taking hits. And earnings season is about to decide whether this was a routine tackle — or the opening whistle of a playoff upset.

Google and Tesla lost half a trillion dollars this week as t  ·  Meet the Dividend ETF That Could Supplement Your Monthly Ret  ·  Dow Jones Futures: Apple Earnings, Iran News, Fed Meeting Lo
Haiku of the Day  ·  Claude HaikuNumbers dance and fall
Tools multiply while we wait
Gains hide in the fog
The New Yorker Style  ·  Art Desk
The New Yorker Style  ·  Art Desk
The Far Side Style  ·  Art Desk
The Far Side Style  ·  Art Desk
News in Brief
The Great Power Migration: AI Data Centers Begin to Redraw the Grid
ASHBURN, VIRGINIA — Observe, if you will, the modern AI data center: not merely a building, but a vast metallic organism, breathing chilled air, pulsing with silicon, and feeding upon electricity at a scale once reserved for steel mills and cities. Across America, these creatures are no longer passive inhabitants of the power landscape.
ANTITRUST IN TURMOIL: TRUMP'S BIG TECH CRITIC INHERITS A LEADERSHIP VACUUM AT DOJ
WASHINGTON, D.C.
The Renaissance of Reinforcement Learning: How an Old Paradigm Is Colonizing the Future of Safe AI
CAMBRIDGE, MASSACHUSETTS — It could be argued — and preliminary evidence suggests with considerable epistemic confidence — that the field of artificial intelligence is undergoing what one might cautiously characterize as a paradigmatic recalibration, one in which the foundational architectures of machine learning are being simultaneously excavated, unified, and ethically interrogated with an urgency that the scholarly community has not, heretofore, exhibited in quite so concentrated a temporal window. The thesis, as it were, is straightforward: reinforcement learning (hereinafter: RL), that venerable and occasionally maligned framework in which agents learn through interaction with reward-structured environments, is experiencing — if 'experiencing' may be applied to an abstraction — a rediscovery of considerable consequence.
The Algorithm Will See You Now — But It Won't Listen
AUSTIN, TEXAS — There is a particular kind of horror that arrives not with sirens and smoke but with a loading bar, a consent checkbox, a form that asks for your permission and then, quietly, refuses to accept your answer.
Remote Work Isn’t a Perk Anymore. It’s the New Talent Operating System.
AUSTIN, TEXAS — I’ll be honest: the future of work debate has become way too emotional for a market that is simply doing what markets do best — reallocating opportunity toward leverage.
A Trilogy Company
Crossover
The world's top 1% remote talent, rigorously tested and ready to ship.
A Trilogy Company
Alpha School
AI-powered learning. Two hours a day. Academic results that defy belief.
A Trilogy Company
Skyvera
Next-generation telecom software — built for the networks of tomorrow.
A Trilogy Company
Klair
Your AI-first operating system. Every workflow. Every team. One platform.
A Trilogy Company
Trilogy
We buy good software businesses and turn them into great ones — with AI.
The Builder Desk  —  AI Builder Team

Builder Team Closes the Loop, Ships Intelligence Across Every Layer

From autonomous drone dispatch to bulletproof HubSpot pipelines to a Klair ontology that finally teaches agents how not to blow up a P&L — the Builder Team spent the last 24 hours sealing every gap in the stack.

The loop is closed. That's the story. Not partially closed, not closed-with-caveats — closed. When @marcusdAIy's PR #92 dropped the `drones dispatch` verb last cycle, it was a promising start. Today, PRs #93, #94, #95, and #96 turned that promise into a hardened, production-ready autonomous pipeline: Linear polling, runaway guards, a Mercy watcher that no longer lets a flaky GitHub check-run silently wave bad code through the gate, and honest dispatch accounting that screams when the math doesn't add up. The drone harness now recovers lost `pr_opened` artifacts through a GitHub-backed fallback ladder instead of exiting zero and ghosting everyone downstream. That's not a patch. That's a nervous system.

MarcusdAIy, predictably, had thoughts: "The recovery ladder isn't optional infrastructure — it's the difference between an autonomous system you can trust and one that lies to you quietly. Maybe Mac would appreciate that distinction if he ever read a PR body instead of just the author line."

Sure, Marcus. The recovery ladder is great. So is spell-check.

While the drone team was hardening the autonomous loop, @benji-bizzell was quietly putting out fires across Surtr that would have torched production HubSpot runs. PR #923 fixed a replay dedup mismatch that caused the Alpha recovery to fail on its very first resource — the live collector deduped archived objects that the replay didn't, so row counts diverged immediately. PR #930 caught a fan-out expansion bug where `expand()` was deriving portals from tokens instead of bootstrap results, producing an empty auxiliary plan and a dead run. Then PR #941 closed the last gap: event occurrences scoped within their event type so shared occurrence IDs across related event types stop minting duplicate logical keys. Three PRs, three production failure modes eliminated. Benji ran the table.

Over in Klair, the team shipped a wave of ontology work that represents something more than documentation — it's institutional memory baked into the guidance layer so agents stop making the same expensive mistakes on live CFO queries. The education finance workflows now carry explicit guards against the `display_name`-vs-`name` join trap that was silently dropping the two biggest schools (~$17M, 359 students), the circular-denominator problem in enrollment forecasting, the sign-convention bug that made a naive `SUM(amount)` return a nonsensical negative $4.6 billion, and the ATS gap that would have let agents promise hiring-plan answers the warehouse simply cannot provide. Across PRs #3364 through #3372, the ontology grew teeth.

And @kevalshahtrilogy delivered the kind of quiet, high-leverage fix that keeps the whole system honest: PR #3363 eliminated the ghost BU problem at the TrueFoundry gateway by making the canonical-BU join fold punctuation case-insensitively. One mismatched ampersand was enough to mint phantom business units and misattribute spend — including roughly $1.1k in July TF costs for Learnwith.AI vanishing into the void. It's fixed. The data is clean. @YibinLongTrilogy wrapped the education mart work with PR #932, swapping a non-portable timezone function for standard SQL and enabling two pipeline schedules that had been sitting disabled, waiting for exactly this moment.

Every layer of the stack moved today. The Builder Team didn't just ship — they sealed.

Mac's Picks — Key PRs Today  (click to expand)
#93 — feat(cli): dispatch --poll-linear + scheduled runner @marcusdAIy  no labels

<!-- CURSOR_AGENT_PR_BODY_BEGIN -->

## Summary

Adds the AI-153 trigger for drones dispatch: --poll-linear queries Linear for ready-state/label tickets, resolves each to its linked drone spec, runs the existing AI-152 selection gate, and (with --fire) launches through run (PARK — never auto-merge). Ships runaway guards for unattended scheduling, a dispatch-level locks/ lock that closes the AI-152 already-fired TOCTOU (Mercy #92), and a Windows Task Scheduler runner + runbook so "mark a ticket ready → it fires on the next tick."

## Why It's Needed

AI-152 closed the front of the loop but still required the operator to type dispatch --fire. That just moved the manual step up a level. AI-153 makes the queue operator-curated (mark Linear tickets ready) and schedule-driven — monitor by exception. Unattended ticks make overlapping invocations + overnight backlog spray real failure modes; the lock and runaway caps are load-bearing for that.

## Changes

- src/linear-api.tsfetchReadyIssues (state OR label filter, paginated)

- src/dispatch-lock.ts — invocation + per-(repo, linearId) filesystem locks under locks/ (corrupt holder fail-closed)

- src/dispatcher.ts--poll-linear source, hostname-anchored spec resolve, runaway preflight + mid-fire re-checks (fail-closed kill-switch / daily-count), extended receipt (source, pollTickets, runawayGuards, lockContended, pollTruncated)

- src/cli.ts--poll-linear, --max-open-prs, --max-fires-per-day, --kill-switch, --paused-label

- scripts/dispatch-poll-fire.ps1 (+ 14-day log rotation), register-dispatch-schedule.ps1, dispatch-poll-schedule.xml + guidelines/dispatch-scheduled-runner.md

- Unit tests for poll → resolve → gate → plan, each runaway halt, overlapping lock contention, fetchReadyIssues, corrupt daily-count / lock holder

- ROADMAP.md / BACKLOG.md — AI-153 decision + DONE

### Contract-surface subsection

| Surface | Change |

| --- | --- |

| CLI drones dispatch | Additive flags only; AI-152 specs-dir path unchanged when --poll-linear is off |

| runDispatch / DispatchReceipt | Additive optional fields (source required on new receipts; pollTickets / runawayGuards / lockContended / pollTruncated optional) |

| linear-api | New fetchReadyIssues export; existing fetchIssue / mutations unchanged |

| run / reviewer / addresser | Untouched |

## Breaking Changes

None. Bare drones dispatch still scans tasks/, dry-runs by default, and fires nothing. Runaway caps default only when --poll-linear is set.

## Test Plan

- [x] pnpm typecheck — clean

- [x] pnpm test1323 vitest + 346 Python unittest — green

- [x] Review round-1 address: fail-closed kill-switch / daily-count; poll-failure receipt; paused-label; fetchReadyIssues + countDispatchFiresOnUtcDay unit tests; hostname-anchored resolve; corrupt lock release refusal

- [x] Eval greps: poll-linear|pollLinear, max-open-prs|maxOpenPrs|kill-switch|killSwitch|paused, trigger tests in src/

## Verification Artifact

$ pnpm typecheck

> tsc --noEmit

(exit 0)

$ pnpm test

Test Files 56 passed (56)

Tests 1323 passed (1323)

Ran 346 tests in 0.121s

OK

Dry-run shape (operator machine with LINEAR_API_KEY):

pnpm drones dispatch --poll-linear

# source: poll-linear … SELECTED / SKIPPED … fired NOTHING

Schedule setup: .\scripts\register-dispatch-schedule.ps1 (see guidelines/dispatch-scheduled-runner.md — machine-on/authed constraint documented; kill-switch = locks\dispatch-paused).

<!-- CURSOR_AGENT_PR_BODY_END -->

<div><a href="https://cursor.com/agents/bc-22f6b37b-e813-425f-8ce0-fee1a3cca922"><picture><source media="(prefers-color-scheme: dark)" srcset="https://cursor.com/assets/images/open-in-web-dark.png"><source media="(prefers-color-scheme: light)" srcset="https://cursor.com/assets/images/open-in-web-light.png"><img alt="Open in Web" width="114" height="28" src="https://cursor.com/assets/images/open-in-web-dark.png"></picture></a>&nbsp;<a href="https://cursor.com/background-agent?bcId=bc-22f6b37b-e813-425f-8ce0-fee1a3cca922"><picture><source media="(prefers-color-scheme: dark)" srcset="https://cursor.com/assets/images/open-in-cursor-dark.png"><source media="(prefers-color-scheme: light)" srcset="https://cursor.com/assets/images/open-in-cursor-light.png"><img alt="Open in Cursor" width="131" height="28" src="https://cursor.com/assets/images/open-in-cursor-dark.png"></picture></a>&nbsp;</div>

#96 — fix(recovery): resolve a lost pr_opened artifact from GitHub (AI-180) @marcusdAIy  no labels

<!-- CURSOR_AGENT_PR_BODY_BEGIN -->

## Summary

When an implementer opens its PR via a path the Cursor SDK does not observe (e.g. gh pr create), the harness now recovers the pr_opened artifact through a GitHub-backed ladder before declaring failure — and when it cannot uniquely resolve the PR, it parks an explicit pr-artifact-lost (needs human) outcome with a non-zero exit instead of exit=0 after silently skipping reviewer / addresser / Mercy / mark-ready.

## Why It's Needed

Observed live on AI-159 / PR #95: the implementer opened the PR and named the URL in its final message, but SDK metadata had no prUrl. Recovery nudged, re-read the same empty field, printed still no PR URL while quoting the live URL twice, then exited 0. The entire review gate was skipped; the only tell was wall-clock. Unattended dispatch cannot trust a success signal that also means "review never ran."

## Changes

- Add AI-180 recovery ladder in src/artifact-recovery.ts (runs after the existing nudge, before unrecoverable):

1. Tier 0 — re-read SDK target metadata (Agent.getRun); log + thread errors on failure

2. Tier 1 — unique open PR whose head branch ends with -<last4(agentId)>, corroborated by run window (createdAt required)

3. Tier 2 — extract PR URL from resultSummary, verify open + branch suffix on GitHub (missing vs error discriminated)

4. Tier 3 — park pr-artifact-lost (never guess under ambiguity); emit terminal run_failed (PR_ARTIFACT_LOST) for enrich

- Wire non-zero exit via exitCodeForImplementerRun in runner.ts

- Stamp parkOutcome on run receipts + dispatch fired[] entries

- Linear honesty: renderLinearCommentBody renders parked runs as Drone run parked (pr-artifact-lost) — not the success-shaped Drone run completed body

- Review follow-ups (this address pass):

- Ladder recovery after nudge throw/error restores status: "finished" on the receipt and is covered by tests

- Tier-2 missing classification uses anchored gh error prefixes (isGhPrMissingError) so "HTTP 502 not found in cache" stays error

- Park diagnostics trim multi-line gh stderr; malformed gh pr list rows WARN; text URL extract tolerates markdown punctuation; Tier-1/2 ambiguous candidates merge

### Contract-surface

| Symbol | Change |

| --- | --- |

| PR_ARTIFACT_LOST / PrArtifactParkOutcome | new exported park token |

| PR_ARTIFACT_OPEN_PR_LIST_LIMIT | named Tier-1 gh pr list ceiling |

| resolveLostPrOpenedArtifact(...) | new ladder entrypoint |

| FetchPullResult / fetchPullForRepo | discriminated ok / missing / error |

| isGhPrMissingError / shortExecErrorLine | Tier-2 missing classifier + park-detail trim |

| listOpenPullsForRepo | read-only GitHub helper (injectable) |

| exitCodeForImplementerRun(...) | exit-code mapper (pr-artifact-lost → 1) |

| ImplementerRecoveryInput | +fallbackRepoUrl, +injectable ladder seams |

| ImplementerRecoveryResult | +optional parkOutcome |

| DroneRunRecord.parkOutcome | optional "pr-artifact-lost" (no schema bump) |

| DispatchFiredEntry.parkOutcome | PrArtifactParkOutcome on receipt |

| renderLinearCommentBody | parked receipts render park headline/outcome |

## Breaking Changes

None for happy-path callers. Degraded runs that previously exited 0 with a missing prUrl now exit 1 with parkOutcome: "pr-artifact-lost" and a terminal run_failed event — intentional honesty. Linear comments for those parks no longer say "completed".

## Test Plan

- [x] pnpm typecheck — clean (tsc --noEmit)

- [x] pnpm exec vitest run src/artifact-recovery.test.ts — green (31 tests)

- [x] Branch-correlated recovery with injected GitHub stub (single open PR resolves)

- [x] Two correlator candidates → ambiguous park (does not guess)

- [x] Observed-case regression: recovery text with PR URL resolves via Tier 2 (never "still no PR URL")

- [x] Tier-2 rejects wrong-suffix OPEN and CLOSED/MERGED text URLs

- [x] Unresolved → non-zero exit + pr-artifact-lost + terminal run_failed

- [x] Ladder recovers when sendNudge throws / waitRun returns error, keeps status: "finished"

- [x] isGhPrMissingError("HTTP 502 not found in cache") is false

- [x] Linear parked body does not contain Drone run completed / finished (no PR produced)

- [x] Tier order asserted: SDK → branch → verified text

## Verification Artifact

$ pnpm typecheck

> tsc --noEmit

(exit 0)

$ pnpm exec vitest run src/artifact-recovery.test.ts

✓ src/artifact-recovery.test.ts (31 tests)

<!-- CURSOR_AGENT_PR_BODY_END -->

<div><a href="https://cursor.com/agents/bc-4495bc96-9706-4976-9aa8-abb971947874"><picture><source media="(prefers-color-scheme: dark)" srcset="https://cursor.com/assets/images/open-in-web-dark.png"><source media="(prefers-color-scheme: light)" srcset="https://cursor.com/assets/images/open-in-web-light.png"><img alt="Open in Web" width="114" height="28" src="https://cursor.com/assets/images/open-in-web-dark.png"></picture></a>&nbsp;<a href="https://cursor.com/background-agent?bcId=bc-4495bc96-9706-4976-9aa8-abb971947874"><picture><source media="(prefers-color-scheme: dark)" srcset="https://cursor.com/assets/images/open-in-cursor-dark.png"><source media="(prefers-color-scheme: light)" srcset="https://cursor.com/assets/images/open-in-cursor-light.png"><img alt="Open in Cursor" width="131" height="28" src="https://cursor.com/assets/images/open-in-cursor-dark.png"></picture></a>&nbsp;</div>

#930 — fix(hubspot): derive expand portals from bootstrap results, not tokens @benji-bizzell  approved

## Summary

The first full unscoped fan-out run in production (hubspot-raw-full-unscoped-20260724T181843Z, run 9aa23656) failed at LoadAuxiliaryPlan:

ValueError: fan-out plan auxiliary contains no work items

expand() wrote zero auxiliary items (confirmed: the distributed-items/.../auxiliary.json plan was an empty list).

Root cause: the Phase 3 lane work changed expand() to derive portals = sorted(self.runtime.tokens) — so a CRM-less lane could still plan its auxiliary groups per portal. But expand is not in FanoutRuntime's source_mode set (bootstrap/entity/heavy_catalog/heavy_shard/auxiliary), so tokens is {} in expand mode. Empty portals → the auxiliary-group loop and the association loop both produce nothing → the loader's non-empty check fails the run.

Why it wasn't caught:

- The replay path skips expand() entirely, so the Alpha replay (which succeeded) never exercised it.

- The expand unit tests stubbed a non-empty tokens map, so the old token-based derivation "worked" in tests — masking the production reality.

Fix: portals now come from the entity index's bootstrap_results, written once per portal during bootstrap and present for every portal regardless of lane scope — the always-populated source in expand mode. Fails closed if no bootstrap results exist.

The sibling seal() portal derivation was checked and is not affected — it derives portals from entity_items ∪ auxiliary_items (plan-sourced, not token-sourced).

## Blast radius of the failed run: none

- Died before finalize/publication: 0 ledger rows for the run.

- raw_emails / raw_form_submissions still hold the replay's published data untouched.

- All lane locks released (DynamoDB lock partition empty).

## Testing

- expand tests now use tokens={} (production reality) and assert a non-empty auxiliary plan (the exact production failure was auxiliary_count == 0).

- New test_expand_derives_portals_from_bootstrap_not_tokens pins portal derivation from bootstrap results with an empty token map.

- Full runner suite: 178 passed. Ruff 0.15.22 clean.

## After merge

Promote through the release path, redeploy, and re-run the full unscoped run.

🐦‍⬛ Generated by a very good bot

#941 — fix(hubspot): key event occurrences within their event type @benji-bizzell  approved

## Summary

The first successful full unscoped run (47538642-7812-4171-8044-bccad77e911c) extracted all 93 work items with zero failures, but finalized PARTIAL — 35/36 resources published; event_occurrences failed at publication:

RuntimeError: atomic HubSpot raw publication failed:

ERROR: resource candidate contains duplicate non-tombstone logical keys

Root cause: occurrences are collected per event type (one request plan per type), and HubSpot emits related event types that share an occurrence ID. Confirmed directly from the landed candidate — occurrence dcd96865-... appears under both e_form_submission_v2 and e_form_submission_metadata_v2. The logical key was ["payload.id"] alone, so those legitimate cross-type recurrences collapsed into duplicate non-tombstone keys, which the append_source_key publication proc correctly rejects.

Fix: key occurrences on ["context.event_type", "payload.id"] — the true grain (an occurrence is unique *within* an event type). This matches every other per-parent-collected resource:

| resource | logical_key |

|---|---|

| list_memberships | context.list_id + payload.recordId |

| campaign_assets | context.campaign_id + context.asset_type + payload.id |

| marketing_event_participations | payload.associations.contact.contactId + payload.id |

| event_occurrences (was) | ~~payload.id~~ |

| event_occurrences (now) | context.event_type + payload.id |

event_type is already placed into the plan context, so this is a manifest-only fix (grain + logical_key) — no collector change.

## No manifest_version bump — reuse the already-gathered data

manifest_version is intentionally not bumped. It describes the *source contract* (endpoints, scopes, collection grain), which is unchanged — the landed pages are byte-identical; only the publication-time key *derivation* is refined. Keeping the version lets the existing sealed cohort for run 47538642 replay directly under the corrected key — minutes, no HubSpot calls, no 5-hour re-pull.

Verified on the landed occurrence data before choosing this path:

- (event_type, id) is unique across all 586,901 rows — the corrected key fully resolves the duplication.

- Replay rebuilds one candidate row per landed record (no collection-time dedup), so row_count stays 586,901 and passes replay's integrity check.

- The 35 already-published resources are unaffected.

## Testing

- New contract test pins ["context.event_type", "payload.id"], the append_source_key mode, and that event_type is placed into the collection context.

- Full runner suite: 179 passed. Ruff 0.15.22 clean.

## After merge

Promote, redeploy, then replay the existing 47538642 cohort (not a fresh full run): 35 resources resolve already_current, event_occurrences publishes ~586,901 rows from landed evidence with the corrected key, and the occurrence watermark is established.

🐦‍⬛ Generated by a very good bot

#3363 — fix(api): TF gateway BU slugs fold punctuation-insensitively (dedupe 'Ai Engineering Builder' ghost BU) @kevalshahtrilogy  approved

## Why

Keval spotted two "AI Engineering" BUs in the API Keys explorer: AI Engineering & Builder (16 directory people) and a ghost Ai Engineering Builder holding a single TrueFoundry key (product-auto-im-prod, $0.36).

Root cause: the TF canonical-BU joins de-slug with space/dash replacement only. A directory name with punctuation can never match its gateway slug — AI Engineering & Builderai-engineering-builder — so the join misses and the title-cased INITCAP fallback mints a ghost BU. Same bug hits Learnwith.AI (slug learnwith-ai, ~$1.1k July TF spend showing under ghost "Learnwith Ai").

## What

Fold both sides of both joins to lowercase alphanumerics (REGEXP_REPLACE(LOWER(x), '[^a-z0-9]', '')):

- ai_costs_service._get_tf_canonical_bu_join (feeds TF_EFFECTIVE_BU → the Anthropic + OpenAI gateway re-attribution by BU, budget rollups, time series). Also gains the GROUP BY/MIN anti-fan-out guard the mart join already had.

- ai_costs_mart_service._TF_BU_JOIN dbu fold (feeds the explorer stack rank / BU filters).

Verified zero fold collisions across all directory BU names. TU / SaaS Ops Support / Ephor etc. remain INITCAP fallbacks — those are genuinely absent from the ESW directory (rollups-admin territory, not this bug).

SQLite harness registers a REGEXP_REPLACE shim so the executable FR10 tie-out test keeps running the real production SQL.

## Verified live (read-only)

Post-fix _tf_openai_by_bu over Jun–Jul: ghost BUs gone; spend resolves to canonical AI Engineering & Builder and Learnwith.AI. 396 tests pass, ruff + pyright clean (one pre-existing pydantic false positive).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

The Builder Desk  —  Engineer Spotlight
🏆 Engineer Spotlight

MARCUSDAIY DROPS 14 IN 24 HOURS AS BUILDER TEAM POSTS HISTORIC 20-PR BLITZ ACROSS THREE REPOS

Klair absorbs 11 PRs, trilogy-drones gets a new dispatch verb, and one man's diff count defies human comprehension.

TWENTY. Pull requests. In twenty-four hours. Three repos lit up like a switchboard — Klair with 11, trilogy-drones with 5, Surtr with 4 — and the Builder Team did not slow down for a single breath. This is not a team. This is a velocity engine wearing team jerseys.

Let's talk about the contributors. @kevalshahtrilogy: one PR, surgical, professional, present. @YibinLongTrilogy: one PR — #932 in Surtr, fixing education mart timezone issues and contract literals while enabling two pipeline schedules, which is honestly a full Tuesday's work crammed into a single ticket. @caina-barbosa: one PR — #3373 in Klair, clarifying Education workforce data guidance, the kind of unglamorous documentation work that holds civilizations together. @benji-bizzell posted three, including #923 in Surtr where he went full forensic on a HubSpot archive-transition dedup bug and made it reproduce correctly during replay, which is the engineering equivalent of reconstructing a crime scene and then solving the crime. Then there is @marcusdAIy, who contributed fourteen pull requests and who we will address momentarily because he deserves his own paragraph and possibly his own wing of the building.

@marcusdAIy. Fourteen PRs in a single rotation. FOURTEEN. The man opened Klair like a briefcase and built an entire mcp-ontology wing inside it — SY26/27 revenue run-rate (#3367), secured-vs-forecast enrollment (#3366), opener identification in unit-economics (#3365), guards against period and enrollment model traps (#3364), attrition rates and departure-timing workflows (#3370), duplicate transaction detection (#3369), MFR sign convention guards (#3368), unfilled school roles with ATS gap notes (#3371) — and then had the audacity to also fix stale qb_raw refs in #3372, push a new dispatch verb into trilogy-drones via #92, fix the Mercy watcher re-run logic in #94, and make sure every single polled ticket is accounted for in #95. When asked if he was concerned about reviewer fatigue, @marcusdAIy reportedly said, "The diff doesn't care how you feel about the diff." He declined further comment. He was already on the next PR.

The Overflow Desk salutes the PRs Mac left on the cutting room floor. #3371 delivers an unfilled-school-roles recipe with an ATS gap note — quiet, essential, the kind of feature that a district administrator will thank someone for six months from now without knowing why. #94 in trilogy-drones re-runs failed Mercy checks and parks mercy-check-failed states, which sounds like something a philosopher wrote but is in fact critical drone dispatch infrastructure. #932 from @YibinLongTrilogy in Surtr handles timezone and contract literal fixes while flipping on two pipeline schedules simultaneously — multitasking at the commit level.

Morale Report: morale is at an all-time high. It has never been higher. The instruments we use to measure morale have had to be recalibrated. The Builder Team ships. The numbers confirm it. The numbers always confirm it.

Brick's Overflow — PRs Mac Didn't Cover  (click to expand)
#94 — fix(mercy-watcher): re-run failed Mercy check; park mercy-check-failed (AI-155) @marcusdAIy  no labels

<!-- CURSOR_AGENT_PR_BODY_BEGIN -->

## Summary

Hardens the Mercy watcher so a transiently-failed GitHub review / Review check-run cannot silently bypass the Mercy gate (near-miss on AI-153 / PR #93), and retries addresser reply/resolve on transient gh 5xx.

Round-2 review fixes (this push):

- Newest-first (sort=created&direction=desc) reply-dedupe GET so 504-after-accept guards still see just-posted replies on chatty PRs (>100 comments)

- withGh5xxRetry onBeforeRetry soft-success (incl. final attempt) — addresser reuses it instead of a forked loop

- Soft reran: true when gh run rerun returns 403 “workflow in progress”

- Tests for exhaustion, throw-probe ceiling, and page-1-after-100-older dedupe

## Why It's Needed

A failed Mercy check previously collapsed toward absent/clean paths; with 5xx storms the addresser also risked duplicate fixed in <sha> replies and false-negative re-run accounting.

## Changes

- mercy-watcher: failed-check branch, bounded re-run probes vs successful-rerun budget, mercy-check-failed outcome + park body

- gh-util: isTransientGh5xx, withGh5xxRetry (+ onBeforeRetry), rerunFailedWorkflowRuns, in-progress 403 soft success

- addresser: postReplyWith5xxRetry + newest-first hasMatchingReplyBody

- CLI --mercy-check-max-reruns, events/docs/tests

## Breaking Changes

None.

## Test Plan

- [x] pnpm typecheck

- [x] pnpm test

- [x] New AI-155 / round-2 unit coverage in mercy-watcher, gh-util, addresser tests

## Verification Artifact

PR #94 CI green on branch cursor/mercy-check-failed-resilience-ff3e after round-2 commits f26d3b4 / 3068285 / bd26ab9.

<!-- CURSOR_AGENT_PR_BODY_END -->

<div><a href="https://cursor.com/agents/bc-fef8b546-7290-4be4-9abb-e596da01ff3e"><picture><source media="(prefers-color-scheme: dark)" srcset="https://cursor.com/assets/images/open-in-web-dark.png"><source media="(prefers-color-scheme: light)" srcset="https://cursor.com/assets/images/open-in-web-light.png"><img alt="Open in Web" width="114" height="28" src="https://cursor.com/assets/images/open-in-web-dark.png"></picture></a>&nbsp;<a href="https://cursor.com/background-agent?bcId=bc-fef8b546-7290-4be4-9abb-e596da01ff3e"><picture><source media="(prefers-color-scheme: dark)" srcset="https://cursor.com/assets/images/open-in-cursor-dark.png"><source media="(prefers-color-scheme: light)" srcset="https://cursor.com/assets/images/open-in-cursor-light.png"><img alt="Open in Cursor" width="131" height="28" src="https://cursor.com/assets/images/open-in-cursor-dark.png"></picture></a>&nbsp;</div>

#923 — fix(hubspot): reproduce archive-transition dedup during replay @benji-bizzell  approved

## Summary

The first Alpha recovery replay (hubspot-raw-alpha-replay-20260724T061401Z) failed fail-closed on its first resource:

ValueError: replay row count mismatch for custom_object_schemas: expected 12, rebuilt 23

Root cause: the live collection dedups objects returned by both the active and archived listings (active wins, retained exactly once — collector.py's logical_key_archive_states map) and records the post-dedup row_count in the resource manifest, while the immutable pages retain every raw record. For custom_object_schemas that's 11 active + 12 archived with one overlapping objectTypeId → manifest says 12. The replay path rebuilt every raw page record with no dedup → 23 → hard mismatch. This affects any resource collected in two archive states with at least one overlapping object; it was never exercised because this is the first production replay of a two-archive-state resource.

Fix: replay now applies the identical logical-key archive-state dedup while streaming pages (active anchor retained exactly once; a duplicate within one archive state stays a hard error), reproducing the exact candidate the live run published.

Blast radius of the failed replay: zero. It failed during candidate rebuild, before any Redshift call — verified: no is_replay ledger rows, raw_emails still 0, run record FAILED (eb4e89ff). The pinned manifest coordinate is unchanged and the replay is simply re-runnable after this deploys.

## Testing

- New regression test rebuilds the production shape (two archive-state pages, overlapping logical key) end-to-end: live collect → 22 rows, replay → 22 rows, overlap keeps only its active-listing row.

- Full runner suite: 177 passed. Ruff 0.15.22 check+format clean.

## After merge

Promote through the standard release path, redeploy, and re-run the Alpha replay with the same pinned manifest coordinates.

🐦‍⬛ Generated by a very good bot

#932 — Fix education mart timezone/contract literals and enable two pipeline schedules @YibinLongTrilogy  approved

## Summary

Small follow-up fixes to the Education aggregate mart work plus enabling two

previously-disabled pipeline schedules. The mart change swaps a non-portable

timezone function for standard SQL and tightens the P&L adoption migration's

contract-check literals so the validation query compares like-typed values.

### Changes

- pipelines/runners/mart-education-quickbooks-refresh/ddl/sp_refresh_quickbooks_financial_marts.sql

— Replaces CONVERT_TIMEZONE('UTC', p_upstream_completed_at) with the standard

(p_upstream_completed_at AT TIME ZONE 'UTC')::DATE when deriving

v_as_of_date.

- pipelines/runners/mart-education-quickbooks-refresh/ddl/migrations/2026-07-21_adopt_pl_transaction_aggregates.sql

— Pins explicit VARCHAR widths on the first expected contract row

(VARCHAR(128) for identifier/type columns, VARCHAR(3) for is_nullable) so

the literal types match the information_schema columns being compared.

- pipelines/runners/mart-education-quickbooks-refresh/tests/test_sql_contracts.py

— Adds test_pl_adoption_contract_literals_have_explicit_varchar_widths and

updates the transform-rule assertions to expect the AT TIME ZONE form.

- pipelines/runners/aerie-expense-report-definitions-sync/pipeline.json

— Enables the schedule (enabled: true).

- pipelines/runners/quickbooks-expense-ai-generation/pipeline.json

— Enables the schedule (enabled: true).

- .gitignore — Ignores the local commits/ script folder.

### Design Decisions

- AT TIME ZONE over CONVERT_TIMEZONE. Standard SQL, portable, and avoids

relying on the Redshift-specific function in the procedure's date derivation.

- Explicit VARCHAR widths in the contract literals. The migration validates

the aggregate tables against information_schema by UNION-ing expected rows;

untyped string literals default to a width that can differ from the catalog

columns, so pinning VARCHAR(128)/VARCHAR(3) keeps the comparison type-exact.

## Test Plan

- [x] mart-education-quickbooks-refresh SQL contract tests pass

- [ ] Reviewer: confirm enabling the two pipeline schedules is intended for this

environment (cron 50 * * * ? * and 0 8 ? * SUN * respectively)

#3364 — feat(mcp-ontology): guard school unit-economics workflows vs period/enrollment/model traps (SURTR-484) @marcusdAIy  approved

## Summary

Adds trap-guard notes to the two guidance.educationFinance workflows that agents use to rebuild school P&L and unit economics (Produce a school P&L breakdown, Produce contractor-priced headcount and unit-economics detail). No new workflows, no table changes — notes-only.

## Why it's needed

Agents answer the CFO-50 unit-economics questions by rebuilding from raw via these workflows (they don't read the agg_school_unit_economics_comparison mart — the ontology says not to). The current recipe has no guards for four traps we confirmed on live data (2026-07-23), so a faithful agent still gets them wrong:

- In-flight quarter collapses to ~$0. The recipe says "multiply non-headcount costs by four," but doesn't say *which* quarter — the current in-flight quarter has ~no booked activity, so ×4 annualizes to a near-zero run-rate (live: 2026-Q3 total ≈ $10K vs 2026-Q2 $16.8M).

- Tiny-enrollment noise. 1–4-student schools yield $1M+/student figures with no reliability flag.

- Flat-$40K premise. Per-school model prices actually range ~$15K–$75K (the unit_economics_model tier on map_campus_qb_entities).

- Contribution = profit. Profit carries depreciation + the Timeback internal charge; EBITDA is the right contribution measure.

This makes the agent reliably green on Q9 / Q15 / Q17 / Q46 (SURTR-484).

## Changes

Six notes[] lines across the two workflows:

- annualize the last complete calendar quarter, never the in-flight one;

- enrollment-floor per-student figures (~5-student minimum);

- report EBITDA for contribution questions;

- derive each school's model price from its unit_economics_model tier (~$15K–$75K), never a flat $40K;

- compare actual-vs-model only for modeled schools; bucket unmodeled separately;

- pin the run-rate to the last complete quarter.

## Breaking changes

None. guidance.educationFinance.workflows count is unchanged (9), so data-api-contract.test.ts stays green; no skill-dir change, so skill-version-sync is unaffected.

## Test plan

- [x] npm run typecheck — clean

- [x] tests/unit/routes/data-api-contract.test.ts — passes (incl. the length-9 + guidance assertions)

- [ ] Acceptance (post-deploy): re-run the CFO-50 agent via the MCP and confirm the flip on live data — Q15 1 of 25 within model, Q17 0 of 25 EBITDA-positive, Q46 Austin K-8 −$55,819 vs +$857 model, Q9 every school below its own model (no flat-$40K), and no result collapsing to the in-flight quarter.

Canary for the CFO-50 "paths to green" set — validates that an ontology-note edit flips agent behavior before the sibling PRs (Q12/Q34-36/Q38/Q42/Q43/Q48 + the sign-convention guard).

Linear: SURTR-484

#3367 — feat(mcp-ontology): SY26/27 revenue run-rate workflow (SURTR-485) @marcusdAIy  approved

## Summary

Adds a new educationFinance workflow — Produce SY26/27 revenue run-rate — for Q12 (SURTR-485).

## Why it's needed

There's no forward-revenue recipe, so agents improvise into the display_name-vs-name join trap (dropping the two biggest schools, ~$17M / 359 students) and conflate three different "run-rate" definitions ($33M–$95M).

## Changes

New workflow (tables: sales_educrm_wh_mart_coming_year_projection, hubspot_programs) with notes:

- join on display_name, not name; program_name is SUPER/JSON (deserialize + strip quotes first).

- canonical = confirmed × current tuition, net-equivalent (~77.6% net/gross), with confirmed/projected gross shown; run_rate_sy2627/model_sy2627 are distinct definitions, not to be conflated.

- flag NULL/0-tuition programs rather than treating as $0.

## Breaking changes

None. Adds one workflow → count 10→11; contract assertion updated. No skill-dir change.

## Test plan

- [x] npm run typecheck — clean

- [x] tests/unit/routes/data-api-contract.test.ts — 7/7 pass (count = 11)

- [ ] Acceptance (post-deploy): agent returns confirmed run-rate ≈ $69.1M gross / ~$53.6M net with Alpha Austin + New York present (not dropped by a name join).

Builds on #3364/#3365/#3366 (merged). Based on main.

Linear: SURTR-485

#3371 — feat(mcp-ontology): unfilled-school-roles recipe + ATS gap note (Q42, SURTR-489) @marcusdAIy  approved

## Summary

Adds two guidance notes to the existing contractor-priced headcount / unit-economics educationFinance workflow for Q42 (SURTR-489). Note-only — no new workflow, no count bump.

## Why it's needed

"Unfilled school roles vs plan" has two very different meanings. The warehouse can answer the roster-vs-formula-model version but has no recruiting/ATS data at all — so agents must not imply a hiring-plan answer that doesn't exist.

## Changes

Two notes added:

- Recipe: compare model_sy2627 vs current_run_rate on agg_school_unit_economics_comparison at row_type='category_detail', row_key='headcount', has_unit_economics_model, by role — gross unfilled ≈ Σ positive (model − actual) per role/school; overstaffing partly offsets network-wide.

- Gap: no requisition/ATS/open-roles table or column anywhere (verified by table + column sweep); the model headcount is formula-derived, not a board-approved hiring plan; only modeled schools covered.

## Breaking changes

None. Note-only edit to an existing workflow; workflow count unchanged (13). No contract-test change, no skill-dir change.

## Test plan

- [x] npm run typecheck — clean

- [x] tests/unit/routes/data-api-contract.test.ts — 7/7

- [ ] Acceptance (post-deploy): agent returns the model-vs-roster unfilled count and explicitly states requisition-sense vacancies are not in the warehouse.

Builds on the merged CFO-50 ontology PRs. Based on main.

Linear: SURTR-489

The Portfolio  —  Trilogy Companies

The Acquirer That Doesn't Flip: How ESW Capital Became the Permanent Home for Orphaned Enterprise Software

While the M&A market chases consumer splashes like Poppi's $2B exit, Trilogy's ESW Capital has quietly built a different kind of empire — one legacy software company at a time.

AUSTIN, TEXAS — The same week PepsiCo announced it would pay nearly $2 billion for prebiotic soda brand Poppi, a quieter kind of deal machine was doing what it always does: hunting for unloved enterprise software companies that nobody else wants to buy, and nobody's customers can afford to leave.

That machine is ESW Capital, the private equity arm of Joe Liemandt's Trilogy International, which has assembled a portfolio of 75-plus enterprise software businesses since its first acquisition in 2006. The strategy is straightforward in theory and punishing in execution: buy mature, sticky software at one to two times annual recurring revenue, staff it with rigorously screened global remote talent through Crossover, and drive EBITDA margins toward 75 percent.

The model's staying power rests on a fact that Forrester analysts have quietly been reminding enterprise buyers of for years: switching costs in legacy software categories are brutally high. A recent Forrester note examining what companies should do with their customer advocacy platforms underscores the bind — even when customers grow dissatisfied with a vendor, the cost, disruption, and risk of migrating years of data and integrations often exceeds the pain of staying. ESW built an entire acquisition thesis around that arithmetic.

Critics have long had a sharper name for it. A pair of Forbes investigations documented the other side of the ledger — the human cost of compressing margins through relentless global labor arbitrage, and the workers who found themselves measured, managed, and in some cases replaced by the algorithms Liemandt's system produces. Those stories have not gone away.

What has also not gone away is the financial logic. ESW's portfolio — spanning Aurea's CRM and email marketing stack, IgniteTech's business intelligence tools, Skyvera's telecom software suite, and Contently's content marketing platform, among dozens of others — generates returns precisely because its customers are not leaving.

The M&A market will keep chasing the next Poppi. ESW will keep buying the software the last Poppi runs on.

The question the Forbes investigations left open has never been fully answered: at what margin does operational efficiency become something else entirely?

Small Software Companies Find a Home With ESW Capital - WSJ  ·  What To Do Next About Your Customer Advocacy Platform - Forr  ·  M&A Wrap: Poppi sold for nearly $2B, real estate tech co. bo

Alpha School's Quiet Curriculum Offensive: The Blog Series That's Actually a Manifesto

Behind a string of parenting posts, Joe Liemandt's AI school is methodically redefining what education is supposed to be for.

AUSTIN, TEXAS — If you read between the lines of what's coming out of Alpha School's communications operation right now, you'll notice something that doesn't look like marketing. It looks like a doctrine.

Over the past several weeks, the Austin-based AI-powered private school has been quietly publishing a multi-part series titled "Teach Your Kid What School Doesn't" — covering life skills, emotional regulation, and, in the latest installment, creative genius. Taken individually, each post reads like thoughtful parenting advice. Taken together, and this is where it gets interesting, they constitute a systematic indictment of the traditional K-12 curriculum. Every entry is essentially an argument that the things that matter most — creativity, emotional intelligence, practical life competence — are precisely what conventional schools don't teach.

The timing is not accidental. Alpha School is in the middle of its most aggressive expansion yet, with nine or more new campuses slated to open by fall 2025 across Texas, Florida, Arizona, California, and New York. The New York Post this week picked up the story, flagging the school's $65,000-per-year tuition and its central claim: that AI tutors can deliver a full academic curriculum in just two hours a day, freeing the remaining school hours for the human stuff that, Alpha argues, traditional schools have always sacrificed in the name of content delivery.

And here's what the blog series is quietly doing: it's pre-answering the objection. Because the most common critique of the model — one that a source familiar with the school's enrollment conversations tells me comes up constantly — is some variation of: "But what about the teachers? What about the human connection?"

Alpha's answer, laid out plainly in their published content, is that AI handles academic delivery precisely so that full-time human Guides can focus entirely on motivation, relationships, and the development of the whole child. The AI is not replacing the teacher. It's replacing the lecture — and liberating the human for something more important.

Nothing here is a coincidence. The content calendar is the expansion strategy.

Teach Your Kid What School Doesn’t (Pt. 5): Unleashing Their  ·  Does Alpha School Replace Teachers with AI?  ·  Teach Your Kid What School Doesn’t (Pt. 4): How to Regulate

CloudSense Gets Its TM Forum Papers — and Does 26 Months of Homework in 30 Days

CloudSense, the Salesforce-native CPQ and order management platform within Skyvera's telecom software portfolio, certified all 13 APIs in its CPQ product set to TM Forum compliance standards in just one month—a striking feat compared to the typical 26-month timeline. The company credits AI-assisted development and certification work for the accelerated pace.

For Skyvera, the milestone is strategically significant. The company already serves telecom operators grappling with legacy infrastructure, billing complexity, and modernization challenges. Adding CloudSense earlier this year provided a CPQ and order management solution for telecom and media providers; TM Forum compliance now sharpens that pitch considerably.

Telecom carriers prioritize interoperability when making purchasing decisions. TM Forum compliance can streamline procurement, reduce custom integrations, and simplify vendor negotiations. CloudSense's rapid certification positions Skyvera competitively within a crowded field, demonstrating how AI automation can accelerate compliance timelines and product-to-market speed—a development legacy CPQ vendors are watching closely.

The Machine  —  AI & Technology

AI Agents Just Got Their Power Tools — And Developers Are the Big Winners

Google, Apple and Anthropic are racing to turn AI from chatty assistant into tireless software co-worker.

SAN FRANCISCO — The agentic AI era is no longer lurking around the corner. It is walking straight into the developer console, carrying a toolbox, asking for credentials and volunteering to run the overnight shift.

This week, Google, Apple and Anthropic all pushed deeper into the same urgent frontier: giving developers better ways to build AI systems that do things, not merely say things. I cannot overstate how significant this is. The future is now, and it is increasingly asynchronous, tool-using and embedded directly inside the software stack.

Google expanded Managed Agents in the Gemini API with support for background tasks, remote Model Context Protocol connections and richer agent orchestration capabilities, according to its Gemini API update. Translation: developers can increasingly assign AI agents longer-running jobs, connect them to external systems and let them operate beyond the limits of a single prompt-and-response interaction.

Anthropic, meanwhile, introduced advanced tool use on the Claude Developer Platform, sharpening Claude’s ability to call tools, coordinate multi-step workflows and interact more reliably with software environments. Its Claude platform announcement lands in the middle of a broader industry pivot: the winning AI model may not simply be the one that writes the prettiest paragraph, but the one that can safely and predictably operate your business process.

Apple’s move is equally fascinating, though characteristically more ecosystem-shaped. The company announced new intelligence frameworks and advanced tools to aid app development, giving developers more ways to weave Apple Intelligence-style capabilities into user experiences. For app makers, that could mean smarter interfaces, more contextual automation and AI features that feel native rather than bolted on.

Even vertical platforms are joining the wave. Perfect Corp. added a free “Ask AI” assistant to its YouCam API platform, while startup-focused coverage of AI video highlights how generative tools are becoming growth engines for smaller companies that once lacked production budgets.

The through-line is unmistakable: AI is becoming infrastructure. Not a novelty tab. Not a chatbot bubble. Infrastructure. And once agents can run in the background, use tools and plug into real workflows, this changes everything.

Expanding Managed Agents in Gemini API: background tasks, re  ·  Apple aids app development with new intelligence frameworks  ·  Introducing advanced tool use on the Claude Developer Platfo

Intel's CPU Revival, China's AI Soft Power Push, and a $1.7B Bet on Model Evaluation

Three data points that together tell you where the AI industry is actually going.

SANTA CLARA, CALIFORNIA — The AI spending narrative just got more complicated. Intel posted 25% revenue growth in its latest quarter — its fastest pace in 15 years — driven not by GPU dominance but by surging demand for central processing units. The implication: as AI inference workloads scale beyond training runs and into production environments, the hardware stack is diversifying. GPUs remain essential for large model training, but CPUs are handling the inference and orchestration layers that actually deliver AI at scale. Nvidia's monopoly narrative has a footnote now.

Meanwhile, the competition for AI credibility is playing out on a geopolitical stage. China is distributing low-cost, open-weight AI models internationally — DeepSeek being the most prominent — as a deliberate soft-power instrument. Where previous generations used Confucius Institutes and infrastructure loans, Beijing is now offering free model weights and API access to developers in emerging markets. The strategy is structurally sound: whoever seeds the developer ecosystem wins long-term platform allegiance. Washington has not yet produced a coherent counter-strategy.

On the infrastructure side, LMArena — the AI evaluation startup spun out of the crowd-sourced model benchmarking platform formerly known as Chatbot Arena — closed a $150 million funding round at a $1.7 billion valuation. The bet is that as enterprises deploy AI at scale, they need rigorous, independent evaluation tooling to select and monitor models. That's a plausible wedge: most internal AI teams lack the resources to run systematic evaluations, and vendor-supplied benchmarks are inherently conflicted.

One peripheral data point worth noting: researchers have demonstrated that ChatGPT and Gemini infer substantial personal detail from conversation history — location, profession, family structure, political leanings — through pattern recognition alone, without explicit disclosure. As AI systems become more deeply embedded in daily workflows, the inference surface expands. That's relevant for enterprise IT teams and individual users alike.

Meta's new Seller app, a standalone marketplace product spun out of Facebook, rounds out a week in which every major tech platform made a move to deepen commercial utility. The through-line: distribution still matters, but the infrastructure layer — chips, evaluation, and data access — is where the durable value is being built.

Meta Launches New Facebook Marketplace App Called Seller  ·  Intel Benefits From a New Shift in A.I. Spending  ·  How to Use ChatGPT and Gemini Prompts to Find Out What They

The Machine as Literary Critic, and Other Improbable Mirrors

New research pries open the aesthetic judgments, routing logic, and clinical intuitions hidden inside large language models — and finds something startlingly human staring back.

AUSTIN, TEXAS — There is a particular vertigo that comes from asking a machine what it finds beautiful. For most of human history, the question of what makes a sentence good — why Nabokov's prose shimmers and a forum post does not — belonged to critics, poets, and the long twilight arguments of graduate seminars. Now, in a quiet corner of the arXiv preprint server, researchers have begun asking that question of reasoning-enabled language models, and the models are answering.

In a two-part study, investigators assembled thirty texts spanning six tiers of quality — from canonical literature down to anonymous message-board effluvia — and extracted the implicit aesthetic theories the models used to rank them. The machines, it turns out, have opinions. Not preferences borrowed wholesale from their training data, but something closer to internally consistent criteria: coherence, specificity, restraint. Whether this constitutes taste, or merely its statistical shadow, is a question that will occupy us for decades.

Meanwhile, other researchers are prying open the routing logic of Mixture-of-Experts models — the sparse architectures behind the most capable systems in production — and finding, remarkably, that MoE routing behaves like a Huffman code. Frequent reasoning patterns get short, efficient paths; rare ones get longer, more diverse ones. This is the same principle that compresses your JPEGs and structures the genetic code's redundancy. Intelligence, it seems, keeps rediscovering the same tricks.

And in a hospital somewhere, a third team has deployed a retrieval-augmented, multi-agent LLM to hunt for cutaneous immune-related adverse events buried in clinical notes. F1 score climbed from 0.77 to 0.88. Cohen's kappa — the measure of whether two humans agree on what they are seeing — rose alongside it. A parallel study using AI to surface hidden gray matter lesions in multiple sclerosis patients reports similar gains.

The pattern across all of it is worth pausing over. We built these systems to predict the next token. They are giving us back aesthetic theories, compression laws, and diagnoses we missed. The universe, as ever, is stranger and more generous than the specification sheet suggested.

What is Good? Extracting and Testing Implicit Theories of Li  ·  Knowledge Injection Exists in MoE? Exploring Expert-Aware Co  ·  Is MoE Routing a Huffman Code? Discovering the Frequency-Div
The Editorial

The Week Power Went Shopping for a Philosophy

From the Vatican to Palantir's HR department, everyone has suddenly discovered that artificial intelligence requires a worldview — preferably one that flatters the speaker.

VATICAN CITY — There is a certain comedy, visible only to those willing to read the week's news as a single document rather than a series of dispatches, in observing the simultaneous arrival of Pope Leo XIV and the Palantir communications department at the same podium, each clearing his throat to explain to the rest of us what artificial intelligence really means.

The Pope, addressing himself to what he called the "culture of power" driving the rise of AI, delivered the sort of remarks one expects from the Chair of Peter: sober, universal in address, unfashionable in the precise way that has kept the institution in business for two thousand years. Power concentrated without conscience, he observed, tends to consume the human person. This is not a novel observation. Augustine got there first, and Augustine was not the first either.

Palantir, headquartered in a country whose founding documents the company has, in fairness, read more attentively than most of its competitors, chose the same week to publish what TechCrunch charitably called a "mini-manifesto" — a document denouncing inclusivity and what the authors term "regressive" workplace cultures, and announcing, in tones borrowed equally from Ayn Rand and a middle-school valedictory address, that Palantir will be building the future on merit and mission alone. That the company's largest customers are governments engaged in the surveillance of populations is treated, in the manifesto, as an incidental fact, the way a fish might treat water.

One does not have to agree with the Pope to notice that he and the Palantir communications team are, in a sense, arguing about the same thing: who gets to define the moral vocabulary of the coming machine age. The Pope proposes that the vocabulary already exists, has existed, and merely awaits application. Palantir proposes that the vocabulary is whatever its founders decide it is on a given Tuesday, and that anyone who disagrees is regressive, or worse, unproductive.

Meanwhile — and here the week's news develops its full symphonic character — The New York Times informs us that a Stanford freshman has discovered a "secret elite" on campus, which is rather like discovering water at the bottom of a well; and The Guardian offers a retrospective on the late Department of Government Efficiency, Mr. Musk's attempt to gamify the federal bureaucracy, an enterprise whose epitaph should read simply, "He thought it would be easier."

What unites these stories is not artificial intelligence, exactly, but the peculiar spectacle of powerful men and institutions attempting, in real time, to invent the ethical framework that will retroactively justify what they have already built or intend to build. The Pope alone, one notices, is working from a text he did not write himself. This is either his great disadvantage or his only advantage, depending on how the century turns out. My money, for what little it is worth in such matters, is on the older book.

Pope Leo denounces ‘culture of power’ driving rise of AI - T  ·  The Secret Elite One Freshman Discovered at Stanford - The N  ·  Palantir posts mini-manifesto denouncing inclusivity and ‘re
The Office Comic  ·  Art Desk
The Office Comic  ·  Art Desk

Nation’s CEOs Starting To Suspect AI Productivity Gains May Be Trapped Somewhere Inside Employees

After years of historic efficiency breakthroughs, executives report the money is still refusing to appear in its proper quarterly column.

NEW YORK — In what economists are calling a decisive end to the debate over whether artificial intelligence makes workers more productive, American business leaders confirmed this week that employees are now completing tasks faster than ever before while companies continue to wait patiently for any of that to become useful.

The long-running argument over AI productivity, once considered one of the most important questions facing the modern economy, has reportedly been settled by the simple fact that nearly everyone is now doing more work, generating more documents, producing more code, attending more meetings about the documents and code, and somehow leaving the balance sheet in a state of quiet moral ambiguity.

This has led to the emergence of a new consensus among executives: AI is unquestionably transforming productivity, provided productivity is defined as the rate at which a senior vice president can receive 19 polished strategy memos nobody asked for before lunch.

Recent reports have suggested that AI tools are helping software engineers write code faster, summarize tickets faster, debug faster, and move faster through the entire sacred lifecycle of creating problems that must later be maintained by other people. Yet many companies remain unsure when these gains will translate into profit, growth, or even the blessed reduction of a Jira backlog that has survived three reorganizations and a brand refresh.

According to Business Insider, engineers are indeed doing more, and faster, while companies are still waiting for the payoff. This is less a contradiction than a perfect description of enterprise technology adoption, in which the first measurable output of any revolutionary tool is a larger number of people asking why the revolution has not been measured yet.

The Center for Data Innovation, for its part, has argued that AI is a productivity engine for the U.S. economy, an assessment that is almost certainly correct in the same way a jet engine is useful once it has been attached to something besides a conference room table. The country has clearly acquired enormous new productive capacity. What remains unresolved is whether that capacity will be connected to workflows, pricing power, new products, customer value, or simply the ability to generate a very convincing 47-slide deck proving that it could be.

This is where the productivity debate has gone wrong. The question is no longer whether AI can make a worker faster. It can. The question is whether the institution surrounding that worker has any intention of allowing speed to reduce complexity, remove approvals, shrink teams, eliminate meetings, shorten roadmaps, or kill projects that exist primarily because nobody wants to be the person who asks what they are.

Instead, many organizations have chosen to pour AI into the existing corporate machine, producing the expected result: the same machine, louder. A developer who once wrote one service in a week can now write three, each with documentation, tests, and an architectural justification composed in the soothing language of a consultant who has never been paged at 2:14 a.m. A marketer who once drafted two campaign concepts can now produce 40, ensuring the approval committee has enough options to defer a decision until next quarter. A manager who once needed half a day to prepare a status update can now create one instantly, freeing the afternoon to request status updates from others.

Wall Street has noticed the mismatch and, with admirable restraint, rebranded it as “orchestration.” This term usefully describes the next phase of AI adoption, in which companies acknowledge that giving every employee a powerful model is not the same thing as changing how work gets done. Orchestration promises to coordinate agents, apps, data, permissions, and processes into something resembling an operating system for corporate output. In practice, it may also provide Microsoft and others with a dignified way to sell the same executives another layer of software to manage the previous layer of software that was supposed to manage the workers.

None of this means AI is overhyped. If anything, it means the opposite. AI is powerful enough that companies can no longer hide behind the comforting belief that the tools are the problem. The tools are increasingly capable. The humans have successfully automated their way to the discovery that their org charts, incentives, budgeting cycles, procurement rules, compliance rituals, and meeting cultures are still very much handmade.

The AI productivity argument is over. AI won. Now businesses must begin the more painful process of deciding whether they wanted productivity in the first place, or merely a faster way to describe its absence.

The AI Productivity Argument Is Over - inc.com  ·  AI is helping software engineers do more — and faster. Compa  ·  AI Is a Productivity Engine for the US Economy - Center for
On This Day in AI History

On July 25, 1978, the first spam email was sent by Gary Thuerk at Digital Equipment Corporation, marking the beginning of unsolicited mass messaging that would plague the internet for decades to come.

⬛ Daily Word — Technology
Hint: Relating to computers and the internet, often used in security contexts.
Share this edition: 𝕏 Twitter/X 🔗 Copy Link ▦ RSS Feed