Theseus Capital · v1.4 · 2026-08-19
Theseus Capital: Institutional Capital thesis, v1.4 (canonical; sources pinned; red-teamed; tokenmaxxing mechanism added)
When intelligence commoditizes, own everything it cannot manufacture: standing, atoms, and meaning.
This document is the standing leg of that sentence. Natural Capital (atoms) and Taste Capital (meaning) are its siblings; the three are one thesis expressed in three asset classes.
Status: canonical. Supersedes thesis.md, the contrarian working notes, and draft_thesis.md as the statement of record; those remain as research lineage. Every claim below survived a structured adversarial process: the original dialectic (Appendix A) plus two independent red-team reviews whose accepted findings are integrated in this version. Every factual figure is pinned to a primary source in the Source Register (Appendix B), verified August 18, 2026; inline tags like [B.5] point to register entries. Supporting apparatus: named public baskets, a scored venture pipeline with deliberate passes, an elasticity evidence review with pre-registered kill criteria, a counsel engagement brief, the two-pager, and the impact methodology, all in apparatus/. The one remaining gate on external use is fund counsel's review of the legal-ownership question (B.9).
As the marginal cost of useful machine cognition falls toward its physical floor, economic rents migrate away from entities that sell intelligence and toward entities that own what intelligence cannot manufacture: regulatory standing, liability-bearing capacity, embedded distribution, proprietary outcome data, and trust as the party of record. The largest single destination of the surplus is the customer, and the expansion of demand that this creates is where new institutions get built.
Three investable corollaries, in descending order of current evidential support:
What this thesis deliberately does not require: artificial superintelligence, a winner-take-all lab, or any particular resolution of the AGI debate. Every essential input is measurable today and tracked quarterly (the Instrument Panel).
The thesis rests on one empirical trend, so the trend must be over-justified. It is not one curve; it is five stacked mechanisms, all pushing the same direction. They are not fully independent: mechanisms two through five share two chokepoints, compute/energy supply and Chinese open-weight policy, and a single decision at either chokepoint can perturb several mechanisms at once (DeepSeek's August 2026 repricing hit distillation economics and the open-weight price cap simultaneously). Both chokepoints are therefore named watch items, not assumptions.
The composite result: the price of a fixed capability tier has fallen roughly two-and-a-half orders of magnitude in four years (≈436x at a uniform 3:1 input:output blend) and continues falling, with the rate varying dramatically by task: Epoch AI measures the decline at 9x to 900x per year depending on the performance milestone [B.1, B.2]. Epoch's own caveat travels with the number: the fastest declines in that range occurred in the most recent year measured, and their persistence is not guaranteed.
Figure 1. The price of a fixed capability tier collapses on a schedule: GPT-3 davinci $60/M (2021) to GPT-5 Nano $0.14/M (2025), log scale; ≈436x decline in ~4 years at a uniform 3:1 blend. Sources pinned in Appendix B.2.
The precise claim, stated so it can be attacked. "Near-zero marginal cost of intelligence" is a slogan, not a unit. The precise version:
For any given workflow, the price of the cheapest model that clears that workflow's quality threshold trends toward the electricity-plus-capital cost of inference.
This makes the economics per-workflow, not per-model. And per-workflow is where the money is. Once any sufficiently cheap model crosses a workflow's required capability, the frontier model's remaining premium is worth only its incremental tail-risk reduction, not its benchmark superiority.
Figure 2. Capability saturation: why the frontier premium collapses per workflow. Stylized mechanism diagram; past the point where the cheapest model crosses the capability the workflow requires, the frontier's quality premium has no marginal economic value.
A workflow is only economically transformed when it crosses four thresholds, not one:
| Threshold | Question |
|---|---|
| Technical | Can the model do the task in evaluation? |
| Economic | Is it cheaper after review, error, and liability costs? |
| Institutional | Is the provider licensed, insured, and trusted to deliver it? |
| Adoption | Do customers actually switch? |
Most AI demos cross the first threshold years before the other three. The fund's entire diligence discipline is refusing to pay for threshold one.
Objection (Jevons / compute scarcity). Token demand is growing faster than compute supply; power and chips are hitting physical ceilings; therefore prices may firm, not collapse. Reply. Conceded, for total spend, which is exactly why aggregate AI capex can boom while this thesis holds. But the thesis is priced on cost per fixed capability, which falls even while total spend explodes (that is what Jevons means). The binding physical constraint migrating to energy and silicon is not a leak in this leg of the stool; it is the Natural Capital leg's entire long case. The stool hedges itself.
Objection (tail risk). Professional services are priced on reliability and tail-risk avoidance, not average performance. A cheap model that fails rare, severe cases is economically worse than an expensive one. Reply. Conceded, and absorbed as a segmentation rather than a refutation: low-risk, easily verified work commoditizes first; liability-bearing work commoditizes only as verification cost falls and reliability at threshold is demonstrated on domain-realistic evals. This ordering is itself a prediction the instrument panel tracks (verification-cost ratio). If verification cost stays proportional to output volume, autonomy never gets cheap, only fast, and the thesis fails. We say so plainly and measure it.
Objection (the plateau). Inference-time reasoning and post-training RL are a new scaling regime; maybe it saturates, capability stalls, and the collapse stops. Reply. A capability plateau does not stop the price collapse: mechanisms 1-5 keep cheapening whatever capability exists (the four-year price series in Fig 1 required no capability miracle, only cost engineering). A plateau does cap which workflows ever cross their thresholds, which shrinks the thesis's addressable set. That is a narrowing scenario, not a kill: the Falsification section carries it as a named falsifier with a dated tripwire.
The consensus AI thesis says: models commoditize, therefore the application layer captures the value. The conclusion does not follow from the premise, because the same force commoditizes both:
And a third mechanism compounds the two: unhobbling. Context, memory, tool use, computer use, retrieval, planning, orchestration (the historical reasons applications existed) keep becoming platform primitives. A useful compression: economically deployable capability E = I x U x A (intelligence x unhobbling x accessibility). Even if I stalls, U and A keep absorbing the application layer's feature surface from below.
The result is not that "the app layer dies." It polarizes:
Objection (the enterprise-software steelman). Most enterprise software value was never code. It is procurement trust, integration, workflow adoption, auditability, someone to sue. Those do not commoditize with tokens. Reply. Correct, and this objection is the thesis, one level down. Verification, absorbed liability, audit trails, and being-the-party-of-record are institutional functions. A vendor renting them out is an institution with worse economics and no license. The steelman does not save the middle layer; it identifies which members of the middle layer are already halfway across the bridge, and it predicts the observed behavior: vertical AI abandoning tool positioning for proprietary data, benchmarks, and outcome pricing (Crosby's RedlineBench; EvenUp's Pre-Litigation-as-a-Service, launched May 2026), and tool vendors pivoting to managed service. Forced vertical integration is the squeeze made visible.
Evidence panel: per-seat AI subscriptions unwinding into metered pricing as true usage costs surface (Shaughnessy; the canonical anecdote: Uber exhausted its 2026 AI coding-tools budget in four months, confirmed by its CTO [B.10]); a market-maker's macro desk writing the same note as AI Twitter (Citadel Securities, Tokenomics, Frank Flight, June 2026: Amazon shutting its internal token leaderboard, Microsoft phasing out Claude Code licenses [B.6]); lab revenue exploding while operating margins stay deeply negative (Anthropic ~$9B to $47B run rate in five months of 2026 at a $965B post-money valuation [B.5]; OpenAI Q1 2026: 39% gross margin, -122% adjusted operating margin [B.5]); and 2025 capital flows showing 53% of AI deal count in vertical AI but only 30% of dollars ($56.2B vertical vs $129.8B non-vertical, the latter dominated by roughly a dozen foundation-model and GPU-infrastructure mega-rounds [B.7]). The gap between where the deals are and where the dollars are is the mispricing this fund trades against; note honestly that the non-vertical dollars are chasing labs and infrastructure more than wrappers, which supports the squeeze (wrapper fundraising is thinning) rather than contradicting it.
This is a Ricardian claim with a Teece engine. When labor (intelligence) becomes abundant, rent accrues to land (the scarce complements); when an innovation is weakly appropriable (and the Engine establishes that machine intelligence is about as weakly appropriable as a technology can be), the rents flow to whoever owns the complementary assets the innovation must pass through (Teece, 1986).
The supply-side ledger (what stays scarce when cognition doesn't):
The demand-side ledger is the half most versions of this thesis omit, and the philosophical core. Humans do not buy cognition. When cognition is free, what do they still pay for?
These demand-side complements are the mirror image of the supply-side list, and together they answer the question "what is an institution?": an institution is a machine for manufacturing legible guarantees.
That definition yields the thesis's trust mechanism, which history supports: trust is institutional, not personal, and it migrates when two conditions hold simultaneously: a large price/convenience gap and a legible guarantee. Deposit insurance moved savings from mattresses to banks; buyer protection and ratings moved commerce from Main Street to strangers on the internet; the index moved wealth from storied stock-pickers to a formula. In every case the incumbent's trust advantage looked unassailable until the guarantee made it portable. Operational implication: an AI-native institution should engineer its guarantees (insured outcomes, money-back accuracy warranties, audited error rates) as deliberately as it engineers its models. The guarantee is the product; the model is the cost line.
Objection (the ledger belongs to the incumbents). Everything on the supply-side list is what JPMorgan, Rocket, and Thomson Reuters already own. Cheap cognition is available to them too, and this technology wave demands unusually little re-architecture: LLMs conform to the organization rather than forcing the organization to conform. The registry's own early reads favor complement-owning incumbents. Reply. Fully conceded, for the served market, and priced into the portfolio rather than argued away: the public-markets leg is long precisely those complement-rich adopters (the Portfolio). The innovator's-dilemma escape hatch is real but narrow: it applies only where incumbent revenue is literally hours sold (billable-hour law, staffing-pyramid consulting, manual bookkeeping), not where incumbents sell balance sheet or risk-bearing (a bank facing cheap cognition experiences pure cost reduction, no revenue conflict, with cheaper funding than any entrant, which is why lending entrants must win on distribution and data, never on "banks are slow"). Where does an entrant face no incumbent complements at all? That is the Expansion, and it is where the venture leg lives.
Every prior draft of this thesis treated the revenue pool as fixed and asked who captures it. That was the largest error in the corpus.
Professional services have spent sixty years under Baumol's cost disease: costs rise where productivity can't, because the binding input was skilled human time. The result is not just expensive services. It is an enormous population priced out of the served market entirely. The evidence of need: 92% of low-income Americans' substantial civil legal problems received no or inadequate professional help (LSC Justice Gap Report, 2022, still the most recent study) [B.8]; roughly 30 million US nonemployer businesses average $58K in revenue, at which a $300/month bookkeeper is arithmetically impossible, and even the dominant accounting software reaches only a fraction of them. Two disciplines the red-team process imposed on this section: the 92% figure is evidence of need, not a TAM (much of that population's market-clearing price is near zero, and TAM claims are restricted to payment-capable segments such as SMBs); and the claim must survive its counter-evidence.
The corrected claim, per the evidence review, requires a distinction sharper than price alone: latent demand unlocks when price falls 5-10x AND the customer's own judgment and labor are actually removed, not merely made cheaper to self-administer, AND distribution already reaches the segment. A cheap tool the customer still operates is not a test of this claim. Wave's free accounting software (1.3% of the nonemployer base), IRS Free File (~3% uptake), and twenty years of $89 online wills (coinciding with falling will ownership) all left the customer doing the same reasoning at a lower price, and their failure to expand the pool says nothing about what happens when the reasoning itself is delegated to an agent; treating them as disconfirmation was an error corrected in apparatus/elasticity-evidence.md. LASIK stands as genuine, category-clean counter-evidence (a 25% price cut met a 30% volume decline) precisely because no judgment was ever delegated on either side of that comparison, so it is informative by contrast: bookkeeping, tax, and basic legal work are compliance necessities, not elective discretionary purchases. Where price fell and labor was genuinely removed and distribution reached the segment, the pool expanded at the extensive margin: zero-commission brokerage added ~30 million accounts in two years, skewing younger and lower-income. LegalZoom's 10% of US LLC formations is retained only as evidence that standardized, judgment-free tasks expand readily once cheap, which is a different and narrower claim than the one this section makes. The honest state of the actual claim is not "supported" or "disconfirmed" but untested: deployable professional-quality autonomous judgment is roughly two years old, so no natural experiment of sufficient scale yet exists. Bench is the closest analogue that genuinely removed labor, and it is retained as real caution (a required calibration case) rather than as clean disconfirmation, because it predates frontier-agent tooling and its collapse is confounded by reported operational failures. Garfield AI and Justpoint are the first live examples of the correct experiment, both agent-native and both serving previously-unserved claimants, both still too small to answer the question at fund scale. Given a genuine evidence vacuum rather than disconfirmed evidence, the wedge is underwritten only on owned-distribution, judgment-removing plays, gated on the pre-registered experiment below.
When all three conditions hold and the price of a service falls 5-10x, the pool is not redistributed. It moves down the demand curve.
Figure 3. A 5-10x price fall does not shrink the pool; it moves it down the demand curve. Surplus passed to existing customers at the incumbent price; the shaded right-hand region at the AI-native price is the unserved market, demand that existed all along, priced out, with no incumbent complements. The venture aperture is that shaded region.
Three consequences, each resolving an objection the fixed-pool framing could not:
Objection (latent demand is a hope). Maybe the priced-out don't want these services at any price. Legal demand is derived from disputes, not from cheap lawyers. And LegalZoom sold near-zero-marginal-cost legal documents for twenty years with massive distribution while the justice gap stayed put; Wave gave away accounting software for a decade and reached 1.3% of the nonemployer base. Reply. These three precedents (LegalZoom, Wave, IRS Free File) do not disconfirm this section, because they tested a different mechanism than the one this section claims. Each is a cheap tool the customer still operates: the customer still does the legal reasoning, the categorization, the interpretation. None of them removed the professional's judgment; they only cut the price of self-service access to a template. This section's mechanism is an AI agent performing the judgment itself, at professional quality, for a customer who does none of the work, which is categorically different from a cheaper form. Citing tool-adoption failures as evidence against agent-adoption is a category error, and treating it as disconfirmation for two years overstated the case against this section (apparatus/elasticity-evidence.md records the correction). The honest consequence is not that the wedge is now supported. It is that no natural experiment has actually tested this claim, because deployable professional-quality autonomous judgment is roughly two years old. What does exist and does count: Bench genuinely removed customer labor and still only reached 12-35K customers before collapsing, which is real caution about SMB services operations (adopted as the Atrium/Bench calibration requirement below) though confounded by pre-frontier tooling and reported operational failures unrelated to demand; Garfield AI (SRA-authorized, winning real county-court judgments for previously-uneconomic sub-£10k claims) and Justpoint are small-n but are the first live examples of the correct experiment. Given a genuine evidence vacuum rather than disconfirmed evidence, unserved-market investments remain restricted to companies with owned distribution into the segment and are gated on the pre-registered price-ladder experiment (apparatus/elasticity-evidence.md), because an untested claim earns a gate, not a presumption.
Carlota Perez's sequence for every major technological revolution (canals, railways, steel, mass production, IT) runs: installation (financial capital floods in, overbuilds, margins negative, valuations detach from cash flow) → turning point (a crash or repricing as financial capital loses conviction) → deployment (production capital diffuses the now-cheap technology through the real economy; the durable fortunes are made here).
The current AI buildout is textbook installation: capability funded by equity rather than cash flow, deeply negative operating margins at the leaders, per-seat pricing subsidized below metered cost, and, the tell, sell-side macro desks beginning to question the leaders' cost trajectory [B.6]. The unwind mechanism is already visible and mundane, and it has a name: "tokenmaxxing." Flat-rate, subsidized-era pricing removed the marginal cost of a token, and several employers responded by gamifying consumption rather than output. Meta ran an internal leaderboard ("Claudeonomics") whose top user reportedly averaged 281 billion tokens over 30 days, at an implied cost above $1.4M, before shutting it down in April 2026. Amazon's internal, employee-built leaderboard ("KiroRank") did the same and was shut down in May 2026 after it "encouraged some staff to perform tasks that didn't necessarily solve problems, just so they could climb the ranks"; Amazon SVP Dave Treadwell told staff, "please don't use AI just for the sake of using AI" [B.6]. Citadel's Frank Flight cited the Amazon episode directly as evidence that "the era of unchecked, free-wheeling token spending is coming to an end" (quoting Pylon's CEO) [B.6]. This is the demand side of the unwind: once metering exposes what tokenmaxxing actually cost, enterprises hit the real bill (Uber's CTO confirmed exhausting the company's 2026 AI coding-tools budget in four months [B.10]) and route to the open-weight model delivering the same measured capability at an order of magnitude lower price. Tokenmaxxing did not cause the commoditization trend in the Engine; it is what made the subsidized installation-period price temporarily invisible, and its abatement is the mechanism by which the real, falling price reasserts itself in enterprise spend.
Perez's framework then makes a claim precisely shaped like this thesis: the crash does not kill the technology. It is the reallocation event that transfers value from the speculative layer to the deploying institutions. The AI-native institution is the deployment-phase asset.
But honesty about Perez cuts our own position: deployment assets get repriced in the crash too. Buying AI-native institutions at 2026 software multiples is paying frenzy prices for deployment assets, the right thesis executed at the wrong price, which in markets is indistinguishable from the wrong thesis. Hence three standing disciplines:
Objection (you'll miss the winners waiting for a crash). The generational institutions may be founded and priced during frenzy; discipline becomes an excuse for absence. And the canonical deployment-phase winners (Amazon 1994, Google 1998) were founded during installation, so a fund absent from frenzy-era formation is absent from the cohort Perez's own history says matters. Reply. The discipline is a price rule, not a calendar rule, and we concede the red-team finding that in the current market it binds like a calendar rule: as of August 2026 nothing in the target set prices at services economics (the two flagship AI-native law firms both raised at software multiples this spring). Three consequences are therefore structural rather than rhetorical: the venture sleeve is separately gated and cannot be called until a signed transaction at services-economics pricing exists; the fund's commitment-period structure must accommodate a late turning point (set with fund counsel); and the fee load on undeployed venture capital is disclosed to LPs as a real cost of the discipline. We also concede that Perez is now consensus-adjacent (a market-maker's macro desk is quoted in this document making part of our argument), which is precisely why the residual edge must come from underwriting discipline and complement selection, not from the periodization itself. If the turning point never comes (lab margins inflect positive, wrapper retention holds), the Falsification section's tripwires force the reassessment explicitly rather than letting the fund drift.
The corpus's natural-experiment registry, scored honestly, currently supports incumbent-adopters (Rocket, Thomson Reuters, Klarna, LegalZoom) more strongly than AI-native entrants (Crosby unproven at scale; Bench a cautionary implosion; Lemonade demonstrating that loss ratios dominate AI; Upstart demonstrating that capital markets dominate models; Cruise demonstrating that capital and regulation can overwhelm a superior cost curve). A thesis that ignores its own scoreboard is dogma. So the portfolio sizes its legs to the evidence each leg has today:
Leg 1, Public markets (most evidence, expressed now). A long-biased spread, named in full in apparatus/public-baskets.md. Long: complement-owning operators adopting AI into hours-free revenue models, with the 2026 de-rating having created the entry (RELX, Wolters Kluwer, and S&P Global fell 27-41% alongside software while reporting accelerating organic growth and expanding margins; the market is pricing the narrative indiscriminately, and complement selection through that de-rating is the active position). Short: a capped sleeve, at most 15% of gross exposure, restricted to liquid names with listed options, expressed via puts wherever short interest exceeds 25% of float, with a pre-set borrow-cost ceiling and a drawdown stop. Two honesty notes from our own research: the short trade partially covered in August 2026 and much of the de-rating has already happened, so late-entry risk on the short side is real and the long side now carries more of the expected spread; and "avoid" is the default expression, with the short sleeve existing for spread capture, never as a standalone profit center. Shorting wrappers into an installation-phase melt-up is a known widow-maker; the sleeve is sized so that scenario is survivable by construction.
Leg 2, Private markets (gated optionality, narrow aperture, strict tests). Honest return framing first, per the red-team reconciliation: the mature AI-native institution is valued at services multiples (8-15x earnings), so entries at services pricing return 3-5x through volume growth, not multiple expansion, and the exit underwrite names no frenzy buyer. This is growth-equity-shaped return with institutional-formation optionality, the sleeve is sized as optionality, and the fund's size is set by the aperture, not the narrative. The aperture is exactly three acquirable complements; everything else on the ledger belongs to incumbents:
Entry criteria, cumulative: (a) the workflow crosses all four thresholds, not just the technical one; (b) incumbent revenue in the vertical is hours-sold or the target market is unserved with owned distribution into it (the Expansion); (c) a contracted guarantee: a carrier term sheet or funded reserve structure, priced into cost per verified outcome and shown to leave a positive sliver, before commitment (an "engineered" guarantee from a seed-stage balance sheet is not legible, per red-team finding B-S2); (d) a cost-stack decomposition of the vertical separating compressible cognition share from incompressible intake, sales, exceptions, liability, and judgment, calibrated against the two named failure cases: Atrium (a16z-backed law-firm-as-a-startup, dead 2020, because the compressible share was overestimated) and Bench (dead December 2024, because the exception tail was the business and its collapse re-poisoned trust for the whole vertical, which we now carry as a portfolio assumption: one implosion per vertical resets trust-creation costs for every entrant); (e) underwritten on ΔRevenue/ΔFTE and cost-per-verified-outcome at services-business pricing. The sleeve is separately gated: it cannot be called until a signed transaction at services-economics pricing exists, the counsel opinion (B.9) is delivered for any regulated-standing play, and the elasticity experiment has passed its pre-registered criteria for any unserved-market play. A scored ten-company pipeline with four deliberate passes lives in apparatus/venture-pipeline.md.
Leg 3, Robustness, not prediction (the ASI question, settled by construction). This fund's returns case requires no position on superintelligence, and its portfolio is built to be robust across the scenarios rather than predicated on one:
| Scenario | What happens | Portfolio result |
|---|---|---|
| Base: trend continues, no ASI | Cost collapse + threshold crossings proceed on the measured curve | Both legs pay as designed |
| Capability plateau | Price collapse continues but few new workflows cross thresholds | Venture leg narrows to already-crossed workflows; publics short leg weakens (wrappers survive as tools); tripwires catch this early |
| ASI reached | Labs vertically integrate downstream; most software theses break | Standing survives by law where software doesn't: complement-owning publics retain the licenses, charters, and liability seats even superintelligent suppliers must transact with; venture leg partially impaired |
| China closes source | Open-weight price cap weakens; frontier pricing power returns | Commoditization slows; squeeze thesis decelerates; named tripwire |
We do not underwrite the ASI branch for the additional reason that our LP base should not be asked to root for racing dynamics: the fund's upside must never depend on the outcome its investors most fear. Values-alignment here is not decoration; it is a risk constraint that happens to also be the only epistemically honest position, since the ASI-gated claim was the least defensible one in the research corpus.
One correlation the fund does not claim to escape: in an AI risk-off event this vehicle draws down with the theme. Its edge is the reallocation entry that follows, funded by the gated sleeve, not immunity from the drawdown; the sibling Natural Capital strategy hedges the family's exposure, not this vehicle's LPs, and this document no longer implies otherwise. Conflict-of-interest structure (recusal rules, a restricted-vertical list maintained with the LPAC, independent third-party valuation of any category-adjacent position) is set out in the fund documents, which govern.
Every metric exists to test a specific claim; a metric that can't kill a claim is decoration. Cost-per-token is a vanity metric everywhere on this panel. The unit of account is the verified outcome. Per the red-team review, the panel is tiered by what a fund of this size can actually measure, and it now measures the fund as well as the world; panel readings are published to LPs semiannually and tripwire determinations are reviewed by the LPAC, not self-graded by the GP.
Tier 1: world metrics (measurable from public data; these adjudicate the thesis).
| Instrument | Definition | Claim it tests | Kills the thesis if... |
|---|---|---|---|
| Open↔closed gap | capability distance and price ratio at the efficient frontier | Engine mechanism 4 | gap widens durably (e.g., China closes source) |
| Cross-vendor price floor | lowest price per fixed capability tier across all vendors, quarterly | the Engine's precise claim (under active test since DeepSeek's capacity-rationed repricing) | the floor stops falling for four consecutive quarters |
| Applications-layer unit economics | gross margin, NRR, CAC payback in public agentic SaaS | the Squeeze (the short sleeve) | wrapper margins and retention hold through two more frontier release cycles |
| Capital flow | venture $ into wrappers vs infrastructure vs AI-native operators | the Clock: the crowd we trade against | the mispricing closes before we deploy |
| Lab gross-margin trajectory | reported gross margins at the frontier labs | frenzy reading vs concentration reading | sustained above 50% for four consecutive quarters: the concentration reading wins and the Perez frenzy read was wrong |
Tier 2: portfolio metrics (small-n, founder-reported, labeled as such; these adjudicate individual investments, with kill thresholds set per deal at entry).
| Instrument | Definition | Kills the investment if... |
|---|---|---|
| Cost per verified outcome | inference + tools + retrieval + review + support + compliance + error correction + insurance + liability reserve + capital cost, per completed workflow | it does not fall materially post-investment |
| ΔRevenue / ΔFTE | incremental revenue per incremental human, after compliance/risk staff | the operator scales like a body shop |
| CAC, churn, LTV/CAC | standard cohort economics for the operator's own customers | below deal-specific thresholds set at entry |
| Measured willingness-to-pay | from the Expansion experiment design (price-ladder or cohort data) | the pre-registered elasticity kill criteria fire |
| Contracted guarantee cost | carrier premium or reserve charge per outcome | the guarantee consumes the sliver |
| Human review ratio | review minutes per completed workflow, over time | flat or rising at constant quality |
Tier 3: fund metrics (these adjudicate the vehicle; absent from v1.1, added per red-team finding that a panel which cannot kill the position sizing is decoration).
| Instrument | Rule |
|---|---|
| Net beta to an AI-deployment factor | capped; measured by quarterly factor decomposition |
| Short-sleeve exposure | ≤15% of gross; borrow-cost ceiling; hard drawdown stop |
| Dry-powder decay clock | undeployed venture capital reported with its fee drag each quarter; commitment-period structure reflects a possibly late turning point |
| Vehicle base rate | stated in LP materials: 9% of thematic funds survived and outperformed over the 15 years to mid-2024 (Morningstar); this fund's design choices exist to escape that base rate and are judged against it |
Demoted from the panel per measurability review: incumbent price pass-through and the economy-wide verification-cost ratio become annual qualitative assessments; a tripwire cannot hang on a series nobody can measure.
A prediction is only honest if its disconfirming observation is written down in advance; otherwise frenzy evidence gets booked as thesis confirmation.
P1 (commoditization velocity), by end of 2027. Open-weight models sit within ~5 points of the closed frontier at under 1/10th the price across major capability tiers, measured on task-level evals, not only aggregate benchmarks.
P2 (institutional formation), by end of 2028. Rewritten after red-team review found the original version unfalsifiable (every outcome confirmed it). New form: at least one AI-native professional-services institution reaches $25M revenue run-rate with ΔRevenue/ΔFTE at least 3x incumbent comparables, at any financing multiple.
P3 (the middle-layer test), rolling from 2026. Agentic-SaaS cohort economics (gross margin, NRR, CAC payback) deteriorate as model releases absorb feature surface and pricing shifts from seats to outcomes.
Standing tripwires (adjudicated with the LPAC against published panel readings, not by the GP alone): portfolio-level human-review ratio flat for 24 months at constant quality (kills autonomy economics; replaces the unmeasurable economy-wide verification-cost ratio); a top-tier incumbent professional firm moving to outcome pricing at scale (incumbent seat holds: reweight leg 1 long, cut leg 2); lab gross margins sustained above 50% for four consecutive quarters (the concentration reading beats the frenzy reading; note honestly that OpenAI's gross margin rose from 33% to 39% year over year, so this series is currently moving toward the tripwire, and that Citadel's Tokenomics itself concludes advanced capability concentrates at balance-sheet-rich firms, which is partially adverse to this thesis rather than supporting it); the cross-vendor price floor failing to fall for four consecutive quarters (the Engine claim fails its live test); model providers acquiring regulated service arms (value-accrual map redraws: labs become the competing institution).
Strip the portfolio away and the thesis makes one claim about the world, and it is worth stating without hedges because everything above is downstream of it.
Machine cognition is becoming infrastructure, following the path of literacy, electricity, and computation itself: from scarce professional attainment, to competitive advantage, to ambient utility whose price asymptotes to its energy cost. Nobody today asks who owns literacy. Within a generation, "who owns intelligence" will sound the same, and the entities built on selling it will look like the companies that once sold electricity door-to-door.
What follows is not the disappearance of firms but their reformation around a different scarce input. For two centuries, institutions were machines for coordinating scarce human cognition; the partnership, the pyramid, the billable hour, the bureaucracy are all artifacts of expensive thought. When thought is cheap, the institution's remaining function is exposed as what it always was underneath: the manufacture of legible guarantees: accountability, standing, absolution, trust. Those are made of law, capital, and human meaning, and no model emits them.
Most of the surplus will go to people, not portfolios: services that were rationed by price for the entire history of the service economy (counsel, bookkeeping, underwriting, advice) become as available as search. That is the largest one-time welfare transfer since mass electrification, it is measurable (the Instrument Panel's market-expansion row), and this fund's structure is built so that its returns come from financing that transfer's institutions, not from taxing it. The residual private fortunes accrue to what intelligence cannot manufacture: the standing to guarantee (this document), the atoms to compute with (Natural Capital), and the meaning humans reserve for each other (Taste Capital).
The thesis in one line, LP-facing: we underwrite to the sliver; society keeps the rest. The next step is unchanged from the research corpus's own closing instruction: not to make this more persuasive, but more measurable.
Claims from the research corpus and what the adversarial process (Claude/Codex roundtable, Aug 2026) did to them:
| Claim | Verdict | Disposition |
|---|---|---|
| Intelligence-per-dollar collapse is real, measurable, ongoing | Survived | the Engine, five mechanisms; hardened with per-workflow threshold framing |
| Rents migrate to scarce complements (Teece) | Survived | the Ledger; extended with the demand-side ledger |
| Thin-wrapper squeeze / double commoditization | Survived, bounded | the Squeeze; boundary condition: polarization, not collapse; the steelman identifies survivors |
| Perez installation/deployment framing | Survived, disciplined | the Clock; converted from citation into price rule + crossover structure |
| "Value accrues to the lab that crosses RSI/ASI" | Killed from the returns case | Self-undermining (an ASI lab integrates downstream); un-measurable; conflicts with LP values; retained only as robustness scenario (the Portfolio) |
| "Incumbents can't restructure (organizational debt)" | Killed as stated | Evidence points the other way (Harvey adoption, bank copilots, Rocket, Klarna); survives only in hours-sold verticals; narrowed into the Ledger reply and leg-2 entry criteria |
| "AI-native structure is an edge" | Killed | Symmetric: fast-followers have the same zero legacy; only the complement endures (was already conceded in thesis.md; now enforced in underwriting) |
| Innovator's dilemma as a general mechanism | Narrowed | Applies only where revenue = hours sold; explicitly fails for balance-sheet businesses (lending) |
| Fixed revenue pool ($1.6T capture framing) | Replaced | the Expansion's demand expansion, the largest correction in this rewrite; resolves the incumbent, trust, and surplus objections simultaneously |
| "Charge $2,000, keep the margin" surplus retention | Narrowed | Competition passes surplus through; retained margin is a sliver on expanded volume, and the pass-through is the impact thesis |
| Codex's "AI lowers minimum efficient scale of institutions" | Adopted, bounded | True and important, but symmetric (also lowers fast-follower scale), so it enlarges the formation opportunity without substituting for a complement |
Second adversarial round (two independent red teams, August 2026; full reports and GP response in apparatus/):
| Claim | Verdict | Disposition |
|---|---|---|
| "Short the wrapper basket" as stated | Killed as unexecutable; rebuilt | Named basket with expressibility limits (apparatus/public-baskets.md); sleeve capped ≤15% gross with borrow ceiling and stop; long-biased spread is the position |
| "The stool hedges itself" (vehicle-level) | Killed | Family-level hedge misrepresented as vehicle-level; deleted; fund-tier risk metrics added (Instrument Panel Tier 3) |
| P2 as originally written | Killed as unfalsifiable; rewritten | No-qualifying-company now disconfirms, full stop; adjudication moved to LPAC |
| "Most upside" venture framing | Killed | Reframed as gated optionality with growth-equity return math at services multiples; exit names no frenzy buyer |
| Unserved wedge on price alone | Narrowed | Three-condition claim (price + labor + distribution) per elasticity evidence; owned-distribution-first; gated on pre-registered experiment; 92% reframed as need, not TAM |
| "Engineered" guarantee | Killed | Guarantee must be contracted (carrier term sheet or funded reserve) and priced into the sliver before commitment |
| "Five largely independent mechanisms" | Narrowed | Two shared chokepoints named (compute/energy; Chinese open-weight policy) |
| Atrium/Bench omission | Corrected | Both added as named calibration cases; cost-stack decomposition required per deal; correlated-trust-risk assumption adopted |
| Principal's conflict as disclosure | Corrected | Structure added: recusal, restricted-vertical list with LPAC, independent valuation; specifics moved to the fund documents |
All entries verified against primary sources on August 18, 2026. Verdicts: CONFIRMED (claim matched source), CORRECTED (thesis text updated to the sourced figure). Corrections have been applied throughout the body; where a widely-circulated draft number was wrong, the register records both.
B.1. Epoch AI efficiency and price trends. CONFIRMED. Algorithmic progress: compute for fixed capability halves ~every 8 months (95% CI 5-14) per Ho, Besiroglu et al., Algorithmic progress in language models, arXiv:2403.05812 (Mar 2024); Epoch's live dashboard (epoch.ai/trends, fetched 2026-08-18) currently estimates ~3.0x/year (~7.6-month doubling). Hardware trends: same dashboard.
B.2. Inference-price decline and the Fig 1 series. CONFIRMED range / CORRECTED series. Range: 9x-900x/year depending on performance milestone, per Epoch AI, LLM inference price trends (Mar 12, 2025), which cautions the fastest declines are recent and may not persist. The draft series mixed price bases; the corrected series uses a uniform 3:1 input:output blend: GPT-3 davinci $60/M flat (2021; The Decoder, Aug 2022, documents the $0.06/1K rate) → GPT-4 $37.50/M blended ($30 in / $60 out at Mar 2023 launch; OpenAI pricing still lists gpt-4-0613) → GPT-4o mini $0.26/M blended ($0.15/$0.60, Jul 2024) → GPT-5 Nano $0.14/M blended ($0.05/$0.40, Aug 2025). Composite: ≈436x in ~4 years (draft said 428x on mixed bases; input-only would be 1,200x, output-only 150x).
B.3. The closed/open pair. CORRECTED, two of four numbers. 80.8 on SWE-bench Verified is Claude Opus 4.6 (Anthropic, Feb 5, 2026); Opus 4.7 scores 87.6 (Anthropic, Apr 16, 2026). Opus output price $25/M confirmed (Anthropic pricing docs). DeepSeek V4-Pro 80.6 confirmed (model card, Apr 26, 2026). V4-Pro pricing: $0.87/M output was preview-period only; since GA (Aug 13, 2026) output is $1.98/M off-peak / $3.96/M peak (DeepSeek pricing; Caixin reported increases up to 1,100%). Current matched-capability gap: ~6.3x-12.6x, not ~28x. Benchmark harnesses differ across evaluators (vals.ai, NIST CAISI); the thesis cites each lab's own reported number and says so.
B.4. Chinese frontier releases. CORRECTED. Four of five shipped within April 2-28, 2026: Qwen 3.6 Plus (Apr 2), GLM-5.1 (Apr 8), Kimi K2.6 (Apr 20-21), DeepSeek V4 preview (Apr 24). MiniMax M2.7 launched Mar 18 (open-sourced ~Apr 12). Honest framing used in the Engine: five releases in a ~six-week window, four within April alone. Sources: MarkTechPost (Apr 12, 2026), dev.to roundup, Yotta Labs, Caixin.
B.5. Lab financials. CONFIRMED. Anthropic: $47B run rate and $65B Series H at $965B post-money per Anthropic's Series H announcement (May 28, 2026); path $9B (end 2025) → $14B (Feb 12, Series G) → $30B (Apr 6) → $47B (May), traced by Simon Willison; Reuters independently confirmed. OpenAI: Q1 2026 revenue $5.7B, gross margin 39% (up from 33% Y/Y), adjusted non-GAAP operating margin -122% (~$6.95B operating loss), per The Information (Sri Muppidi), summarized by Ed Zitron (May 22, 2026). Caveat carried in the Squeeze: Q1-only, non-GAAP.
B.6. Citadel Securities, Tokenomics, and "tokenmaxxing." CONFIRMED via secondary reporting; primary text not independently fetchable (citadelsecurities.com returns 403 to automated retrieval). Author Frank Flight, Global Macro Strategy, June 2026, confirmed via Yahoo Finance (Jun 15, 2026) and MSN syndication. Argues the binding constraint shifted from capability to input cost/scarcity, and explicitly cites Amazon's leaderboard shutdown as evidence; also quotes Pylon CEO Marty Kausas that "the era of unchecked, free-wheeling token spending is coming to an end." The Microsoft "phasing out Claude Code licenses" detail is single-sourced to this coverage and not independently confirmed. On tokenmaxxing itself: Amazon's leaderboard was an internal, employee-built dashboard called "KiroRank", shut down May 2026 per Business Insider (May 29, 2026) because it "encouraged some staff to perform tasks that didn't necessarily solve problems, just so they could climb the ranks"; SVP Dave Treadwell: "please don't use AI just for the sake of using AI." This is a distinct incident, roughly six weeks later, from Meta's separate internal leaderboard ("Claudeonomics," shut down ~April 8-9, 2026; top user ~281 billion tokens/30 days, ~$1.4M+ implied cost), per Yahoo Finance and Inc. (Apr 2026); the two should not be conflated. No single coiner of the term "tokenmaxxing" was identifiable; mainstream usage clusters in the April 8-11, 2026 coverage window. Additional corroborating anecdotes (informational, not separately pinned): Uber's CTO citing the "tokenmaxxing era" ending (TheNextWeb, Aug 2026); Chamath Palihapitiya warning that "CEOs and CFOs... probably have no idea how much tokenmaxxing is going on inside of their organizations" (CNBC, Jul 14, 2026); YC-backed Weave raising $13.5M explicitly to curb developer tokenmaxxing (Business Insider, Jul 2026).
B.7. 2025 capital flows. CONFIRMED with reframe. Euclid Ventures, The Vertical Report 2026 (Apr 6, 2026; PitchBook + Harmonic data, US/Canada, $1M+ financings): vertical AI = 53% of deal count but 30% of dollars ($56.2B vs $129.8B non-vertical). Reframe applied in the Squeeze: the source's category is "non-vertical," not "wrapper"; the dollar gap is driven by ~12 foundation-model/GPU mega-rounds (excluding them, vertical took 51% of capital).
B.8. Justice gap. CONFIRMED. "Low-income Americans received no or inadequate legal help for 92% of their substantial civil legal problems": LSC, The Justice Gap (with NORC, April 2022; press release). Still the most recent study as of Aug 2026; cited as 2022 in the Expansion.
B.9. Legal-ownership structures. CONFIRMED / counsel review still required. ABA Model Rule 5.4 (official text; non-binding template, adopted state-by-state). Arizona ABS: ER 5.4 eliminated Aug 27, 2020; licensing effective Jan 1, 2021; 114 active licensed entities as of Dec 31, 2024 (program page; April 2025 annual report). Utah sandbox: pilot extended to seven years, through Aug 2027, permanence under review; 7 currently authorized entities (Utah Innovation Office); five-year data review: Stanford Law, Jun 2025. Whether these structures make the legal vertical investable remains a fund-counsel determination.
B.10. Uber. CONFIRMED, scope narrowed. Uber exhausted its 2026 AI coding-tools budget (not its entire AI budget) in four months; CTO Praveen Neppalli Naga confirmed to The Information. Coverage: Forbes (May 17), Fortune (May 26), TechCrunch (Jun 2: subsequent $1,500/employee/month cap).
Theseus Capital, Institutional Capital leg, v1.4. Prepared August 2026; all sources pinned August 18, 2026; red-teamed by two independent adversarial reviews with accepted findings integrated (apparatus/); the Expansion corrected August 19, 2026 to distinguish tool-based from agent-based historical precedents (apparatus/elasticity-evidence.md); the Clock and B.6 expanded August 19, 2026 with the tokenmaxxing demand-side mechanism and direct evidence. Fund counsel's review of the legal-ownership question (B.9) remains an open item; until resolved, legal-vertical examples are illustrations, not investments.
Share