Theseus Capital · v1 · 2026-08-20

Neo-Firms Explainer

The question it answers

The consensus view of the AI investment landscape runs: foundation models will commoditize, so value will accrue to the application layer built on top of them, because applications own the customer workflow while the model becomes interchangeable infrastructure underneath. This thesis argues that conclusion does not follow from the premise. If model commoditization also commoditizes the applications built on models (because cheap AI makes cheap software), then the money is not safely parked one layer up. The real question becomes: once cognition itself is abundant and cheap, what remains scarce, and who owns it?

The core claim

As the marginal cost of useful machine cognition falls toward its physical floor (electricity plus capital), economic rents migrate away from entities that sell intelligence and toward entities that own what intelligence cannot manufacture: regulatory standing, liability-bearing capacity, embedded distribution, proprietary outcome data, and trust as the party of record. The largest single destination of the productivity surplus this creates is the customer, not the producer, and the expansion of demand that follows from cheaper, better service is where new institutions get built.

One sentence: when intelligence commoditizes, own everything it cannot manufacture: standing, atoms, and meaning. ("Standing" is this thesis; "atoms" and "meaning" are sibling legs of the fund, compute/energy and taste/relationship respectively, not developed here.)

The engine: why cognition keeps getting cheaper

This is the empirical foundation, and it is deliberately built to need no speculative assumption about artificial general intelligence or superintelligence. The price of a fixed capability tier, that is, the cost to get a model to perform at a specific benchmark level, has fallen roughly two and a half orders of magnitude in about four years, and continues falling anywhere from 9x to 900x per year depending on the task, per Epoch AI's tracking. Five distinct mechanisms drive this, and they would all have to stop simultaneously to kill the trend:

  1. Algorithmic efficiency. The compute needed to reach a given capability level has historically halved roughly every seven to eight months, faster than Moore's Law and independent of hardware gains. The training recipe itself keeps improving.
  2. Hardware price-performance. FLOPs per dollar keeps falling through accelerator competition and inference-specific silicon.
  3. Distillation. Whatever a frontier, expensive model can do this year, a much cheaper model can do next year, because large models are used to train smaller ones. This converts a temporary capability lead into a permanent, falling capability floor available to everyone.
  4. Open-weight diffusion. Open and open-weight models within reach of the frontier cap what any closed lab can charge, because the buyer's cheap alternative is one API call away. As of mid-2026, a closed frontier model and a leading open-weight model performed almost identically on a coding benchmark (SWE-bench Verified, roughly 80.6 to 80.8) while differing in price by roughly an order of magnitude. Two honest complications: the current closed frontier has since pulled seven points ahead, so the price cap operates with a lag; and the open-weight vendor itself raised prices sharply after that comparison, showing commoditization is not a smooth, monotonic line for any single vendor even while the cross-vendor floor keeps falling.
  5. Competitive structure. At least five Western labs and five Chinese labs now ship frontier-tier models, with several major Chinese releases landing within weeks of each other in spring 2026. No lab has demonstrated it can durably protect a capability lead; research and talent both diffuse quickly.

The precise version of the claim, stated so it can be tested rather than merely asserted: for any given workflow, the price of the cheapest model that clears that workflow's quality bar trends toward the underlying cost of compute, not toward whatever a market leader can currently charge. This reframes the question from "how good are frontier models" to "how good does a model need to be for this specific job, and how much does the cheapest model that clears that bar cost." A workflow only actually gets economically transformed once it clears four separate thresholds: it must be technically possible for the model, cheaper after accounting for review and error costs, delivered by a provider legally allowed and trusted to do it, and something customers will actually switch to. Most AI demonstrations clear only the first threshold. Refusing to pay for capability that has not cleared all four is a central discipline of the thesis.

The squeeze: why the application layer is not automatically safe

The consensus thesis assumes model commoditization does not touch the application layer built on top. But the same forces that make models cheap also make it cheap to build competing software, because AI increasingly writes the code, and increasingly performs natively (through longer context, tool use, and better reasoning) tasks that used to require a dedicated application to orchestrate. This is a double commoditization: models get cheaper, and building software that wraps them gets cheaper too, which erodes the second-mover cost of copying whatever a thin application does.

The result is not that every application dies. It polarizes. Products whose entire value is repackaged model access, prompt wrappers, thin orchestration layers, per-seat copilots, get squeezed because their differentiation decays on the release schedule of the underlying model. Products that already own a system of record, a compliance workflow, proprietary outcome data, or genuine distribution survive, because those assets do not evaporate when a better model ships. Live 2026 evidence for the squeeze side: several formerly high-flying per-seat SaaS companies (in categories like sales intelligence, IT research subscriptions, project-management software, and consumer subscription products with AI-inference-heavy features) took severe stock declines, guidance cuts, and layoffs specifically attributed to AI eroding their reason to exist as standalone products.

The ledger: where the money actually goes

This part borrows from two pieces of economic theory. Ricardo's rent theory says that when one input (here, cognition, standing in for labor) becomes abundant, the economic rent shifts to whatever remains scarce (here, the "land," meaning the complementary assets). Teece's theory of profiting from innovation says that when an innovation is easy to copy (weak appropriability, which cheap and diffusing AI clearly is), the profits flow not to the innovator but to whoever owns the scarce complementary assets the innovation has to pass through to reach a customer.

The supply-side list of what stays scarce: regulatory and licensing standing (a bank charter, a bar admission, underwriting authority), liability-bearing capacity and balance sheet, proprietary data that is actually tied to verified outcomes (not just usage logs, which are not a moat), embedded distribution and transaction flow, and trust as the party of record.

The more original half of the argument is the demand side, which most versions of this thesis skip: what do humans keep paying for once cognition itself is free? Three things. Accountability, someone to blame, sue, and be made whole by, which is why liability absorption is itself a sellable product (insurance was the first institution built entirely around this; an AI-native institution generalizes it). Absolution, the transfer of decision-anxiety to a party of record, which is most of what a client is actually buying from a lawyer or a wealth manager. And status or relationship, the social ritual of being advised by someone, which does not commoditize because its scarcity is the point.

Put together, this yields a compact definition: an institution is a machine for manufacturing legible guarantees. History suggests trust migrates from an incumbent to a new provider only when two things happen together: a large price or convenience gap, and a legible guarantee. Deposit insurance moved savings from mattresses into banks; buyer protection and ratings moved commerce onto strangers on the internet; the index fund moved money away from storied stock-pickers. In each case the incumbent's trust advantage looked unbeatable until the guarantee became portable. The practical implication: a new AI-native institution should build its guarantee (insured outcomes, money-back accuracy warranties, audited error rates) as deliberately as it builds its model, because the guarantee is the actual product being sold and the model is just a cost line underneath it.

The obvious objection is that the supply-side list above describes assets incumbents (banks, insurers, established law and accounting firms) already own, and cheap cognition is available to them too. This is largely conceded. The mechanism only favors a new entrant where the incumbent's revenue model is literally selling human hours (so cheap cognition creates a direct conflict with its own economics, the classic innovator's dilemma), not where the incumbent sells balance sheet or risk-bearing capacity (a bank facing cheap cognition is not conflicted, it is simply handed a cost reduction, and it has cheaper funding than any startup could get). This is why the strongest version of the entrant argument lives somewhere else entirely: markets nobody currently serves at all.

The expansion: the unserved market, and a correction worth understanding carefully

Professional services have spent decades under what economists call Baumol's cost disease: costs rise wherever productivity cannot improve, because the bottleneck is expensive, skilled human time. One consequence, largely ignored by prior versions of this thesis, is that an enormous population is priced out of professional services entirely, not merely paying a high price for them. In the United States, roughly 92% of low-income Americans' substantial civil legal problems receive no or inadequate professional help, and the median very small business (of which there are roughly 30 million) earns under $60,000 a year, at which even a modest $300-a-month bookkeeper is an arithmetic impossibility, not a preference. Cheap cognition is the first plausible reversal of this dynamic in the service economy's history: when a service's price falls five to tenfold, the customer pool does not simply pay less for the same total revenue, it moves down the demand curve into people who were never customers of anyone.

Here is the correction. An earlier draft of this argument cited LegalZoom, Wave (free accounting software), and the IRS's free e-filing program as evidence against the demand-expansion claim, on the reasoning that all three offered dramatically cheaper access to professional-adjacent services for years without meaningfully closing these gaps. That reasoning contains a category error. LegalZoom, Wave, and Free File are all cheap tools that the customer still has to operate themselves: the customer still does the legal reasoning to fill in the form correctly, still categorizes their own transactions, still interprets their own tax situation. None of them removed the professional's judgment; they only lowered the price of access to a template. This thesis's actual mechanism is an AI agent performing the judgment itself, at professional quality, for a customer who does none of the work. Citing the failure of a cheaper tool as evidence against the success of judgment-removal is comparing two different products.

Correcting for this, the true claim requires three conditions together, not price alone: the price must fall five to tenfold, the customer's own labor and judgment must actually be removed rather than merely made cheaper to self-administer, and distribution must actually reach the previously unserved segment. Historical cases that had all three genuinely did expand the pool: zero-commission stock trading added roughly 30 million new brokerage accounts in two years, skewing toward younger and lower-income people who had never invested before; LegalZoom's cheap templates did reach 10% of all new US business formations, a real instance of a standardized, judgment-free task expanding once cheap. A one-time price cut with no labor removal (LASIK surgery got about 25% cheaper in the early 2000s and volume fell, not rose) shows discretionary, elective purchases behave differently from compliance-driven necessities like bookkeeping or basic legal work.

The honest state of the corrected claim is not proven, and it is not disproven either. It is genuinely untested, because deployable, professional-quality autonomous judgment from AI agents is roughly two years old, so no experiment of real scale has actually been run yet. The closest historical analogue that did remove customer labor, an AI-assisted bookkeeping startup that burned over $100 million reaching only a small fraction of the addressable market before collapsing in late 2024, is real caution about the operational difficulty of serving this segment (customer acquisition cost, exception handling, adverse selection), but it also predates today's frontier-quality agents and its collapse is tangled with reported internal mismanagement unrelated to demand. The first live examples of the actually-correct experiment, a fully AI-run law firm in the UK winning real court judgments for previously-uneconomic small claims, and an AI-native firm originating mass-tort cases nobody else found, exist now but are too early and too small to answer the question definitively.

The clock: timing discipline, and a named mechanism worth knowing

Carlota Perez's framework for technological revolutions describes a repeating sequence: an installation phase (financial capital floods in, overinvestment, negative margins, valuations detached from cash flow), a turning point (a crash or repricing as confidence in the installation-phase leaders breaks), and a deployment phase (the technology, now cheap and proven, actually diffuses through the economy and the durable fortunes get made). The 2026 AI buildout looks like textbook installation: capability funded by equity rather than revenue, deeply negative margins at the frontier labs even as their revenue explodes, and enterprise AI pricing subsidized well below its true metered cost.

The specific mechanism by which this subsidy became visible has an actual name in 2026: "tokenmaxxing." Under flat-rate, effectively-free-at-the-margin pricing, some companies' employees began maximizing their own token consumption for its own sake rather than for real output. Amazon built and then shut down an internal leaderboard ("KiroRank") after it was found to be encouraging employees to perform pointless tasks just to climb the rankings; Meta did the same with a separate internal leaderboard ("Claudeonomics") whose top user reportedly ran up roughly 281 billion tokens in a month, an implied cost over a million dollars. A Citadel Securities research note in June 2026 cited the Amazon episode directly as evidence that "the era of unchecked, free-wheeling token spending is coming to an end." The practical implication: as metering exposes what this behavior actually cost, enterprises hit a real bill (Uber's own CTO confirmed the company burned through its entire year's AI coding-tools budget in four months) and respond by switching to whichever open-weight model delivers the same measured capability for an order of magnitude less, which is the concrete, already-observable version of the commoditization mechanism described in the engine section above.

The practical discipline this creates: never pay software-company valuation multiples for a company that has not yet secured its scarce complement, because doing so is betting on the same installation-phase frenzy the thesis expects to eventually correct. This is treated as a pricing rule applied at the moment of investment, not as a prediction of when any crash will happen; Perez's framework describes a sequence, not a calendar.

The portfolio: what this actually means for where capital goes

The thesis translates into three positions, deliberately sized to how much evidence currently supports each rather than to how exciting each one sounds:

A public-markets spread, long-biased: own operators that already control a scarce complement and are adopting AI aggressively (companies with regulated licenses, proprietary authoritative data, embedded transaction flow, or balance sheet, several of which were unfairly de-rated in 2026 alongside pure AI-software companies despite reporting accelerating growth), and short or avoid a capped, carefully named basket of companies whose entire product is repackaged model access. This leg does not require any new institution to succeed; it profits from the squeeze happening at all.

A venture sleeve, deliberately framed as optionality rather than the primary source of returns, because a mature AI-native institution is likely to be valued like an ordinary services business (modest multiples of earnings), not like software. It targets companies that can plausibly acquire one of exactly three things: distribution already embedded inside a transaction flow, a proprietary data loop tied to verified outcomes, or acquirable regulatory standing (a real but currently narrow set of legal structures exist for this in places like Arizona and Utah). This sleeve is deliberately gated: it will not commit capital until a target company can be bought at services-business pricing rather than software pricing, until any liability guarantee it makes is backed by an actual contracted insurer rather than merely asserted, and until the unserved-market demand claim above has cleared a real, pre-designed experiment rather than being assumed.

Explicit robustness, not a bet, on how artificial intelligence's endgame plays out. The thesis does not require superintelligence, a single winning lab, or any particular resolution of the debate about where AI capability tops out; every input it actually relies on is measured quarterly using data available today. If a lab does eventually achieve something like recursive self-improvement, the portfolio is designed to survive that scenario (the same scarce complements, licenses, balance sheets, and trust, still cannot be manufactured by a smarter model) rather than depending on it for its return.

How the thesis checks its own work

The thesis is built around a set of measurable instruments (grouped into things measurable from public data, things measurable only from a small number of portfolio companies, and things that measure the fund's own risk rather than the world) and a small number of dated, specific predictions, each written together with the exact observation that would prove it wrong. For example: if, by the end of 2028, no AI-native professional-services company has reached meaningful revenue scale with dramatically better revenue-per-employee than an incumbent competitor, at any valuation, the institutional-formation part of the thesis is considered disconfirmed and new venture commitments stop, with no escape clause available.

The thesis has already been through two rounds of deliberate, independent attempts to kill it, one from a public-markets skeptic and one from a services-operations skeptic who studied prior failed attempts at AI-native professional-services companies. Both concluded that the underlying economic argument survives scrutiny, but that the original fund construction around it did not: the short book was not concrete enough to actually trade, one of the core predictions could not have been proven wrong by any outcome, and the venture criteria did not yet require the discipline (a real contracted guarantee, real comparison to companies that already failed attempting something similar) that would keep it honest. Every one of those findings has since been incorporated directly into the current version of the thesis rather than argued away, which is why the current document differs meaningfully, and for the better, from any earlier draft that might exist elsewhere.

Share