← Back to Signal
Routing & OrchestrationTokenomicsPosition on Market Signal

Solow among the thinking machines: the illusion of the silicon moat

Robert Solow's productivity residual explains why raw GPU stockpiles are no longer a moat. As compute commoditizes and depreciates, the orchestration layer, runtime routing across the token stream, is where defensible margin now lives.

George Fortin · June 17, 2026 · 8 min

Key takeaways
  1. 01Hyperscaler capex hits $660B to $690B in 2026, at 45% to 57% of revenue, a utility grade ratio for what was sold as a software business.
  2. 02Depreciation schedules are reversing. Amazon shortened server life back to five years, and one estimate puts sector under-reporting at $176B through 2028.
  3. 03Inference cost is now exceeding payroll for some teams. Microsoft pulled engineers off Claude Code, and Uber burned its 2026 AI coding budget in four months.
  4. 04A static engine spends the same compute on a comma as on a hard deduction. Roughly 85% of tokens are trivial. The remaining 15% carries the real reasoning.
  5. 05Runtime orchestration routes by token difficulty, breaking the linear lockstep between hardware spend and reasoning output. Prizmal measures 2x to 4x throughput gains.

00 · The frame

Darwin among the machines, revisited.

"Man will have become to the machine what the horse and the dog are to man."
Samuel Butler · Darwin Among the Machines · 1863

Writing on a boat from New Zealand to England in 1863, Butler looked at the fast pace of early industrial machinery and argued that engines were evolving at a pace that, following Darwin, would result in machines becoming the superior race. He saw, over 160 years ago, that humans were already sacrificing a growing portion of our labor and capital just to keep the hardware clean, organized, and lubricated.

Many read his work as an early warning of an existential, science fiction apocalypse. I have always looked at it differently. As an economist, an asset allocator, and an entrepreneur, I look at the hundreds of billions of dollars currently being funneled into artificial intelligence infrastructure and see that we have walked right into the trap Butler anticipated. We are sacrificing our energy grids, our capital, and our natural resources to sustain a brute force infrastructure that misaligns with basic production economics. We are treating raw compute as a permanent competitive moat, and we are about to hit an inevitable economic wall.

01 · The residual

Solow's A is the alpha in compute.

How do we measure the productivity increase from hundreds of billions of dollars in AI infrastructure? In a knowledge economy, reasoning for both low and complex tasks is what we pay people for. For the sake of this argument, a unit of reasoning is equivalent to a unit of production.

To trace where the financial alpha will be captured, strip away the marketing and look at the foundation of modern growth theory: Robert Solow's total factor productivity residual. Output (Y) represents reasoning quality and quantity, Capital (K) represents the hardware stack, and Labor (L) represents the raw token workload. The residual, A, is everything that turns the same capital and the same labor into more output. For the past three years, venture funds and technology monopolies have disproportionately allocated capital to a single linear bet: if you want more cognitive output, you buy more GPU.

That brute force thesis is already showing diminishing marginal return. Blackstone's joint venture with Google puts up an initial five billion dollars in equity to fund computing power and custom TPUs directly, intentionally decoupling the silicon from traditional data center real estate. Traditional asset managers have financialized compute into its own sovereign asset class, treating it like a standard utility.

The paradox for asset allocators

If raw compute is financialized, scaled, and distributed by private equity players, owning a large pile of chips is no longer a competitive moat. The hardware layer becomes a race to the bottom on margins, and Solow's A is quite literally the alpha in the compute market.

02 · The cracks

The scale of the brute force bet.

2026 capex
$660B

To $690B across the five largest hyperscalers.

Capital intensity
45-57%

Of revenue. A utility ratio.

Under-reported
$176B

Estimated depreciation gap through 2028.

Revenue per head
$2.8M

OpenAI, against a $450K industry average.

The first crack is the accelerating depreciation schedule. By 2023 and 2024, hyperscalers had quietly stretched the assumed useful life of their servers from three or four years out to six, a move that collectively pulled an estimated $18B out of annual depreciation expense and flattered earnings accordingly. Then the schedules started reversing. In early 2025, Amazon shortened the life of a subset of its servers back to five years, citing the accelerating pace of AI hardware development, and booked a corresponding hit to net income. In November 2025, Michael Burry went public with the claim that the sector was understating depreciation by as much as $176B through 2028. When the hardware itself depreciates faster every quarter, owning a large pile of it stops looking like a moat and starts looking like a melting asset.

The second crack sits on the labor side of the ledger, and it cuts against the narrative that AI replaces headcount. As of mid 2025, AI native firms were posting revenue per employee that dwarfs the incumbents: OpenAI near $2.8M per head and Anthropic near $2.5M, against a tech industry average closer to $450,000. That is the prize everyone is chasing. But the chase is turning out to be expensive. By 2026, an Nvidia research executive was saying plainly that for his team the cost of compute now runs higher than the cost of the employees. Microsoft pulled thousands of engineers off Claude Code licenses once the bill became hard to justify, and by April 2026 Uber's CTO reported burning through the company's entire 2026 budget for AI coding tools in only four months. The promise was that a token would be cheaper than a salary. For a growing number of firms, it is not.

Input variableThe old moatThe new paradigm
Capital (K)Accumulating proprietary GPU clustersCommoditized, private equity funded utilities
Labor (L)Massive, homogeneous token volumeSegmented, heterogeneous agentic loops
Residual (A)Static, un-optimized model depthDynamic, runtime token orchestration

Read the matrix down the right hand column and the conclusion writes itself. If the capital is commoditizing, and the labor it was supposed to replace is now more expensive than the compute, the only column left with a defensible margin is the residual: the orchestration layer.

03 · The operating budget

Where the CFO enters the argument.

This stops being an abstract infrastructure debate the moment it touches the operating budget. A CFO needs to forecast and control expenses. Headcount is a known cost, budgeted every year with precision. Compute is a wild card. The cost of an agentic workflow is not a function of users and it is not a function of revenue. It is a function of how much compute the engine decides to spend, and a static engine spends the same compute on a comma as it does on a hard deduction.

The trivial 85%

Most of the token stream is formatting, boilerplate, and predictable continuation. A dense engine runs every layer and every attention head on all of it, at full depth.

The load bearing 15%

A minority of tokens carries the actual reasoning. That is where full compute depth earns its cost, and where accuracy is genuinely at risk if you cut.

04 · The residual, made real

Orchestration as the defensible layer.

Instead of treating the token stream as a uniform block of labor, an intelligent runtime layer reads token signals in real time. It transforms the inference engine from a blunt instrument into a dynamic router, and in doing so it breaks the linear lockstep between hardware spend and reasoning output.

Think of how junior employees tend to do the simpler tasks, while more senior contributors plan, oversee, and handle the most complex parts of a task. This is how AI models should behave, and this is how PrizmalSwitch treats workflows.

2x to 4x efficiency gain in token throughput

That is the Solow residual made real: more output without another dollar added to the hardware bill. It is also how we avoid the Butlerian prediction, by putting the machines back at the service of mankind and not the other way around.

Sources · Futurum Group, Introl, CreditSights, Epoch AI, CNBC, Tom's Hardware, Yahoo Finance, Google Blog.

Run the numbers on your own traffic

The R/E Audit measures reasoning per unit of energy on your real workload, read only and reversible.