The counter
Roughly 1.2 billion tokens per second, as a lower bound.
Basis: derived from a public statement of approximately 3.2 quadrillion tokens processed per month at one hyperscaler in 2025. That is 3.2e15 divided by 2.6 million seconds, or about 1.235 billion tokens per second. Industry-wide volume is materially higher, so treat this as a conservative lower bound.
Force gauges
Where each force stands.
Fabs contributing late 2027.
At the frontier.
The RAM side of the stack.
The processor side, governed by PrizmalRun.
What changed
Dated events behind the curves.
| Date | Event | Source |
|---|---|---|
| 2026-06-04 | Major cloud provider disclosed HBM allocation as a binding 2027 constraint on inference capacity. | Public earnings commentary, Q2 2026 |
| 2026-05-28 | Hybrid routing entered the per-call lane; per-token routing remains the structurally cheaper plane. | Computex 2026 keynotes |
| 2026-05-12 | 66% of large enterprises reported at least one AI workload in production. | Deloitte TMT Predictions 2026 |
| 2026-04-22 | DRAM contract prices up 90 to 95% quarter over quarter, the largest jump in a decade. | Counterpoint Research and TrendForce, April 2026 |
| 2026-03-15 | Independent analysis put the served-inference gross-margin ceiling near 50% under current cost structure. | Sacra research note, March 2026 |
| 2026-02-08 | Leading-edge foundry capacity reported sold out through 2027; new fabs begin contributing late 2027. | Industry foundry roadmap disclosures, 2026 |
Three days, your traffic, a guaranteed result. The audit runs in your environment at $25,000, reserved with a $1,000 refundable deposit. At least 15% inference savings identified or the full amount is refunded. On a $10M bill, that floor is $1.5M.