01 · The frame
The token is the unit. Almost all of it is unoptimized.
Jensen Huang announced it at GTC: engineers will earn tokens as compensation, on top of salary. Chamath Palihapitiya is already enforcing token budgets to avoid running out of money. Every serious enterprise will consume 10 billion tokens per employee per year by 2027. The question is no longer whether you use AI. It is how much of that compute is wasted.
Enterprise consumption in 2026 as agentic workflows scale.
Every LLM runs at full depth on every token, trivial or complex.
Reported by Jason Calacanis for agents at partial capacity.
FLOPs reduction at inference time. 99% reasoning retained.
02 · The problem
Every token runs at full cost. Most of them do not need to.
Jensen Huang told GTC in March 2026 that he would give engineers half their salary again in AI tokens. That statement codifies what enterprise finance teams are already discovering. Tokens are a commodity with a cost, and the volume is accelerating faster than anyone modeled. By mid-2025, OpenRouter alone was routing 100 trillion tokens per year, with over 1 trillion tokens processed in a single day.
The structural problem is not the volume. It is the waste built into how transformers execute. Every LLM today runs at full capacity on every single token, whether the task requires deep reasoning or a trivial lookup. Agents that route correctly outperform brute-force frontier calls at scale, because the cost per token is not the issue. The issue is that most tokens cost too much for what they actually deliver.
"I'm going to give engineers probably half their base pay on top of their salary as tokens. Tokens are becoming one of the recruiting tools in Silicon Valley."
"I've been forced to institute token budgets for my top developers. Without them, I'll run out of money. The usage far exceeded every projection."
"Enterprise spending on generative AI hit $37 billion in 2025, a 3.2x increase from 2024. 98% of organizations now actively manage AI spend. The era of all-you-can-eat AI pricing is structurally unsustainable."
03 · The math
10 billion tokens per employee. Here is what that costs.
Model this at current mid-tier API pricing of $2.50 per million input tokens. A 1,000-person enterprise consuming 10 billion tokens per employee per year generates 10 trillion tokens annually. That is a $25 million per year inference bill at full compute. With PrizmalRun cutting 60% of FLOPs while retaining 99% of reasoning quality, the same intelligence output costs $10 million. $15 million saved, without retraining a single model or modifying a single weight.
| Line | Without PrizmalRun | With PrizmalRun |
|---|---|---|
| Employees | 1,000 | 1,000 |
| Tokens per employee per year | 10B | 10B |
| Total tokens per year | 10T | 10T |
| Average price per M tokens | $2.50 | $2.50 |
| Annual inference cost | $25M | $10M |
| Reasoning quality | 100% | 99% |
| Compute wasted | ~90% | Governed per token |
| Annual savings | None | $15M |
Of 10 billion tokens consumed per employee per year, SpeCo classifies and routes each one in real time. Roughly 45% are low complexity, lookups, formatting, and short completions, which skip 60 to 80% of layers. Another 35% are mid complexity, summarization, analysis, and structured generation, which skip 20 to 40% of layers. Only 20% are high complexity, multi-step reasoning, code generation, and chain of thought, which run at full depth.
04 · Measured performance
The numbers are not projected. They are measured.
All results measured under controlled A/B conditions using NVML energy instrumentation on Llama-2 7B and Mistral 7B. Full methodology in the technical appendix.
| Metric | Without PrizmalRun | With PrizmalRun | Business impact |
|---|---|---|---|
| Energy per inference | 95.4 J | 71.5 J | 25.1% reduction in power and cooling cost |
| Reasoning per joule (R/E) | 0.275 RPJ | 0.965 RPJ | +251% more intelligence per energy dollar |
| GSM8K reasoning score | 28.0% | 55.0% | +96.4% on complex multi-step tasks |
| CoQA coherence | 93.0% | 92.6% | Quality preserved across conversational tasks |
| Tokens per second | 38.2 t/s | 46.9 t/s | +22.8% throughput on the same hardware |
| Compute cost (FLOPs) | Baseline | -60% | $15M saved annually at 1,000 employees |
A system that produces 5,000 trivial tokens per second scores identically to one that produces 5,000 tokens of substantive reasoning. The metric optimizes for throughput. In agentic deployments where agents run for hours, coordinate across tools, and reason dynamically, throughput is not the constraint. Intelligence per joule is.
05 · How it works
Token-level compute modulation.
SpeCo scores each token for difficulty in real time and adjusts layer depth, attention head count, and numeric precision before that token resolves. Simple lookups skip layers. Complex reasoning chains get the full model.
KV cache depth control, sparse KV retention, and reduced context fetch width shrink the active memory footprint per inference. Models that previously exceeded device DRAM now fit and run stably.
More correct answers delivered per unit of energy than any static approach. GRIFFIN reaches 72%. CATS reaches 104% on its training distribution and collapses outside it.
Integrates directly into ONNX Runtime, ExecuTorch, and vendor stacks from Qualcomm, Intel, and AMD. No changes to training pipelines. Runtime control only.
It is the control system for AI compute. Simple tokens get less compute. Complex reasoning tokens get full depth. Every token, every inference. No retraining, no fine-tuning, no model modification.