RESEARCH
Mid-query, token-based decisions
A Prizmal answer is written by several models, moved between them without loss, and governed token by token. Here is what we can show, and what we keep.
The pieces
RASP makes open models radically cheaper per token without losing the reasoning. Training free: the principal structure stays dense, the residual is pruned.
These are 7B to 9B class results. The bet, and the patents, are about what happens above.
Mid-answer, the writing model can change. The full working context moves with it: every fact, every constraint, every open thread of the reasoning. Nothing is summarized away, nothing is re-asked, and the sentence lands as if one mind wrote it. In a routed answer the model can change between one word and the next, and the seam is not visible.
Moving a live context between two different models without loss is the hardest problem on this page, and the one we say the least about.
What we publish and what we keep
We publish results, benchmark scores, recipe outcomes, and audit formats. We keep the estimator, the switch policy, and the transfer method.
Six patent families
Less compute, same answer
Identifies, per token, computation whose removal leaves the output unchanged at zero divergence. The removable set is knowable before the work is done.
How much model a token gets
A generating model rarely needs all of itself. This family sets, live, how much of the network each token engages. Deciding costs less than the compute it saves.
The recovery reflex
Aggressive efficiency requires a failure detector. When output confidence degrades, the compute envelope widens inside the same answer, then narrows once the risk passes.
Error, bounded by design
Approximation error compounds with depth. This family caps the error each layer may contribute, so the total stays engineered rather than emergent.
Speed on the metal
Sparsity pays only when hardware can skip what theory removed. This family turns pruned structure into kernels that keep their speedup on real accelerators.
The move without loss
A live generation carries state: facts, constraints, the open threads of a reasoning. This family moves that state across models mid-answer, nothing summarized and nothing re-derived. The seam does not appear in the text.
Six USPTO patent families, the first filed December 2025. Legal titles publish with the filings.
Application serials on file include US 63/928,790, US 63/928,818, and US 64/072,126.
If you work on these problems, the interesting question is not which model to call. It is when a token stops being cheap, and how a thought survives a change of mind.
BenchAlign note
BenchAlign v5 is benchlm.ai's reasoning suite, scored 0 to 100. A routed Prizmal answer is scored the way any single model is: the suite sees answers, not the mix. Prizmal's figure is 84, July 2026.
The rest is in the audit trail.