Product
One control layer
PrizmalSwitch decides. PrizmalRun delivers. Warden governs. The bulk of your work runs on open models, and frontier is the escalation path, on your own keys, only when the work requires it.
01 · Routing
PrizmalSwitch decides per request
One OpenAI compatible key in front of every request. It reads each one and keeps deciding while the answer is written, so routine spans run on the best capable open class and frontier is engaged only for the spans that require it.
Provider outages do not become your outages. Failover preserves the session: state, context, and tools survive the reroute. 99.9% SLA, contractual, with credits.
02 · Runtime
PrizmalRun governs compute token by token
The runtime that operates the open models. Inside each one, live decoder signals decide how much compute the answer actually needs: dense where reasoning is at stake, light where it is not. Built on RASP.
RASP · Llama-3 8B · 50% sparsity
RASP
NVML device level
At 50% sparsity
Compression governs RAM. PrizmalRun governs compute. They compose. They do not compete.
03 · Method
The science, in brief
R/E means reasoning per unit of energy: evaluation suite score per joule at the device, NVML for energy, standard reasoning suites for quality. Prizmal prices and reports inference in R/E because tokens per dollar ignores how much intelligence each token actually buys.
- Models tested: Llama-3 8B, Gemma-2 9B, DeepSeek-R1 7B
- Sparsity setting: 50%
- Evaluation suites: Terminal-Bench v2, SWE-bench, HumanEval, DeepSWE
- Baseline: dense inference, same weights, same hardware
- Ahead of academic sparsity baselines: +7.93 vs WANDA, +7.33 vs GRIFFIN, +2.10 vs TEAL on Gemma-2 9B IT at 50% sparsity
RASP paper published June 2026. The method is patented, not open sourced. Five USPTO provisional patent families, filed December 2025, PCT in progress: Isocline Pruning, Draft-Guided Router, Disagreement Control Loop, Layer Sensitivity Caps, Masked Matrix Multiply. The moat is on the mechanism, not the weights.
04 · Governance
Warden governs every token
Policy, identity, and an immutable record on every token, with a kill switch that works at runtime.
The R/E Audit
3 days in shadow mode on your real traffic. At least 15% in savings identified, or the $25,000 fee is refunded, deposit included.