Product

One control layer

PrizmalSwitch decides. PrizmalRun delivers. Warden governs. The bulk of your work runs on open models, and frontier is the escalation path, on your own keys, only when the work requires it.

PrizmalSwitchPrizmalRunWarden

01 · Routing

PrizmalSwitch decides per request

One OpenAI compatible key in front of every request. It reads each one and keeps deciding while the answer is written, so routine spans run on the best capable open class and frontier is engaged only for the spans that require it.

Step 1
Request
Your app calls api.prizmal.ai with model set to auto.
Step 2
PrizmalSwitch
Per request decision on live decoder signal, not static rules.
Step 3
Open classes
Default to the best capable class for the prompt.
Step 4
Frontier
Escalate to your own frontier keys only when the work requires it.

Provider outages do not become your outages. Failover preserves the session: state, context, and tools survive the reroute. 99.9% SLA, contractual, with credits.

02 · Runtime

PrizmalRun governs compute token by token

The runtime that operates the open models. Inside each one, live decoder signals decide how much compute the answer actually needs: dense where reasoning is at stake, light where it is not. Built on RASP.

R/E vs dense
+251%

RASP · Llama-3 8B · 50% sparsity

Less compute
60%

RASP

Less energy
~25%

NVML device level

Reasoning kept
99%

At 50% sparsity

Compression governs RAM. PrizmalRun governs compute. They compose. They do not compete.

03 · Method

The science, in brief

R/E means reasoning per unit of energy: evaluation suite score per joule at the device, NVML for energy, standard reasoning suites for quality. Prizmal prices and reports inference in R/E because tokens per dollar ignores how much intelligence each token actually buys.

How the numbers are made
  • Models tested: Llama-3 8B, Gemma-2 9B, DeepSeek-R1 7B
  • Sparsity setting: 50%
  • Evaluation suites: Terminal-Bench v2, SWE-bench, HumanEval, DeepSWE
  • Baseline: dense inference, same weights, same hardware
  • Ahead of academic sparsity baselines: +7.93 vs WANDA, +7.33 vs GRIFFIN, +2.10 vs TEAL on Gemma-2 9B IT at 50% sparsity
Provenance

RASP paper published June 2026. The method is patented, not open sourced. Five USPTO provisional patent families, filed December 2025, PCT in progress: Isocline Pruning, Draft-Guided Router, Disagreement Control Loop, Layer Sensitivity Caps, Masked Matrix Multiply. The moat is on the mechanism, not the weights.

Methods questions: research@prizmal.ai

Read the research

04 · Governance

Warden governs every token

Policy, identity, and an immutable record on every token, with a kill switch that works at runtime.

G-01
Policy and guardrails
Per tenant, per team, per task policies. Allowed models, allowed data classes, allowed regions. Enforced at the token, not the dashboard.
G-02
Identity aware QoS
SSO and SCIM. Quotas, caps, and priority bound to identity. The CFO seat does not pay for the intern's runaway loop.
G-03
Sovereignty
On-prem, sovereign cloud, regional residency. Choose where the prompt lives and where the answer is computed.
G-04
Kill switch at runtime
One control plane decision halts a model, a tenant, a route, or a class within seconds. Not a ticket, a switch.
Before Warden
AI spend is an opaque line item. Prompts in provider logs, budgets discovered at invoice, no kill switch, audit by screenshot.
After Warden
Every token classified and policied. Budgets enforced mid-stream. One switch halts anything. Audit is an export, not an archaeology.

The R/E Audit

3 days in shadow mode on your real traffic. At least 15% in savings identified, or the $25,000 fee is refunded, deposit included.