A second opinion for every action.
One call turns any text in 100+ languages into a calibrated verdict: act, review, escalate, block.
One call answers every question.
Department, urgency, churn, refund. Scored together, routed by language before the model ever sees it.
Try it in the consoleLanguages with a router that picks the right checkpoint per request.
A single decision on one GPU. Batched, about a millisecond each.
Thresholds you set. ≥ 0.85 moves on its own. Everything else finds a human.
Self-hosted weights, Apache-2.0. Safety stops being a line item.
Calibrated, not merely confident.
Probabilities trained with proper scoring rules, then temperature-fit on held-out data. When the gate says 0.89, it means it.
See live metricsExpected calibration error after fitting. Down from 0.466 raw.
Languages clearing 3× random, routed automatically.
Faster than hosted LLM judges at zero marginal cost.
Confidence you can wire to automation.
From policy to verdict in three moves.
Write the policy
A few typed questions in YAML: choice, score, or yes-or-no. No training, no prompts to babysit.
Make one call
Send any state in any language. The router picks the checkpoint; every question scores in a single forward pass.
Act on the verdict
High confidence moves on its own. Everything else lands in front of a human with the reason attached.
Four policies cover nearly everything.
LLM firewall
Jailbreaks, injections and leaks get blocked before your model sees them.
Verdicts, fresh from the gate.
“Billed twice for March. Refund it today or we cancel.”
Ship calibrated doubt.
One container, one policy file, every stack you already use.