Guardrails — catch it before it reaches, or leaves, the model
Content guardrails scan both directions of every call —
what's sent to the model and what comes back — against built-in threat
packs and rules your own admins write. Applied inline, streaming included,
with independent control down to the individual billing node.
Detection
Three built-in threat packs
Prompt-injection, jailbreak, and PII patterns, ready to turn on —
protection from day one, no security hire required.
Write your own rules
Any tenant admin can add a custom regex-based rule to cover risks
unique to the business — leaked project codenames, confidential terms,
anything a pattern can catch.
Input and output, not just input
Rules scan what users send and what the model sends back —
stopping leaks going out, not only attacks coming in.
Control
Four states, not two
Every rule is Disabled
Flag Redact
Block — match the response to the risk
instead of an all-or-nothing switch.
Assign by billing node
Scope a rule to one team's node, several nodes, or the whole account
— different products, different rules, one account.
Scope to one key or teammate
Restrict a rule to a single API key or member, so a new rule can be
trialled safely before it's rolled out to everyone.
Operations
One-click kill switch
A single toggle suspends every guardrail, account-wide — undo a
misfiring rule instantly, with zero downtime.
Every rule counts its hits
Live match counters on every rule, built-in and custom, so you can
tune false positives with real traffic data, not guesswork.
Safe even mid-stream
Redaction works correctly on live streamed responses without
corrupting the stream — production-grade for real chat products.
Guardrails are one layer of defense-in-depth, configured per tenant from
the admin console — they run alongside, not instead of, your own
application-level checks.