Guardrails
The Guardrails block checks text against safety and policy rules before your workflow acts on it. Use it to filter incoming user messages, or to check AI responses before they are sent out.
It is a branching block with two outputs:
- Pass — every enabled check passed. Continue the normal flow.
- Fail — at least one check failed. Route to a fallback (e.g. a polite refusal message or human review).
Two positions: Put a Guardrails block right after a trigger to screen user messages, or right after an AI block to screen AI responses before delivery.
Key Features
- Pass / Fail Branching: Connect each outcome to its own flow
- Input and Output Modes: Choose whether you are checking user messages or AI responses
- Built-in Checks: Toggle ready-made guardrails — no prompts to write
- Custom Guardrails: Add your own rules in plain English
- Result Variables: Save the overall result and per-check details into variables
Configuration
| Parameter | Type | Required | Description |
|---|---|---|---|
| Response type | Dropdown | Yes | User Messages (input) or AI Responses (output) |
| Text to check | Text | Yes | The text to evaluate, e.g. {{user_message}} or {{ai_response}} |
| Default guardrails | Toggles | No | The built-in checks listed below |
| Custom guardrails | List | No | Your own plain-English rules |
| Response mapping | List | No | Save results into workflow variables |
Default guardrails — User Messages (input)
| Toggle | What it does | Extra settings |
|---|---|---|
| Block API keys & passwords | Blocks messages containing secrets like API keys or passwords | — |
| Block prompt injection | Detects attempts to override your bot's instructions | — |
| Block jailbreak attempts | Detects jailbreak-style prompts | Detection threshold (0–1) |
| Block prompt leaking | Blocks attempts to make the bot reveal its system prompt | — |
| Block offensive language | Filters toxic or abusive messages | — |
| Stay on topic | Only allows messages about your listed topics | Allowed topics |
| Language filter | Only allows messages in the languages you list | Languages |
| Block spam & gibberish | Filters random characters and incoherent messages | Detection threshold (0–1) |
| Limit message length | Blocks messages over a token limit (protects against context-stuffing; ~4 characters per token) | Max tokens (min 100) |
Default guardrails — AI Responses (output)
| Toggle | What it does | Extra settings |
|---|---|---|
| Prevent hallucinations | Checks the response against a trusted document for factual accuracy | Reference document |
| Block harmful responses | Filters toxic, hateful, or harmful content before users see it | — |
| Enforce response format | Fails responses that are not in the expected format | Expected format: JSON, Markdown, Plain Text, or Custom (with your own format description) |
| Brand voice check | Checks the response against your tone and style rules | Brand voice guidelines |
| Block competitor mentions | Fails responses that mention listed competitors | Competitor names |
| Enforce company policies | Checks the response against your written policies | Your policy rules |
| Require minimum confidence | Fails when the AI's confidence in its own response is below the threshold | Minimum confidence score (0–1) |
Custom guardrails
Click Add guardrail to write your own rule. No coding needed — each rule has:
| Field | Description |
|---|---|
| Guardrail name | Short label, e.g. No Discount Over 20% |
| Rule | The rule in plain English, e.g. "Never suggest a competitor product" |
Setting up the block
- Add the Guardrails block where you want the check (after the trigger for input, after the AI block for output).
- Set Response type to User Messages or AI Responses.
- Set Text to check to the variable holding the text.
- Toggle the default guardrails you need and fill in their extra settings.
- Add any custom guardrails.
- Connect the Pass output to the normal flow and the Fail output to your fallback flow.
Example: Safe support bot
Trigger → Guardrails (User Messages: Block prompt injection + Block offensive language + Stay on topic: "product support") → LLM Agent → Guardrails (AI Responses: Block harmful responses + Block competitor mentions) → Send reply.
If either check fails, the Fail branch sends "Sorry, I can't help with that" instead.
Save answer
Use Response mapping to store results in variables:
| Value | Description |
|---|---|
| Overall result (true/false) | true when every enabled check passed |
| All check results (JSON) | The result of every check that ran |
| Failed checks only (JSON) | Only the checks that failed, with reasons |