Guide
AI guardrails explained
AI guardrails are automated safety checks that screen the inputs to and outputs from an AI model, blocking or redacting content that violates policy (such as personal data, secrets, or prompt-injection attempts) before it reaches a user. They run independently of the model itself, so they apply consistently even when the underlying model changes. Guardrails turn loosely defined usage policies into enforceable, machine-checked rules.
What risks do guardrails address?
Guardrails exist because language models will faithfully repeat whatever they are given and generate whatever they are asked, with no built-in notion of what an organization considers acceptable. They close that gap by inspecting traffic in both directions.
- PII, secret, and credential leakage: personal data, API keys, passwords, or tokens being sent into or returned from a model.
- Prompt injection and jailbreaks: crafted inputs that try to override system instructions or bypass safety rules.
- Off-policy content: responses that fall outside what an organization permits for a given team or use case.
- Unverified output: answers that are released without being checked against policy or expected formats.
Where guardrails run
Effective guardrails run on both input and output: every prompt and every response. On the input side they catch sensitive data and injection attempts before the model sees them. On the output side they screen what the model produces before it reaches the user, redacting or blocking anything that breaches policy.
Checking only one direction leaves a gap: input-only filtering misses unsafe generations, and output-only filtering lets sensitive data and attacks reach the model unchecked. Bidirectional screening is what makes the control reliable.
How ChatLite applies guardrails
ChatLite enforces organization policies on every prompt and response. It detects and redacts PII, secrets, and credentials, and blocks prompt injection and jailbreak attempts before they reach the model.
Content policies can be tuned per team or department, so different parts of an organization can operate under rules that fit their needs. ChatLite keeps a full audit trail of what was screened, redacted, or blocked, and lets admins enable or disable specific detectors per organization.
Learn more about ChatLite's guardrails and how they fit into the broader security model.