Engineering
What running AI guardrails in production taught us
On a slide, guardrails are a box between the user and the model: screen the input, screen the output, block what violates policy. In production, that box sits on the critical path of every single message, and the engineering is in the details the slide leaves out. Here is what operating input and output screening on real traffic taught us.
Latency is the first constraint, not the last
Users tolerate a model taking time to think. They do not tolerate a blank screen before the first token appears. Every check you add to the input path (prompt-injection detection, PII scanning, policy classification) happens before the model even starts. We learned to treat screening as a strict latency budget: run checks in parallel, keep the fast deterministic ones (patterns, secrets formats) ahead of the slower model-based ones, and measure the p95, not the average. Output screening has a different shape: responses stream, so checks have to work on text as it arrives rather than waiting for the full reply.
False positives cost more than false negatives
A missed violation is a risk; a wrongly blocked legitimate request is a lesson every user remembers. Block a lawyer for asking about personal data in a contract clause, a completely normal work task, and they conclude the tool is useless and route around it. We tune for precision on blocking, and prefer softer interventions where policy allows: redacting the specific token rather than refusing the message, or flagging for review rather than stopping the conversation. Every block message says why, in plain language. “Request blocked by policy” with no explanation generates a support ticket and an annoyed user; a specific reason usually generates a rephrased, compliant request.
Decisions need an audit trail
The moment guardrails act on real traffic, someone will ask: what was blocked, why, and was the decision right? Every screening decision (pass, redact, block) needs a log entry with the rule that fired and enough context to review it later. That audit trail turned out to be as valuable as the blocking itself: it is how you prove to a security team what the system caught, how you find rules that misfire, and how you demonstrate compliance without arguing from theory.
Guardrails must survive a model swap
The most important architectural decision we made was keeping guardrails independent of the models they wrap. Screening runs in its own layer in ChatLite, so when a new model is added, or a team switches mid-conversation from one provider to another, the same policies apply with zero migration. If your safety checks are implemented inside one vendor’s ecosystem, every model change is also a security project. Decoupling them is what makes a multi-model workspace governable at all.