Where I Put Safety Checks in an LLM Serving Path
A practical way to place LLM safety checks without turning every request into another full model inference.
Table of Contents
When people draw an LLM safety architecture, it often becomes a chain of boxes: validate the input, run a classifier, ask a second model, filter the output, log everything.
The boxes are reasonable. The difficult part is deciding which requests visit which boxes—and what the system does when a safety dependency is slow or unavailable.
My default is a risk-based serving path: apply deterministic checks to every request, reserve expensive semantic checks for traffic that needs them, and inspect the output independently from the input.
request
-> protocol and schema validation
-> deterministic policy checks
-> risk scoring
-> low risk: model inference
-> elevated risk: semantic safety check -> model inference
-> output policy and data-loss checks
-> response or controlled refusal
Start with failures that do not require a model
Malformed encodings, unsupported content types, oversized payloads, invalid tool schemas, and impossible parameter combinations should fail before any model is involved.
This layer protects both safety and capacity. An input that can be rejected deterministically should not consume accelerator time or a safety-model budget.
I also keep explicit policies here when they can be expressed precisely: allowed tools, maximum argument size, tenant permissions, and known data classifications. A language model should not make a decision that an access-control system can make exactly.
Use classifiers to route, not to declare truth
A fast classifier or heuristic score is useful for deciding how much scrutiny a request needs. It is less useful as an unquestionable final verdict.
Every classifier has a threshold, and every threshold represents a trade-off between false positives and false negatives. That threshold should be chosen from the product’s risk tolerance and measured on representative traffic—not copied from an offline benchmark.
I would usually let the score select among three paths:
- continue normally;
- send the request through a stronger semantic check;
- reject immediately when a deterministic policy is already violated.
This keeps the common path short without pretending that all traffic has the same risk.
Check outputs separately
A clean prompt does not guarantee a clean response. The model can reveal sensitive data, produce disallowed content, or generate unsafe tool arguments without an obviously malicious input.
Output checks therefore need their own policy. For text, that may include data-loss detection, policy classification, or redaction. For tool use, I care more about typed arguments, authorization, and the effect of the proposed action.
The system should distinguish between blocking, redacting, asking for confirmation, and returning a safe fallback. Treating every violation as the same generic refusal makes the product harder to use and the incidents harder to investigate.
Failure policy is a product decision
“Fail closed” sounds safe, but applied without context it can turn a minor dependency outage into a full product outage.
I prefer to define failure behavior per capability:
- a read-only summarization feature may degrade to a restricted mode;
- an external side effect may require the safety service and authorization path to be healthy;
- a high-risk request may be refused when its semantic check times out;
- an internal low-risk request may continue with tighter output controls.
The important part is that this behavior is deliberate and testable. A timeout should not accidentally choose the safety policy.
Measure the routing system
Aggregate block rate is not enough. I want to know:
- how much traffic each layer receives;
- latency added by each layer at p50, p95, and p99;
- disagreement rates between the fast and deep checks;
- false-positive reports by policy and tenant;
- timeout and fallback frequency;
- and which output actions were blocked after the input passed.
Those measurements tell us whether the layered design is actually selective. If nearly every request reaches the expensive check, the routing layer is not doing useful work. If nothing reaches it, either the product is unusually safe or the threshold is wrong.
The design rule I keep
LLM safety is not one model placed in front of another model. It is a serving policy that combines deterministic controls, statistical judgment, authorization, and explicit failure behavior.
Put exact checks where exact checks are possible. Spend model inference where context is necessary. Then measure the disagreements, because that is where the policy will need to evolve.