Skip to main content
AI Safety Case Study / Architecture · Delivery · Operations

Sentinel A Supervision Layer for Safer LLM Applications

A case study in enforcing content safety, compliance, and output quality without putting every request through the slowest and most expensive model.

Role
Architecture, implementation, and API product design
Ownership
Independent project
Status
Reference implementation; hosted API inactive
Validation
Local functional checks and browser simulation
AI Safety Policy Enforcement FastAPI Cloud Run API Product Design

3 layers

Tiered supervision

Deterministic filters, lightweight classifiers, and an LLM judge cover different risk and latency profiles.

Formerly public

RapidAPI deployment

The API was served through RapidAPI with a Google Cloud Run backend before being deactivated for cost.

Local validation

Evidence boundary

Functionality was exercised locally; no claim is made about sustained production traffic or production SLOs.

My Role

I owned the safety contract end to end.

Sentinel was an independent build. I defined the product boundary, implemented the API and detection pipeline, shaped the verdict contract, deployed the former hosted version, and documented the operational model.

Owned

  • Policy and verdict semantics
  • Tiered detection architecture
  • FastAPI and deployment packaging
  • Simulation and technical narrative

Evidence Boundary

  • No sustained production-traffic claim
  • No published production SLO
  • No benchmarked model-accuracy claim
  • Hosted API is currently inactive

The Problem

Model output is untrusted until policy says otherwise.

LLM applications can leak sensitive data, follow prompt injections, generate unsafe advice, or violate domain policy even when the underlying model is generally capable.

Sentinel inserts an independent supervision boundary between generation and delivery. It evaluates the prompt, draft output, and application context before a response reaches the user.

Architecture

Escalate uncertainty, not every request.

A layered pipeline resolves obvious cases quickly and spends semantic compute only when the decision remains ambiguous.

01 · INPUT

Prompt + Draft

Application context and active policies travel with the request.

02 · FAST PATH

Deterministic

Schemas, regex, PII patterns, and explicit deny rules.

03 · CLASSIFY

Lightweight ML

Toxicity, injection, and domain-risk classifiers.

04 · ESCALATE

LLM Judge

Contextual reasoning for difficult or conflicting signals.

05 · DECIDE

Allow · Redact · Block

Return a verdict, reasons, policy IDs, and audit signals.

Policy plane: versioned rules select checks, thresholds, and enforcement behavior.
Telemetry plane: latency, verdicts, policy hits, and redaction events feed operations.

Key Decisions

Separate policy intent from detection mechanics.

Teams define what must be protected; Sentinel chooses the cheapest reliable mechanism that can enforce it.

DECISION 01

Use tiered supervision instead of one universal judge.

Fast deterministic checks handle known patterns; ambiguity escalates through classifiers to semantic review.

DECISION 02

Return structured verdicts, not opaque scores.

Policy IDs, reasons, evidence, and actions make decisions auditable and easier to integrate.

DECISION 03

Support redaction as a first-class action.

Not every violation requires a hard block; targeted transformation can preserve usefulness while removing risk.

DECISION 04

Keep supervision outside the application runtime.

A dedicated service centralizes enforcement, telemetry, rollout, and SaaS gateway concerns.

Trade-offs

Safety quality is an operating curve.

Thresholds, latency, cost, and user friction move together. Sentinel exposes those choices instead of hiding them behind a single safety score.

Recall vs precision

Aggressive thresholds catch more harmful content but create more review and false-positive cost.

Depth vs latency

Semantic judges improve contextual coverage but belong off the common request path.

Central policy vs autonomy

Shared enforcement reduces drift, while domain teams still need configurable rules and rollout control.

Explainability vs flexibility

Structured reasons constrain implementation but make downstream handling and audit far more reliable.

Evidence

Separate implementation evidence from adoption claims.

The project demonstrates a working product contract and a previous deployment path. It does not establish production scale or detector quality on representative traffic.

Hosted History

RapidAPI + Cloud Run

Proves the service was packaged and exposed through a public API path before deactivation for cost.

Does not prove sustained traffic, availability, or an SLO.

Functional Validation

Local policy behavior

Exercises configured detection, redaction, and verdict flows in the reference implementation.

Does not prove production accuracy on representative datasets.

Interactive Artifact

Browser simulation

Makes PASS, FIX, and FAIL semantics inspectable without depending on an active backend.

Illustrates the contract; it is not the production detector.

Illustrative Local Simulation

Sentinel Decision Contract

Runs deterministic browser-side examples of PASS, FIX, and FAIL behavior. The former RapidAPI and Cloud Run deployment is currently inactive.

Run the illustrative check to inspect a structured supervision verdict.

This interaction demonstrates the product contract, not model accuracy or production performance.

Lesson and Next Step

A safety layer is a decision system, not a filter.

The architecture is only credible when policy quality, failure behavior, rollout, and appeals are measurable. The next iteration would establish reproducible benchmarks first, then evolve a repeatable safety release process.

Adversarial regression suitesVersion jailbreak, PII, toxicity, and domain-policy datasets alongside code and policy changes.
Shadow evaluation and staged rolloutCompare new detectors and thresholds against live traffic before enforcement changes.
Decision-quality monitoringTrack appeals, overrides, false positives, and policy drift as operational SLOs.

Continue the Conversation

Need a technical leader who connects architecture to operations?

I work across AI infrastructure, distributed systems, safety, capacity, and reliability—and I stay close enough to implementation to make the trade-offs concrete.