zeph-sanitizer
Content sanitization, exfiltration guard, PII filtering, and quarantine for Zeph — untrusted input isolation before LLM context injection.
Overview
Implements a multi-stage security pipeline that processes all external data before it enters the LLM context window. The pipeline detects prompt injection patterns, wraps content in spotlighting XML delimiters, optionally routes high-risk sources through an isolated quarantine LLM call, and guards outbound paths against data exfiltration. Memory retrieval sources are classified via MemorySourceHint to suppress false positive injection flags on recalled user conversations and LLM-generated summaries.
Key types
| Type | Description |
|---|---|
ContentSanitizer |
4-step pipeline: truncate → strip control chars → detect injections → spotlighting XML wrap |
ContentTrustLevel |
Trusted / LocalUntrusted / ExternalUntrusted |
ContentSourceKind |
Source category (tool output, web scrape, document, etc.) |
SanitizedContent |
Output with body, source, injection_flags, and was_truncated |
InjectionFlag |
Detected injection pattern (pattern_name, byte_offset, matched_text) |
pii::PiiFilter |
Regex PII scrubber (email, phone, SSN, credit card; opt-in name heuristic) |
guardrail::GuardrailFilter |
LLM-based pre-screener at the input boundary |
quarantine::QuarantinedSummarizer |
Dual LLM pattern — routes high-risk content through an isolated, tool-less LLM call |
response_verifier::ResponseVerifier |
Post-LLM response scanner |
exfiltration::ExfiltrationGuard |
Three outbound guards: markdown image tracking, tool URL cross-validation, memory write suppression |
memory_validation::MemoryWriteValidator |
Structural write guards for the memory store |
causal_ipi::TurnCausalAnalyzer |
Behavioral deviation detection at tool-return boundaries |
nli::NliSanitizer |
Probabilistic NLI entailment check for injected instructions |
secret_mask::SecretMaskRegistry |
Vault-secret placeholder masking at the LLM boundary |
ipi_filter::IpiFilter / IpiVerdict |
Indirect prompt injection filter and verdict type |
ContentSource |
Source metadata with ContentSourceKind and optional MemorySourceHint for memory retrieval classification |
MemorySourceHint |
ConversationHistory / LlmSummary / ExternalContent — classifies memory retrieval sources to suppress false positive injection flags on recalled user text and LLM-generated summaries |
media::MediaSanitizer |
Validation pipeline for MCP-sourced images: magic-byte vs declared-MIME check, format allowlist, byte-size cap, and decoded-dimension/pixel caps (decompression-bomb defense) before an image is attached as a native MessagePart::Image |
media::MediaRejected |
Typed rejection reason (SizeExceeded, DimensionExceeded, MimeMismatch, DecodeFailed, format-not-allowed); the text placeholder always remains as a fallback |
Architecture
The crate is a layered defense-in-depth pipeline; each layer is independently configurable and optional except layer 1:
| Layer | Type | Description |
|---|---|---|
| 1 | ContentSanitizer |
Regex-based injection detection + spotlighting |
| 2 | pii::PiiFilter |
Regex PII scrubber (email, phone, SSN, credit card) |
| 3 | guardrail::GuardrailFilter |
LLM-based pre-screener at the input boundary |
| 4 | quarantine::QuarantinedSummarizer |
Isolated LLM fact extractor |
| 5 | response_verifier::ResponseVerifier |
Post-LLM response scanner |
| 6 | exfiltration::ExfiltrationGuard |
Outbound channel guards (markdown images, tool URLs) |
| 7 | memory_validation::MemoryWriteValidator |
Structural write guards for the memory store |
| 8 | causal_ipi::TurnCausalAnalyzer |
Behavioral deviation detection at tool-return boundaries |
| 9 | nli::NliSanitizer |
Probabilistic NLI entailment check for injected instructions |
| 10 | secret_mask::SecretMaskRegistry |
Vault-secret placeholder masking at the LLM boundary |
[!NOTE]
media::MediaSanitizeris a separate, image-specific validation pipeline for MCP tool-result passthrough ([mcp.media]) — it does not sit in the text-content layer chain above and is invoked directly byzeph-mcp's tool executor when a server hasmedia_passthroughenabled.
Sanitization pipeline (layer 1 detail)
External data
↓ 1. Truncate to max_content_size
↓ 2. Strip null bytes and control characters
↓ 3. Detect 17 injection patterns (OWASP variants + encoding)
↓ 4. Wrap in spotlighting XML delimiters
<tool-output>…</tool-output> (local sources)
<external-data>…</external-data> (external sources)
Usage
use ;
use ContentIsolationConfig;
let config = default;
let sanitizer = new;
let source = new;
let result = sanitizer.sanitize;
// result.body contains the wrapped, injection-cleaned text
// result.injection_flags contains any detected patterns (advisory — content is never removed)
for flag in &result.injection_flags
Configuration
[]
= true
= 65536 # bytes; content truncated before injection detection
[]
= true
= ["web_scrape", "a2a_message"] # source kinds routed through quarantine
= "claude-haiku-4-5-20251001" # optional; defaults to primary provider
= 2048
[]
= true
= true
= true
= true
[]
= true # default: true — scrubs email/phone/SSN/credit-card before LLM context and debug dumps
= false # opt-in: higher-recall, lower-precision name heuristic
[]
= true # default: true — vault-resolved secrets replaced with placeholders before outbound LLM calls
[!NOTE]
PiiFilterConfigandSecretMaskingConfigboth default toenabled = true. An operator's explicitenabled = falsein an existingconfig.tomlis always respected.
Features
zeph-memory (and transitively zeph-db) needs exactly one backend selected to compile; sqlite is the default so the crate builds in isolation.
| Feature | Default | Description |
|---|---|---|
sqlite |
yes | SQLite backend for zeph-memory |
postgres |
no | PostgreSQL backend for zeph-memory |
classifiers |
no | ML-backed injection detection (classify_injection) and NER-based PII detection (detect_pii); requires an attached classifier backend via with_classifier / with_pii_detector |
Security metrics
ContentSanitizer exposes metrics via the shared MetricsSnapshot:
| Metric | Description |
|---|---|
sanitizer_runs |
Total sanitization invocations |
sanitizer_injection_flags |
Cumulative injection pattern detections |
sanitizer_truncations |
Content truncations applied |
quarantine_invocations |
Quarantine LLM calls triggered |
quarantine_failures |
Quarantine LLM call failures (falls back to direct sanitization) |
exfiltration_images_blocked |
Markdown image pixel-tracking attempts blocked |
exfiltration_tool_urls_flagged |
Tool URLs cross-validated against untrusted sources |
exfiltration_memory_guards |
Memory write suppression events |
Installation
Documentation
Full documentation: https://bug-ops.github.io/zeph/
License
Licensed under either of MIT or Apache License, Version 2.0 at your option.