Back to Prompt Guard

Synthetic detector evaluation

Scanner 2026-09-28.1; 43 hand-authored synthetic cases. These results are reproducible engineering evidence, not a real-world accuracy estimate, security certification or guarantee.

Each row checks whether the intended detector recognizes the labeled value. The sample deliberately includes obfuscated and truncated sensitive values that this scanner misses. False negatives are retained in the report. Overlapping rules can share a redaction span; these isolated examples do not measure combinations.

DetectorTrue positivesFalse negativesFalse positivesTrue negativesMiss rateFalse-positive rate
api_key410220.0%0.0%
private_key210233.3%0.0%
database_url310225.0%0.0%
email210233.3%0.0%
us_ssn210233.3%0.0%
phone210233.3%0.0%
payment_card210233.3%0.0%
custom_term210233.3%0.0%

Known misses and limits

The corpus exposes misses for split API keys, truncated PEM blocks, encoded database URLs, obfuscated emails, unformatted SSNs/phones, punctuated cards and zero-width custom terms. Names, addresses, medical facts and arbitrary proprietary code are outside comprehensive coverage. A clean scan means only that no supported pattern matched. Never treat it as proof that a prompt is safe.

Local timing

1,000 short fixture scans on linux, v24.19.0: median 0.004 ms, p95 0.009 ms, maximum 0.663 ms. This excludes all network, TLS, storage and provider time. It is neither hosted request latency nor a service-level commitment.

Reproduce and evaluate your policy

Download synthetic corpus ยท Download every result

The repository script scripts/tools/evaluate-prompt-guard.mjs regenerates this report using the shipped scanner. Use the playground's shadow evaluation to inspect representative text locally before choosing a production policy. Shadow evaluation never sends data to a provider and never changes the hosted policy; protected hosted traffic continues to block or redact.