TypeSafe Jev · System One demo
Moderation Guard
Type text, get parallel Noul hazard probabilities + severity Score from one TypeSafe request. Verdict computed in code via policy thresholds.
74/1000
Ready
policy strict · severity 0.00/3 · awaiting check
Profanity · Cursing / obscene language0.0%
Hate · Identity attack / hate0.0%
Sexual · Explicit sexual content0.0%
Violence · Threats / gore0.0%
Illegal · Crime / illicit help0.0%
Self-harm · Suicidal / self-injury0.0%
Spam / Jailbreak · Scam, spam, prompt injection0.0%
Thresholds: review ≥0.35, action ≥0.70, severity blocks at ≥2.0. Hit Check to screen the text on the left.