TypeSafe LabJev · System One
connecting

04 · Guardrail

Prompt pre-flight

The testA guardrail in front of an image generator: one call, ten questions, before any credits are spent or a provider sees the request. The questions and thresholds come from the NanoStudioPro analysis (148 real prompts, 23 people) in Experiments.
How to read itThe verdict applies four rules in order: block + strike (risk > 0.75, or an explicit / undress / minors category above 70%), allow but log (grey-zone risk or any flagged category above 50%), images-only notice in quick mode (reads like a chat message or refers to a previous result), otherwise OK. The strike counter below is per session.
What to watchRead the category and the risk score: on edge cases they disagree. Try a blunt request, a euphemism, a request phrased as a question, a Spanish one, and "the same but in blue" with no image attached.
Typed in
Request · exactly what was sent to the model
Full response · every probability · raw JSON