Moderation policy
Category catalogue and decision thresholds for this deployment. Overridable via POLICY_* env at deploy time.
| Category | Action | Threshold | State |
|---|---|---|---|
Adult / NSFW nsfw | Block (fail) | ≥ 70% | Enabled |
Graphic violence violence | Block (fail) | ≥ 75% | Enabled |
Hate speech hate | Block (fail) | ≥ 60% | Enabled |
Self-harm self_harm | Flag (review) | ≥ 60% | Enabled |
Harassment harassment | Flag (review) | ≥ 65% | Enabled |
Illegal content illegal | Block (fail) | ≥ 60% | Enabled |
Profanity / foul language profanity | Flag (review) | ≥ 80% | Enabled |
Spam / scam spam | Flag (review) | ≥ 85% | Enabled |
Block categories fail a submission outright; review categories flag it for a human. A submission passes only when no enabled category meets its threshold.