What Is AI Moderation, and Does Discord Need It?
How Airwavy's 14-category AI Moderation layer works, what it actually catches that deterministic filters miss, and its real limits.
Quick answer
AI Moderation is a Premium, opt-in second-opinion layer that reviews message candidates deterministic signals already flagged — scam-like links, configured phrases, spam, mass mentions, suspicious attachments, or high-risk channels — across 14 safety categories. It never bans, kicks, or times out anyone; it only alerts staff or deletes a confirmed-unsafe message, layered on top of Airwavy's regular moderation tools, not replacing them.
What you'll learn
- What AI Moderation actually reviews, and what triggers a review in the first place
- The 14 categories it can be configured to cover
- What it can and can't do — and why that limit is deliberate
- How to set it up without over-relying on it
- How it differs from Anti-Scam, which also uses automated detection
Why deterministic filters alone aren't enough
Keyword lists, pattern matching, and rate limits catch a lot — but they're built to catch what someone anticipated in advance. A message that's ambiguous, uses phrasing nobody thought to block, or combines several small red flags that individually wouldn't trip anything, can slip through a purely rule-based system.
What Discord's own AutoMod does about ambiguous content
Discord's native AutoMod evaluates message content directly against rules you define — keyword lists, a preset spam filter, mention-spam limits. It has no concept of reviewing a genuinely ambiguous case with broader context; a message either matches a rule or it doesn't.
How AI Moderation works
Deterministic signals run first — AI Moderation only reviews candidates those signals already flagged as worth a second look: scam-like links, configured phrases, spam, mass mentions, suspicious attachments, or activity in designated high-risk channels.
- Covers 14 categories: violence and threats; scams, fraud, and illegal activity; sexual harm; child safety; extortion; high-impact advice; privacy exposure; piracy; dangerous weapons; hate or harassment; self-harm; sexual content; election misinformation; and malware or cyber abuse.
- Choose Alert staff only for an advisory workflow, or Delete confirmed unsafe messages for confirmed destructive action — with optional escalation requiring stronger confirmation on serious or ambiguous results before deletion.
- Configure a private alert channel, the staff roles allowed to use alert actions, candidate phrases, high-risk channels, ignored channels or roles, and scoring thresholds.
- It never bans, kicks, or times out members — a delete action still requires the configured confirmation path and Discord permission.
Set it up
- Confirm Server Premium is active — AI Moderation is opt-in and Premium-only.
- Open Moderation in the dashboard and enable AI Moderation.
- Choose the safety categories relevant to your community — you don't have to enable all 14 at once.
- Set candidate phrases, high-risk channels, and any ignored channels or roles.
- Choose Alert staff only to start, review real results for a while, and only move to Delete confirmed unsafe messages once you trust the configuration.
Common mistakes
- Expecting it to ban, kick, or time out — it never does; pair it with /mod or Warning Points for destructive action.
- Enabling Delete mode on day one — start with Alert staff only so you can validate accuracy before anything gets removed automatically.
- Treating a flag as certain — automated safety analysis can be wrong or lack context; keep staff review and appeal procedures appropriate to your community.
- Enabling all 14 categories without considering which are actually relevant to your community's real risk profile.
Troubleshooting
I enabled it and nothing seems to be happening
That can be correct — it only reviews candidates that deterministic signals already flagged. It's a second opinion on edge cases, not a general content scanner running on every message.
How is this different from Anti-Scam?
Anti-Scam is a dedicated, always-on pipeline built specifically for scam messages, links, behaviour, and images. AI Moderation is a separate, general-purpose safety classifier covering 14 broader categories, one of which is scams — the two run together without conflicting.
A flagged message didn't actually seem unsafe
Lower confidence in a specific category by adjusting the enabled categories and candidate phrases, and consider switching that category to alert-only rather than delete.
Related guides
How to Stop Spam in Discord
Why Discord spam happens, what Discord's own tools miss, and how to stop it automatically with Airwavy's Anti-Scam detection and AI Moderation.
Read guideWhat Is a Discord AI Bot?
What actually makes a bot 'AI-powered,' and how Airwavy's two separate AI systems — AI Chatbot and AI Moderation — differ.
Read guide