Content Moderation Pipeline#
Problem statement (interviewer prompt)
Design a content-moderation pipeline for user-generated images + videos + text. Auto-classify with ML (NSFW, violence, hate speech, CSAM), route ambiguous cases to human reviewers, support appeals + audit trails, and meet regulatory deadlines (e.g. 24h take-downs).
flowchart LR
UP[Upload / Post]
ML([ML classifiers])
RULES[Policy rules]
Q[[Review queue]]
HUM[Human moderators]
ACT[Action]
UP --> ML --> RULES --> ACT
RULES --> Q --> HUM --> ACT
classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
class UP,RULES,HUM,ACT service;
class Q queue;
class ML compute;
flowchart TB
subgraph Inputs
TXT[Text posts]
IMG[Images]
VID[Videos]
LIVE[Live streams]
AUD[Audio]
end
subgraph ML[ML classifiers]
NSFW([NSFW image])
VIOL[Violence]
HATE[Hate speech]
SPAM[Spam]
CHILD[CSAM hash match]
COPY[Copyright fingerprint]
DEEPF[Deepfake detection]
OCR([OCR text])
ASR[Audio transcription]
end
subgraph Hashes
PD[PhotoDNA hash]
PERC[Perceptual hash dedup]
MEDIA_GRAPH[Known-bad media graph]
end
subgraph Policy
RULE[Policy / rules]
GEO[Region-specific laws]
AGE[Age gating]
end
subgraph Queue
HIGH[[High-priority queue]]
MED[Medium]
LOW[Low]
SLA[SLA tracking]
end
subgraph Human[Human review]
MOD[Moderators]
TOOL[Mod tooling]
TRAIN[Training + welfare]
ESC[Escalation]
end
subgraph Actions
REMOVE
LIMIT[Reduce reach]
LABEL[Warn label]
BAN[Ban / strike]
LEGAL[Legal hold / reporting]
end
subgraph Loop
APPEAL([User appeal])
FEED[Decision feedback into ML]
AUDIT[Audit log + transparency report]
end
Inputs --> ML --> Policy --> Queue --> Human --> Actions
Hashes --> ML
Loop --- Actions
classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
class APPEAL client;
class TXT,IMG,VID,LIVE,AUD,VIOL,HATE,SPAM,CHILD,COPY,DEEPF,ASR,PD,PERC,MEDIA_GRAPH,RULE,GEO,AGE,MED,LOW,SLA,MOD,TOOL,TRAIN,ESC,LIMIT,LABEL,BAN,LEGAL,FEED service;
class HIGH queue;
class NSFW,OCR compute;
class AUDIT obs;
Glossary & fundamentals#
| Concept | What it is | Fundamentals |
|---|---|---|
| Probabilistic structures | perceptual hash, MinHash | probabilistic-data-structures |
| Pub/Sub | media events into pipeline | pub-sub-pattern |
| Idempotency | safe retries | idempotency-retries |
| CDC | decisions feed search/index | change-data-capture |
| Observability | SLA + appeal SLOs | observability |
Quick reference#
Functional#
- Detect & act on disallowed content (text/image/video/audio).
- Human review queue for ambiguous.
- Appeals.
- Compliance reports & legal hold.
Non-functional#
- Latency: live streams need second-scale; posts can be minutes.
- High recall on CSAM, terrorism (zero-tolerance categories).
- Worker welfare considerations critical.
Trade-offs#
- Auto-action vs human-review: high precision auto-act, ambiguous → queue.
- Region-specific policies force per-region pipelines.
- Appeals friction vs abuse of appeals.
Refs#
- Meta Transparency Reports, Trust & Safety Professional Association papers.
- "Behind the Screen" Sarah T. Roberts (book).
- PhotoDNA whitepaper.
FAQ#
How does content moderation work at scale?#
New posts run through hash matchers (PhotoDNA, PDQ) for known bad content, then ML classifiers for NSFW, violence, and hate. Borderline scores go to a human review queue with SLA tracking.
Why use hash-based detection plus ML?#
Hash matching catches known violating media in milliseconds with near-zero false positives. ML covers novel content but is probabilistic, so hashes act as a cheap first filter.
How are human reviewers integrated?#
Ambiguous items land in a priority queue sorted by severity and recency. Reviewers see context, take an action, and the label feeds back into model training and policy tuning.
How do you tune moderation thresholds?#
Each content type and surface (image vs comment) has its own confidence threshold tuned to balance recall and reviewer cost. Thresholds are versioned and shadow-tested before rollout.
How do appeals and audit work?#
Every decision logs the model version, score, and reviewer ID. Users can appeal, which reopens the case for a senior reviewer and updates the audit trail without rewriting history.