Skip to content

Content Moderation Pipeline#

Problem statement (interviewer prompt)

Design a content-moderation pipeline for user-generated images + videos + text. Auto-classify with ML (NSFW, violence, hate speech, CSAM), route ambiguous cases to human reviewers, support appeals + audit trails, and meet regulatory deadlines (e.g. 24h take-downs).

flowchart LR
  UP[Upload / Post]
  ML([ML classifiers])
  RULES[Policy rules]
  Q[[Review queue]]
  HUM[Human moderators]
  ACT[Action]
  UP --> ML --> RULES --> ACT
  RULES --> Q --> HUM --> ACT

    classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
    classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
    classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
    classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
    classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
    classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
    classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
    classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
    classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
    classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
    class UP,RULES,HUM,ACT service;
    class Q queue;
    class ML compute;
flowchart TB
  subgraph Inputs
    TXT[Text posts]
    IMG[Images]
    VID[Videos]
    LIVE[Live streams]
    AUD[Audio]
  end

  subgraph ML[ML classifiers]
    NSFW([NSFW image])
    VIOL[Violence]
    HATE[Hate speech]
    SPAM[Spam]
    CHILD[CSAM hash match]
    COPY[Copyright fingerprint]
    DEEPF[Deepfake detection]
    OCR([OCR text])
    ASR[Audio transcription]
  end

  subgraph Hashes
    PD[PhotoDNA hash]
    PERC[Perceptual hash dedup]
    MEDIA_GRAPH[Known-bad media graph]
  end

  subgraph Policy
    RULE[Policy / rules]
    GEO[Region-specific laws]
    AGE[Age gating]
  end

  subgraph Queue
    HIGH[[High-priority queue]]
    MED[Medium]
    LOW[Low]
    SLA[SLA tracking]
  end

  subgraph Human[Human review]
    MOD[Moderators]
    TOOL[Mod tooling]
    TRAIN[Training + welfare]
    ESC[Escalation]
  end

  subgraph Actions
    REMOVE
    LIMIT[Reduce reach]
    LABEL[Warn label]
    BAN[Ban / strike]
    LEGAL[Legal hold / reporting]
  end

  subgraph Loop
    APPEAL([User appeal])
    FEED[Decision feedback into ML]
    AUDIT[Audit log + transparency report]
  end

  Inputs --> ML --> Policy --> Queue --> Human --> Actions
  Hashes --> ML
  Loop --- Actions

    classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
    classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
    classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
    classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
    classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
    classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
    classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
    classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
    classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
    classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
    class APPEAL client;
    class TXT,IMG,VID,LIVE,AUD,VIOL,HATE,SPAM,CHILD,COPY,DEEPF,ASR,PD,PERC,MEDIA_GRAPH,RULE,GEO,AGE,MED,LOW,SLA,MOD,TOOL,TRAIN,ESC,LIMIT,LABEL,BAN,LEGAL,FEED service;
    class HIGH queue;
    class NSFW,OCR compute;
    class AUDIT obs;

Glossary & fundamentals#

Concept What it is Fundamentals
Probabilistic structures perceptual hash, MinHash probabilistic-data-structures
Pub/Sub media events into pipeline pub-sub-pattern
Idempotency safe retries idempotency-retries
CDC decisions feed search/index change-data-capture
Observability SLA + appeal SLOs observability

Quick reference#

Functional#

  • Detect & act on disallowed content (text/image/video/audio).
  • Human review queue for ambiguous.
  • Appeals.
  • Compliance reports & legal hold.

Non-functional#

  • Latency: live streams need second-scale; posts can be minutes.
  • High recall on CSAM, terrorism (zero-tolerance categories).
  • Worker welfare considerations critical.

Trade-offs#

  • Auto-action vs human-review: high precision auto-act, ambiguous → queue.
  • Region-specific policies force per-region pipelines.
  • Appeals friction vs abuse of appeals.

Refs#

  • Meta Transparency Reports, Trust & Safety Professional Association papers.
  • "Behind the Screen" Sarah T. Roberts (book).
  • PhotoDNA whitepaper.

FAQ#

How does content moderation work at scale?#

New posts run through hash matchers (PhotoDNA, PDQ) for known bad content, then ML classifiers for NSFW, violence, and hate. Borderline scores go to a human review queue with SLA tracking.

Why use hash-based detection plus ML?#

Hash matching catches known violating media in milliseconds with near-zero false positives. ML covers novel content but is probabilistic, so hashes act as a cheap first filter.

How are human reviewers integrated?#

Ambiguous items land in a priority queue sorted by severity and recency. Reviewers see context, take an action, and the label feeds back into model training and policy tuning.

How do you tune moderation thresholds?#

Each content type and surface (image vs comment) has its own confidence threshold tuned to balance recall and reviewer cost. Thresholds are versioned and shadow-tested before rollout.

How do appeals and audit work?#

Every decision logs the model version, score, and reviewer ID. Users can appeal, which reopens the case for a senior reviewer and updates the audit trail without rewriting history.