Open Problems

Low-Resource Languages & Dialects

A Benchmark and Labeling Framework for Community-Contingent Hate Speech Annotations, Replacing Single Gold Labels with Distributions of Plausible Ones

Effect to explainPartly addressed
Possible candidate · 3/5 runs8 papers report this25% from 2025+

Generated automatically from the limitations stated in 8 papers (EMNLP, TheWebConference, NAACL), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Every current hate speech dataset collapses a genuinely plural judgment — what counts as hateful depends on annotator identity, community norms, reclamation practices, and exposure to implicit forms — into one label, then trains and evaluates models against it. This means models cannot represent the disagreement they were built from, audits of pretraining data and model outputs in low-resource settings have no trustworthy ground truth to measure against, and moderation systems calibrated on one community's thresholds misfire (or stay silent) on another's counter-speech, quotes, and reclaimed slurs. Structurally, nobody today can state, for a given utterance, how hateful-label assignments vary across the populations that matter — so "biased model" and "ambiguous data" are permanently confounded.

Why it matters

Moderation and auditing systems that can be tuned to a named community's threshold with known error bars, and a clean separation of model bias from label ambiguity in any future hate speech evaluation.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Re-annotate an existing multilingual/multimodal hate corpus with structured disagreement: multiple annotator pools stratified by community identity and region, each instance receiving a distribution over labels plus free-text rationale, with inter-pool divergence measured per item. Deliverable: a public dataset where the unit of truth is the distribution, and metrics (pool-divergence scores, agreement-by-subgroup) computed against it.

  2. 2

    Train and evaluate label-distribution models — models outputting P(hateful | annotator profile) instead of a point label — and compare their calibration against single-label baselines on the misclassification cases already known to be systematic (disguised memes, communal words, reclaimed slurs). Measure calibration error and per-subgroup error gaps.

  3. 3

    Audit downstream use: rerun a toxicity audit of one pretraining corpus under distributional labels versus gold labels, quantifying how much of the documented bias signal (e.g., occupation associations) survives when label subjectivity is explicitly modeled. Measure the change in flagged-content rates and which flagged items flip.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

Stratified annotator pools may not actually capture "community" norms — self-identified membership may predict labeling behavior far more weakly than the framework assumes, collapsing the distributions into noise and leaving the single-label status quo justified.

Sub-problems

  • Learning Nuanced and Implicit Hate Speech Classifiers Under Weak Auxiliary Supervision

    Current high-performing moderation models for implicit and target-specific hate speech depend on auxiliary metadata—such as explicitness tags, target-group annotations, and fine-grained annotator demographic profiles—that are absent in standard production corpora. When forced to train solely on aggregate binary labels, models conflate communal identity terms with toxic intent and fail on contextual nuances like reclaimed slurs or quoted counter-speech. Furthermore, hard-coded aggregation thresholds treat subjective socio-cultural disagreement as random noise, rendering classifiers non-transferable across diverse online communities.

Evidence

Each paper's own statement of the limitation, verbatim.

Nearest existing work

Related open problems

Low-Resource Languages & Dialects

Effect to explainOpen

Reliability-Calibrated Annotation Protocols for Subjective and Low-Resource labeling Tasks

Across many NLP annotation efforts, gold standards rest on so few annotators—and on tasks so subjective—that reported inter-annotator agreement is often moderate at best (κ 0.18–0.46), and in some cases reliability fails outright (agreement too weak to include in evaluation at all). Because papers rarely distinguish "the task is ambiguous" from "the pool was too small," downstream users cannot judge how trustworthy the gold labels are, and resources with single-annotator or single-community provenance silently embed one worldview. Structurally, there is no standard way to report, audit, or correct for annotation reliability when it is low or unmeasured.

Possible candidate · 2/5 runs19 papers report this60% from 2025+

Low-Resource Languages & Dialects

Barrier to removeOpen

Metric Inflation and Reliability Limits in IPA-Space Generation Evaluation

Multilingual systems increasingly evaluate generations in International Phonetic Alphabet (IPA) space to bridge script disparities, yet standard character-level metrics like chrF inflate scores because IPA relies heavily on shared Latin base characters. Furthermore, automatic grapheme-to-phoneme tools (such as Epitran) introduce systematic transcription errors and inventory mismatches that are masked when evaluating against synthetic or machine-translated benchmarks. Researchers currently cannot distinguish whether high IPA-level chrF gains reflect genuine semantic translation quality or superficial phonetic transcription artifacts.

Possible candidate · 2/5 runs3 papers report this100% from 2025+
Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.