Low-Resource Languages & Dialects
Reliability-Calibrated Annotation Protocols for Subjective and Low-Resource labeling Tasks
Generated automatically from the limitations stated in 19 papers (ACL, EMNLP, NAACL), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Across many NLP annotation efforts, gold standards rest on so few annotators—and on tasks so subjective—that reported inter-annotator agreement is often moderate at best (κ 0.18–0.46), and in some cases reliability fails outright (agreement too weak to include in evaluation at all). Because papers rarely distinguish "the task is ambiguous" from "the pool was too small," downstream users cannot judge how trustworthy the gold labels are, and resources with single-annotator or single-community provenance silently embed one worldview. Structurally, there is no standard way to report, audit, or correct for annotation reliability when it is low or unmeasured.
Why it matters
Trustworthy evaluation in low-resource and subjective domains becomes possible: datasets can be compared on evidence quality, not just size, and papers can claim reliability instead of assuming it.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Reliability auditing of existing datasets: take 15–20 public datasets with published agreement statistics and re-estimate per-item label confidence using bootstrap over annotator subsamples, measuring how downstream model rankings change when only high-confidence items are used. Deliverable: a "reliability report card" format for datasets.
- 2
Minimum-annotator/minimum-agreement guidelines: run controlled annotation studies on the same subjective task with pools of 1, 3, 5, and 10 annotators, plotting measured agreement stability and downstream metric variance as a function of pool size, to produce evidence-based thresholds for when a gold standard is fit for purpose.
Partly donePaper [9] establishes empirical and bootstrap methods for measuring reliability as a function of rater pool size ($k$-rater reliability) on WordSim-353, but does not evaluate downstream metric variance or provide general fit-for-purpose thresholds across subjective tasks.
Other papers read (9)
- [1]Efficient Online Scalar Annotation with Bounded SupportACL 2018
- [2]The “Problem” of Human Label Variation: On Ground Truth in Data, Modeling and EvaluationEMNLP 2022
- [3]Noise Correction on Subjective DatasetsACL 2024
- [4]Noise or additional information? Leveraging crowdsource annotation item agreement for natural language tasks.EMNLP 2015
- [5]Cross-replication Reliability - An Empirical Approach to Interpreting Inter-rater ReliabilityACL 2021
- [6]The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human LabelsNAACL 2024
- [7]Beyond Averages: Learning with Annotator Disagreement in STSEMNLP 2025
- [8]Nonsense!: Quality Control via Two-Step Reason Selection for Annotating Local Acceptability and Related Attributes in News EditorialsEMNLP 2019
- [10]Best-Worst Scaling More Reliable than Rating Scales: A Case Study on Sentiment Intensity AnnotationACL 2017
Checked 9/26/2026.
- 3
Disagreement-preserving labels: pilot releasing distributions or per-annotator labels (rather than adjudicated gold) for one contested task, measuring whether models trained on soft labels match human judgment distributions better than those trained on collapsed gold.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
If re-analysis shows model rankings are robust to low agreement—i.e., noisy gold labels don't actually change conclusions—then the audit protocol solves a problem with little practical consequence.
Sub-problems
Benchmarking Subjective Content Moderation Under Persistent Multi-Annotator Disagreement
Current computational social science and content moderation benchmarks collapse subjective judgments into a single adjudicated "gold" label, despite observed inter-annotator agreement metrics falling into poor or moderate ranges (e.g., Fleiss' $\kappa < 0.45$, Gwet's AC1 $< 0.30$). These labels are routinely produced by small, demographically narrow pools (often 2–4 in-person annotators, the paper authors, or single-region crowdsource workers), baking localized cultural and political worldviews directly into ground truth. As a result, standard evaluation pipelines cannot distinguish between a model that fails at content understanding and one that simply aligns with a valid alternative perspective absent from the original annotator pool.
Evidence
Each paper's own statement of the limitation, verbatim.
- CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web DataACL 2026
Only 12.9% of lines have multiple annotators, so a full inter-annotator agreement study was impossible; multiple valid labels per line (macro/micro codes, mixed-language lines) remain inherent ambiguities.
- CAPC-CG: A Large-Scale, Expert-Directed LLM-Annotated Corpus of Adaptive Policy Communication in ChinaACL 2026
Gold standard annotated by only three domain experts (~6,000 paragraphs), which may capture a narrower range of interpretations than a larger annotator pool
- SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?EMNLP 2025
Three annotated language pairs (eng-ibo, eng-zul, fra-lin) had inter-annotator agreement too weak for dev/test inclusion (e.g., eng-zul Spearman 0.107), and Igbo, Swahili, and Sesotho were dropped entirely (Spearman < 0.20), so annotation reliability itself fails for some languages
- Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in KazakhACL 2025
Full manual verification required 12 expert native-speaker annotators, each working the equivalent of five full working days, which is costly and limits scalability to other languages
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationACL 2025
Only 42 languages; community annotator participation was uneven across languages, skewing coverage and annotator diversity
- Discourse-Driven Code-Switching: Analyzing the Role of Content and Communicative Function in Spanish-English Bilingual SpeechEMNLP 2025
Single annotator per task with no full inter-annotator agreement (spot-check κ=0.81 for NEs, κ=0.69 for DAs on small samples of 457 and 2345 instances), creating reliability concerns
- Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive SummarizationACL 2025
Annotation counts and rater pools vary widely across languages — Japanese had only one worker and 188 annotations, so per-language correlations rest on very different evidence bases
- Are Large Language Models Chronically Online Surfers? A Dataset for Chinese Internet Meme ExplanationEMNLP 2025
Annotation subjectivity from few annotators; Fleiss' kappa of 0.442 indicates only moderate inter-annotator agreement
Show all 19 papers
- MAFALDA: A Benchmark and Comprehensive Study of Fallacy Detection and ClassificationNAACL 2024
Annotation was done by the authors themselves, so collective cultural/political bias may be baked into the gold standard despite their diversity and template-based justification process
- Leveraging Conflicts in Social Media Posts: Unintended Offense DatasetEMNLP 2024
Annotation subjectivity: AMT non-expert raters from only USA/UK/Australia with averaging-based agreement, potentially introducing cultural bias; only ~80% of ratings had confidence ≥50
- Harnessing Toulmin’s theory for zero-shot argument explicationACL 2024
Evaluation datasets are small (ARCT: 445 pairs; MCT: 112 arguments) and English-only; annotator agreement on warrant acceptability was low (kappa 0.18-0.31), reflecting task subjectivity
- Extracting Lexical Features from Dialects via Interpretable Dialect ClassifiersNAACL 2024
Evaluation requires annotators proficient in both dialects, so human evaluation covered only CN-TW (3 annotators) and SCN-IT (1 annotator), limiting coverage
- Error Analysis of Multilingual Language Models in Machine Translation: A Case Study of English-Amharic TranslationEMNLP 2024
Inter-annotator agreement was only fair (Fleiss' Kappa mostly 0.21-0.40), indicating substantial annotator disagreement
- Beyond Recognising Entailment: Formalising Natural Language Inference from an Argumentative PerspectiveACL 2024
Annotator agreement on full argument type is low (Gwet's AC1 24.7, 'fair'), and argument form only 58.3, indicating the annotation scheme itself is ambiguous; one annotator had to be dropped due to agreement drift
- Aya Dataset: An Open-Access Collection for Multilingual Instruction TuningACL 2024
Contribution distribution is heavily skewed: a long tail of annotators contributed one or two examples, and some languages depend on one or two prolific contributors, risking single-viewpoint cultural bias (e.g., French data reflecting France rather than Quebec, Togo, Senegal)
- Your spouse needs professional help: Determining the Contextual Appropriateness of Messages through Modeling Social RelationshipsACL 2023
Annotations come from few in-person annotators (4 in Phase 1, 2 in Phase 2), so judgments reflect those annotators' cultural/personal values; initial raw agreement was moderate (α=0.46 appropriateness before adjudication)
- Towards standardizing Korean Grammatical Error Correction: Datasets and AnnotationACL 2023
Inter-annotator agreement for the evaluation itself is only moderate (Fleiss' kappa 0.4386), and INS/DEL edits are left unrefined (subject omission treated as one type)
- NormBank: A Knowledge Bank of Situational Social NormsACL 2023
Only about half of gold annotations are judged fully correct by third-party evaluators, with disagreements traced to annotators' cultural and personal worldviews
- A Multi-Task Dataset for Assessing Discourse Coherence in Chinese Essays: Structure, Theme, and Logic AnalysisEMNLP 2023
Manual expert annotation is labor-intensive and subject to annotator subjectivity and potential bias (IAA as low as 52.23% for inappropriate clause identification)
Nearest existing work
- The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human LabelsNAACL 2024
- Further Investigation into Reference Bias in Monolingual Evaluation of Machine TranslationEMNLP 2017
- Towards Semi-Automatic Generation of Proposition Banks for Low-Resource LanguagesEMNLP 2016
- Building Better: Avoiding Pitfalls in Developing Language Resources when Data is ScarceACL 2025
- The “Problem” of Human Label Variation: On Ground Truth in Data, Modeling and EvaluationEMNLP 2022
- k-Rater Reliability: The Correct Unit of Reliability for Aggregated Human AnnotationsACL 2022
- Nonsense!: Quality Control via Two-Step Reason Selection for Annotating Local Acceptability and Related Attributes in News EditorialsEMNLP 2019
- A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for SummarizationACL 2023
- STABLEVAL: Disagreement-Aware and Stable Evaluation of AI SystemsICML 2026
- Noise or additional information? Leveraging crowdsource annotation item agreement for natural language tasks.EMNLP 2015
- Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators’ DisagreementEMNLP 2021
- Modelling Instance-Level Annotator Reliability for Natural Language Labelling TasksNAACL 2019
- NUTMEG: Separating Signal From Noise in Annotator DisagreementEMNLP 2025
- Anchoring and Agreement in Syntactic AnnotationsEMNLP 2016
- VariErr NLI: Separating Annotation Error from Human Label VariationACL 2024
Related open problems
Low-Resource Languages & Dialects
A Benchmark and Labeling Framework for Community-Contingent Hate Speech Annotations, Replacing Single Gold Labels with Distributions of Plausible Ones
Every current hate speech dataset collapses a genuinely plural judgment — what counts as hateful depends on annotator identity, community norms, reclamation practices, and exposure to implicit forms — into one label, then trains and evaluates models against it. This means models cannot represent the disagreement they were built from, audits of pretraining data and model outputs in low-resource settings have no trustworthy ground truth to measure against, and moderation systems calibrated on one community's thresholds misfire (or stay silent) on another's counter-speech, quotes, and reclaimed slurs. Structurally, nobody today can state, for a given utterance, how hateful-label assignments vary across the populations that matter — so "biased model" and "ambiguous data" are permanently confounded.
Low-Resource Languages & Dialects
Metric Inflation and Reliability Limits in IPA-Space Generation Evaluation
Multilingual systems increasingly evaluate generations in International Phonetic Alphabet (IPA) space to bridge script disparities, yet standard character-level metrics like chrF inflate scores because IPA relies heavily on shared Latin base characters. Furthermore, automatic grapheme-to-phoneme tools (such as Epitran) introduce systematic transcription errors and inventory mismatches that are masked when evaluating against synthetic or machine-translated benchmarks. Researchers currently cannot distinguish whether high IPA-level chrF gains reflect genuine semantic translation quality or superficial phonetic transcription artifacts.