Computational Social Science & Content Moderation
Benchmarking Systematic Biases and Reproducibility in LLM Pragmatic Discourse Annotation
Generated automatically from the limitations stated in 3 papers (ACL, EMNLP), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Computational social science and online moderation research increasingly substitute human coding with LLM annotation for subjective pragmatic phenomena such as conversational roles, implicit toxicity, and discourse dynamics. However, existing studies rely heavily on proprietary models like GPT-4, which introduce non-determinism, reproducibility barriers, and systematic biases toward surface stylistic features like politeness. Because smaller open-weight models frequently fail or skip messages on complex pragmatic schemas, researchers currently lack a validated, reproducible protocol or understanding of how LLM annotation errors propagate across subjective discourse settings.
Why it matters
Enables computational social science and moderation pipelines to run auditable, reproducible, and open-source pragmatic annotations without inheriting stylistic confounds or closed-API dependencies.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Conduct a cross-model benchmarking study comparing open-weight models (7B-70B) against proprietary models on established pragmatic discourse benchmarks (e.g., conversational role labeling, advice identification, implicit derailment), measuring inter-annotator agreement with human gold standards and parse failure rates.
- 2
Quantify stylistic confounding by constructing controlled perturbation suites (e.g., altering message politeness, tone, or dialect while preserving underlying pragmatic intent) and measuring the shift in predicted pragmatic role labels.
- 3
Evaluate lightweight fine-tuning and calibration strategies (such as rationale-distillation from reasoning models into open-weight architectures) to measure whether small open models can achieve stable pragmatic annotation without skipping or defaulting to surface style heuristics.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
If rapid prompt-engineering recipes or standard commercial API model updates immediately resolve pragmatic inconsistency and style bias without requiring architectural or methodological changes.
Evidence
Each paper's own statement of the limitation, verbatim.
- Who Plays Which Role When? Communication Role Dynamics for Peer Recognition and Team Performance PredictionACL 2026
Best LLM annotation depends on prompting choices and proprietary reasoning models; model biases toward language style may propagate (e.g., polite messages more likely labeled with constructive roles)
- A Theoretically Grounded Approach to Summarizing Conversation Dynamics for Forecasting the Derailment of Online ConversationsACL 2026
Automatic pragmatic annotation is error-prone and depends on proprietary GPT-4, raising annotation bias, reproducibility, and model dependence concerns; smaller open models (Qwen2.5-1.5B) failed, consistently skipping messages
- Evaluation and Facilitation of Online Discussions in the LLM Era: A SurveyEMNLP 2025
LLM annotators remain unreliable on subjective dimensions (pragmatics, humor, irony) and exhibit biases (sociodemographic, political) that can be exacerbated over a discussion; closed-source models add non-determinism
Nearest existing work
- A Theoretically Grounded Approach to Summarizing Conversation Dynamics for Forecasting the Derailment of Online ConversationsACL 2026
- Pragmatic Inference Chain (PIC) Improving LLMs’ Reasoning of Authentic Implicit Toxic LanguageEMNLP 2025
- Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and ChallengesACL 2025
- Rethinking Pragmatics in Large Language Models: Towards Open-Ended Evaluation and Preference TuningEMNLP 2024
- Should a Chatbot be Sarcastic? Understanding User Preferences Towards Sarcasm GenerationACL 2022
- PerspectiveMod: A Perspectivist Resource for Deliberative ModerationEMNLP 2025
- Comparing human and LLM politeness strategies in free productionEMNLP 2025
- Conversations Gone Awry: Detecting Early Signs of Conversational FailureACL 2018
- It Is Not Only the Negative that Deserves Attention! Understanding, Generation & Evaluation of (Positive) ModerationNAACL 2025
- Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals’ Subjective Text PerceptionsACL 2025
- Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful BeliefsACL 2026
- The Importance of Modeling Social Factors of Language: Theory and PracticeNAACL 2021
- A Multidimensional Lexicon for Interpersonal StancetakingACL 2017
- Can Language Model Moderators Improve the Health of Online Discourse?NAACL 2024
- TalkDown: A Corpus for Condescension Detection in ContextEMNLP 2019