Open Problems

Computational Social Science & Content Moderation

Benchmarking Systematic Biases and Reproducibility in LLM Pragmatic Discourse Annotation

Barrier to removeOpen
Strong candidate · 4/5 runs3 papers report this100% from 2025+

Generated automatically from the limitations stated in 3 papers (ACL, EMNLP), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Computational social science and online moderation research increasingly substitute human coding with LLM annotation for subjective pragmatic phenomena such as conversational roles, implicit toxicity, and discourse dynamics. However, existing studies rely heavily on proprietary models like GPT-4, which introduce non-determinism, reproducibility barriers, and systematic biases toward surface stylistic features like politeness. Because smaller open-weight models frequently fail or skip messages on complex pragmatic schemas, researchers currently lack a validated, reproducible protocol or understanding of how LLM annotation errors propagate across subjective discourse settings.

Why it matters

Enables computational social science and moderation pipelines to run auditable, reproducible, and open-source pragmatic annotations without inheriting stylistic confounds or closed-API dependencies.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Conduct a cross-model benchmarking study comparing open-weight models (7B-70B) against proprietary models on established pragmatic discourse benchmarks (e.g., conversational role labeling, advice identification, implicit derailment), measuring inter-annotator agreement with human gold standards and parse failure rates.

  2. 2

    Quantify stylistic confounding by constructing controlled perturbation suites (e.g., altering message politeness, tone, or dialect while preserving underlying pragmatic intent) and measuring the shift in predicted pragmatic role labels.

  3. 3

    Evaluate lightweight fine-tuning and calibration strategies (such as rationale-distillation from reasoning models into open-weight architectures) to measure whether small open models can achieve stable pragmatic annotation without skipping or defaulting to surface style heuristics.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

If rapid prompt-engineering recipes or standard commercial API model updates immediately resolve pragmatic inconsistency and style bias without requiring architectural or methodological changes.

Evidence

Each paper's own statement of the limitation, verbatim.

Nearest existing work

Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.