Open Problems

Word Embeddings & Lexical Semantics

Robustness Benchmarking of Lexical Semantics Methods Under Language Model Scarcity and Distribution Shift

Barrier to removeOpen
Strong candidate · 5/5 runs3 papers report this0% from 2025+

Generated automatically from the limitations stated in 3 papers (EMNLP), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Current lexical semantics and semantic change methods assume the availability of high-quality, off-the-shelf pretrained language models trained on massive, well-matched text distributions. When applied to historical texts, low-resource languages, or niche domains where such models do not exist or perform poorly, practitioners have no empirical evidence on how severely these methods degrade. This dependence leaves lexical semantics largely untested across ancient corpora, morphologically non-standard texts, and non-mainstream model architectures.

Why it matters

Establishes empirical operating boundaries for lexical semantics methods across non-standard linguistic domains and provides clear guidelines on when static or lightweight methods should be preferred over poorly matched language models.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Multi-domain and historical stress-testing: Evaluate standard LM-based lexical representation and semantic change detection methods across ancient, historical, and low-resource corpora alongside static embedding baselines, measuring semantic shift detection accuracy and nearest-neighbor stability as LM suitability varies.

  2. 2

    Architecture and context-length scaling benchmark: Run a controlled comparison of lexical semantic extraction across diverse LM families, context window lengths, and parameter scales to measure the exact point of performance degradation when moving from high-capacity modern LMs to restricted or smaller models.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

Broad multilingual foundation models may improve historical and low-resource coverage rapidly enough to dissolve the practical gaps identified across target domains.

Evidence

Each paper's own statement of the limitation, verbatim.

Nearest existing work

Related open problems

Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.