Open Problems

Recommender Systems

Cross-Domain and Interaction-Density Robustness of Sequential Recommender Systems

Scope to testOpen
Possible candidate · 2/5 runs9 papers report this67% from 2025+

Generated automatically from the limitations stated in 9 papers (SIGIR, RecSys, ICML), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Current sequential and language-model-based recommender algorithms are overwhelmingly developed and evaluated on 5-core filtered subsets of Amazon product reviews (e.g., Books, Electronics, Beauty). Consequently, it is unknown whether reported performance advantages hold across non-e-commerce domains (such as news browsing, streaming video, or social feeds), under varying interaction densities, or at larger catalog scales. Researchers cannot determine whether recent architectural improvements reflect general algorithmic progress or overfitting to the statistical idiosyncrasies of filtered Amazon review datasets.

Why it matters

Provides the first verified mapping of which sequential recommendation architectures generalize across distinct domain properties, catalog scales, and sparsity levels.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Multi-domain benchmark evaluation: Standardize evaluation across public datasets from diverse domains (e.g., MIND for news, KuaiRand for short video, Steam for gaming, and Reddit for content interaction) alongside Amazon subsets, evaluating representative sequential and PLM-based models under identical split and metric protocols (NDCG@K, Recall@K).

  2. 2

    Sparsity and core-filtering stress tests: Systematically vary interaction density from raw, long-tail logs (1-core/2-core) to dense 5-core and 10-core regimes to measure performance degradation curves and determine whether relative model rankings flip under sparse conditions.

  3. 3

    Modality and catalog scale sweeps: Evaluate the sensitivity of text-enhanced and ID-based sequential models across differing catalog sizes and metadata configurations (pure IDs, categorical attributes, and rich textual metadata) to measure computational cost relative to performance gains across domains.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

If comprehensive cross-domain benchmark surveys are already under review by major recommender systems venues, or if scaling the compute-heavy PLM/LLM baselines across massive external datasets exceeds available student computing resources.

Evidence

Each paper's own statement of the limitation, verbatim.

Show all 9 papers

Nearest existing work

Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.