Recommender Systems
Cross-Domain and Interaction-Density Robustness of Sequential Recommender Systems
Generated automatically from the limitations stated in 9 papers (SIGIR, RecSys, ICML), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Current sequential and language-model-based recommender algorithms are overwhelmingly developed and evaluated on 5-core filtered subsets of Amazon product reviews (e.g., Books, Electronics, Beauty). Consequently, it is unknown whether reported performance advantages hold across non-e-commerce domains (such as news browsing, streaming video, or social feeds), under varying interaction densities, or at larger catalog scales. Researchers cannot determine whether recent architectural improvements reflect general algorithmic progress or overfitting to the statistical idiosyncrasies of filtered Amazon review datasets.
Why it matters
Provides the first verified mapping of which sequential recommendation architectures generalize across distinct domain properties, catalog scales, and sparsity levels.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Multi-domain benchmark evaluation: Standardize evaluation across public datasets from diverse domains (e.g., MIND for news, KuaiRand for short video, Steam for gaming, and Reddit for content interaction) alongside Amazon subsets, evaluating representative sequential and PLM-based models under identical split and metric protocols (NDCG@K, Recall@K).
- 2
Sparsity and core-filtering stress tests: Systematically vary interaction density from raw, long-tail logs (1-core/2-core) to dense 5-core and 10-core regimes to measure performance degradation curves and determine whether relative model rankings flip under sparse conditions.
- 3
Modality and catalog scale sweeps: Evaluate the sensitivity of text-enhanced and ID-based sequential models across differing catalog sizes and metadata configurations (pure IDs, categorical attributes, and rich textual metadata) to measure computational cost relative to performance gains across domains.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
If comprehensive cross-domain benchmark surveys are already under review by major recommender systems venues, or if scaling the compute-heavy PLM/LLM baselines across massive external datasets exceeds available student computing resources.
Evidence
Each paper's own statement of the limitation, verbatim.
- Learning Decomposed Contextual Token Representations from Pretrained and Collaborative Signals for Generative RecommendationSIGIR 2026
Evaluated only on Amazon Review subsets with 5-core filtering and max sequence length 20; generalization to other domains or sparse-interaction settings is untested
- Adaptive Mix Preference Optimization for Generative RecommendationSIGIR 2026
Evaluated primarily on a single industrial e-commerce platform plus Amazon reviews; generalization to other domains (e.g., content recommendation with likes/dislikes) is asserted but not demonstrated
- MELON: Learning Multi-Aspect Modality Preferences for Accurate Multimedia RecommendationSIGIR 2025
Evaluated only on four Amazon datasets (Women/Men Clothing, Toys & Games, Sports) with users/items having ≥5 interactions; generalization to other domains, denser data, or other modality types is untested
- Improving LLMs for Recommendation with Out-Of-Vocabulary TokensICML 2025
Evaluated only on Amazon review datasets (Sports, Beauty, Toys) with a limited number of clusters (G=100, 200 for Toys); scalability to billions of items, as in industrial systems, is not demonstrated
- Diversity-aware Dual-promotion Poisoning Attack on Sequential RecommendationSIGIR 2025
Evaluated only on Amazon review datasets (Beauty, Sports, Toys) with relatively sparse interactions; results may not transfer to denser industrial settings
- DIFF: Dual Side-Information Filtering and Fusion for Sequential RecommendationSIGIR 2025
Evaluated only on four 5-core e-commerce/review datasets with item-related attributes (category, brand, position); no evidence for other domains or attribute types
- The Elephant in the Room: Rethinking the Usage of Pre-trained Language Model in Sequential RecommendationRecSys 2024
Experiments limited to Amazon review datasets with a single PLM (Longformer) and single PLM-based SR model (RECFORMER), so generalization to other PLMs, LLMs, or industrial-scale settings is unverified
- Data Augmentation using Reverse Prompt for Cost-Efficient Cold-Start RecommendationRecSys 2024
Evaluated only on tiny Amazon subsets (570-3000 users, ~1000-6000 training interactions, 100 warm-start and 100 cold items); no large-scale validation
Show all 9 papers
- A Model-Agnostic Framework for Recommendation via Interest-aware Item EmbeddingsRecSys 2023
Evaluated only on two Amazon review datasets (Electronics, Books) with CTR-style tasks; no industrial-scale or online evaluation despite claims of industrial relevance
Nearest existing work
- Scaling Sequential Recommendation Models with TransformersSIGIR 2024
- SAER: Scalable Assessment of E-commerce Recommendations using Large Language ModelsSIGIR 2026
- LLM-EDT: Large Language Models Enhanced Cross-domain Sequential Recommendation with Dual-phase TrainingSIGIR 2026
- Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic EncodersACL 2026
- Leveraging Large Language Models for Sequential RecommendationRecSys 2023
- Bridge the Domains: Large Language Models Enhanced Cross-domain Sequential RecommendationSIGIR 2025
- Amazon-KG: A Knowledge Graph Enhanced Cross-Domain Recommendation DatasetSIGIR 2024
- Personalized Review Recommendation based on Implicit dimension miningNAACL 2024
- Wukong: Towards a Scaling Law for Large-Scale RecommendationICML 2024
- Adaptive Domain Scaling for Personalized Sequential Modeling in RecommendersSIGIR 2025
- X-Cross: Dynamic Integration of Language Models for Cross-Domain Sequential RecommendationSIGIR 2025
- Yambda-5B — A Large-Scale Multi-Modal Dataset for Ranking and RetrievalRecSys 2025
- LLM-RecG: A Semantic Bias-Aware Framework for Zero-Shot Sequential RecommendationRecSys 2025
- TLRec: A Transfer Learning Framework to Enhance Large Language Models for Sequential Recommendation TasksRecSys 2024
- When Large Vision Language Models Meet Multimodal Sequential Recommendation: An Empirical StudyTheWebConference 2025