ACM TORS Volume 4, Issue 3: From SIGIR’24 Highlights to Reproducible Recommendation Research

The September 2026 issue of ACM Transactions on Recommender Systems (vol. 4, no. 3) follows two movements in recommender-systems research at once: models are incorporating more signals, languages, behaviors, and objectives, while evaluation is being asked to carry more of the scientific load.

The new issue contains 14 contributions, across article numbers [34e] and [34-46]: the editorial by Dietmar Jannach and Li Chen [34e], six papers in the “Highlights of SIGIR’24” section [34-39], and seven regular research articles [40-46]. Together, they comprise 459 article pages and 47 byline appearances by 45 distinct authors. The average contribution is 32.8 pages long, the median is 31 pages, and the 13 research articles average 34.2 pages. The six SIGIR’24 highlights contribute 192 pages, or 41.8% of the issue; the seven regular papers contribute 252 pages; and the editorial contributes 15 pages. The 69-page study by Benigni et al. [46] occupies 15.0% of the issue and is exactly three times as long as its shortest research article, the 23-page paper by De Kerpel and Benoit [40]. Reproducibility, in this issue, is given room to show its work.

Contents of Volume 4, Issue 3

[34e] Dietmar Jannach and Li Chen. 2026. “Improving Methodological Standards in Recommender Systems Offline Evaluation.” ACM Transactions on Recommender Systems 4, 3, Article 34e, 15 pages, 34e:1–34e:15. DOI: 10.1145/3800587.

Highlights of SIGIR’24

[34] Yi Yu, Kazunari Sugiyama, and Adam Jatowt. 2026. “Beyond Recommendations: Sequential Recommendation with Collaborative Explanation.” ACM Transactions on Recommender Systems 4, 3, Article 34, 31 pages, 34:1–34:31. DOI: 10.1145/3731458.

[35] Yang Li, Junpeng Du, Chenzhan Wang, Zunlong Liu, Xiaomin Zhu, and Chen Lin. 2026. “CROSS: Feedback-Oriented Multi-Modal Dynamic Alignment in Recommendation Systems.” ACM Transactions on Recommender Systems 4, 3, Article 35, 24 pages, 35:1–35:24. DOI: 10.1145/3734527.

[36] Qi Xiao, Jing Xiao, Weike Pan, and Zhong Ming. 2026. “MBASR: A Generic Framework for Multi-Behavior Data Augmentation in Sequential Recommendation.” ACM Transactions on Recommender Systems 4, 3, Article 36, 28 pages, 36:1–36:28. DOI: 10.1145/3749998.

[37] Theresia Veronika Rampisela, Maria Maistro, Tuukka Ruotsalo, Falk Scholer, and Christina Lioma. 2026. “Relevance-aware Individual Item Fairness Measures for Recommender Systems: Limitations and Usage Guidelines.” ACM Transactions on Recommender Systems 4, 3, Article 37, 45 pages, 37:1–37:45. DOI: 10.1145/3765624.

[38] Patrik Dokoupil and Ladislav Peska. 2026. “SM-RS 2.0: User-perceived Qualities of Single- and Multi-Objective Recommender Systems.” ACM Transactions on Recommender Systems 4, 3, Article 38, 28 pages, 38:1–38:28. DOI: 10.1145/3754459.

[39] Andreea Iana, Goran Glavaš, and Heiko Paulheim. 2026. “Multilinguality in MIND: Advancing Cross-lingual News Recommendation with a Multilingual Dataset.” ACM Transactions on Recommender Systems 4, 3, Article 39, 36 pages, 39:1–39:36. DOI: 10.1145/3777415.

Regular papers

[40] Lukas De Kerpel and Dries F. Benoit. 2026. “A Reward-Informed Semi-Personalized Bandit Approach for Enhancing Accuracy and Serendipity in Online Slate Recommendations.” ACM Transactions on Recommender Systems 4, 3, Article 40, 23 pages, 40:1–40:23. DOI: 10.1145/3771931.

[41] Kirandeep Kaur, Vinayak Gupta, Manya Chadha, and Chirag Shah. 2026. “Efficient and Responsible Adaptation of Large Language Models for Robust Top-k Recommendations.” ACM Transactions on Recommender Systems 4, 3, Article 41, 31 pages, 41:1–41:31. DOI: 10.1145/3774778.

[42] Eranjana Kathriarachchi, Shafiq Alam, and Salman Rashid. 2026. “Overcoming Hazards of E-commerce Recommender Systems for Social Good.” ACM Transactions on Recommender Systems 4, 3, Article 42, 34 pages, 42:1–42:34. DOI: 10.1145/3785355.

[43] Diego Corrêa da Silva and Dietmar Jannach. 2026. “Calibrated Recommendations: Survey and Future Directions.” ACM Transactions on Recommender Systems 4, 3, Article 43, 32 pages, 43:1–43:32. DOI: 10.1145/3789266.

[44] Célina Treuillier, Sylvain Castagnos, Evan Dufraisse, Özlem Özgöbek, and Armelle Brun. 2026. “Shift your Focus for the Greater Good: Improving Fairness at no cost for Accuracy and Diversity in News Recommender Systems.” ACM Transactions on Recommender Systems 4, 3, Article 44, 36 pages, 44:1–44:36. DOI: 10.1145/3790097.

[45] Jaime Hieu Do, Trung-Hoang Le, and Hady Wirawan Lauw. 2026. “Compositions of Variant Experts for Integrating Short-Term and Long-Term Preferences.” ACM Transactions on Recommender Systems 4, 3, Article 45, 27 pages, 45:1–45:27. DOI: 10.1145/3795520.

[46] Michael Benigni, Maurizio Ferrari Dacrema, and Dietmar Jannach. 2026. “Diffusion Recommender Models and the Illusion of Progress: A Concerning Study of Reproducibility and a Conceptual Mismatch.” ACM Transactions on Recommender Systems 4, 3, Article 46, 69 pages, 46:1–46:69. DOI: 10.1145/3795792.

Evaluation as part of the contribution

The issue opens with the editorial by Dietmar Jannach and Li Chen [34e], which calls for stronger methodological standards in offline recommender-systems evaluation. The editorial places baseline selection and tuning, reproducibility, transparent reporting, and the full experimental pipeline inside the scientific contribution. A claimed improvement depends on the evidence around it: preprocessing choices, comparison methods, optimization effort, repeated runs, statistical analysis, and the artifacts required to inspect or reproduce the result.

Benigni et al. [46] provide the issue’s largest empirical demonstration of why these standards matter. Their 69-page investigation studies four diffusion-recommendation papers from SIGIR 2023 and SIGIR 2024, covering nine diffusion-based algorithms and 18 baseline models. Only 25% of the reported experimental results were fully reproducible. Repeated executions produced variations reaching 18% in some settings, including differences of up to 14% in NDCG@10 and 18% in Recall@10 for one examined model. Once conventional methods were tuned consistently, every examined paper had at least one baseline that matched or surpassed the diffusion approach. The finding connects methodological practice to the interpretation of progress: architectural complexity becomes persuasive when it survives controlled comparison.

Corrêa da Silva and Jannach [43] examine a related gap between an idea’s technical literature and its operational evidence. Their review retrieves 57 papers from the ACM Digital Library and 356 from Google Scholar, then retains 51 works and adds two earlier related studies, yielding 53 papers for analysis. More than 80% of these studies appeared in conference or workshop proceedings, and 58% primarily proposed technical methods. The authors found one study that had applied calibration in a production recommender. Their survey organizes calibration as a way of aligning recommendation distributions with user interests and other target distributions, while identifying field studies, domain coverage, and impact assessment as routes for the next stage of research.

Rampisela et al. [37] turn the same scrutiny toward measurement. They analyze relevance-aware measures of individual item fairness, test them on real and synthetic data, identify at least one limitation for every measure under consideration, and provide corrections or reformulations where appropriate. Their 45-page contribution—the second-longest paper in the issue—shows that the choice of fairness metric can change the apparent behavior of a recommender, particularly when relevance and exposure are brought into the same calculation.

Dokoupil and Peska [38] add the user’s assessment to this methodological picture. SM-RS 2.0 records propensities for four objectives—relevance, diversity, novelty, and exploration—and supports six benchmark tasks, from propensity estimation and selections-aware reranking to perceived-quality and satisfaction prediction. Reproduced baselines reach Kendall’s τ = 0.196 for perceived diversity, while τ is 0.039 for serendipity and 0.045 for satisfaction. These figures quantify how difficult it is to infer subjective qualities from observable recommendation interactions: an offline metric and a user’s experience are related variables, rather than interchangeable ones.

Learning from more than one signal

Yu et al. [34] formulate sequential recommendation and explanation as a joint prediction problem. The model produces two linked outputs—the next item and its collaborative explanation—and uses a purpose-built learning objective to align recommendation and explanation learning. Explanation therefore participates in learning and evaluation rather than being attached after ranking. The framework also establishes a quantitative task on which recommendation and explanation quality can be assessed together.

Li et al. [35] study the alignment of collaborative feedback with multimodal item information. CROSS combines dynamic item-level alignment with multi-grained collaborative alignment and is evaluated on four datasets with five conventional collaborative-filtering models and six multimodal recommenders: 11 backbones in total. Across the five collaborative-filtering backbones, the reported average improvements range from 21.52% to 70.78%. Across the six multimodal backbones, they range from 8.70% to 20.73%, while comparisons with FETTLE yield gains between 3.82% and 5.24%. The magnitude of these results suggests that alignment is a central modeling operation when visual, textual, and interaction signals enter through different representation spaces.

Xiao et al. [36] approach sparse sequential data from the data side. MBASR supplies five behavior-aware augmentation operations, an additional operation that combines two augmentation mechanisms, and two position-based sampling strategies. The framework is tested on four real-world datasets and can be connected to existing multi-behavior sequential recommenders without changing their internal architecture. The journal version expands the SIGIR work from three augmentation operations and three datasets to five operations and four datasets, turning a conference method into a broader framework for interactions such as viewing, adding to cart, favoriting, and purchasing.

Do et al. [45] focus on the timescale of preference. Their variant experts model short-term and long-term interests separately and then compose the experts for ranking. Across Diginetica, RetailRocket, and Cosmetics, the largest relative gains occur on Diginetica: 38.03% in MRR, 32.66% in NDCG@10, 32.10% in NDCG@20, 17.64% in Recall@10, and 16.22% in Recall@20. On RetailRocket, the corresponding MRR improvement is 26.76%, while the Cosmetics improvements range from 4.12% to 5.32% across the five ranking measures. The variation across datasets also shows that the value of preference composition depends on session density and on how much historical evidence exists for each user.

Coverage across languages, users, and objectives

Iana et al. [39] extend news recommendation across 14 languages, 13 language families, six scripts, five low-resource languages, and five geographic macro-areas. xMIND builds on a source dataset containing about one million users, 130,379 distinct news articles, and more than 24 million clicks. The dataset’s typological-diversity score is 0.42, compared with 0.05 and 0.31 for two multilingual resources used in the paper’s comparison; its language-family diversity reaches 0.93, and its geographic entropy reaches 1.13. A translation assessment uses 50 sampled news items, with two annotators evaluating the translations for each language. Together, these design choices make multilinguality an experimental variable rather than a change of interface language.

Kaur et al. [41] address coverage at the user level. Their two-stage framework first identifies users for whom a conventional recommender produces weak rankings and then allocates those cases to an LLM-based reranker. The evaluation covers eight recommendation algorithms, three LLMs, and three datasets. The authors report an approximately 12% reduction in the number of weak users. The task is carried out under substantial sparsity: Amazon Software records only 0.49% of possible user–item interactions, while Amazon Video Games records 0.68%. Directing LLM inference toward the users who benefit from it links robustness to resource allocation rather than applying the most expensive component uniformly.

Treuillier et al. [44] treat accuracy, diversity, and fairness as three linked objectives in news recommendation. Their ADF framework makes fairness a constraint on diversification, exposing users to a broader range of opinions while maintaining accuracy and preserving the distribution of their interests. This contributes an algorithmic answer to a recurring multi-objective question: fairness-constrained diversification can improve diversity while limiting its impact on accuracy and maintaining fairness across recommendation models. The work also complements the user-perceived objectives in Dokoupil and Peska [38] and the distributional perspective on calibration in Corrêa da Silva and Jannach [43].

Online decisions and social consequences

De Kerpel and Benoit [40] study online slate recommendation through semi-personalized bandits. Their experiments span MovieLens, FinancialNews, and ZOZOTOWN and are repeated across five runs. TreeTS records the lowest average regret on all three datasets: 1.4429 on MovieLens, 0.0717 on FinancialNews, and 0.0851 on ZOZOTOWN. On FinancialNews, 0.0717 represents a 39.0% reduction relative to the best competing value of 0.1176. On ZOZOTOWN, TreeTS reaches a serendipity score of 0.3303, approximately 21.3% above the next score of 0.2723. The dataset comparison is itself informative: the available action space is 84.6% dense for FinancialNews, 20.9% for ZOZOTOWN, and 4.8% for MovieLens, giving the proposed hierarchy three different online environments in which to trade reward against exploration.

Kathriarachchi et al. [42] widen the unit of analysis from ranking outcomes to consequences for customers and society. Their review starts with 68 records and retains 11 articles for detailed synthesis. The hazards include biased product recommendations, privacy breaches, cold-start effects, and recommendation processes that fail to serve customer needs. The proposed responses fall into three connected layers: technological solutions, customer awareness, and laws and regulations. This three-layer structure places the recommender inside an e-commerce system that also contains interfaces, organizations, policies, and informed users.

The same systems perspective runs through the fairness work of Rampisela et al. [37] and Treuillier et al. [44]. One paper examines whether item exposure is measured in a relevance-aware manner; the other constrains diversification by fairness to broaden the range of opinions while preserving users’ interest distributions. Together with the hazard review [42], they show how item-level measures, ranking algorithms, and institutional safeguards can be studied as parts of the same recommendation process.

Highlights of SIGIR 2024

ACM SIGIR is the Association for Computing Machinery’s Special Interest Group on Information Retrieval and one of the central research communities for search, ranking, recommendation, and related forms of information access. Its organizational history reaches back to 1971, while the first conference in the annual SIGIR series was held in Rochester, New York, in 1978. SIGIR 2024 was the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. It took place in Washington, D.C., from July 14 to 18, 2024, and brought together work on retrieval models, recommendation, evaluation, datasets, reproducibility, user interaction, and the growing role of large language models in information access.

This issue of ACM Transactions on Recommender Systems includes six extended papers selected as highlights of SIGIR 2024. Their topics range from explanations and multimodal alignment to multi-behavior learning, fairness measurement, user-perceived recommendation quality, and multilingual news recommendation. The collection shows how recommender-systems research is broadening both the signals used for prediction and the criteria used for evaluation.

Yu, Sugiyama, and Jatowt [34] connect sequential recommendation with collaborative explanation. Their framework predicts a user’s next item while also generating an explanation based on patterns found in the behavior of related users. Recommendation and explanation are learned jointly, allowing the explanatory component to influence the representation used for ranking. The work therefore treats explanations as part of the predictive task rather than as an additional text-generation step after the recommendation has been produced.

Li et al. [35] study how collaborative feedback can be aligned with visual and textual item information. Their CROSS framework combines dynamic item-level alignment with multi-grained collaborative alignment and is evaluated on four datasets using five collaborative-filtering models and six multimodal recommenders. Across the five collaborative-filtering backbones, the reported average improvements range from 21.52% to 70.78%. Across the six multimodal backbones, they range from 8.70% to 20.73%, while comparisons with FETTLE yield gains between 3.82% and 5.24%. These results indicate how strongly recommendation quality can depend on the way interaction signals and multimodal representations are connected.

Xiao et al. [36] address sparse sequential data through multi-behavior data augmentation. MBASR provides five behavior-aware augmentation operations, an additional operation that combines two augmentation mechanisms, and two position-based sampling strategies. The framework is evaluated on four real-world datasets and can be attached to existing multi-behavior sequential recommenders without changing their internal architectures. By modeling interactions such as viewing, favoriting, adding to cart, and purchasing, the paper shows how auxiliary behaviors can provide useful evidence when the target behavior alone is limited.

Rampisela et al. [37] examine the measurement of individual item fairness. Their study analyzes relevance-aware fairness measures on both real and synthetic data and identifies at least one limitation for every measure considered. The authors also propose corrections, reformulations, and usage guidelines. Their results demonstrate that fairness conclusions depend substantially on how exposure and relevance are combined, making metric selection part of the substantive research design rather than a final reporting choice.

Dokoupil and Peska [38] investigate recommendation quality from the user’s perspective. SM-RS 2.0 records user propensities related to relevance, diversity, novelty, and exploration and supports six benchmark tasks, including propensity estimation, selections-aware reranking, and the prediction of perceived quality and satisfaction. Reproduced baselines reach Kendall’s τ = 0.196 for perceived diversity, while the corresponding values are 0.039 for serendipity and 0.045 for satisfaction. The low correlations for several subjective outcomes illustrate how difficult it remains to infer user experience from standard behavioral or offline signals.

Iana, Glavaš, and Paulheim [39] expand news recommendation across languages through the xMIND dataset. The resource covers 14 languages, 13 language families, six scripts, five low-resource languages, and five geographic macro-areas. It builds on a source collection containing about one million users, 130,379 distinct news articles, and more than 24 million clicks. The dataset reaches a typological-diversity score of 0.42, a language-family diversity score of 0.93, and a geographic entropy of 1.13. Its translation assessment uses 50 sampled news items, with two annotators evaluating each language. The result is a benchmark that makes cross-lingual variation a measurable part of recommendation research.

Taken together, the six SIGIR 2024 highlights show a field working simultaneously on richer models and richer evidence. The papers incorporate explanations, images, text, multiple behavior types, subjective judgments, fairness considerations, and linguistic diversity. They also extend evaluation beyond a single accuracy score by examining user perceptions, metric behavior, cross-lingual coverage, and performance across multiple model families and datasets. In that sense, the collection reflects a broader development within SIGIR: recommendation is increasingly studied as an information-access problem involving representation, interaction, evaluation, and social context at the same time.

Add a Comment

Your email address will not be published. Required fields are marked *