What Survives the Stress Test? 18 New Papers from ACM TORS (Vol. 4; Issue 4; December)

The latest issue of ACM Transactions on Recommender Systems brings 18 new papers, including 13 “Highlights of RecSys’24.” Volume 4, Issue 4 covers a broad range of topics: multimodal recommendation, sequential behavior, fairness, calibration, news, music, sustainability, quantum optimization, and research infrastructure. More importantly, many of the papers share a common pattern: they test how robust our conclusions are when the conditions change. Sequences are shuffled, modalities disappear, models move from hot start to cold start, fairness methods are tested on denser graphs, and offline assumptions meet users, organizations, power meters, and hardware.

The timing is convenient. There are still about three weeks before the recommender-systems community meets in Minneapolis for RecSys 2026. That leaves just enough time to read a few of these papers before the conference—and, in the case of the RecSys’24 highlights, perhaps meet some of their authors again in Minneapolis. With the 20th RecSys approaching, the issue also raises a fitting question: after two decades of recommender-systems research, which of our models, metrics, datasets, and conclusions actually survive a stress test? The full list of papers, authors, affiliations, and DOI links is at the end of this post.

Stress Test 1: What Happens When We Disturb the Data?

The most literal stress test comes from Klenitskiy et al. [9] in An Analysis of Sequential Patterns in Datasets for Evaluation of Sequential Recommendations. They take 19 datasets used for sequential recommendation and shuffle the order of user interactions. With SASRec+, NDCG@10 changes by +4% on Yelp and falls by only 3% on RetailRocket, 4% on Foursquare, and 9% on Gowalla. At the other end, MovieLens-20M loses 58%, Yambda-50M 60%, Zvuk 61%, and 30Music 86%. The same procedure thus finds anything from weak to dominant sequential structure. The paper also shows that preprocessing can change the result. Before comparing increasingly elaborate sequential models, checking whether the benchmark cares about sequence at all seems like a sensible experiment.

Zhao et al. [3] and Alshabanah et al. [6] disturb two other assumptions about interaction data. In Unlocking the Unlabeled Data, Zhao et al. introduce a neutral state between positive and negative feedback rather than forcing all unobserved interactions into the negative class. With NGCF on Yelp, Recall@20 rises from 0.0738 to 0.0977 and NDCG@20 from 0.0447 to 0.0613; the average gain across six metrics is 34.59%. In The Upside of Bias, Alshabanah et al. deliberately sample more heavily from rare interactions. On BookCrossing, mean synthesis improves overall HR by 59.8% and tail NDCG by 101.7%; the attention variant improves tail NDCG by 105.8%. Both papers obtain gains by reconsidering which parts of sparse interaction data carry the most information.

Music provides another test of what we treat as signal. Tran et al. [7] report in PISA: Combining Transformers and ACT-R for Repeat-Aware Sequential Listening Session Recommendation that 72.37% of target songs on Last.fm and 89.10% on Deezer are repeats. PISA explicitly models repetition and exploration and produces the best result in 10 of the 12 central NDCG/Recall comparisons. On Deezer, overall NDCG reaches 11.20%, while its predicted repeat ratio of 85.16% is close to the observed 89.10%. This fits the broader complexity of music recommendation, where discovery and familiarity coexist rather than compete as simple opposites.

Kheiri and Ding [17] broaden the data question to emotion. Their Emotion-Aware Recommender Systems: A Comprehensive Review, Challenges, and Future Directions synthesizes 196 papers published from 2014 to 2024. Individual studies report substantial effects—for example, one multimodal video recommender moves from 78% to 89% accuracy after adding affective features. Yet the review also documents the complications that arrive with emotion as a signal: subjective labels, cultural differences, noisy sensing, privacy, and multimodal fusion. More context can improve recommendation, but the context itself has to survive measurement.

Stress Test 2: Do the Models Still Work When the Conditions Change?

Missing information provides a direct robustness test for multimodal recommenders. Ganhör et al. [5], in Single-Branch Network Architectures to Close the Modality Gap in Multimodal Recommendation, compare shared single-branch encoders with conventional multi-branch systems. Under missing modalities, the single-branch systems are significantly better in 16 of 31 tested configurations, with NDCG@10 gains reaching about 40%. On MovieLens-1M with only one modality available, the basic multi-branch model reaches NDCG@10=0.0866; the strongest single-branch configuration reaches 0.1851. This is precisely the kind of failure condition discussed in our earlier post on challenges in modern multimodal recommendation.

Yi and Ounis [1] test Unifying Isolated Processes for Enhanced Multi-Modal Recommendations Using a Graph Transformer under several shifts. Their UGT model improves Recall@20 over the strongest competing result by 6.60% on Amazon Sports, 13.97% on Clothing, and 6.41% on Baby. Under cold start on Clothing, NDCG@20 improves by 20.30% over FREEDOM. The model is then moved from e-commerce to micro-video recommendation, where its out-of-domain NDCG@20 gains range from 5.80% to 8.30%. The strongest evidence in this paper is therefore not only another benchmark win, but that part of the gain survives a change of domain.

Tamm and Aljanaki [13] ask whether pretrained audio representations survive a change of task. In Adopting State-of-the-Art Pretrained Audio Representations for Music Recommender Systems, BERT4Rec with MusiCNN reaches NDCG@50=0.0520 in hot start, while the purely collaborative ELSA baseline reaches 0.0631. In cold start, the ordering changes: MERT has the highest HitRate@20 at 0.2829, while MusiCNN produces the best NDCG@20 at 0.0496. A representation that performs well on a Music Information Retrieval task still has to prove itself again when the downstream task becomes recommendation.

Sato [12] performs a similar test on calibration. Calibrating the Predictions of User Preferences in Top-Ranked Items asks whether scores that look calibrated globally remain calibrated among the items users actually see. On KuaiRec, Top-N-Focused calibration reduces ECE@20 from 0.165 to 0.016 and rank-discounted calibration error from 0.220 to 0.022. On sparse Modcloth, however, average non-parametric ECE worsens from 0.087 to 0.122. Calibration therefore needs a location: calibrated where?

Lee et al. [2] stress-test the design choices behind LightGCN. Their LightGCN++ makes degree scaling, neighbor weighting, and layer pooling more flexible while retaining linear message passing. NDCG@20 rises from 0.2549 to 0.2747 on LastFM, from 0.1415 to 0.1764 on CiteULike, and from 0.0397 to 0.0513 on Alibaba. The largest reported NDCG@10 gain is 29.38%. LightGCN++ is also less sensitive when feature transformations and nonlinearities are added. The intervention is small; the result is a version of LightGCN whose simplicity is less brittle.

Stress Test 3: What Happens When the Benchmark Meets the World?

Tomita and Yokoyama [10] and Boratto et al. [11] put fairness methods under two different kinds of pressure. In Balancing Fairness and High Match Rates in Reciprocal Recommender Systems, the efficiency-oriented method produces 111.37 expected matches on Japanese dating data, together with 434 male-side and 331 female-side envy violations. Nash Social Welfare produces 90.39 expected matches with only 31 and 14 violations. On the Speed Dating dataset, envy violations fall from 4,417/4,043 to 17/18, while expected matches fall from 160.04 to 131.99. The trade-off becomes visible rather than disappearing into one aggregate metric.

Boratto et al. [11] vary graph scale and density in Graph Augmentation for Intersectional Unfairness Mitigation. On KRECS, HMLET’s NDCG rises from 3.94 to 7.51 while intersectional disparity worsens from 0.15 to 0.47. On LFM1M, the same family of interventions can reduce disparity—for HMLET from 0.50 to 0.27—while leaving NDCG almost unchanged. On the much larger KRECB graph, XSimGCL moves from NDCG 4.93/disparity 0.10 to 8.76/0.09. Fairness results depend strongly on the graph on which they are produced.

Timmers et al. [14] test whether diversity in a recommendation list survives contact with user choice. PI-adaptDiv adapts diversification according to the diversity users actually consume. In the online experiment, mean consumed-diversity change is -0.015 for SVD, +0.054 for PI-adaptDiv, and +0.131 for the aggressive MMR-EDC condition. PI-adaptDiv also produces a mean item-rating change of +0.113. Its within-condition increase in consumed diversity is significant at p=0.006, while pairwise Holm-corrected differences against SVD and MMR remain inconclusive. Recommended diversity and consumed diversity are related measures, but the user sits between them.

News recommendation provides two complementary encounters with organizational reality. Vandenbroucke, Michiels, and Smets [8] conduct 22 interviews across four news organizations in Welcome to the Metrics Jungle. Editorial, commercial, product, and technical stakeholders connect recommendation to different combinations of engagement, reach, conversion, retention, pluralism, and editorial goals. Holzleitner et al. [16] then test controlled personalization directly at Aftenposten in a 34-day production A/B test with about 28,000 established paying mobile subscribers per condition. CTR rises from 0.524 to 0.601, or 14.60%, while daily Average Recommendation Popularity falls by 49.29% and Average Click Popularity by 44.04%. CTR here is clicks divided by visible article impressions, and the reported analysis is restricted to articles clicked at least once during the test; including never-clicked articles would reduce the absolute CTR values. The section-impression Gini also falls from 0.650 to 0.615, while changes in reading depth and duration are much smaller. Together, the two papers show why news recommendation is difficult to evaluate with one KPI.

Wegmeth et al. [4] add energy to the evaluation. Green Recommender Systems studies 23 models on 13 datasets and five computers, with energy measured directly. The reconstructed deep-learning research pipeline is estimated to emit around 42 times more CO2-equivalent than the traditional-model pipeline, with an estimated average of 2,909 kg CO2e per deep-learning-based paper. Dataset choice alone changes DGCF from roughly 0.005 kWh on Hetrec-LastFM to 6.6 kWh on Yelp-2018. This develops a line of work we discussed earlier in Green Recommender Systems: A Call for Attention. Disclosure: I am a co-author of [4].

Niu et al. [15] take a different system all the way to physical hardware. In Performance-Driven QUBO for Recommender Systems on Quantum Annealers, PDQUBO reaches NDCG@10=0.1140 versus 0.1021 for the strongest competing QUBO method in one Item-KNN setting, a gain of 11.7%. The direct quantum-hardware experiment is more revealing. At 50 features, quantum annealing returns QUBO energy 2.361 while simulated annealing and hybrid optimization reach about -1.76. At 150 features, direct quantum annealing deteriorates to 164.041 while the other two remain near -7.97 and -8.21. The formulation scales more gracefully than the current hardware.

Heitz et al. [18] address another barrier between offline benchmarks and empirical research. Informfully is an open-source platform for running recommender-system user studies with external algorithms, experimental groups, surveys, multimodal content, and interaction logging. Its four reported deployments include 151, 941, 143, and 283 participants. The largest ran for 183 days with 941 participants, 41,340 items, five algorithms, and about 1.21 million recorded interactions. We have also discussed POPROX, which tackles a related problem through a live news-recommendation environment with its own participant pool. Both efforts reduce the amount of infrastructure researchers need before they can test what happens when real users enter the loop.

Twenty Years In, Stress Tests Matter More

Across these papers, robustness rarely has a single meaning. Sometimes it means surviving shuffled data. Sometimes it means functioning with missing modalities or new items. Elsewhere, a result has to survive a change in graph density, another optimization objective, an organization with competing goals, or the physical hardware on which the method is supposed to run. Several papers become most informative exactly where a familiar conclusion stops transferring cleanly.

That seems appropriate just before RecSys returns to Minneapolis for its twentieth edition. A mature research field should keep producing better models. It should also become increasingly good at knowing which conclusions remain true when we shuffle the history, remove a modality, move from hot start to cold start, change the graph, introduce another objective, or leave the benchmark altogether. Sometimes almost nothing happens. Sometimes NDCG drops by 86%. Both outcomes tell us something.

The 18 Papers in ACM TORS Volume 4, Issue 4

The list number below is the reference number used throughout this post. Papers [1]–[13] belong to the “Highlights of RecSys’24” section; [14]–[18] are regular papers.

  1. Unifying Isolated Processes for Enhanced Multi-Modal Recommendations Using a Graph Transformer. Zixuan Yi — University of Glasgow, United Kingdom; Iadh Ounis — School of Computing Science, University of Glasgow, United Kingdom. ACM TORS Article 47.
  2. Revisiting LightGCN: Unexpected Inflexibility, Inconsistency, and A Remedy Towards Improved Recommendation. Geon Lee, Kyungho Kim, Fanchen Bu, Langzhang Liang, and Kijung Shin — KAIST, Republic of Korea. ACM TORS Article 48.
  3. Unlocking the Unlabeled Data: Enhancing Recommendations with Neutral Samples and Uncertainty. Yuhan Zhao — Harbin Engineering University, China, and Hong Kong Baptist University, Hong Kong; Rui Chen — Harbin Engineering University, China; Qilong Han — Harbin Engineering University, China; Hongtao Song — Harbin Engineering University, China; Li Chen — Hong Kong Baptist University, Hong Kong. ACM TORS Article 49.
  4. Green Recommender Systems: Understanding and Minimizing the Carbon Footprint of AI-Powered Personalization. Lukas Wegmeth — Intelligent Systems Group, University of Siegen, Germany; Tobias Vente — Intelligent Systems Group, University of Siegen, Germany; Alan Said — University of Gothenburg, Sweden; Joeran Beel — Intelligent Systems Group, University of Siegen, Germany. ACM TORS Article 50.
  5. Single-Branch Network Architectures to Close the Modality Gap in Multimodal Recommendation. Christian Ganhör — Institute of Computational Perception, Johannes Kepler University Linz, Austria; Marta Moscati — Institute of Computational Perception, Johannes Kepler University Linz, Austria; Anna Hausberger — Institute of Computational Perception, Johannes Kepler University Linz, Austria; Shah Nawaz — Institute of Computational Perception, Johannes Kepler University Linz, Austria; Markus Schedl — Institute of Computational Perception, Johannes Kepler University Linz, and AI Lab, Linz Institute of Technology, Austria. ACM TORS Article 51.
  6. The Upside of Bias: Personalizing Long-Tail Item Recommendations with Biased Sampling. Abdulla Alshabanah — University of Southern California, United States; Keshav Balasubramanian — University of Southern California, United States; Elan Markowitz — University of Southern California, United States; Greg Ver Steeg — University of California Riverside, United States; Murali Annavaram — University of Southern California, United States. ACM TORS Article 52.
  7. PISA: Combining Transformers and ACT-R for Repeat-Aware Sequential Listening Session Recommendation. Viet Anh Tran — Deezer Research, France; Guillaume Salha-Galvan — SPEIT, Shanghai Jiao Tong University, China, with part of the work conducted at Deezer Research; Bruno Sguerra — Deezer Research, France; Romain Hennequin — Deezer Research, France. ACM TORS Article 53.
  8. Welcome to the Metrics Jungle: Organizational Stakeholder Perspectives on Evaluation of News Recommender Systems in Industry. Hanne Vandenbroucke — imec-SMIT, Vrije Universiteit Brussel, Belgium; Lien Michiels — imec-SMIT, Vrije Universiteit Brussel, and Adrem Data Lab, Universiteit Antwerpen, Belgium; Annelien Smets — imec-SMIT, Vrije Universiteit Brussel, Belgium. ACM TORS Article 54.
  9. An Analysis of Sequential Patterns in Datasets for Evaluation of Sequential Recommendations. Anton Klenitskiy — Sber AI Lab, Russia; Anna Volodkevich — Sber AI Lab and Skoltech, Russia; Anton Pembek — Sber AI Lab and Lomonosov Moscow State University, Russia; Alexey Vasilev — Sber AI Lab and HSE University, Russia. ACM TORS Article 55.
  10. Balancing Fairness and High Match Rates in Reciprocal Recommender Systems: A Nash Social Welfare Approach. Yoji Tomita — CyberAgent, Inc., Japan; Tomohiko Yokoyama — School of Information Science and Technology, The University of Tokyo, Japan. ACM TORS Article 56.
  11. Graph Augmentation for Intersectional Unfairness Mitigation: A Study across Dataset Scales and Interaction Densities. Ludovico Boratto — University of Cagliari, Italy; Francesco Fabbri — Spotify AB, Spain; Gianni Fenu — University of Cagliari, Italy; Mirko Marras — University of Cagliari, Italy; Giacomo Medda — University of Cagliari, Italy. ACM TORS Article 57.
  12. Calibrating the Predictions of User Preferences in Top-Ranked Items. Masahiro Sato — FUJIFILM, Japan. ACM TORS Article 58.
  13. Adopting State-of-the-Art Pretrained Audio Representations for Music Recommender Systems. Yan-Martin Tamm — University of Tartu, Estonia; Anna Aljanaki — University of Tartu, Estonia. ACM TORS Article 59.
  14. PI-adaptDiv: An Adaptive Algorithm to Prevent and Escape Online Filter Bubbles. Colin Timmers — UCLouvain, Belgium; François Fouss — UCLouvain, Belgium; Corentin Vande Kerckhove — UCLouvain, Belgium. ACM TORS Article 60.
  15. Performance-Driven QUBO for Recommender Systems on Quantum Annealers. Jiayang Niu — RMIT University, Australia; Jie Li — School of Computing Technologies, RMIT University, Australia; Ke Deng — School of Computing Technologies, RMIT University, Australia; Mark Sanderson — School of Computing Technologies, RMIT University, Australia; Nicola Ferro — Department of Information Engineering, University of Padua, Italy; Yongli Ren — School of Computing Technologies, RMIT University, Australia. ACM TORS Article 61.
  16. Controlled Personalization in Legacy Media Online Services: A Case Study in News Recommendation. Marlene Holzleitner — University of Klagenfurt, Austria; Stephan Leitner — University of Klagenfurt, Austria; Hanna Lind Jorgensen — Schibsted, Norway; Christoph Schmitz — Schibsted, Norway; Jacob Welander — Schibsted, Norway; Dietmar Jannach — University of Klagenfurt, Austria, and University of Bergen, Norway. ACM TORS Article 62.
  17. Emotion-Aware Recommender Systems: A Comprehensive Review, Challenges, and Future Directions. Kiana Kheiri — Computer Science, Toronto Metropolitan University, Canada; Chen Ding — Computer Science, Toronto Metropolitan University, Canada. ACM TORS Article 63.
  18. Informfully Research Platform – An Open Source Project for Conducting Empirical Research with Recommender Systems. Lucien Heitz — Department of Informatics and Digital Society Initiative, University of Zurich, Switzerland; Julian A. Croci — Department of Informatics, University of Zurich, Switzerland; Madhav Sachdeva — Department of Informatics, University of Zurich, Switzerland; Abraham Bernstein — Department of Informatics, University of Zurich, Switzerland. ACM TORS Article 64.

Add a Comment

Your email address will not be published. Required fields are marked *