When Defaults Become Research Questions: 10 New ACM TORS Papers

Recommender systems have accumulated a comfortable collection of conventions. Give every user the same number of interest vectors. Evaluate explanations using the first recommendation. Use overlapping users as the bridge between domains. Generate item identifiers token by token. Let one LLM reconcile several objectives. Treat the user simulator as experimental infrastructure. Most of these choices have sensible origins. But every now and then it is worth asking a simple question: Why do we do it this way?

The latest ten papers accepted at ACM Transactions on Recommender Systems (TORS) ask that question surprisingly often. Some revisit a modeling convention. Others question an evaluation protocol or system architecture. A few examine assumptions so deeply embedded in the experimental pipeline that we rarely describe them as assumptions at all. This is a useful continuation of our previous announcement of nine newly accepted TORS papers and of the journal’s recent push for stronger reproducibility in offline evaluation. Reproducibility asks whether another researcher can obtain the same result. Several papers here go one step upstream: why was the representation, metric, simulator, or architecture designed that way in the first place?

The ten newly accepted papers

[1] Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation
Elena Epure — Idiap Research Institute, Martigny, Switzerland; Yashar Deldjoo — Polytechnic University of Bari, Bari, Italy; Bruno Sguerra — Deezer Research, Paris, France; Markus Schedl — Johannes Kepler University Linz, Linz, Austria; Manuel Moussallam — Deezer Research, Paris, France. Epure, Deldjoo, and Sguerra contributed equally.

[2] JR4CE: Job Recommendation for Career Exploration
Yosuke Saito — Kyoto University, Kyoto, Japan; Kazunari Sugiyama — Osaka Seikei University, Osaka, Japan.

[3] SC-VUG: Fairness-Aware Cross-Domain Recommendation for Non-Overlapping Users via Sparsity-Calibrated Virtual User Generation
Yuhan Zhao — Hong Kong Baptist University, Hong Kong, China; Weixin Chen — Hong Kong Baptist University, Hong Kong, China; Li Chen — Hong Kong Baptist University, Hong Kong, China; Weike Pan — Shenzhen University, Shenzhen, China. Zhao and Chen contributed equally.

[4] Measuring What Matters: Consistency and Compactness in Evaluation of Counterfactual Explanations
Amir Reza Mohammadi — University of Innsbruck, Innsbruck, Austria; Andreas Peintner — University of Innsbruck, Innsbruck, Austria; Michael Mueller — University of Innsbruck, Innsbruck, Austria; Eva Zangerle — University of Innsbruck, Innsbruck, Austria.

[5] One Size Does Not Fit All: Adaptive Interest Representation Learning for Multi-interest Recommendation
Yaokun Liu — University of Illinois Urbana-Champaign, Champaign, United States; Yifan Liu — University of Illinois Urbana-Champaign, Champaign, United States; Ruichen Yao — University of Illinois Urbana-Champaign, Champaign, United States; Zelin Li — University of Illinois Urbana-Champaign, Champaign, United States; Dong Wang — University of Illinois Urbana-Champaign, Urbana, United States.

[6] Diffusion Models in Recommendation Systems: A Survey
Ting-Ruen Wei — Santa Clara University, Santa Clara, United States; Yi Fang — Santa Clara University, Santa Clara, United States.

[7] Recommendation-as-Experience: A Framework for Context-Sensitive Adaptation in Conversational Recommender Systems
Raj Mahmud — University of Technology Sydney, Sydney, Australia; Shlomo Berkovsky — Australian Institute of Health Innovation, Macquarie University, Sydney, Australia; Mukesh Prasad — University of Technology Sydney, Sydney, Australia; A. Baki Kocaballi — University of Technology Sydney, Sydney, Australia.

[8] SETRec++: Scaling Order-agnostic Identifier for Large Language Model-based Generative Recommendation
Xinyu Lin — National University of Singapore, Singapore; Chuanyu Zhang — University of Science and Technology of China, Hefei, China; Yufan Liu — University of Science and Technology of China, Hefei, China; Haihan Shi — University of Science and Technology of China, Hefei, China; Wenjie Wang — University of Science and Technology of China, Hefei, China; Fuli Feng — University of Science and Technology of China, Hefei, China; Qifan Wang — Meta AI, Menlo Park, United States; See-Kiong Ng — National University of Singapore, Singapore; Tat-Seng Chua — National University of Singapore, Singapore.

[9] Collab-REC: An LLM-based Agentic Framework for Balancing Recommendations in Tourism
Ashmi Banerjee — Technical University of Munich, Munich, Germany; Adithi Satish — Technical University of Munich, Munich, Germany; Fitri Nur Aisyah — Technical University of Munich, Munich, Germany; Wolfgang Wörndl — Technical University of Munich, Munich, Germany; Yashar Deldjoo — Polytechnic University of Bari, Bari, Italy.

[10] Comparative Analysis and Unified Evaluation Framework of User Simulators for Realistic Behaviour Modelling
Md Faisal Ahmed — RMIT University, Melbourne, Australia, and Noakhali Science and Technology University, Bangladesh; Estrid He — RMIT University, Melbourne, Australia; Chenglong Ma — RMIT University, Melbourne, Australia; Jeffrey Chan — RMIT University, Melbourne, Australia.

Adaptive representations

Liu et al. [5] question one of the cleanest conveniences in multi-interest recommendation: choosing one number K and giving every user K interest vectors. AirRec estimates a user-specific number instead. Across Retail Rocket, Gowalla, and Amazon Books, it improves NDCG@20 by about 15% and NDCG@50 by about 25% on average over the compared state-of-the-art methods. The diagnostics are more telling than another accuracy table: over 75% of sampled low-diversity Amazon Books users receive two or three interests, while higher-diversity users concentrate around five or six.

Zhao et al. [3] examine another convenient representation: overlapping users as the bridge in cross-domain recommendation. SC-VUG creates probabilistic source-domain representations for target users who do not occur in the source domain. On Epinions with BiTGCF, the resulting NDCG@20 is .0863 and the group disparity measured by UGF is reported as essentially zero at the paper’s precision. More importantly, the subgroup analysis indicates that the gap closes mainly by improving non-overlapping users rather than degrading the overlapping group. The authors also vary the available overlap from 25% to 100%, directly testing the condition behind the proposed method.

Mahmud et al. [7] move adaptation into the conversation. Their Recommendation-as-Experience framework distinguishes educative, explorative, and affective interaction aims. In a vignette-based study with 168 participants across ten domains, perceived item value is strongly associated with all three aims, with correlations around .8. Domain also changes what participants want from the interaction. The study provides evidence for adapting conversational behavior to the situation, while longitudinal effects in a deployed recommender remain open.

Evaluation and evidence under scrutiny

Mohammadi et al. [4] examine the convention of evaluating counterfactual explanations through the first recommended item. Their experiments span three datasets, two recommenders, and six explanation methods. Aggregating evaluation over the first five recommendations produces a Spearman rank correlation of .971 and Rank Variance of .063, making method rankings substantially more stable. Compactness changes the comparison as well: methods that perform strongly with large perturbation sets can lose their advantage when explanations are restricted to a few interactions. There is a small reporting inconsistency between prose and table for one top-1 result, but the broader top-k pattern is clear.

Ahmed et al. [10] inspect another part of the evaluation stack that is easy to treat as neutral: the simulated user. Their comparison covers LLM-based simulators such as RecAgent and Agent4Rec and RL environments including RecSim and KuaiSim. LLM-based simulation performs strongly on denser data, but performance falls by about 44–58% under sparse and negative-feedback conditions such as those constructed from Book-Crossing. The RL environments are stronger at reproducing some longer-term effects such as user conformity. Simulator choice can therefore become an experimental assumption rather than mere infrastructure.

Epure et al. [1] reach a related conclusion for LLM-based music recommendation. This is a survey and position paper rather than a new benchmark study. It reviews how LLMs enter user modeling, item modeling, natural-language recommendation, and evaluation, and argues for assessing dimensions such as grounding, personalization, discovery, cultural coverage, and hallucination. This connects well to our recent post on the ACM TORS Special Issue on Music Recommender Systems, where evaluation, context, LLMs, fairness, and user experience already appeared as central problems for the domain.

Wei and Fang [6] provide the map for another fast-growing line. Their survey manually curates 188 diffusion-based recommender-system papers and organizes them by recommendation task, modality and domain, and trustworthy objectives. It complements our July discussion of ACM TORS Volume 4, Issue 3. That issue included the reproducibility study by Benigni et al., which found only 25% of the examined reported experimental results fully reproducible. One paper maps how large diffusion recommendation has become; the earlier one asks how much of its evidence survives closer inspection.

Decomposing the LLM recommender

Lin et al. [8] move surprisingly far down the generative-recommendation stack. SETRec replaces ordered item-token sequences with sets of order-agnostic collaborative and semantic tokens, while SETRec++ adds mechanisms to disentangle their information. On Toys with Qwen-1.5B, cold-item Recall@10 rises from .0883 with SETRec to .1055 with SETRec++. More strikingly, simultaneous token generation yields reported inference speedups of roughly 8× to 18× across the four evaluated datasets.

The gains depend on the backbone. SETRec++ is less consistently beneficial with T5, making SETRec the simpler option when resources or architecture do not favor the additional disentanglement machinery. This is a useful result because the paper does more than propose a more elaborate identifier: it also shows where the elaboration pays off.

Banerjee et al. [9] decompose the LLM recommender at a higher level. Collab-REC assigns personalization, popularity, and sustainability to three specialist agents and uses a deterministic moderator for catalog grounding and aggregation. The study evaluates 900 tourism queries over 200 European cities. With Claude, iterative refinement raises the grounded success score from .465 to .657 and catalog coverage from 66% to 81.5%. Most gains arrive by roughly four to five rounds.

The deterministic moderator is the more interesting design choice here. The generative components propose candidates, while catalog validity and stakeholder trade-offs remain explicit and reproducible. This also makes the cost of additional rounds visible instead of hiding coordination inside one large prompt.

Diversity with a purpose

Saito and Sugiyama [2] ask what diversity should accomplish in job recommendation. A list can be diverse while being useless for career exploration. JR4CE therefore combines interaction data with explicit preferences, current employment information, and data augmentation from suitable role-model users to recommend alternative career paths that remain related to the user’s direction.

The setting is difficult: the evaluated job-interaction matrices are more than 99.9% sparse. On GLIT-2022, JR4CE reaches NDCG@5=.0370, compared with .0245 for KGRec. Applying MMR to KGRec increases diversity but lowers its NDCG@5 to .0137. The absolute values are small, as one would expect on data this sparse, but the comparison captures the main result: diversity that follows a career-exploration objective preserves considerably more relevance than generic diversification.

The paper also makes the trade-off visible in its ablation. On GLIT-2021, removing JR4CE’s Diversity Data Augmentation component slightly increases NDCG@5 from .0407 to .0414, while ILAD@5 falls from .3444 to .3136. The harder question is longitudinal: whether these diversified lists actually improve career exploration rather than its offline proxies.

A useful question for the methods section

Previous ACM TORS round-ups on RS_c have documented how quickly recommender-systems research is adding objectives, modalities, generative models, and evaluation criteria. Our December 2025 collection of ten TORS papers, for example, already covered evaluation metrics, user behavior, LLM adaptation, and responsible recommendation. The present batch has a slightly different character. It repeatedly turns choices that normally sit in the implementation or evaluation details into explicit objects of study.

That does not make every default a mistake. Defaults are useful precisely because researchers cannot reopen every decision in every experiment. But once a convention affects which method wins, which users benefit, or what conclusion an evaluation supports, it becomes scientifically interesting in its own right. Sometimes progress begins with a new architecture. Sometimes it begins with a much shorter question: Why do we do it this way?

Add a Comment

Your email address will not be published. Required fields are marked *