Technical and Regulatory Perspectives on Information Retrieval and Recommender Systems
Technical and Regulatory Perspectives on Information Retrieval and Recommender Systems
- Conference Article
59
- 10.1145/2600428.2610382
- Jul 3, 2014
The development and evaluation of Information Retrieval and Recommender Systems has traditionally focused on the relevance and accuracy of retrieved documents and recommendations, respectively. However, there is an increasing realization that accuracy alone might be a sub-optimal strategy for a successful user experience. Properties such as novelty and diversity have been explored in both fields for assessing and enhancing the usefulness of search results and recommendations. In this doctoral research we study the assessment and enhancement of both properties in the confluence of Information Retrieval and Recommender Systems.
- Research Article
5
- 10.1145/3458537.3458545
- Jun 1, 2019
- ACM SIGIR Forum
Information retrieval addresses the information needs of users by delivering relevant pieces of information but requires users to convey their information needs explicitly. In contrast, recommender systems offer personalized suggestions of items automatically. Ultimately, both fields help users cope with information overload by providing them with relevant items of information. This thesis aims to explore the connections between information retrieval and recommender systems. Our objective is to devise recommendation models inspired in information retrieval techniques. We begin by borrowing ideas from the information retrieval evaluation literature to analyze evaluation metrics in recommender systems [2]. Second, we study the applicability of pseudo-relevance feedback models to different recommendation tasks [1]. We investigate the conventional top-N recommendation task [5, 4, 6, 7], but we also explore the recently formulated user-item group formation problem [3] and propose a novel task based on the liquidation of long tail items [8]. Third, we exploit ad hoc retrieval models to compute neighborhoods in a collaborative filtering scenario [9, 10, 12]. Fourth, we explore the opposite direction by adapting an effective recommendation framework to pseudo-relevance feedback [13, 11]. Finally, we discuss the results and present our conclusions. In summary, this doctoral thesis adapts a series of information retrieval models to recommender systems. Our investigation shows that many retrieval models can be accommodated to deal with different recommendation tasks. Moreover, we find that taking the opposite path is also possible. Exhaustive experimentation confirms that the proposed models are competitive. Finally, we also perform a theoretical analysis of some models to explain their effectiveness. Advisors : Álvaro Barreiro and Javier Parapar. Committee members : Gabriella Pasi, Pablo Castells and Fidel Cacheda. The dissertation is available at: https://www.dc.fi.udc.es/~dvalcarce/thesis.pdf.
- Book Chapter
12
- 10.1007/978-3-319-97556-6_5
- Sep 20, 2018
This chapter provides a brief introduction to two of the most common applications of data science methods in e-commerce: information retrieval and recommender systems. First, a brief overview of the systems is presented followed by details on some of the most commonly applied models used for these systems and how these systems are evaluated. The chapter ends with an overview of some of the application areas in which information retrieval and recommender systems are typically developed.
- Book Chapter
10
- 10.1007/978-3-642-23014-1_15
- Jan 1, 2011
The powerful and democratic activity of social tagging allows the wide set of Web users to add free annotations on resources. Tags express user interests, preferences and needs, but also automatically generate folksonomies. They can be considered as gold mine, especially for e-commerce applications, in order to provide effective recommendations. Thus, several recommender systems exploit folksonomies in this context. Folksonomies have also been involved in many information retrieval approaches. In considering that information retrieval and recommender systems are siblings, we notice that few works deal with the integration of their approaches, concepts and techniques to improve recommendation. This paper is a first attempt in this direction. We propose a trail through recommender systems, social Web, e-commerce and social commerce, tags and information retrieval: an overview on the methodologies, and a survey on folksonomy-based information retrieval from recommender systems point of view, delineating a set of open and new perspectives.
- Conference Article
8
- 10.1109/cbmi.2019.8877420
- Sep 1, 2019
Multimedia search is an emerging area in information retrieval (IR) and recommender systems (RS) research. However, there is a lack of standardized audiovisual datasets that include rich content descriptors, which are a necessity in content-based IR and RS. The contributions of this paper are twofold: First, we present a new multimedia dataset of movie clips, named MFVCD-7K Multifaceted Video Clip Dataset, that comes with low-level and semantic multimodal descriptions of their content (textual, audio, and visual). In addition, we showcase the use of this dataset for a novel content-based video clip retrieval and result diversification task we introduce. We investigate baseline algorithms for retrieval and diversification, and provide experimental results according to relevance and diversity measures. We believe that both dataset and baseline results constitute an important asset for the IR, RS, and multimedia communities.
- Dissertation
- 10.18122/td.2238.boisestate
- May 1, 2024
Evaluation of recommender systems is key to ensuring that we are making progress by only promoting proposed algorithms that actually outperform the state-of-the-art. Evaluation can be performed offline using logged historic data, or online such as A/B tests. Offline evaluation is the most popular evaluation paradigm used in recent research publications because of its accessibility. It has been an area of study since its earliest days, and is still currently an active area of research. Even though significant work has been done to improve the process and metrics of offline evaluation, it still faces fundamental difficulties. Given the significant role of offline evaluation in recommender system evaluation, it is important that the difficulties that embody it be mitigated. This dissertation aims to improve offline evaluation of recommender systems by identifying gaps in existing practices in two specific components of the offline evaluation protocol: candidate set sampling – which is the selection of the set of candidate items that the recommendation system is expected to rank for each user in an experiment; and statistical inference techniques – which are used to analyze the evaluation results in order to make inferences about the effectiveness of a proposed system relative to a baseline. This dissertation addresses gaps in candidate set sampling by showing that uniform sampling of the candidate set exacerbates popularity bias, while popularity-weighted sampling mitigates this bias. Additionally, it demonstrated that candidate set sampling improves the accuracy of effectiveness performance estimation in top-N recommender systems. With respect to statistical inference, this dissertation identified a lack of rigorous statistical analysis in evaluations within RecSys and revealed that the Wilcoxon and Sign tests display higher-than-expected Type-1 error rates for large sample sizes, recommending their discontinuation in recommender system experiments. It demonstrates that in Top-N recommendation and large search evaluation data, most tests are likely to yield statistically significant results, emphasizing the need to prioritize effect size for practical or scientific significance. Additionally, the dissertation found that the Benjamini-Yekutieli test exhibited the lowest error rate and greater power than the Bonferroni test, recommending it as the default correction for comparing multiple systems in information retrieval and recommender system experiments. By addressing these gaps, the dissertation significantly contributes to the improvement of offline evaluation, equipping recommender system researchers with evidence-based knowledge to make informed decisions when configuring their evaluation experiments.
- Conference Article
6
- 10.1109/aiccsa.2016.7945626
- Nov 1, 2016
The emergence of social networks and the communication facilities they offer have generated an enormous informational mass. This social content is used in several research and industrial works and has had a great impact in different processes. In this paper, we present an overview of social information use in Information Retrieval (IR) and Recommendation systems. We first describe several user profile models using social information. A special attention is given to the following points: the analysis of the different user profiling models incorporating social content in Information Retrieval (IR) and in social recommendation methods. We distinguish between the models using social signals and relations, and the models using temporal information. We also present current and future challenges and research directions to enhance IR and recommendation process. We then describe our proposed model of social polarized and temporal user profile building and use in social recommendation context. Our proposal tries to address open challenges and establish a new model of user profile that fits information needs in recommender systems.
- Conference Article
9
- 10.1145/3583780.3615010
- Oct 21, 2023
Information Retrieval (IR) and Recommender Systems (RSs) tasks are moving from computing a ranking of final results based on a single metric to multi-objective problems. Solving these problems leads to a set of Pareto-optimal solutions, known as Pareto frontier, in which no objective can be further improved without hurting the others. In principle, all the points on the Pareto frontier are potential candidates to represent the best model selected with respect to the combination of two, or more, metrics. To our knowledge, there are no well-recognized strategies to decide which point should be selected on the frontier in IR and RSs. In this paper, we propose a novel, post-hoc, theoretically-justified technique, named "Population Distance from Utopia" (PDU), to identify and select the one-best Pareto-optimal solution. PDU considers fine-grained utopia points, and measures how far each point is from its utopia point, allowing to select solutions tailored to user preferences, a novel feature we call "calibration". We compare PDU against state-of-the-art strategies through extensive experiments on tasks from both IR and RS, showing that PDU combined with calibration notably impacts the solution selection.
- Dissertation
15
- 10.18297/etd/2744
- Oct 5, 2017
Websites and online services thrive with large amounts of online information, products, and choices, that are available but exceedingly difficult to find and discover. This has prompted two major paradigms to help sift through information: information retrieval and recommender systems. The broad family of information retrieval techniques has given rise to the modern search engines which return relevant results, following a user's explicit query. The broad family of recommender systems, on the other hand, works in a more subtle manner, and do not require an explicit query to provide relevant results. Collaborative Filtering (CF) recommender systems are based on algorithms that provide suggestions to users, based on what they like and what other similar users like. Their strength lies in their ability to make serendipitous, social recommendations about what books to read, songs to listen to, movies to watch, courses to take, or generally any type of item to consume. Their strength is also that they can recommend items of any type or content because their focus is on modeling the preferences of the users rather than the content of the recommended items. Although recommender systems have made great strides over the last two decades, with significant algorithmic advances that have made them increasingly accurate in their predictions, they suffer from a few notorious weaknesses. These include the cold-start problem when new items or new users enter the system, and lack of interpretability and explainability in the case of powerful black-box predictors, such as the Singular Value Decomposition (SVD) family of recommenders, including, in particular, the popular Matrix Factorization (MF) techniques. Also, the absence of any explanations to justify their predictions can reduce the transparency of recommender systems and thus adversely impact the user's trust in them. In this work, we propose machine learning approaches for multi-domain Matrix Factorization (MF) recommender systems that can overcome the new user cold-start problem. We also propose new algorithms to generate explainable recommendations, using two state of the art models: Matrix Factorization (MF) and Restricted Boltzmann Machines (RBM). Our experiments, which were based on rigorous cross-validation on the MovieLens benchmark data set and on real user tests, confirmed that our proposed methods succeed in generating explainable recommendations without a major sacrifice in accuracy.
- Research Article
29
- 10.1145/3636341.3636351
- Jun 1, 2023
- ACM SIGIR Forum
This report documents the program and the outcomes of Dagstuhl Seminar 23031 "Frontiers of Information Access Experimentation for Research and Education", which brought together 38 participants from 12 countries. The seminar addressed technology-enhanced information access (information retrieval, recommender systems, natural language processing) and specifically focused on developing more responsible experimental practices leading to more valid results, both for research as well as for scientific education. The seminar featured a series of long and short talks delivered by participants, who helped in setting a common ground and in letting emerge topics of interest to be explored as the main output of the seminar. This led to the definition of five groups which investigated challenges, opportunities, and next steps in the following areas: reality check, i.e. conducting real-world studies, human-machine-collaborative relevance judgment frameworks, overcoming methodological challenges in information retrieval and recommender systems through awareness and education, results-blind reviewing, and guidance for authors. Date: 15--20 January 2023. Website: https://www.dagstuhl.de/23031.
- Research Article
1782
- 10.1145/3285029
- Feb 25, 2019
- ACM Computing Surveys
With the growing volume of online information, recommender systems have been an effective strategy to overcome information overload. The utility of recommender systems cannot be overstated, given their widespread adoption in many web applications, along with their potential impact to ameliorate many problems related to over-choice. In recent years, deep learning has garnered considerable interest in many research fields such as computer vision and natural language processing, owing not only to stellar performance but also to the attractive property of learning feature representations from scratch. The influence of deep learning is also pervasive, recently demonstrating its effectiveness when applied to information retrieval and recommender systems research. The field of deep learning in recommender system is flourishing. This article aims to provide a comprehensive review of recent research efforts on deep learning-based recommender systems. More concretely, we provide and devise a taxonomy of deep learning-based recommendation models, along with a comprehensive summary of the state of the art. Finally, we expand on current trends and provide new perspectives pertaining to this new and exciting development of the field.
- Research Article
- 10.1145/3722449.3722464
- Dec 1, 2024
- ACM SIGIR Forum
IIR 2024, the 14th Italian Information Retrieval Workshop, served as the annual event for the Information Retrieval (IR) and Recommender Systems (RS) communities both in Italy and collaborating with Italian research institutions. This year's event spanned two days and featured studies on various topics within IR, RS, and Large Language Models (LLMs). Key focus areas included enhanced retrieval models, personalized information systems, conversational interfaces and user-centric systems, comparative evaluations and metrics, and practical applications in specific fields. IIR 2024 was jointly organized by the University of Udine and the University of Milano-Bicocca and was held in Udine, Italy. Date : September 5--6, 2024. Website : https://iir2024.uniud.it/.
- Research Article
75
- 10.1007/s10462-020-09892-9
- Aug 19, 2020
- Artificial Intelligence Review
With the growth of online information, varying personalization drifts and volatile behaviors of internet users, recommender systems are effective tools for information filtering to overcome the information overload problem. Recommender systems utilize rating prediction approaches i.e. predicting the rating that a user will give to a particular item, to generate ranked lists of items according to the preferences of each user in order to make personalized recommendations. Although previous recommendation systems are effective in creating attired recommendations, however, they still suffer from different types of challenges such as accuracy, scalability, cold-start, and data sparsity. In the last few years, deep learning has attained substantial interest in various research areas such as computer vision, speech recognition, and natural language processing. Deep learning based approaches are vigorous in not only performance improvement but also to feature representations learning from the scratch. The impact of deep learning is also prevalent, recently validating its efficacy on information retrieval and recommender systems research. In this study, a comprehensive review of deep learning-based rating prediction approaches is provided to help out new researchers interested in the subject. More concretely, the classification of deep learning-based recommendation/rating prediction models is provided and articulated along with an extensive summary of the state-of-the-art. Lastly, new trends are exposited with new perspectives pertaining to this novel and exciting development of the field.
- Conference Article
- 10.1145/3170427.3173036
- Apr 20, 2018
Due to the enormous amount of information being carried over online systems today, no user can access all such information. Therefore, to help the users, all major online organizations deploy information retrieval (content recommendation, search or ranking) systems to find important information. Current information retrieval systems have to make certain design choices. For example, news recommendation systems need to decide on the quality of recommended news stories, how much emphasis to give to a story's long-term importance over its recency or freshness etc. Similarly, recommendation systems over user generated contents (e.g., in social media like Facebook and Twitter) need to take into account the content posted by heterogeneous user groups. However, such design choices can introduce unintended biases in the contents presented to the users. For example, the recommended contents may have poor quality or less news value, or the news discourse may get hijacked by hyper-active demographic groups. In this thesis, we want to systematically measure the effect of such design choices in the content recommendation systems, and build alternate recommendation systems that mitigate the biases in the recommendation output.
- Conference Article
- 10.1145/3726302.3730369
- Jul 13, 2025
We present GENNEXT, a workshop dedicated to exploring the integration of language agents, generative models, and conversational AI within information retrieval (IR) and recommender systems (RS). Building on the success of our recent RecSys'24 workshop, GENNEXT aims to advance discussions on the applications of language agents powered by Large Language Models (LLMs). The workshop will focus on enhancing interactivity between users and systems through multi-turn dialogues, improving creative content generation, advancing personalization, and enabling multifaceted, context-aware decision-making. For example, a language agent could respond to a query like ''Suggest an eco-friendly food tour for a weekend in my city'' by using a recommendation API to identify eateries specializing in sustainable or organic cuisine and a pollution API to ensure the selected routes have low air pollution levels.