Faculty of Information · University of Toronto

Information systems that are effective, robust, and fair.

LS3 is a research laboratory studying how search, ranking and language models actually behave — where they break, whose assumptions they carry, and what they do to the people who depend on them.

Two LS3 researchers working at a desk with laptops, data dashboards on the wall behind them and a server rack alongside.

What we work on

All research →

Recent publications

All 329 →
2027 Conference

Emotional and Informational Trajectories of Immigrants: A Longitudinal Study of Reddit Communities

Zarif Masud, Abhijit Paul, Naimul Khan, Syed Ishtiaque Ahmed, Ebrahim Bagheri

ICWSM 2027 21st International Conference on Web and Social Media (ICWSM 2027)

Abstract

Migration is often studied through surveys, administrative records, or clinical data, yet these sources offer limited visibility into how settlement unfolds in everyday life before and after arrival. We introduce a longitudinal analysis of self-reported immigrants to Canada on Reddit, treating migration as a temporally anchored life-course transition rather than a single event. Using keyword-based retrieval and an LLM-assisted temporal inference pipeline, we reconstruct pre-arrival, arrival-year, and post-arrival timelines for 9,834 immigrant Reddit users who arrived between 2014 and 2024, comprising 4,861,954 posts and comments. Combining LIWC-based psycholinguistic analysis with embedding-based topic clustering, we find that overt affective change around arrival is modest: contrary to expectations from migration mental-health literature, negative emotion barely shifts, while positive emotion declines reliably from arrival to post-arrival with a small effect. The stronger linguistic signal is a reorientation of self-positioning: after arrival, newcomers write less about themselves and more about families, communities, and external institutions. Topical attention, in contrast, shifts sharply and structurally. Around arrival, leisure, technology, and media-oriented discussion give way to settlement geography, immigration procedures, jobs, and housing; after arrival, acute procedural concerns recede while infrastructural ones such as cost of living, transport, and finance remain elevated. These findings show that migration reshapes Reddit discourse less through dramatic emotional change than through a durable reallocation of attention toward the relational, institutional, and material work of settlement. Methodologically, our study offers a scalable framework for temporally anchoring social media data around life-course events; substantively, it reveals when newcomer information needs emerge, with implications for the timing of settlement support.

2026 Conference

A Benchmark & Dataset for Detecting AI-Manipulated Visual Evidence in the Court System

Kelly McConvey, Sajad Ebrahimi, Nima Jamali, Jalehsadat Mahdavimoghaddam, Matina Mahdizadeh Sani, Maksym Taranukhin, Wentao Zhang, Jaquelyn Burkell, Yuntian Deng, Karen Eltis, Maura Grossman, Vered Shwartz, Ebrahim Bagheri

CIKM 2026 ACM International Conference on Information and Knowledge Management (CIKM 2026)

Abstract

Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not well equipped to evaluate. Surveillance frames, dashcam stills, and phone photographs may be used to establish presence, sequence, causation, damage, or identity, yet contemporary generative systems allow non-experts to alter or fabricate such images through ordinary prompt-based interfaces. Existing image-forensics benchmarks provide important resources for face manipulation, classical tampering, and general synthetic-image detection, but they are not organized around the forms of visual evidence submitted in courts, the localized edits that can change what an exhibit appears to prove, or the consumer-tool threat model now facing the justice system. We introduce the CIFAR Synthetic Evidence Corpus for Detecting AI-Manipulated Images, a benchmark for evidentiary image authentication in court and justice-system contexts. The corpus contains 1,505 photographic items, including 720 authentic controls and 785 manipulated or fabricated images, spanning surveillance, dashcam, and consumer-photo imagery. Manipulations are organized into scene-condition edits, localized element edits, and full fabrications produced with contemporary generative systems. Each item is released with structured metadata covering source provenance, manipulation tier, subtype, generator, prompt template, and scene attributes, enabling controlled evaluation beyond aggregate binary detection. We also establish state-of-the-art baselines with publicly available image-manipulation detectors, showing that current systems exhibit error profiles that remain problematic for evidentiary use. The dataset, prompts, metadata manifest, code, and baseline evaluation scripts are released to support research on visual evidence authentication, information integrity, and trustworthy AI for the justice system.

2026 Conference

A Large-Scale Dataset for Gender-Fair Query Reformulations

Hai Son Le, Shirin Seyedsalehi, Morteza Zihayat, Ebrahim Bagheri

SIGIR 2026 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2026)

Abstract

Information Retrieval systems can amplify societal inequalities when training data and ranking algorithms encode biases towards certain gender identities. Query reformulation methods are known to improve retrieval effectiveness, yet these methods seldom consider retrieval fairness as one of their criteria for succesfull retrieval. In this paper, we introduce Refairmulate, an open resource for gender fair query reformulation that uses a multi-objective procedure to balance retrieval effectiveness and gender bias. The offered resource contains three dataset subsets. The Optimal subset consists of 112,261 query pairs including the original query and its reformulated variant where the reformulated variant enjoys a perfect RR@10 equal to 1 and a complete reduction of gender bias. The Effective subset includes 209,343 query pairs where the reformulated variant is guaranteed to have both better retrieval effectiveness and less gender bias compared to the original query. The Fair subset contains 321,604 query pairs but only guarantees that the reformulated variant enjoys a less gender bias regardless of its retrieval effectiveness. We benchmark Refairmulate with BM25 and several dense retrievers, including SPLADE, SBERT, TCT ColBERT, and ANCE, all of which show consistent gains, with up to 76.0% relative improvement in MRR@10 and up to 48.5% reduction in gender bias. To our knowledge, Refairmulate is the first large scale benchmark designed for gender fairness aware query reformulation, and it enables reproducible evaluationand model development.

2026 Journal

A Regularization Framework for Gender Bias Mitigation in Dense Neural Rankers

Shirin SeyedSalehi, Morteza Zihayat, Ebrahim Bagheri

MLJ Machine Learning Journal

Abstract

Dense neural retrievers have improved retrieval effectiveness but can also amplify social biases in ranked results. This paper investigates gender bias in retrieval systems and introduces a fairness-aware training approach that regularizes standard ranking losses with bias and fairness terms. The formulation applies penalty or reward signals at the document level within pairwise objectives, enabling a tunable trade-off between effectiveness and fairness. We evaluate the approach on MS MARCO-derived benchmarks using two encoders (BERT-mini and ELECTRA-small) and two query sets (gender-neutral and socially sensitive). Across ARaB, LIWC, and NFaiRR, our best configurations substantially reduce gender bias while preserving MRR@10 within small to moderate deltas, and in some cases improving effectiveness. We also compare against fairness-aware baselines such as adversarial and neutrality-regularized rankers and find competitive or superior bias reduction under comparable effectiveness. The findings are empirical and scoped to binary gender bias in English on the evaluated datasets and models, without claims of broader generality.

2026 Conference

A Reproducibility Study of LLM-Based Query Reformulation

Amin Bigdeli, Radin Hamidi Rad, Le Hai Son, Mert Incesu, Negar Arabzadeh, Charles L. A. Clarke, Ebrahim Bagheri

SIGIR 2026 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2026)

Abstract

Large Language Models (LLMs) have recently been widely adopted for query reformulation and expansion in Information Retrieval, with numerous studies reporting substantial effectiveness gains. However, these results are typically obtained under heterogeneous experimental conditions, making it difficult to assess which findings are reproducible and which depend on specific implementation choices. In this work, we present a systematic reproducibility and comparative study of ten representative LLM-based query reformulation methods under a unified and strictly controlled experimental framework. We evaluate methods across two architectural LLM families at two parameter scales, three retrieval paradigms (lexical, learned sparse, and dense), and nine benchmark datasets spanning TREC Deep Learning and BEIR. All decoding parameters, preprocessing pipelines, indexing configurations, and evaluation protocols are held constant to isolate the effect of the reformulation strategy itself. Our results show that reformulation gains are strongly conditioned on the retrieval paradigm, that improvements observed under lexical retrieval do not consistently transfer to neural retrievers, and that larger LLMs do not uniformly yield better downstream performance. These findings clarify the stability and limits of reported gains in prior work and provide a controlled benchmark to support future evaluation and replication efforts. To support transparency and future extensibility, we release all prompts, evaluation scripts, and experimental artifacts.

2026 Journal

A Social Media Lens on the Needs and Concerns of Information Workers

Jalehsadat Mahdavimoghaddam, Koustuv Saha, Julie Hui, Tawanna Dillahunt, Ebrahim Bagheri

TSC ACM Transactions on Social Computing

Abstract

Technological advancements have greatly impacted labor market dynamics, leaving a psychological impact on workers. Although some studies have explored such labor market changes and their effects on workers, they are limited to self-reported data, such as surveys and questionnaires. In this paper, we propose a new approach for identifying information workers’ challenges and their impact on workers’ emotional well-being using large-scale, inexpensive, and near-real-time online social network data. Our research is among the first to utilize statistical methods, machine learning techniques, and natural language analysis to show how labor market-related issues faced by information workers can be modeled and to what extent linguistic differences in online social content can be observed depending on gender and age strata. We find and systematically report highly discussed topics by information workers in areas related to education, skill development, job applications/job search, and employment/job concerns. We show that workers from different gender and age groups disclose their needs in notably different ways. In terms of the implications of our work, we discuss how online social network data can serve as a reliable lens, allowing us to better understand workers’ challenges and well-being in the knowledge economy and how online social platforms can be vehicles for providing effective peer support to improve workers’ well-being