Research theme

Responsible Information Access

When a ranking model learns from human data, whose assumptions does it carry forward — and can we take them back out?

Retrieval systems are trained on text people wrote, and they inherit the biases in it. A dense retriever asked a neutral query can return a systematically gendered picture of who does a job; a query reformulation model can quietly narrow what a user is allowed to find.

We work on making those effects measurable and then correctable. That means building the datasets and metrics that let bias be observed in the first place, and designing training-time interventions — regularization objectives, fair reformulation strategies — that reduce it without destroying the retrieval effectiveness the system exists for.

The work extends past the model itself. We study how AI ethics is actually taught and assessed, and how retrieval systems shape public narratives about groups of people who have little say in how they are represented.

Representative work

  • A Regularization Framework for Gender Bias Mitigation in Dense Neural Rankers — Machine Learning Journal, 2026
  • A Large-Scale Dataset for Gender-Fair Query Reformulations — SIGIR 2026
  • AI ethics education: A scoping review of pedagogy, curriculum, and assessment — IP&M, 2026

Publications in this theme

27
2026 Conference

A Benchmark & Dataset for Detecting AI-Manipulated Visual Evidence in the Court System

Kelly McConvey, Sajad Ebrahimi, Nima Jamali, Jalehsadat Mahdavimoghaddam, Matina Mahdizadeh Sani, Maksym Taranukhin, Wentao Zhang, Jaquelyn Burkell, Yuntian Deng, Karen Eltis, Maura Grossman, Vered Shwartz, Ebrahim Bagheri

CIKM 2026 ACM International Conference on Information and Knowledge Management (CIKM 2026)

Abstract

Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not well equipped to evaluate. Surveillance frames, dashcam stills, and phone photographs may be used to establish presence, sequence, causation, damage, or identity, yet contemporary generative systems allow non-experts to alter or fabricate such images through ordinary prompt-based interfaces. Existing image-forensics benchmarks provide important resources for face manipulation, classical tampering, and general synthetic-image detection, but they are not organized around the forms of visual evidence submitted in courts, the localized edits that can change what an exhibit appears to prove, or the consumer-tool threat model now facing the justice system. We introduce the CIFAR Synthetic Evidence Corpus for Detecting AI-Manipulated Images, a benchmark for evidentiary image authentication in court and justice-system contexts. The corpus contains 1,505 photographic items, including 720 authentic controls and 785 manipulated or fabricated images, spanning surveillance, dashcam, and consumer-photo imagery. Manipulations are organized into scene-condition edits, localized element edits, and full fabrications produced with contemporary generative systems. Each item is released with structured metadata covering source provenance, manipulation tier, subtype, generator, prompt template, and scene attributes, enabling controlled evaluation beyond aggregate binary detection. We also establish state-of-the-art baselines with publicly available image-manipulation detectors, showing that current systems exhibit error profiles that remain problematic for evidentiary use. The dataset, prompts, metadata manifest, code, and baseline evaluation scripts are released to support research on visual evidence authentication, information integrity, and trustworthy AI for the justice system.

2026 Conference

A Large-Scale Dataset for Gender-Fair Query Reformulations

Hai Son Le, Shirin Seyedsalehi, Morteza Zihayat, Ebrahim Bagheri

SIGIR 2026 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2026)

Abstract

Information Retrieval systems can amplify societal inequalities when training data and ranking algorithms encode biases towards certain gender identities. Query reformulation methods are known to improve retrieval effectiveness, yet these methods seldom consider retrieval fairness as one of their criteria for succesfull retrieval. In this paper, we introduce Refairmulate, an open resource for gender fair query reformulation that uses a multi-objective procedure to balance retrieval effectiveness and gender bias. The offered resource contains three dataset subsets. The Optimal subset consists of 112,261 query pairs including the original query and its reformulated variant where the reformulated variant enjoys a perfect RR@10 equal to 1 and a complete reduction of gender bias. The Effective subset includes 209,343 query pairs where the reformulated variant is guaranteed to have both better retrieval effectiveness and less gender bias compared to the original query. The Fair subset contains 321,604 query pairs but only guarantees that the reformulated variant enjoys a less gender bias regardless of its retrieval effectiveness. We benchmark Refairmulate with BM25 and several dense retrievers, including SPLADE, SBERT, TCT ColBERT, and ANCE, all of which show consistent gains, with up to 76.0% relative improvement in MRR@10 and up to 48.5% reduction in gender bias. To our knowledge, Refairmulate is the first large scale benchmark designed for gender fairness aware query reformulation, and it enables reproducible evaluationand model development.

2026 Journal

A Regularization Framework for Gender Bias Mitigation in Dense Neural Rankers

Shirin SeyedSalehi, Morteza Zihayat, Ebrahim Bagheri

MLJ Machine Learning Journal

Abstract

Dense neural retrievers have improved retrieval effectiveness but can also amplify social biases in ranked results. This paper investigates gender bias in retrieval systems and introduces a fairness-aware training approach that regularizes standard ranking losses with bias and fairness terms. The formulation applies penalty or reward signals at the document level within pairwise objectives, enabling a tunable trade-off between effectiveness and fairness. We evaluate the approach on MS MARCO-derived benchmarks using two encoders (BERT-mini and ELECTRA-small) and two query sets (gender-neutral and socially sensitive). Across ARaB, LIWC, and NFaiRR, our best configurations substantially reduce gender bias while preserving MRR@10 within small to moderate deltas, and in some cases improving effectiveness. We also compare against fairness-aware baselines such as adversarial and neutrality-regularized rankers and find competitive or superior bias reduction under comparable effectiveness. The findings are empirical and scoped to binary gender bias in English on the evaluated datasets and models, without claims of broader generality.

2026 Journal

AI ethics education: A scoping review of pedagogy, curriculum, and assessment

Calvin Hillis, Maushumi Bhattacharjee, Batool AlMousawi, Riley Martens, Tarik Eltanahy, Sara Ono, Ba’ Pham Marcus Hui, Michelle Swab, Gordon V. Cormack, Maura R. Grossman, Ebrahim Bagheri, Zack Marshall

IP&M Information Processing and Management

Abstract

Background: Artificial intelligence (AI) is increasingly embedded in social and institutional decision-making, creating a growing demand for ethically literate practitioners. Universities have responded by introducing AI ethics instruction; however, the structure, content, pedagogy, and evaluation of these efforts remain unevenly documented. Objective: This study aims to map and synthesize research on university-level AI ethics education by characterizing course design, pedagogy, ethical themes, and assessment methods, while also identifying evidence gaps that limit knowledge consolidation and instructional refinement. Methods: We conducted a scoping review using Continuous Active Learning to screen 50,766 records published up to 2024. A total of 43 studies met the inclusion criteria following title, abstract, and full-text review. We coded instructional design, curricular themes, pedagogical methods, and evaluation approaches using descriptive frequency counts and qualitative synthesis. Results: Most included studies were conceptual or descriptive, with relatively few empirical evaluations. Instruction was concentrated in computing and engineering disciplines and primarily targeted undergraduate learners. Ethics content was more often embedded within technical courses rather than offered as standalone courses. Reported pedagogical approaches relied heavily on lectures and case-based discussions, with fewer studies describing participatory formats such as simulations or role-play. Curricular emphasis was concentrated on bias, fairness, and privacy, with comparatively less attention given to governance, explainability, and trust. Evaluation methods most commonly relied on self-report and reflective approaches, while validated instruments and performance-based assessments were less common, and behavioral or applied outcomes were rarely assessed. Conclusions: The literature suggests that the field is oriented more toward awareness-building than the development of measurable ethical competence. Clearer competency definitions, stronger transparency in assessment, and better alignment between instructional design and evaluation would improve comparability across studies and support evidence-informed course development.

2026 Conference

How Information Retrieval Systems Construct and Amplify Immigration Narratives

Zarif Masud, Abhijit Paul, Syed Ishtiaque Ahmed, Ebrahim Bagheri

ECIR 2026 48th European Conference on Information Retrieval (ECIR 2026)

Abstract

Information retrieval systems play a central role in how peo- ple access and understand information about complex social issues, in- cluding immigration. Yet little is known about how the datasets that underpin these systems represent migrants or structure public narratives about migration. In this paper, we investigate how immigration is framed within a widely used IR benchmark and how ranking models shape the visibility of those frames. Using MS MARCO as our data source, we cu- rate immigration-related queries and annotate retrieved passages using a migration-specific framing taxonomy grounded in social-science research. Our goal is to identify which narratives dominate and to measure how different retrieval models influence their exposure. We find that legality and security frames are far more common than humanitarian or inclu- sive ones, and that neural reranking amplifies exclusionary portrayals compared to sparse retrieval.

2026 Conference

QueryGym: A Toolkit for Reproducible LLM-Based Query Reformulation

Amin Bigdeli, Radin Hamidi Rad, Mert Incesu, Negar Arabzadeh, Charles L. A. Clarke, Ebrahim Bagheri

ACM The Web Conference (TheWebConf 2026)

Abstract

We present QueryGym, a lightweight, extensible Python toolkit that supports large language model (LLM)-based query reformulation. This is an important tool development since recent work on llm-based query reformulation has shown notable increase in retrieval effectiveness. However, while different authors have sporadically shared the implementation of their methods, there is no unified toolkit that provides a consistent implementation of such methods, which hinders fair comparison, rapid experimentation, consistent benchmarking and reliable deployment. QueryGym addresses this gap by providing a unified framework for implementing, executing, and comparing llm-based reformulation methods. The toolkit offers: (1) a Python API for applying diverse LLM-based methods, (2) a retrieval-agnostic interface supporting integration with backends such as Pyserini and PyTerrier, (3) a centralized prompt management system with versioning and metadata tracking, (4) built-in support for benchmarks like BEIR and MS MARCO, and (5) a completely open-source extensible implementation available to all researchers. QueryGym is publicly available at https://github.com/radinhamidi/QueryGym.

2026 Conference

Reliable Evaluation of AI Assisted Peer Review

Negar Arabzadeh, Sajad Ebrahimi, Shakiba Amirshahi, Seyed Mohammad Hosseini, Hai Son Le, Mahdi Bashari, Ebrahim Bagheri

ACM International Conference on Information and Knowledge Management (CIKM 2026)

Abstract

Large language models (LLMs) are rapidly moving from experimental prototypes into operational peer review workflows, where they are used to polish, rewrite, expand, synthesize, screen, and in some cases generate reviews. This shift creates an urgent challenge for the organizations that depend on expert assessment, including publishers, conferences, funding agencies, research institutions, and enterprise research teams. Rather than asking only whether AI can improve peer review, this talk argues that we must first confront a more fundamental evaluation question. Can we reliably measure what makes a review useful, fair, accurate, and decision-relevant? The talk examines this question through the lens of AI-assisted peer review and shows that commonly used signals, including recommendation agreement, similarity to reference reviews, writing fluency, and single score LLM-as-a-judge evaluation, often fail to capture substantive review quality. Drawing on real world deployment experience at Reviewerly, we present case studies in which evaluation metrics reward surface-level improvements such as polish, organization, and LLM like wording while overlooking the evidentiary content, specificity, correctness, and editorial usefulness of the review. We then discuss practical lessons for building reliable evaluation systems in real peer review workflows, including stress testing review metrics, separating surface sensitivity from practical robustness, comparing judge models under realistic perturbations, and designing layered evaluation pipelines that are interpretable, evidence-grounded, and aligned with editorial needs. Although peer review is the central case study, the broader message applies to many industry and organizational settings where AI is being introduced into expert judgment workflows. Responsible automation requires more than stronger generation models. It requires evaluation methods that measure the qualities we actually want AI systems to improve.

2026 Conference

Self-Paced Fair Ranking with Loss as a Proxy for Bias

Shirin Seyedsalehi, Hai Son Le, Morteza Zihayat, Ebrahim Bagheri

International Conference on Web Search and Data Mining (WSDM 2026)

Abstract

Neural rankers often mirror societal biases in training data, such as gender. Prior work typically requires access to protected-attribute labels or model changes, limiting applicability. We propose a simple, model-agnostic approach that uses the model’s own loss values as a proxy for bias. Through self-paced learning, our method first prioritizes lower-loss (and less biased) examples, then gradually introduces harder (more biased) ones. This loss-aware curriculum reduces reliance on biased samples without demographic annota- tions. We prove the gender loss gap decreases monotonically during training and show on MS MARCO that our method reduces bias while maintaining or improving ranking effectiveness, outperform- ing strong baselines.

2025 Conference

Bias-Aware Curriculum Sampling For Fair Ranking

Shirin Seyedsalehi, Hai Son Le, Morteza Zihayat, Ebrahim Bagheri

SIGIR 2025 The 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2025)

Abstract

No abstract available for this paper.

2025 Journal

Exploring the methodological quality and risk of bias in 200 systematic reviews: A comparative study of ROBIS and AMSTAR-2 tools,

Carole Lunny, Nityanand Jain, Tina Nazari, Melodi Kosaner-Kließ, Lucas Santos, Ian Goodman, Alaa AM Osman, Stefano Berrone, Mohammad Dadam, Connor Brenna, Heba Hussein, Gioia Dahdal, Diana Cespedes A., Nicola Ferri, Salmaan Kanji, Yuan Chi, Dawid Pieper, Beverly Shea, Amanda Parker, Dipika Neupane, Paul Khan, Daniella Rangira, Kat Kolaski, Ben Ridley, Amina Berour, Kevin Sun, Radin Hamidi Rad, Zihui Ouyang, Emma K Reid, Iván Pérez-Neri, Sanabel Barakat, Silvia Bargeri, Silvia Gianola, Greta Castellini, Sera Whitelaw, Adrienne Stevens, Shailesh Kolekar, E. Bagheri, Kristy Wong, Paityn Major, Andrea Tricco

Research Synthesis Methods

Abstract

Background: AMSTAR-2 and ROBIS are tools used to assess the methodological quality and the risk of bias in systematic review (SRs). Methods: We applied AMSTAR-2 and ROBIS to a sample of 200 published SRs. We investigated the overlap in their methodological constructs, responses by item and overall, percentage agreement, direction of effect, and timing of assessments. Results: AMSTAR-2 contains 16 items and ROBIS contains 24 items. Three items in AMSTAR-2 and nine items in ROBIS did not overlap in construct. Of the 200 SRs, 73% were low or critically low quality using AMSTAR-2 and 81% had high risk of bias using ROBIS. The median time to complete AMSTAR-2 and ROBIS was 51 and 64 minutes, respectively. When assessment times were calibrated to the number of items in each tool, each item took on average of 3.2 minutes for AMSTAR-2 compared to 2.7 minutes for ROBIS. Nine percent of SRs had opposed ratings (i.e., AMSTAR-2 was high quality while ROBIS was high risk). In both tools, three-quarters of items showed more than 70% agreement between senior and junior raters after extensive training and piloting. Conclusions: The tools are not exchangeable due to their unique items and differences in underlying concepts. While AMSTAR-2 only considers the methodological quality of the results, ROBIS considers the bias in the results and conclusions. Additionally, ROBIS invites reviewers to assess the external validity, which is absent from AMSTAR-2. AMSTAR-2 may be more appropriate where faster assessments are prioritised. ROBIS may be more appropriate when a comprehensive bias assessment is sought.

2025 Journal

Partisan Perspectives on AI and Immigration: An Analysis of Canadian Parliamentary Discourse (2014–2024)

Eleyan Sawafta, Zarif Masud, Yasmeen Abu-Laban, Syed Ishtiaque Ahmed, Ebrahim Bagheri, Geoffrey Rockwell

Canadian Ethnic Studies

Abstract

In 2017, Canada became the first country to advance a national artificial intelligence (AI) strategy. In a now global race to adopt this technology, various stakeholders, including researchers and civil society organizations, have questioned its fairness in immigration, highlighting colonial legacies disguised as efficiency. These debates are particularly intense in the context of immigration decision-making due to the ethical and human rights implications, such as life-altering consequences for migrants and refugees. This article critically examines partisan discourse on AI and the immigration decisions on visa applications in Canada by analyzing parliamentary discussions. Specifically, extracting House of Commons Debates (Hansard) between 2014 and 2024, we analyze this data using Critical Discourse Analysis with the help of computational text analysis tools, including Voyant. We argue that while partisan differences exist in how AI is framed in Canada’s legislative debates on immigration, these debates overall fail to systematically engage with the full range of ethical and human rights implications, and continuing colonial legacies, embedded in the concerns raised by many researchers and stakeholders about immigration and artificial intelligence. Beyond the main argument, we observe that although some parties, like the Conservative Party, engaged in criticism, their interventions resembled protest more than a policy response of a “government-in-waiting” prepared to govern.

2025 Workshop

Responsible AI Day

Ebrahim Bagheri, Faezeh Ensan, Calvin Hillis, Reihaneh Rabbany, Robin Cohen, Benjamin C. M. Fung, Sébastien Gambs

KDD 2025 ACM SIGKDD Conference on Knowledge Discovery and Data Mining

Abstract

No abstract available for this paper.

2025 Journal

The Role of Protocol Papers, Scoping Reviews, and Systematic Reviews in Responsible AI Research

Calvin Hillis, Ebrahim Bagheri, Zack Marshall

IEEE Technology and Society Magazine

Abstract

Research protocol papers, scoping reviews, and systematic reviews are important tools for ensuring transparency, methodological rigor, accountability, and reproducibility in scholarly research. While widely adopted in disciplines such as medicine and social sciences, these practices remain underutilized in computer science and engineering, particularly within Responsible AI research. This paper explores the critical role that scoping and systematic reviews play in synthesizing knowledge, identifying gaps, and guiding future research in Responsible AI. It also examines how research protocol papers—especially those developed for structured reviews—can improve the credibility, clarity, and replicability of Responsible AI studies. By outlining the benefits and challenges of these methods, we argue for their greater adoption in AI ethics and Responsible AI research. We conclude with recommendations to support the institutional, editorial, and cultural shifts necessary to integrate these rigorous methodological tools into Responsible AI scholarship.

2025 Journal

Understanding and Mitigating Gender Bias in Information Retrieval Systems

Shirin Seyedsalehi, Amin Bigdeli, Negar Arabzadeh, Batool AlMousawi, Zack Marshall, Morteza Zihayat, Ebrahim Bagheri

Foundations and Trends® in Information Retrieval (FnTIR)

Abstract

Gender bias is a pervasive issue that continues to influence various aspects of society, including the outcomes of information retrieval (IR) systems. As these systems become increasingly integral to accessing and navigating the vast amounts of information available today, the need to understand and mitigate gender bias within them is paramount. This book provides a comprehensive examination of the origins, manifestations, and consequences of gender bias in IR systems, as well as the current methodologies employed to address these biases. Theoretical frameworks surrounding gender and its representation in artificial intelligence (AI) systems are explored, particularly focusing on how traditional gender binaries are perpetuated and reinforced through data and algorithmic processes. Metrics and methodologies used to identify and measure gender bias within IR systems are then analyzed, offering a detailed evaluation of existing approaches and their limitations. Subsequent chapters address the sources of gender bias, including biased input queries, retrieval methods, and gold standard datasets. Various data-driven and method-level debiasing strategies are presented, including techniques for debiasing neural embeddings and algorithmic approaches aimed at reducing bias in IR system outputs. The book concludes with a discussion of the challenges and limitations faced by current debiasing efforts and provides insights into future research directions that could lead to more equitable and inclusive IR systems. This book serves as a valuable resource for researchers, practitioners, and students in the fields of information retrieval, artificial intelligence, and data science, providing the knowledge and tools needed to address gender bias and contribute to the development of fair and unbiased information systems.

2024 Journal

Gender Disentangled Representation Learning in Neural Rankers

Shirin Seyedsalehi, Sara Salamat, Negar Arabzadeh, Sajad Ebrahimi, Morteza Zihayat, Ebrahim Bagheri

Machine Learning Journal

Abstract

Recent studies have demonstrated that while neural ranking methods excel in retrieval effectiveness, they also tend to amplify stereotypical biases, especially those related to gender. Current mitigation strategies often focus on adjusting training methods, like adversarial techniques or data balancing, but typically overlook explicit consideration of gender as an attribute. In this paper, we introduce a systematic approach that treats gender as a distinct component within neural ranker representations. Our neural disentanglement method separates content semantics from gender information, enabling the neural ranker to evaluate document relevance based on content alone, without the interference of gender-related information during retrieval. Our extensive experiments demonstrate that: (1) our disentanglement approach matches the effectiveness of baseline models and offers more consistent performance across queries of different gender affiliations; (2) isolating gender within the representations allows the neural ranker to produce an unbiased list of documents, not favoring any specific gender; and (3) the disentangled gender component effectively and concisely captures gender information independently from the semantic content.

2023 Conference

De-Biasing Relevance Judgements for Fair Ranking

Amin Bigdeli, Negar Arabzadeh, Shirin Seyedsalehi, Bhaskar Mitra, Morteza Zihayat, Ebrahim Bagheri

ECIR 2023 The 45th European Conference on Information Retrieval (ECIR 2023)

DOI
Abstract

The objective of this paper is to show that it is possible to significantly reduce stereotypical gender biases in neural rankers without modifying the ranking loss function, which is the current approach in the literature. We systematically de-bias gold standard relevance judgement datasets with a set of balanced and well-matched query pairs. Such a de-biasing process will expose neural rankers to comparable queries from across gender identities that have associated relevant documents with compatible degrees of gender bias. Therefore, neural rankers will learn not to associate varying degrees of bias to queries from certain gender identities. Our experiments show that our approach is able to (1) systematically reduces gender biases associated with different gender identities, and (2) at the same time maintain the same level of retrieval effectiveness.

2022 Conference

A Light-weight Strategy for Restraining Gender Biases in Neural Rankers

Amin Bigdeli, Negar Arabzadeh, Shirin SeyedSalehi, Morteza Zihayat, Ebrahim Bagheri

ECIR 2022 44th European Conference on IR Research (ECIR 2022)

DOI
Abstract

In light of recent studies that show neural retrieval methods may intensify gender biases during retrieval, the objective of this paper is to propose a simple yet effective sampling strategy for training neural rankers that would allow the rankers to maintain their retrieval effectiveness while reducing gender biases. Our work proposes to consider the degrees of gender bias when sampling documents to be used for training neural rankers. We report our findings on the MS MARCO collection and based on different query datasets released for this purpose in the literature. Our results show that the proposed light-weight strategy is able to show competitive (or even better) performance compared to the state of the art neural architectures specifically designed to reduce gender biases.

2022 Conference

Addressing Gender-related Performance Disparities in Neural Rankers

Shirin Seyedsalehi, Negar Arabzadeh, Amin Bigdeli, Morteza Zihayat, Ebrahim Bagheri

SIGIR 2022 The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2022)

DOI
Abstract

While neural rankers continue to show notable performance improvements over a wide variety of information retrieval tasks, there have been recent studies that show such rankers may intensify certain stereotypical biases. In this paper, we investigate whether neural rankers introduce retrieval effectiveness (performance) disparities over queries related to different genders. We specifically study whether there are significant performance differences between male and female queries when retrieved by neural rankers. Through our empirical study over the MS MARCO collection, we find that such performance disparities are notable and that the performance disparities may be due to the difference between how queries and their relevant judgements are collected and distributed for different gendered queries. More specifically, we observe that male queries are more closely associated with their relevant documents compared to female queries and hence neural rankers are able to more easily learn associations between male queries and their relevant documents. We show that it is possible to systematically balance relevance judgment collections in order to reduce performance disparity between different gendered queries without negatively compromising overall model performance.

2022 Conference

Bias-aware Fair Neural Ranking for Addressing Stereotypical Gender Biases

Shirin Seyedsalehi, Amin Bigdeli, Negar Arabzadeh, Bhaskar Mitra, Morteza Zihayat, Ebrahim Bagheri

EDBT 2022 EDBT/ICDT 2022

DOI
Abstract

Research has shown that neural rankers can pick up and intensify gender biases. The expression of stereotypical gender biases in retrieval systems can lead to their reinforcement in users' beliefs. As such, the objective of this paper is to propose a bias-aware fair ranker that explicitly incorporates a notion of gender bias and hence controls how bias is expressed in documents that are retrieved. The proposed approach is designed such that it learns the notion of relevance between the document and the query from the relevant sampled documents while incorporating the notion of gender bias by penalizing irrelevant biased sampled documents. We show that unlike the state of the art, our approach reduces bias while maintaining retrieval effectiveness over different query sets.

2022 Conference

Gender Fairness in Information Retrieval Systems

Amin Bigdeli, Negar Arabzadeh, Shirin Seyedsalehi, Morteza Zihayat, Ebrahim Bagheri

SIGIR 2022 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2022) (tutorial)

DOI
Abstract

Recent studies have shown that it is possible for stereotypical gender biases to find their way into representational and algorithmic aspects of retrieval methods; hence, exhibit themselves in retrieval outcomes. In this tutorial, we inform the audience of various studies that have systematically reported the presence of stereotypical gender biases in Information Retrieval (IR) systems. We further classify existing work on gender biases in IR systems as being related to (1) relevance judgement datasets, (2) structure of retrieval methods, and (3) representations learnt for queries and documents. We present how each of these components can be impacted by or cause intensified biases during retrieval. Based on these identified issues, we then present a collection of approaches from the literature that have discussed how such biases can be measured, controlled, or mitigated. Additionally, we introduce publicly available datasets that are often used for investigating gender biases in IR systems as well as evaluation methodology adopted for determining the utility of gender bias mitigation strategies.

2022 Conference

On the Characteristics of Ranking-based Gender Bias Measures

Anja Klasnja, Negar Arabzadeh, Mahbod Mehrvarz, Ebrahim Bagheri

WebSci 2022 WebSci'22

DOI
Abstract

With increased recent awareness on the possible impact of retrieval techniques on intensifying gender biases, researchers have embarked on defining quantifiable gender bias metrics that can provide the means to concretely measure such biases in practice. While successful in allowing for identifying possible sources of gender bias, there has been little work that systematically explores the characteristics of these metrics. This paper argues that effective future works on gender biases in information retrieval require a careful understanding of the bias metrics in terms of their consistency, robustness, sensitivity and also their relation with psychological characteristics and what they actually measure. Through our experiments, we show that more rigorous work on gender bias metrics need to be pursued as existing metrics may not necessarily be consistent and robust and often capture differing psychological characteristics.

2021 Conference

Exploring Gender Biases in Information Retrieval Relevance Judgement Datasets

Amin Bigdeli, Negar Arabzadeh, Morteza Zihayat, Ebrahim Bagheri

ECIR 2021 43rd European Conference on IR Research (ECIR 2021)

DOI
Abstract

Recent studies in information retrieval have shown that gender biases have found their way into representational and algorithmic aspects of computational models. In this paper, we focus specifically on gender biases in information retrieval gold standard datasets, often referred to as relevance judgements. While not explored in the past, we submit that it is important to understand and measure the extent to which gender biases may be present in information retrieval relevance judgements primarily because relevance judgements are not only the primary source for evaluating IR techniques but are also widely used for training end-to-end neural ranking methods. As such, the presence of bias in relevance judgements would immediately find its way into how retrieval methods operate in practice. Based on a fine-tuned BERT model, we show how queries can be labeled for gender at scale based on which we label MS MARCO queries. We then show how different psychological characteristics are exhibited within documents associated with gendered queries within the relevance judgement datasets. Our observations show that stereotypical biases are prevalent in relevance judgement documents.

2021 Conference

On the Orthogonality of Bias and Effectiveness in Ad hoc Retrieval

Amin Bigdeli, Negar Arabzadeh, Shirin Seyedsalehi, Morteza Zihayat, Ebrahim Bagheri

The 44th International ACM SIGIR Conference. on Research and Development in Information Retrieval (SIGIR 2021)

DOI
Abstract

Various researchers have recently explored the impact of different types of biases on information retrieval tasks such as ad hoc retrieval and question answering. While the impact of bias needs to be controlled in order to avoid increased prejudices, the literature has often viewed the relationship between increased retrieval utility (effectiveness) and reduced bias as a tradeoff where one can suffer from the other. In this paper, we empirically study this tradeoff and explore whether it would be possible to reduce bias while maintaining similar retrieval utility. We show this would be possible by revising the input query through a bias-aware pseudo-relevance feedback framework. We report our findings based on four widely used TREC corpora namely Robust04, Gov2, ClueWeb09 and ClueWeb12 and using two classes of bias metrics. The findings of this paper are significant as they are among the first to show that decrease in bias does not necessarily need to come at the cost of reduced utility.

2021 Conference

On the Orthogonality of Bias and Utility in Ad hoc Retrieval

Amin Bigdeli, Negar Arabzadeh, Shirin Seyedsalehi, Morteza Zihayat, Ebrahim Bagheri

SIGIR 2021 International ACM SIGIR Conference on Research and Development in Information Retrieval

Abstract

No abstract available for this paper.

2018 Journal

Online social network response to studies on antidepressant use in pregnancy

Simone N Vigod, Ebrahim Bagheri, Fattane Zarrinkalam, Hillary K Brown, Muhammad Mamdani, Joel G Ray

Journal of Psychosomatic Research

DOI
Abstract

Background: About 8% of U.S women are prescribed antidepressant medications around the time of pregnancy. Decisions about medication use in pregnancy can be swayed by the opinion of family, friends and online media, sometimes beyond the advice offered by healthcare providers. Exploration of the online social network response to research on antidepressant use in pregnancy could provide insight about how to optimize decision-making in this complex area. Methods: For all 17 research articles published on the safety of antidepressant use in pregnancy in 2012, we sought to explore online social network activity regarding antidepressant use in pregnancy, via Twitter, in the 48 hours after a study was published, compared to the social network activity in the same period 1 week prior to each article’s publication. Results: Online social network activity about antidepressants in pregnancy quickly doubled upon study publication. The increased activity was driven by studies demonstrating harm associated with antidepressants, by lower-quality studies, and studies where abstracts presented relative versus absolute risks. Implications: These findings support a call for leadership from medical journals to consider how to best incentivize and support a balanced and clear translation of knowledge around antidepressant safety in pregnancy to their readership and the public.

2010 Book section

A Framework for the Manifestation of Tacit Critical Infrastructure Knowledge

Ebrahim Bagheri, Ali A Ghorbani

Sustainable and Resilient Critical Infrastructure Systems

Abstract

Critical infrastructure systems are tightly-coupled socio-technical systems with complicated behavior. They have emerged as an important focal point of research due to both their vital role in the normal conduct of societal activities as well as their inherent appealing complications for researchers. In this chapter, we will report on our experience in developing techniques, tools and algorithms for revealing and interpreting the hidden intricacies of such systems. The chapter will include the description of several of our technologies that allow for the guided understanding of the current status quo of infrastructure systems through the Astrolabe methodology, the formal profiling of infrastructure systems using the UML-CI meta-modeling mechanism, and also observing the emergent behavior of these complex systems through the application of the agent-based AIMS simulation suite.