Research theme

Information, People & Society

Information systems are studied as engineering. What do they look like when studied as social infrastructure?

The same techniques that improve a ranking can be turned outward, onto the question of what these systems are doing to people. This strand of the lab’s work is empirical and social: it uses computational methods to study communities, labour, and institutions under the pressure of algorithmic mediation.

Recent studies follow immigrants’ emotional and informational trajectories across years of Reddit activity, examine the needs and concerns information workers voice on social platforms, test whether decentralized autonomous organizations are meaningfully decentralized in their labour, and build benchmarks for detecting AI-manipulated visual evidence entering the court system.

The through-line is a refusal to treat retrieval as a purely technical object. Somebody is always on the other side of the query.

Representative work

  • Emotional and Informational Trajectories of Immigrants: A Longitudinal Study of Reddit Communities — ICWSM 2027
  • A Social Media Lens on the Needs and Concerns of Information Workers — ACM TSC, 2026
  • A Benchmark & Dataset for Detecting AI-Manipulated Visual Evidence in the Court System — CIKM 2026
  • Decentralized in Name Only: The Centralization of DAO Labor — TheWebConf 2026

Publications in this theme

71
2027 Conference

Emotional and Informational Trajectories of Immigrants: A Longitudinal Study of Reddit Communities

Zarif Masud, Abhijit Paul, Naimul Khan, Syed Ishtiaque Ahmed, Ebrahim Bagheri

ICWSM 2027 21st International Conference on Web and Social Media (ICWSM 2027)

Abstract

Migration is often studied through surveys, administrative records, or clinical data, yet these sources offer limited visibility into how settlement unfolds in everyday life before and after arrival. We introduce a longitudinal analysis of self-reported immigrants to Canada on Reddit, treating migration as a temporally anchored life-course transition rather than a single event. Using keyword-based retrieval and an LLM-assisted temporal inference pipeline, we reconstruct pre-arrival, arrival-year, and post-arrival timelines for 9,834 immigrant Reddit users who arrived between 2014 and 2024, comprising 4,861,954 posts and comments. Combining LIWC-based psycholinguistic analysis with embedding-based topic clustering, we find that overt affective change around arrival is modest: contrary to expectations from migration mental-health literature, negative emotion barely shifts, while positive emotion declines reliably from arrival to post-arrival with a small effect. The stronger linguistic signal is a reorientation of self-positioning: after arrival, newcomers write less about themselves and more about families, communities, and external institutions. Topical attention, in contrast, shifts sharply and structurally. Around arrival, leisure, technology, and media-oriented discussion give way to settlement geography, immigration procedures, jobs, and housing; after arrival, acute procedural concerns recede while infrastructural ones such as cost of living, transport, and finance remain elevated. These findings show that migration reshapes Reddit discourse less through dramatic emotional change than through a durable reallocation of attention toward the relational, institutional, and material work of settlement. Methodologically, our study offers a scalable framework for temporally anchoring social media data around life-course events; substantively, it reveals when newcomer information needs emerge, with implications for the timing of settlement support.

2026 Conference

A Benchmark & Dataset for Detecting AI-Manipulated Visual Evidence in the Court System

Kelly McConvey, Sajad Ebrahimi, Nima Jamali, Jalehsadat Mahdavimoghaddam, Matina Mahdizadeh Sani, Maksym Taranukhin, Wentao Zhang, Jaquelyn Burkell, Yuntian Deng, Karen Eltis, Maura Grossman, Vered Shwartz, Ebrahim Bagheri

CIKM 2026 ACM International Conference on Information and Knowledge Management (CIKM 2026)

Abstract

Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not well equipped to evaluate. Surveillance frames, dashcam stills, and phone photographs may be used to establish presence, sequence, causation, damage, or identity, yet contemporary generative systems allow non-experts to alter or fabricate such images through ordinary prompt-based interfaces. Existing image-forensics benchmarks provide important resources for face manipulation, classical tampering, and general synthetic-image detection, but they are not organized around the forms of visual evidence submitted in courts, the localized edits that can change what an exhibit appears to prove, or the consumer-tool threat model now facing the justice system. We introduce the CIFAR Synthetic Evidence Corpus for Detecting AI-Manipulated Images, a benchmark for evidentiary image authentication in court and justice-system contexts. The corpus contains 1,505 photographic items, including 720 authentic controls and 785 manipulated or fabricated images, spanning surveillance, dashcam, and consumer-photo imagery. Manipulations are organized into scene-condition edits, localized element edits, and full fabrications produced with contemporary generative systems. Each item is released with structured metadata covering source provenance, manipulation tier, subtype, generator, prompt template, and scene attributes, enabling controlled evaluation beyond aggregate binary detection. We also establish state-of-the-art baselines with publicly available image-manipulation detectors, showing that current systems exhibit error profiles that remain problematic for evidentiary use. The dataset, prompts, metadata manifest, code, and baseline evaluation scripts are released to support research on visual evidence authentication, information integrity, and trustworthy AI for the justice system.

2026 Journal

A Social Media Lens on the Needs and Concerns of Information Workers

Jalehsadat Mahdavimoghaddam, Koustuv Saha, Julie Hui, Tawanna Dillahunt, Ebrahim Bagheri

TSC ACM Transactions on Social Computing

Abstract

Technological advancements have greatly impacted labor market dynamics, leaving a psychological impact on workers. Although some studies have explored such labor market changes and their effects on workers, they are limited to self-reported data, such as surveys and questionnaires. In this paper, we propose a new approach for identifying information workers’ challenges and their impact on workers’ emotional well-being using large-scale, inexpensive, and near-real-time online social network data. Our research is among the first to utilize statistical methods, machine learning techniques, and natural language analysis to show how labor market-related issues faced by information workers can be modeled and to what extent linguistic differences in online social content can be observed depending on gender and age strata. We find and systematically report highly discussed topics by information workers in areas related to education, skill development, job applications/job search, and employment/job concerns. We show that workers from different gender and age groups disclose their needs in notably different ways. In terms of the implications of our work, we discuss how online social network data can serve as a reliable lens, allowing us to better understand workers’ challenges and well-being in the knowledge economy and how online social platforms can be vehicles for providing effective peer support to improve workers’ well-being

2026 Journal

AI ethics education: A scoping review of pedagogy, curriculum, and assessment

Calvin Hillis, Maushumi Bhattacharjee, Batool AlMousawi, Riley Martens, Tarik Eltanahy, Sara Ono, Ba’ Pham Marcus Hui, Michelle Swab, Gordon V. Cormack, Maura R. Grossman, Ebrahim Bagheri, Zack Marshall

IP&M Information Processing and Management

Abstract

Background: Artificial intelligence (AI) is increasingly embedded in social and institutional decision-making, creating a growing demand for ethically literate practitioners. Universities have responded by introducing AI ethics instruction; however, the structure, content, pedagogy, and evaluation of these efforts remain unevenly documented. Objective: This study aims to map and synthesize research on university-level AI ethics education by characterizing course design, pedagogy, ethical themes, and assessment methods, while also identifying evidence gaps that limit knowledge consolidation and instructional refinement. Methods: We conducted a scoping review using Continuous Active Learning to screen 50,766 records published up to 2024. A total of 43 studies met the inclusion criteria following title, abstract, and full-text review. We coded instructional design, curricular themes, pedagogical methods, and evaluation approaches using descriptive frequency counts and qualitative synthesis. Results: Most included studies were conceptual or descriptive, with relatively few empirical evaluations. Instruction was concentrated in computing and engineering disciplines and primarily targeted undergraduate learners. Ethics content was more often embedded within technical courses rather than offered as standalone courses. Reported pedagogical approaches relied heavily on lectures and case-based discussions, with fewer studies describing participatory formats such as simulations or role-play. Curricular emphasis was concentrated on bias, fairness, and privacy, with comparatively less attention given to governance, explainability, and trust. Evaluation methods most commonly relied on self-report and reflective approaches, while validated instruments and performance-based assessments were less common, and behavioral or applied outcomes were rarely assessed. Conclusions: The literature suggests that the field is oriented more toward awareness-building than the development of measurable ethical competence. Clearer competency definitions, stronger transparency in assessment, and better alignment between instructional design and evaluation would improve comparability across studies and support evidence-informed course development.

2026 Conference

Decentralized in Name Only: The Centralization of DAO Labor

Lingling Zhang, Morteza Zihayat, Ebrahim Bagheri

TheWebConf 2026 ACM The Web Conference (TheWebConf 2026)

Abstract

This paper studys how decentralized the Decentralized Autonomous Organizations (DAOs) truly are in their everyday labor practices. We collect task assignment data from Dework and construct a cleaned analysis subset of 266 DAOs with sufficient activity to support meaningful decentralization measurement. We measure decentralization through inequality, diversity, dominance, coverage, and combined them into a Decentralization Diversity Index (DDI). We find that operational work is highly concentrated: a small group of contributors completes most tasks. A fine-tuned T5 model trained for contributor recommendation reproduces these patterns, amplifying exposure to dominant members. To intervene, we propose a decentralization-aware reranking method that penalizes historically overrepresented contributors. Experiments reveal a tunable trade-off between relevance and decentralization, with small but consistent DDI gains at top ranks. Our findings show that DAO labor is far from decentralized and that lightweight post-hoc adjustments can broaden contributor exposure.

2026 Conference

How Information Retrieval Systems Construct and Amplify Immigration Narratives

Zarif Masud, Abhijit Paul, Syed Ishtiaque Ahmed, Ebrahim Bagheri

ECIR 2026 48th European Conference on Information Retrieval (ECIR 2026)

Abstract

Information retrieval systems play a central role in how peo- ple access and understand information about complex social issues, in- cluding immigration. Yet little is known about how the datasets that underpin these systems represent migrants or structure public narratives about migration. In this paper, we investigate how immigration is framed within a widely used IR benchmark and how ranking models shape the visibility of those frames. Using MS MARCO as our data source, we cu- rate immigration-related queries and annotate retrieved passages using a migration-specific framing taxonomy grounded in social-science research. Our goal is to identify which narratives dominate and to measure how different retrieval models influence their exposure. We find that legality and security frames are far more common than humanitarian or inclu- sive ones, and that neural reranking amplifies exclusionary portrayals compared to sparse retrieval.

2025 Conference

Few-Shot Adversarial Attacks against Neural Ranking Models

Amin Bigdeli, Negar Arabzadeh, Ebrahim Bagheri, Charles L. A. Clarke

International ACM SIGIR Conference on Information Retrieval in the Asia Pacific (SIGIR AP 2025)

Abstract

Neural ranking models have become the backbone of modern information retrieval systems, yet they remain vulnerable to adversarial manipulation. This paper introduces Few-Shot Adversarial Prompting (FSAP), a novel framework that leverages large language models (LLMs) to generate harmful, high-ranking adversarial documents without access to model gradients or internal states. Unlike prior attacks that modify existing documents or rely on handcrafted templates, FSAP exploits in-context learning to synthesize realistic adversarial documents conditioned on a small support set of previously seen harmful examples. We propose two variants: FSAPIntraQ, which uses examples from the same query, and FSAPInterQ, which transfers adversarial patterns across unrelated topics. Through comprehensive evaluation on the TREC 2020 and TREC 2021 Health Misinformation Tracks and across four neural rankers, we show that FSAP achieves superior attack effectiveness, strong stance fidelity, and high undetectability. Our findings demonstrate that FSAP generalizes across different LLMs, posing a transferable and scalable threat model for neural retrieval systems.

2025 Journal

Jointly Learning Content-Network Representations for Collaborative Expert Discovery

Roohollah Etemadi, Morteza Zihayat Kermani, Kuan Feng, Jason Adelman, Fattane Zarrinkalam, Ebrahim Bagheri

Machine Learning Journal

Abstract

The success of community question answering (CQA) platforms is in part dependent on how well questions are matched with the best experts that can effectively answer them in a timely manner. As the complexity of questions on CQA platforms increases, necessitating the collaboration of multiple experts, existing expert-finding methods are inadequate because they are designed to evaluate the suitability of individual experts for specific questions rather than collaborative expertise. This paper introduces a novel approach to identify collaborative teams of experts by employing graph neural networks and kernel pooling techniques, trained end-to-end. Our approach not only predicts potential individual experts but also forms teams that collectively enhance the answer quality. Through extensive experiments on real-life datasets, we show our proposed model is able to (1) find more qualified experts for new questions by at least 4.4% and 6.7% superior ranks in terms of ranking metrics NDCG and MAP, respectively; (2) form a collaborative team of experts with 4.7% higher skill coverage.

2025 Journal

Learning Context-aware Term Importance for Query Performance Prediction

Abbas Saleminezhad, Negar Arabzadeh, Soosan Beheshti, Ebrahim Bagheri

ACM Transactions on Intelligent Systems and Technology (TIST)

Abstract

Ad hoc retrieval, a cornerstone task in Information Retrieval (IR), aims to rank documents in response to a user’s query, often without prior knowledge of the user’s specific information need. While transformer-based neural rankers have achieved state-of-the-art performance in ad hoc retrieval, their effectiveness varies significantly across queries. Certain queries—commonly referred to as hard queries—remain particularly challenging, highlighting critical gaps in retrieval models. Identifying these hard queries is essential for improving retrieval systems, motivating the task of Query Performance Prediction (QPP), which aims to estimate the effectiveness of a query without requiring access to relevance judgments. In this paper, we propose Context-Aware Query Performance Prediction (CA–QPP), a novel post-retrieval QPP method, which builds on the foundations of perturbation-based QPP methods that hypothesize a relationship between query sensitivity to small perturbations and query retrieval effectiveness. Building on this foundation, our approach exposes the given query to perturbations by constructing two query variations: an effective variation emphasizing terms that enhance retrieval and an ineffective variation accentuating terms that hinder it. By contrasting the retrieval outcomes of these variations using a cross-encoder model, CA–QPP captures the interplay of term contributions and predicts the performance for the given query. We evaluate CA–QPP on the widely used MS MARCO datasets and their associated query sets, including TREC DL 2019, TREC DL 2020, DL-Hard, TREC DL 2021, and TREC DL 2022, which feature extensive human-labeled relevance judgments. Our experiments demonstrate that CA–QPP consistently outperforms traditional and neural-based QPP baselines across standard correlation metrics, including Pearson’s ρ, Kendall’s τ, and Spearman’s ρ. Through a detailed case study, we further illustrate the mechanics of CA–QPP and provide empirical evidence for its ability to model the contextual impact of individual query terms, making it a robust framework for query performance prediction.

2025 Conference

Say the Task, Build the Team: Prompt-Based Team Formation

Lingling Zhang, Radin Hamidi Rad, Morteza Zihayat, Ebrahim Bagheri

ASONAM 2025 The 17th International Conference on Advances in Social Networks Analysis and Mining (ASONAM 2025)

Abstract

The problem of assembling effective expert teams based on project needs is central to expert networks such as LinkedIn. However, current team formation methods typically depend on keyword-matching techniques that fail to capture the nuanced semantics of natural lan- guage project descriptions. This results in inadequate modeling of re- quired expertise and suboptimal team selection. Addressing this gap, we propose a contextual, prompt-driven framework for team formation that infers latent expertise from rich textual descriptions of project goals. Our approach fine-tunes a T5-Large sequence-to-sequence model to trans- late project prompts into expert team compositions by benefiting from enhanced expertise annotations. To facilitate this task, we curate, and publicly release, a dataset based on DBLP V14 collection, augmented with high-confidence expertise labels generated by large language mod- els. Experimental results across multiple evaluation metrics show that our proposed model outperforms existing state-of-the-art baselines, un- derscoring the importance of contextualized representations in expert discovery and team assembly.

2025 Journal

The Role of Protocol Papers, Scoping Reviews, and Systematic Reviews in Responsible AI Research

Calvin Hillis, Ebrahim Bagheri, Zack Marshall

IEEE Technology and Society Magazine

Abstract

Research protocol papers, scoping reviews, and systematic reviews are important tools for ensuring transparency, methodological rigor, accountability, and reproducibility in scholarly research. While widely adopted in disciplines such as medicine and social sciences, these practices remain underutilized in computer science and engineering, particularly within Responsible AI research. This paper explores the critical role that scoping and systematic reviews play in synthesizing knowledge, identifying gaps, and guiding future research in Responsible AI. It also examines how research protocol papers—especially those developed for structured reviews—can improve the credibility, clarity, and replicability of Responsible AI studies. By outlining the benefits and challenges of these methods, we argue for their greater adoption in AI ethics and Responsible AI research. We conclude with recommendations to support the institutional, editorial, and cultural shifts necessary to integrate these rigorous methodological tools into Responsible AI scholarship.

2024 Journal

Exploring decision-makers’ challenges and strategies when selecting multiple systematic reviews: insights for AI decision support tools in healthcare

Carole Lunny, Sera Whitelaw, Emma K Reid, Yuan Chi, Nicola Ferri, Jia He (Janet) Zhang, Dawid Pieper, Salmaan Kanji, Areti-Angeliki Veroniki, Beverley Shea, Jasmeen Dourka, Clare Ardern, Ba Pham, Ebrahim Bagheri, Andrea C Tricco

BMJ Open

DOI
Abstract

Background Systematic reviews (SRs) are being published at an accelerated rate. Decision-makers may struggle with comparing and choosing between multiple SRs on the same topic. We aimed to understand how healthcare decision-makers (eg, practitioners, policymakers, researchers) use SRs to inform decision-making and to explore the potential role of a proposed artificial intelligence (AI) tool to assist in critical appraisal and choosing among SRs. Methods We developed a survey with 21 open and closed questions. We followed a knowledge translation plan to disseminate the survey through social media and professional networks. Results Our survey response rate was lower than expected (7.9% of distributed emails). Of the 684 respondents, 58.2% identified as researchers, 37.1% as practitioners, 19.2% as students and 13.5% as policymakers. Respondents frequently sought out SRs (97.1%) as a source of evidence to inform decision-making. They frequently (97.9%) found more than one SR on a given topic of interest to them. Just over half (50.8%) struggled to choose the most trustworthy SR among multiple. These difficulties related to lack of time (55.2%), or difficulties comparing due to varying methodological quality of SRs (54.2%), differences in results and conclusions (49.7%) or variation in the included studies (44.6%). Respondents compared SRs based on the relevance to their question of interest, methodological quality, and recency of the SR search. Most respondents (87.0%) were interested in an AI tool to help appraise and compare SRs. Conclusions Given the identified barriers of using SR evidence, an AI tool to facilitate comparison of the relevance of SRs, the search and methodological quality, could help users efficiently choose among SRs and make healthcare decisions.

2024 Conference

It Takes a Team to Triumph: Collaborative Expert Finding in Community QA Networks

Roohollah Etemadi, Morteza Zihayat, Kuan Feng, Jason Adelman, Fattane Zarrinkalam, Ebrahim Bagheri

International ACM SIGIR Conference on Information Retrieval in the Asia Pacific (SIGIR AP 2024)

Abstract

The increasing complexity and multidisciplinary nature of queries on Community Question Answering (CQA) platforms have rendered the traditional model of individual expert response inadequate. This paper tackles the challenge of identifying a group of experts whose combined expertise can effectively address such complex inquiries collaboratively, leading to more accepted answers. Our approach jointly learns topological and textual information extracted from the CQA environment in an end-to-end fashion. Extensive experiments on several real-life datasets indicate that our approach improves the quality of expert ranks on average 4.6% and 7.1% in terms of NDCG and MAP, respectively, compared to the best baseline. The results also reveal that groups formed by our approach are more collaborative and on average 61.6% of members recommended by our approach are among the true answerers of questions which is around 6.1 times improvement compared to the baselines.

2024 Journal

Predicting Users’ Future Interests on Social Networks: A Reference Framework

Fattane Zarrinkalam, Havva Alizadeh Noughabi, Zeinab Noorian, Hossein Fani, Ebrahim Bagheri

Information Processing and Management

Abstract

Predicting users’ interests on social networks is gaining attention due to its potential to cater customized information and services to the end users. Although previous works have extensively explored how users’ interests can be modeled on social networks, there has been limited investigation into the prediction of users’ future interests. The objective of our work in this paper is to empirically study the effectiveness of different sets of features based on users’ past social interactions, historical interests and their temporal dynamics to predict their interests over a collection of future-yet-unobserved topics. More specifically, we introduce and formalize the features for interest prediction in four categories: user-based, topical, explicit user-topic engagement, and friends’ influence. We further explore the influence of temporality by augmenting features with information pertaining to users’ historical interests and social connections. We model the task of future interest prediction as a learning-to-rank problem where different features and their related categories are ranked based on their relevance and performance in interest prediction, and investigate the efficiency of different features individually and comparatively for predicting the future interest of users with different activity levels in social networks over on unobserved topics. After conducting experiments on a real-world dataset sourced from Twitter, we have identified several noteworthy findings: 1) relevance feature in the category of past explicit user-topic engagement is the strongest indicator for predicting user’s future interest across all user groups, with an observed 8.57% decrease in NDCG and an 8.95% decrease in MAP when it is removed in the ablation study. 2) the observation of an 8.06% decrease in NDCG and a 7.3% decrease in MAP, when topical features such as popularity, freshness, and coherence are removed in the ablation study, highlights their significance as among the strongest indicators for users’ future interest, particularly for low-activity users. 3) although temporal features show a clear positive impact across user groups with varying levels of activity (resulting in a 4.5% decrease in NDCG and a 7.3% decrease in MAP when removed in the ablation study), the temporal topical features do not demonstrate a significant positive effect, and 4) The removal of user-specific characteristics such as influence and personality traits in the ablation study reveals their significant impact in predicting future interest over cold topics, reflected by a 5.49% decrease in NDCG and a 5.72% decrease in MAP. Our findings make significant contributions to the field of future interest prediction, offering valuable insights and practical implications for various applications in social network analysis.

2022 Journal

Embedding-based Team Formation for Community Question Answering

Roohollah Etemadi, Morteza Zihayat, Kuan Feng, Jason Adelman, Ebrahim Bagheri

Information Sciences

DOI
Abstract

Finding a qualified individual who can independently answer a question on a community question answering platform is becoming more challenging due to the increasing multidisciplinary nature of posted questions. As such, finding a group of experts to collaboratively answer the questions is of paramount importance. To this end, our proposed method forms teams of experts who can collectively answer new questions. Our approach, called team2box, learns neural embedding representations based on the content of the posted questions, experts’ engagement with these questions, and past expert collaboration history in order to form a team to answer the posted question. It embeds experts and questions as points and existing teams as regions within the embedding space. Such an approach allows team2box to form a team whose members (1) collectively cover the knowledge required to answer a question, (2) have successful past experience in jointly answering similar questions, and (3) can work efficiently together to answer the question. Extensive experiments on real-life datasets from Stack Exchange show that team2box outperforms the state-of-the-art by discovering teams with on average 38.97% more covering the skills required to answer new questions and employing experts with collectively a high expertise level.

2022 Journal

Exploring the Utility of Social Content for Understanding Future In-Demand Skills

Jalehsadat Mahdavimoghaddam, Ayush Bahuguna, Ebrahim Bagheri

CSCW 2022 Proc. ACM Hum. Comput. Interact. (CSCW)

DOI
Abstract

Rapid technological innovations, especially in the Information Technology space, demand the workforce to be vigilant by acquiring new skills to remain relevant and employable. The workforce needs to be engaged in a continuous lifelong learning process by educating themselves about skills that will be in-demand in the future. To do so, it is important for students, job seekers, and even recruiters to know which skills will be in demand in the future to invest time and resources in these skills. On this basis, the main objective of this paper is to investigate whether social content can offer insight into potential future in-demand skills in the IT job market. Based on the analysis of social content from Reddit and Job Posting data from Dice and Monster websites, we find that social content related to job skills are strong indicators for future in-demand skills. We further find that specific social content associated with recruitment-related topics are stronger indicators of future skills. Our findings encourage learners and job seekers to pay close attention to online social content to forward plan new skills and maximize their employability.

2022 Journal

Feature-based Question Routing in Community Question Answering Platforms

Soroosh Sorkhani, Amin Bigdeli, Roohollah Etemadi, Morteza Zihayat, Ebrahim Bagheri

Information Sciences

DOI
Abstract

Community question answering (CQA) platforms are receiving increased attention and are becoming an indispensable source of information in different domains ranging from board games to physics. The success of these platforms dependent on how efficiently new questions are assigned to community experts, known as called question routing. In this paper, we address the problem of question routing by adopting a learning to rank approach over five CQA websites in the context of which we introduce 74 features and systematically classify them into content-based and social-based categories. Our extensive experiments on datasets from five real online question answering websites indicate that content-based features related to tags and topics as well as social features that are related to user characteristics and user temporality are effective for question routing. Our work shows the ability to improve performance compared to the state-of-the-art neural matchmaking methods that lack the interpretability offered by our work. The improvement can be as high as on average 2.74% and20 2.27% in terms of common ranking metrics, Normalized Discounted Cumulative Gain (NDCG) and Mean Average Precision (MAP) respectively, compared to our best baselines.

2022 Journal

Social Alignment Contagion in Online Social Networks

Amin Mirlohi, Jalehsadat Mahdavimoghaddam, Jelena Jovanovic, Feras Al-Obeidat, Ali A Ghorbani, Ebrahim Bagheri

IEEE Transactions on Computational Social Systems

DOI
Abstract

Researchers have already observed social contagion effects in both in-person and online interactions. However, such studies have primarily focused on users’ beliefs, mental states, and interests. In this article, we expand the state of the art by exploring the impact of social contagion on social alignment, i.e., whether the decision to socially align oneself with the general opinion of the users on the social network is contagious to one’s connections on the network or not. The novelty of our work in this article includes: 1) unlike earlier work, this article is among the first to explore the contagiousness of the concept of social alignment on social networks; 2) our work adopts an instrumental variable approach to determine reliable causal relations between observed social contagion effects on the social network; and 3) our work expands beyond the mere presence of contagion in social alignment and also explores the role of population heterogeneity on social alignment contagion. Based on the systematic collection and analysis of data from two large social network platforms, namely, Twitter and Foursquare, we find that a user’s decision to socially align or distance from social topics and sentiments influences the social alignment decisions of their connections on the social network. We further find that such social alignment decisions are significantly impacted by population heterogeneity.

2021 Journal

Learning to Rank Implicit Entities on Twitter

Hawre Hosseini, Ebrahim Bagheri

Information Processing and Management

DOI
Abstract

Linking textual content to entities from the knowledge graph has received increasing attention in the context of which surface form representations of entities, e.g., terms or phrases, are disambiguated and linked to appropriate entities. This allows textual content, e.g., social user-generated content, to be interpreted and reasoned on at a higher semantic level. However, recent research has shown that at least 15% of social user-generated content do not have explicit surface form representation of entities that they discuss. In other words, the subject of the content is only implied. For such cases, existing entity linking methods, known as explicit entity linking, cannot perform linking because entity surface form is missing. In this paper, we investigate how implicit entities within social content can be identified and linked. The contributions of our work include (1) modeling the problem of implicit entity linking as a learn to rank problem where knowledge graph entities are ranked based on their relevance to the input tweet, (2) the introduction and systematic classification of appropriate features for identifying implicit entities, (3) extensive evaluation of the proposed approach in comparison with existing state of the art as well as performing feature analysis over proposed features, and (4) the qualitative assessment of the root causes for mislabeled instances in our experiments and careful discussion on how mislabeled entity links can be addressed as a part of future work. In our experiments, we show that our proposed features are able to improve the state of the art over the standard Precision at 1 (P@1) metric.

2021 Journal

On the Congruence Between Online Social Content and Future IT Skill Demand

Jalehsadat Mahdavimoghaddam, Niranjan Krishnaswamy, Ebrahim Bagheri

CSCW 2021 Proc. ACM Hum. Comput. Interact.

DOI
Abstract

The speed of digital transformation has resulted in new challenges for job seekers to become lifelong learners and to develop new skills faster than before. In this paper, our main objective is to examine how online content can serve as indicators for changes to the Information Technology (IT) industry and its in-demand skills. To study this relationship, we collect Reddit posts to represent social media content and job postings to reflect the IT industry based on which we explore possible correlations between them. Further, we propose a methodology to quantitatively estimate the predictive power of social media content for future in-demand skills. Our results show that the frequency of skill-related conversations on Reddit correlates with the popularity of skills in job posting data. Additionally, our findings indicate that the number of social posts dedicated to a specific skill can be a strong indicator for future job requirements. This is an important finding because identifying what skills the labor force should acquire will assist job seekers to plan their lifelong learning objectives to (a) maximize their employability, (b) continuously update their skills to remain in demand, and (c) be informed and actively engaged in defining knowledge trends, rather than reactively becoming informed of the latest information.

2020 Journal

Extracting, Mining and Predicting Users' Interests from Social Media

Fattane Zarrinkalam, Stefano Faralli, Guangyuan Piao, Ebrahim Bagheri

Foundations and Trends in Information Retrieval (FnTIR)

DOI
Abstract

The abundance of user generated content on social media provides the opportunity to build models that are able to accurately and effectively extract, mine and predict users' interests with the hopes of enabling more effective user engagement, better quality delivery of appropriate services and higher user satisfaction. While traditional methods for building user profiles relied on AI-based preference elicitation techniques that could have been considered to be intrusive and undesirable by the users, more recent advances are focused on a non-intrusive yet accurate way of determining users' interests and preferences. In this paper, we will cover five important subjects related to the mining of user interests from social media: (1) the foundations of social user interest modeling, such as information sources, various types of representation models and temporal features, (2) techniques that have been adopted or proposed for mining user interests, (3) different evaluation methodologies and benchmark datasets, (4) different applications that have been taking advantage of user interest mining from social media platforms, and (5) existing challenges, open research questions and exciting opportunities for further work.

2020 Conference

Mining User Interests from Social Media

Fattane Zarrinkalam, Stefano Faralli, Guangyuan Piao, Ebrahim Bagheri

The 29th ACM International Conference on Information and Knowledge Management (tutorial), (CIKM2020)

DOI
Abstract

The abundance of user generated content on social media provides the opportunity to build models that are able to accurately and effectively extract, mine and predict users’ interests with the hopes of enabling more effective user engagement, better quality delivery of appropriate services and higher user satisfaction. While traditional methods for building user profiles relied on AI-based preference elicitation techniques that could have been considered to be intrusive and undesirable by the users, more recent advances are focused on a non-intrusive yet accurate way of determining users’ interests and preferences. In this tutorial, we will cover five important aspects related to the effective mining of user interests: we will introduce (1) the information sources that are used for extracting user interests, (2) the variety of types of user interest profiles that have been proposed in the literature, (3) techniques that have been adopted or proposed for mining user interests, (4) the scalability and resource requirements of the state of the art methods and, finally (5)the evaluation methodologies that are adopted in the literature for validating the appropriateness of the mined user interest profiles.We will also introduce existing challenges, open research questions and exciting opportunities for further work.

2020 Conference

Neural Embedding-based Metrics for Pre-Retrieval Query Performance Prediction

Negar Arabzadeh, Fattane Zarrinkalam, Jelena Jovanovic, Ebrahim Bagheri

42nd European Conference on IR Research (ECIR 2020)

DOI
Abstract

Query Performance Prediction (QPP) is concerned with estimating the effectiveness of a query within the context of a retrieval model. It allows for operations such as query routing and segmentation, leading to improved retrieval performance. Pre-retrieval QPP methods are oblivious to the performance of the retrieval model as they predict query difficulty prior to observing the set of documents retrieved for the query. Since neural embedding-based models are showing wider adoption in the information retrieval community, in this paper, we propose a set of pre-retrieval QPP metrics based on the properties of pre-trained neural embeddings and show that such metrics are more effective for performance prediction compared to the widely known QPP metrics such as SCQ, PMI and SCS. We report our findings based on Robust04, ClueWeb09 and Gov2 corpora and their associated TREC topics.

2020 Journal

On The Causal Relation Between Real World Activities and Emotional Expressions of Social Media Users

Seyed Amin Mirlohi Falavarjani, Jelena Jovanovic, Hossein Fani, Ali A Ghorbani, Zeinab Noorian, Ebrahim Bagheri

Journal of the Association for Information Science and Technology (JASIST)

DOI
Abstract

Social interactions through online social media have become a daily routine of many, and the number of those whose real-world (offline) and online lives have become intertwined is continuously growing. As such, the interplay of individuals' online and offline activities has been the subject of numerous research studies, the majority of which explored the impact of people's online actions on their offline activities. The opposite direction of impact - the effect of real-world activities on online actions - has also received attention but to a lesser degree. To contribute to the latter form of impact, this paper reports on a quasi-experimental design study that examined the presence of causal relations between real-world activities of online social media users and their online emotional expressions. To this end, we have collected a large dataset (over 17K users) from Twitter and Foursquare, and systematically aligned user content on the two social media platforms. Users' Foursquare check-ins provided information about their offline activities, whereas the users' expressions of emotions and moods were derived from their Twitter posts. Since our study was based on a quasi-experimental design, to minimise the impact of covariates, we applied an innovative model of computing propensity scores. Our main findings can be summarised as follows: (i) users' offline activities do impact their affective expressions, both of emotions and moods, as evidenced in their online shared textual content; (ii) the impact depends on the type of offline activity and if the user embarks on or abandons the activity. Our findings can be used to devise a personalised recommendation mechanism to help people better manage their online emotional expressions.

2020 Conference

ReQue: A Configurable Workflow and Dataset Collection for Query Refinement

Mahtab Tamannaee, Hossein Fani, Fattane Zarrinkalam, Jamil Samouh, Samad Paydar, Ebrahim Bagheri

The 29th ACM International Conference on Information and Knowledge Management, (CIKM2020)

DOI
Abstract

In this paper, we implement and publicly share a configurable software workflow and a collection of gold standard datasets for training and evaluating supervised query refinement methods. Existing datasets such as AOL and MS MARCO, which have been extensively used in the literature for this purpose, are based on the weak assumption that users’ input queries improve gradually within a search session, i.e., the last query where the user ends her information seeking session is the best reconstructed version of her initial query. In practice, such an assumption is not necessarily accurate for a variety of reasons, e.g., topic drift. The objective of our work is to enable researchers to build gold standard query refinement datasets without having to rely on such weak assumptions. Our software workflow, which generates such gold standard query datasets, takes three inputs: (1) a dataset of queries along with their associated relevance judgements (e.g. TREC topics), (2) an information retrieval method (e.g., BM25), and (3) an evaluation metric (e.g., MAP), and outputs a gold standard dataset. The produced gold standard dataset includes a list of revised queries for each query in the input dataset, each of which effectively improves the performance of the specified retrieval method (e.g., BM25) in terms of the desirable evaluation metric (e.g., MAP). Since our workflow can be used to generate gold standard datasets for any input query set, in this paper, we have generated and publicly shared gold standard datasets for TREC queries associated with Robust04, Gov2, ClueWeb09, and ClueWeb12. The source code of our software workflow, the generated gold datasets, and benchmark results for three state-of-the-art supervised query refinement methods over these datasets are made publicly available for reproducibility purposes.

2020 Conference

Temporal Latent Space Modeling for Community Prediction

Hossein Fani, Ebrahim Bagheri, Weichang Du

42nd European Conference on IR Research (ECIR 2020)

DOI
Abstract

We propose a temporal latent space model for user community prediction in social networks, whose goal is to predict future emerging user communities based on past history of users' topics of interest. Our model assumes that each user lies within an unobserved latent space, and similar users in the latent space representation are more likely to be members of the same user community. The model allows each user to adjust its location in the latent space as her topics of interest evolve over time. Empirically, we demonstrate that our model, when evaluated on a Twitter dataset, outperforms existing approaches under two application scenarios, namely news recommendation and user prediction on a host of metrics such as mrr, ndcg as well as precision and f-measure.

2019 Conference

Neural Embedding Features for Point-of-Interest Recommendation

Alireza Pourali, Fattane Zarrinkalam, Ebrahim Bagheri

ASONAM 2019 IEEE/ACM International Conference on Social Networks Analysis and Mining (ASONAM 2019)

DOI
Abstract

The focus of point-of-interest recommendation techniques is to suggest a venue to a given user that would match the users' interests and is likely to be adopted by the user. Given the multitude of venues and the sparsity of user check-ins, the problem of recommending venues has shown to be a difficult task. Existing literature has already explored various types of features such as geographical distribution, social structure and temporal behavioral patterns to make a recommendation. In this paper, we propose a new set of features derived based on the neural embeddings of venues and users. We show how the neural embeddings for users and venues can be jointly learnt based on the prior check-in sequence of users and then be used to define three types of features, namely user, venue, and user-venue interaction features. These features are integrated into a feature-based matrix factorization model. Our experiments show that the features defined over the user and venue embeddings are effective for venue recommendation.

2019 Conference

On the Causal Relation between Users' Real-World Activities and their Affective Processes

Seyed Amin Mirlohi Falavarjani, Ebrahim Bagheri, Ssu Yu Zoe Chou, Jelena Jovanovic, Ali A Ghorbani

ASONAM 2019 IEEE/ACM International Conference on Social Networks Analysis and Mining (ASONAM 2019)

DOI
Abstract

Research in social network analytics has already extensively explored how engagement on online social networks can lead to observable effects on users' real-world behavior (e.g., changing exercising patterns or dietary habits), and their psychological states. The objective of our work in this paper is to investigate the flip-side and examine whether engaging in or disengaging from real-world activities would reflect itself in users' affective processes such as anger, anxiety, and sadness, as expressed in users' posts on online social media. We have collected data from Foursquare and Twitter and found that engaging in or disengaging from a real-world activity, such as frequenting at bars or stopping going to a gym, have direct impact on the users' affective processes. In particular, we report that engaging in a routine real-world activity leads to expressing less emotional content online, whereas the reverse is observed when users abandon a regular real-world activity.

2019 Journal

The Reflection of Offline Activities on Users’ Online Social Behavior: An Observational Study

Seyed Amin Mirlohi Falavarjani, Fattane Zarrinkalam, Jelena Jovanovic, Ebrahim Bagheri, Ali A Ghorbani

Information Processing and Management

DOI
Abstract

The ever increasing presence of online social networks in users’ daily lives has led to the interplay between users’ online and offline activities. There have already been several works that have studied the impact of users’ online activities on their offline behavior, e.g., the impact of interaction with friends on an exercise social network on the number of daily steps. In this paper, we consider the inverse to what has already been studied and report on our extensive study that explores the potential causal effects of users’ offline activities on their online social behavior. The objective of our work is to understand whether the activities that users are involved with in their real daily life, which place them within or away from social situations, have any direct causal impact on their behavior in online social networks. Our work is motivated by the theory of normative social influence, which argues that individuals may show behaviors or express opinions that conform to those of the community for the sake of being accepted or from fear of rejection or isolation. We have collected data from two online social networks, namely Twitter and Foursquare, and systematically aligned user content on both social networks. On this basis, we have performed a natural experiment that took the form of an interrupted time series with a comparison group design to study whether users’ socially situated offline activities exhibited through their Foursquare check-ins impact their online behavior captured through the content they share on Twitter. Our main findings can be summarised as follows (1) a change in users’ offline behaviour that affects the level of users’ exposure to social situations, e.g., starting to go to the gym or discontinuing frequenting at bars, can have a causal impact on users’ online topical interests and sentiment; and (2) the causal relations between users’ socially situated offline activities and their online social behavior can be used to build effective predictive models of users’ online topical interests and sentiments.

2019 Journal

User Community Detection via Embedding of Social Network Structure and Temporal Content

Hossein Fani, Eric Jiang, Ebrahim Bagheri, Feras Al-Obeidat, Weichang Du, Mehdi Kargar

Information Processing and Management

DOI
Abstract

Identifying and extracting user communities is an important step towards understanding social network dynamics from a macro perspective. For this reason, the work in this paper explores various aspects related to the identification of user communities. To date, user community detection methods employ either explicit links between users (link analysis), or users' topics of interest in posted content (content analysis), or in tandem. Little work has considered temporal evolution when identifying user communities in a way to group together those users who share not only similar topical interests but also similar temporal behavior towards their topics of interest. In this paper, we identify user communities through textitmultimodal feature learning (embeddings). Our core contributions can be enumerated as (a) we propose a new method for learning neural embeddings for users based on their temporal content similarity; (b) we learn user embeddings based on their social network connections (links) through neural graph embeddings; (c) we systematically interpolate temporal content-based embeddings and social link-based embeddings to capture both social network connections and temporal content evolution for representing users, and (d) we systematically evaluate the quality of each embedding type in isolation and also when interpolated together and demonstrate their performance on a Twitter dataset under two different application scenarios, namely textitnews recommendation and textituser prediction. We find that (1) content-based methods produce higher quality communities compared to link-based methods; (2) methods that consider temporal evolution of content, our proposed method in particular, show better performance compared to their non-temporal counter-parts; (3) communities that are produced when time is explicitly incorporated in user vector representations have higher quality than the ones produced when time is incorporated into a generative process, and finally (4) while link-based methods are weaker than content-based methods, their interpolation with content-based methods leads to improved quality of the identified communities.

2018 Conference

Causal Dependencies for Future Interest Prediction on Twitter

Negar Arabzadeh, Hossein Fani, Fattaneh Zarrinkalam, Ahmed Navivala, Ebrahim Bagheri

CIKM 2018 The 27th ACM International Conference on Information and Knowledge Management (CIKM 2018)

DOI
Abstract

The accurate prediction of users' future topics of interests on social networks can facilitate content recommendation and platform engagement. However, researchers have found that future interest prediction, especially on social networks such as Twitter, is quite challenging due to the rapid changes in community topics and evolution of user interactions. In this context, temporal collaborative filtering methods have already been used to perform user interest prediction, which benefit from similar user behavioral patterns over time to predict how a user's interests might evolve in the future. In this paper, we propose that instead of considering the whole user base within a collaborative filtering framework to predict user interests, it is possible to much more accurately predict such interests by only considering the behavioral patterns of the most influential users related to the user of interest. We model influence as a form of causal dependency between users. To this end, we employ the concept of Granger causality to identify causal dependencies. We show through extensive experimentation that the consideration of only one causally dependent user leads to much more accurate prediction of users' future interests in a host of measures including ranking and rating accuracy metrics.

2018 Journal

Detecting Life Events From Twitter based on Temporal Semantic Features

Maryam Khodabakhsh, Mohsen Kahani, Ebrahim Bagheri, Zeinab Noorian

Knowledge-based Systems

DOI
Abstract

The wide adoption of social networking and microblogging platforms by a large number of users across the globe has provided a rich source of unstructured information for understanding users' behaviors, interests and opinions at both micro and macro levels. An active area in this space is the detection of important real-world events from user-generated social content. The works in this area identify instances of events that impact a large number of users. However, a more nuanced form of an event, known as life event, is also of high importance, which in contrast to real-world events, does not impact a large number of users and is limited to at most a few people. For this reason, life events, such as marriage, travel, and career change, among others, are more difficult to detect for several reasons: i) they are specific to a given user and do not have a wider reaching reflection; ii) they are often not reported directly and need to be inferred from the content posted by individual users; and iii) many users do not report their life events on social platforms, making the problem highly class-imbalanced. In this paper, we propose a semantic approach based on word embedding techniques to model life events. We then use word mover's distance to measure the similarity of a given tweet to different types of life events, which are used as input features for a multi-class classifier. Furthermore, we show that when a sequence of tweets that have appeared before and after a given tweet of interest (temporal stacking) are considered, the performance of the life event detection task improves significantly.

2018 Journal

Entity Linking of Tweets based on Dominant Entity Candidates

Yue Feng, Fattane Zarrinkalam, Ebrahim Bagheri, Hossein Fani, Feras Al-Obeidat

Social Network Analysis and Mining

DOI
Abstract

Entity linking, also known as semantic annotation, of textual content has received increasing attention. Recent works in this area have focused on entity linking on text with special characteristics such as search queries and tweets. The semantic annotation of tweets is specially proven to be challenging given the informal nature of the writing and the short length of the text. In this paper, we propose a method to perform entity linking on tweets built based on one primary hypothesis. We hypothesize that while there are formally many possible entity candidates for an ambiguous mention in a tweet, as listed on the disambiguation page of the corresponding entity on Wikipedia, there are only few entity candidates that are likely to be employed in the context of Twitter. Based on this hypothesis, we propose a method to identify such dominant entity candidates for each ambiguous mention and use them in the annotation process. Particularly, our proposed work integrates two phases i) dominant entity candidate detection, which applies community detection methods for finding the dominant candidates of ambiguous mentions; and ii) named entity disambiguation that links a tweet to entities in Wikipedia by only considering the identified dominant entity candidates. Our investigations show that: 1) there are only very few entity candidates for each ambiguous mention in a tweet that need to be considered when performing disambiguation. This helps us limit the candidate search space and hence noticeably reduce the entity linking time; 2) limiting the search space to only a subset of disambiguation options will not only improve entity linking execution time but will also lead to improved accuracy of the entity linking process when the main entity candidates of each mention are mined from a temporally aligned corpus. We show that our proposed method offers competitive results with the state-of-the-art methods in terms of precision and recall on widely-used gold standard datasets while significantly reducing the time for processing each tweet. year = 2018

2018 Conference

Implicit Entity Linking through Ad-hoc Retrieval

Hawre Hosseini, Tam T Nguyen, Ebrahim Bagheri

ASONAM 2018 IEEE/ACM International Conference on Social Networks Analysis and Mining (ASONAM 2018)

DOI
Abstract

The systematic linking of explicitly-observed phrases within a document to entities of a knowledge base has already been explored in a process known as entity linking. The objective of this paper, however, is to identify and entity link those entities that are not mentioned but are implied within a document, more specifically within a tweet. This process is referred to as implicit entity linking. Unlike prior work that build a representation for each entity based on its related content in the knowledge base, we propose to perform implicit entity linking by determining how a tweet is related to user-generated content posted online and as such indirectly perform entity linking. We formulate this problem as an ad-hoc document retrieval process where the input query is the tweet, which needs to be implicitly linked and the document space is the set of user-generated content related to the entities of the knowledge base. We systematically compare our work with the state-of-the-art baseline and show that our method is able to provide statistically significant improvements.

2018 Journal

Mining User Interests over Active Topics on Social Networks

Fattane Zarrinkalam, Mohsen Kahani, Ebrahim Bagheri

Information Processing and Management

DOI
Abstract

Inferring users' interests from their activities on social networks has been an emerging research topic in the recent years. Most existing approaches heavily rely on the explicit contributions (posts) of a user and overlook users' implicit interests, i.e., those potential user interests that the user did not explicitly mention but might have interest in. Given a set of active topics present in a social network in a specified time interval, our goal is to build an interest profile for a user over these topics by considering both explicit and implicit interests of the user. The reason for this is that the interests of free-riders and cold start users who constitute a large majority of social network users, cannot be directly identified from their explicit contributions to the social network. Specifically, to infer users' implicit interests, we propose a graph-based link prediction schema that operates over a representation model consisting of three types of information: user explicit contributions to topics, relationships between users, and the relatedness between topics. Through extensive experiments on different variants of our representation model and considering both homogeneous and heterogeneous link prediction, we investigate how topic relatedness and users' homophily relation impact the quality of inferring users' implicit interests. Comparison with state-of-the-art baselines on a real-world Twitter dataset demonstrates the effectiveness of our model in inferring users' interests in terms of perplexity and in the context of retweet prediction application. Moreover, we further show that the impact of our work is especially meaningful when considered in case of free-riders and cold start users

2018 Journal

Online social network response to studies on antidepressant use in pregnancy

Simone N Vigod, Ebrahim Bagheri, Fattane Zarrinkalam, Hillary K Brown, Muhammad Mamdani, Joel G Ray

Journal of Psychosomatic Research

DOI
Abstract

Background: About 8% of U.S women are prescribed antidepressant medications around the time of pregnancy. Decisions about medication use in pregnancy can be swayed by the opinion of family, friends and online media, sometimes beyond the advice offered by healthcare providers. Exploration of the online social network response to research on antidepressant use in pregnancy could provide insight about how to optimize decision-making in this complex area. Methods: For all 17 research articles published on the safety of antidepressant use in pregnancy in 2012, we sought to explore online social network activity regarding antidepressant use in pregnancy, via Twitter, in the 48 hours after a study was published, compared to the social network activity in the same period 1 week prior to each article’s publication. Results: Online social network activity about antidepressants in pregnancy quickly doubled upon study publication. The increased activity was driven by studies demonstrating harm associated with antidepressants, by lower-quality studies, and studies where abstracts presented relative versus absolute risks. Implications: These findings support a call for leadership from medical journals to consider how to best incentivize and support a balanced and clear translation of knowledge around antidepressant safety in pregnancy to their readership and the public.

2018 Journal

Predicting Future Personal Life Events on Twitter via Recurrent Neural Networks

Maryam Khodabakhsh, Mohsen Kahani, Ebrahim Bagheri

Journal of Intelligent Information Systems

DOI
Abstract

Social network users publicly share a wide variety of information with their followers and the general public ranging from their opinions, sentiments and personal life activities. There has already been significant advance in analyzing the shared information from both micro (individual user) and macro (community level) perspectives, giving access to actionable insight about user and community behaviors. The identification of personal life events from user's profiles is a challenging yet important task, which if done appropriately, would facilitate more accurate identification of users' preferences, interests and attitudes. For instance, a user who has just broken his phone, is likely to be upset and also be looking to purchase a new phone. While there is work that identifies tweets that include mentions of personal life events, our work in this paper goes beyond the state of the art by predicting a future personal life event that a user will be posting about on Twitter solely based on the past tweets. We propose two architectures based on recurrent neural networks, namely the classification and generation architectures, that determine the future personal life event of a user. We evaluate our work based on a gold standard Twitter life event dataset and compare our work with the state of the art baseline technique for life event detection. While presenting performance measures, we also discuss the limitations of our work in this paper.

2018 Conference

Predicting Personal Life Events from Streaming Social Content

Maryam Khodabakhsh, Hossein Fani, Fattane Zarrinkalam, Ebrahim Bagheri

CIKM 2018 The 27th ACM International Conference on Information and Knowledge Management (CIKM 2018)

DOI
Abstract

Researchers have shown that it is possible to identify reported instances of personal life events from users' social content, e.g., tweets. This is known as personal life event detection. In this paper, we take a step forward and explore the possibility of predicting users' next personal life event based solely on the their historically reported personal life events, a task which we refer to as personal life event prediction. We present a framework for modeling streaming social content for the purpose of personal life event prediction and describe how various instantiations of the framework can be developed to build a life event prediction model. In our extensive experiments, we find that (i) historical personal life events of a user have strong predictive power for determining the user's future life event; (ii) the consideration of sequence in historically reported personal life events shows inferior performance compared to models that do not consider sequence, and (iii) the number of historical life events and the length of the past time intervals that are taken into account for making life event predictions can impact prediction performance whereby more recent life events show more relevance for the prediction of future life events.

2018 Conference

The Impact of Foursquare Checkins on Users' Emotions on Twitter

Seyed Amin Mirlohi Falavarjani, Hawre Hosseini, Ebrahim Bagheri

International Workshop on Social Aspects in Personalization and Search collocated with the 40th European Conference on IR Research, ECIR 2018, Grenoble, France, March 26-29, 2018

DOI
Abstract

Performing observational studies based on social network content has recently gained attraction where the impact of various types of interruptions has been studied on users’ behavior. There has been recent work that have focused on how online social network behavior and activity can impact users’ offline behavior. In this paper, we study the inverse where we focus on whether users’ offline behavior captured through their check-ins at different venues on Foursquare can impact users’ online emotion expression as depicted in their tweets. We show that users’ offline activity can impact users’ online emotions; however, the type of activity determines the extent to which a user’s emotions will be impacted.

2018 Journal

Topic and Sentiment Aware Microblog Summarization for Twitter

Syed Muhammad Ali, Zeinab Noorian, Ebrahim Bagheri, Chen Ding, Feras Al-Obeidat

Journal of Intelligent Information Systems

DOI
Abstract

Recent advances in microblog content summarization has primarily viewed this task in the context of traditional multi-document summarization techniques where a microblog post or their collection form one document. While these techniques already facilitate information aggregation, categorization and visualization of microblog posts, they fall short in two aspects: i) when summarizing a certain topic from microblog content, not all existing techniques take topic polarity into account. This is an important consideration in that the summarization of a topic should cover all aspects of the topic and hence taking polarity into account (sentiment) can lead to the inclusion of the less popular polarity in the summarization process. ii) Some summarization techniques produce summaries at the topic level. However, it is possible that a given topic can have more than one important aspect that need to have representation in the summarization process. Our work in this paper addresses these two challenges by considering both topic sentiments and topic aspects in tandem. We compare our work with the state of the art Twitter summarization techniques and show that our method is able to outperform existing methods on standard metrics such as ROUGE-1.

2018 Conference

Topic-Association Mining for User Interest Detection

Anil Kumar Trikha, Fattane Zarrinkalam, Ebrahim Bagheri

ECIR 2018 Advances in Information Retrieval: 40th European Conference on IR Research, ECIR 2018, Grenoble, France, March 26-29, 2018, Proceedings

DOI
Abstract

The accurate identification of user interests on Twitter can lead to more efficient procurement of targeted content for the users. While the analysis of user content has engaged with on Twitter is a rich source for detecting the user’s interests, prior research have shown that it may not be sufficient. There have been work that attempt to identify a user’s implicit interests, i.e., those topics that could interest the user but the user has not engaged with them in the past. Prior work has shown that topic semantic relatedness is an important feature for determining users’ implicit interests. In this paper, we explore the possibility of identifying users’ implicit interests solely based on topic association through frequent pattern mining without regard for the semantics of the topics. We show in our experiments that topic association is a strong feature for determining users’ implicit interests.

2018 Journal

User Interest Prediction over Future Unobserved Topics on Social Networks

Fattane Zarrinkalam, Mohsen, Kahani, Ebrahim Bagheri

Information Retrieval Journal

DOI
Abstract

The accurate prediction of users' future interests on social networks allows one to perform future planning by studying how users will react if certain topics emerge in the future. It can improve areas such as targeted advertising and the efficient delivery of services. Despite the importance of predicting user future interests on social networks, existing works mainly focus on identifying user current interests and little work has been done on the prediction of user potential interests in the future. There have been work that attempt to identify a user future interests, however they cannot predict user interests with regard to new topics since these topics have never received any feedback from users in the past. In this paper, we propose a framework that works on the basis of temporal evolution of user interests and utilizes semantic information from knowledge bases such as Wikipedia to predict user future interests and overcome the cold item problem. Through extensive experiments on a real-world Twitter dataset, we demonstrate the effectiveness of our approach in predicting future interests of users compared to state-of-the-art baselines. Moreover, we further show that the impact of our work is especially meaningful when considered in case of cold items.

2017 Book section

Community detection in social networks

Hossein Fani, Ebrahim Bagheri

Encyclopedia with Semantic Computing and Robotic Intelligence

DOI
Abstract

Online social networks have become a fundamental part of the global online experience. They facilitate different modes of communication and social interactions, enabling individuals to play social roles that they regularly undertake in real social settings. In spite of the heterogeneity of the users and interactions, these networks exhibit common properties. For instance, individuals tend to associate with others who share similar interests, a tendency often known as homophily, leading to the formation of communities. This entry aims to provide an overview of the definitions for an online community and review different community detection methods in social networks. Finding communities are beneficial since they provide summarization of network structure, highlighting the main properties of the network. Moreover, it has applications in sociology, biology, marketing and computer science which help scientists identify and extract actionable insight.

2017 Conference

Estimating the Effect of Exercising on Users' Online Behavior booktitle = International Workshop on Observational Studies Through Social Media (OSSM 2017) collocated with International AAAI Conference on Web and Social Media (ICWSM)

Seyed Amin Mirlohi Falavarjani, Hawre Hosseini, Zeinab Noorian, Ebrahim Bagheri

International Workshop on Observational Studies Through Social Media (OSSM 2017) collocated with International AAAI Conference on Web and Social Media (ICWSM)

DOI
Abstract

This study aims to estimate the influence of offline activity on users’ online behavior, relying on a matching method to reduce the effect of confounding variables. We analyze activities of 850 users who are active on both Twitter and Foursquare social networks. Users’ offline activity is extracted from Foursquare posts and users’ online behavior is extracted from Twitter posts. Users’ interests, representing their online behavior, are extracted with regards to a set of topics in several subsequent time intervals. The shift of users’ interests across different time intervals is taken as a measure of user behavior change on the social network. On the other hand, we employ user check-ins at a gym or fitness center as a sign of exercise and consider it to be an offline activity. In order to find the effect of exercise on online behavior, we identify users who did not go to the gym for at least two months but did so at least nine times in the next three months. We show that shift in interest reduces significantly for users after they start exercising, which implies that the offline activity of exercising can influence how users’ interests are shaped and change on the social network over time.

2017 Book section

Event identification in social networks

Fattane Zarrinkalam, Ebrahim Bagheri

Encyclopedia with Semantic Computing and Robotic Intelligence

DOI
Abstract

Social networks enable users to freely communicate with each other and share their recent news, ongoing activities or views about different topics. As a result, they can be seen as a potentially viable source of information to understand the current emerging topics/events. The ability to model emerging topics is a substantial step to monitor and summarize the information originating from social sources. Applying traditional methods for event detection which are often proposed for processing large, formal and structured documents, are less effective, due to the short length, noisiness and informality of the social posts. Recent event detection techniques address these challenges by exploiting the opportunities behind abundant information available in social networks. This article provides an overview of the state of the art in event detection from social networks.

2017 Conference

Predicting Users’ Future Interests on Twitter

Fattane Zarrinkalam, Hossein Fani, Ebrahim Bagheri, Mohsen Kahani

ECIR 2017 Advances in Information Retrieval: 39th European Conference on IR Research, ECIR 2017, Aberdeen, UK, April 8-13, 2017, Proceedings

DOI
Abstract

In this paper, we address the problem of predicting future interests of users with regards to a set of unobserved topics in microblogging services which enables forward planning based on potential future interests. Existing works in the literature that operate based on a known interest space cannot be directly applied to solve this problem. Such methods require at least a minimum user interaction with the topic to perform prediction. To tackle this problem, we integrate the semantic information derived from the Wikipedia category structure and the temporal evolution of user’s interests into our prediction model. More specifically, to capture the temporal behaviour of the topics and user’s interests, we consider discrete intervals and build user’s topic profile in each time interval separately. Then, we generalize users’ interests that have been observed over several time intervals by transferring them over the Wikipedia category structure. Our approach not only allows us to generalize users’ interests but also enables us to transfer users’ interests across different time intervals that do not necessarily have the same set of topics. Our experiments illustrate the superiority of our model compared to the state of the art.

2017 Journal

Query Expansion Using Pseudo Relevance Feedback on Wikipedia

Andisheh Keikha, Faezeh Ensan, Ebrahim Bagheri

Journal of Intelligent Information Systems

DOI
Abstract

One of the major challenges in Web search pertains to the correct interpretation of users’ intent. Query Expansion is one of the well-known approaches for determining the intent of the user by addressing the vocabulary mismatch problem. A limitation of the current query expansion approaches is that the relations between the query terms and the expanded terms is limited. In this paper, we capture users’ intent through query expansion. We build on earlier work in the area by adopting a pseudo-relevance feedback approach; however, we advance the state of the art by proposing an approach for feature learning within the process of query expansion. In our work, we specifically consider the Wikipedia corpus as the feedback collection space and identify the best features within this context for term selection in two supervised and unsupervised models. We compare our work with state of the art query expansion techniques, the results of which show promising robustness and improved precision.

2017 Conference

Temporally Like-minded User Community Identification through Neural Embeddings

Hossein Fani, Ebrahim Bagheri, Weichang Du

CIKM 2017 The 26th ACM International Conference on Information and Knowledge Management (CIKM)

DOI
Abstract

We propose a neural embedding approach to identify temporally like-minded user communities, i.e., those communities of users who have similar temporal alignment in their topics of interest. Like-minded user communities in social networks are usually identified by either considering explicit structural connections between users (link analysis), users' topics of interest expressed in their posted contents (content analysis), or in tandem. In such communities, however, the users' rich temporal behavior towards topics of interest is overlooked. Only few recent research efforts consider the time dimension and define like-minded user communities as groups of users who share not only similar topical interests but also similar temporal behavior. Temporal like-minded user communities find application in areas such as recommender systems where relevant items are recommended to the users at the right time. In this paper, we tackle the problem of identifying temporally like-minded user communities by leveraging unsupervised feature learning (embeddings). Specifically, we learn a mapping from the user space to a low-dimensional vector space of features that incorporate both topics of interest and their temporal nature. We demonstrate the efficacy of our proposed approach on a Twitter dataset in the context of three applications: news recommendation, user prediction and community selection, where our work is able to outperform the state-of-the-art on important information retrieval metrics.

2016 Conference

Inferring Implicit Topical Interests on Twitter

Fattane Zarrinkalam, Hossein Fani, Ebrahim Bagheri, Mohsen Kahani

ECIR 2016 Advances in Information Retrieval - 38th European Conference on IR Research, ECIR 2016, Padua, Italy, March 20-23, 2016. Proceedings

DOI
Abstract

Inferring user interests from their activities in the social network space has been an emerging research topic in the recent years. While much work is done towards detecting explicit interests of the users from their social posts, less work is dedicated to identifying implicit interests, which are also very important for building an accurate user model. In this paper, a graph based link prediction schema is proposed to infer implicit interests of the users towards emerging topics on Twitter. The underlying graph of our proposed work uses three types of information: user’s followerships, user’s explicit interests towards the topics, and the relatedness of the topics. To investigate the impact of each type of information on the accuracy of inferring user implicit interests, different variants of the underlying representation model are investigated along with several link prediction strategies in order to infer implicit interests. Our experimental results demonstrate that using topics relatedness information, especially when determined through semantic similarity measures, has considerable impact on improving the accuracy of user implicit interest prediction, compared to when followership information is only used.

2016 Journal

Overview of Text Annotation with Pictures

Kent Poots, Ebrahim Bagheri

IEEE IT Professional

Abstract

The vast array of information available on the Web makes it a challenge for readers to quickly browse through and decide about the importance and relevance of content. Interpreting large-volumes of data is particularly demanding for users with handheld devices in the social media and micro-blogging sphere. Various approaches address this challenge through text summarization, content ranking and personalized recommendation. We describe a family of techniques that help users understand text by automatically annotating text with pictures, referred to as text picturing . The objective is to find a set of pictures that cover the main concepts in a textual snippet. We provide an overview of text picturing, its constituent steps such as knowledge extraction, map ping, scene rendering, as well as application areas. We give a picturing-related literature overview, and list use-cases that offer IT professionals insight into how picturing techniques can be successfully incorporated into real world applications.

2016 Conference

Query Expansion Using Pseudo Relevance Feedback on Wikipedia

Ebrahim Bagheri Andisheh Keikha Faezeh Ensan

International Workshop on Query Understanding and Reformulation for Mobile and Web Search collocated with The 9th ACM International Conference on Web Search and Data Mining (WSDM 2016)

DOI
Abstract

One of the major challenges in Web search pertains to the correct interpretation of users’ intent. Query Expansion is one of the well-known approaches for determining the intent of the user by addressing the vocabulary mismatch problem. A limitation of the current query expansion approaches is that the relations between the query terms and the expanded terms is limited. In this paper, we capture users’ intent through query expansion. We build on earlier work in the area by adopting a pseudo-relevance feedback approach; however, we advance the state of the art by proposing an approach for feature learning within the process of query expansion. In our work, we specifically consider the Wikipedia corpus as the feedback collection space and identify the best features within this context for term selection in two supervised and unsupervised models. We compare our work with state of the art query expansion techniques, the results of which show promising robustness and improved precision.

2016 Conference

Time-Sensitive Topic-Based Communities on Twitter

Hossein Fani, Fattane Zarrinkalam, Ebrahim Bagheri, Weichang Du

Advances in Artificial Intelligence - 29th Canadian Conference on Artificial Intelligence, Canadian AI 2016, Victoria, BC, Canada, May 31 - June 3, 2016. Proceedings

DOI
Abstract

This paper tackles the problem of detecting temporal content-based user communities from Twitter. Most existing contentbased community detection methods consider the users who share similar topical interests to be like-minded and use this as a basis to group the users. However, such approaches overlook the potential temporality of users’ interests. In this paper, we propose to identify time-sensitive topic-based communities of users who have similar temporal tendency with regards to their topics of interest. The identification of such communities provides the potential for improving the quality of community-level studies, such as personalized recommendations and marketing campaigns that are sensitive to time. To this end, we propose a graph-based framework that utilizes multivariate time series analysis to represent users’ temporal behavior towards their topics of interest in order to identify like-minded users. Further, Topic over Time (TOT) topic model that jointly captures keyword co-occurrences and locality of those patterns over time is utilized to discover users’ topics of interest. Experimental results on our Twitter dataset demonstrates the effectiveness of our proposed temporal approach in the context of personalized news recommendation and timestamp prediction compared to non-temporal community detection methods.

2015 Conference

Lexical Semantic Relatedness for Twitter Analytics

Yue Feng, Hossein Fani, Ebrahim Bagheri, Jelena Jovanovic

27th IEEE International Conference on Tools with Artificial Intelligence, ICTAI 2015, Vietri sul Mare, Italy, November 9-11, 2015

DOI
Abstract

Existing work in the semantic relatedness literature has already considered various information sources such as WordNet, Wikipedia and Web search engines to identify the semantic relatedness between two words. We will show that existing semantic relatedness measures might not be directly applicable to microblogging content such as tweets due to i) the informality and short length of microblogging content, which can lead to shift in the meaning of words when used in microblog posts, ii) the presence of non-dictionary words that have their semantics defined/evolved by the Twitter community. Therefore, we propose the Twitter Space Semantic Relatedness (TSSR) technique that relies on the latent relation hypothesis to measure semantic relatedness of words on Twitter. We construct a graph representation of terms in tweets and apply a random walk procedure to produce a stationary distribution for each word, which is the basis for relatedness calculation. Our experiments examine TSSR from three different perspectives and show that TSSR is better suited for Twitter analytics compared to the standard semantic relatedness techniques.

2015 Conference

Semantics-Enabled User Interest Detection from Twitter

Fattane Zarrinkalam, Hossein Fani, Ebrahim Bagheri, Mohsen Kahani, Weichang Du

IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology, WI-IAT 2015, Singapore, December 6-9, 2015 - Volume I

DOI
Abstract

Social networks enable users to freely communicate with each other and share their recent news, ongoing activities or views about different topics. As a result, user interest detection from social networks has been the subject of increasing attention. Some recent works have proposed to enrich social posts by annotating them with unambiguous relevant ontological concepts extracted from external knowledge bases and model user interests as a bag of concepts. However, in the bag of concepts approach, each topic of interest is represented as an individual concept that is already predefined in the knowledge base. Therefore, it is not possible to infer fine-grained topics of interest, which are only expressible through a collection of multiple concepts or emerging topics, which are not yet defined in the knowledge base. To address these issues, we view each topic of interest as a conjunction of several concepts, which are temporally correlated on Twitter. Based on this, we extract active topics within a given time interval and determine a users inclination towards these active topics. We demonstrate the effectiveness of our approach in the context of a personalized news recommendation system. We show through extensive experimentation that our work is able to improve the state of the art.

2015 Conference

Social Media in Human-Robot Interaction

Jacky Au Duong Mager Alanna Ebrahim Bagheri Frauke Zeller David Harris Smith, Frank Rudzicz

Social Media & Society Conference (#SMSociety15)

Abstract

No abstract available for this paper.

2009 Journal

Can reputation migrate? On the propagation of reputation in multi-context communities

Ebrahim Bagheri, Reza Zafarani, M Barouni-Ebrahimi

Knowl.-Based Syst.

Abstract

As e-communities grow in both quality and quantity, their online users require more appropriate tools to suite their needs in such environments. Many such tools are not explicitly needed in real-world communities where humans directly interact with each other. Trust making and reputation ascription are among the most important examples of such tools. Humans often build trust relationships through interaction or recommendation, and are therefore able to ascribe relevant reputation to those they interact with. However, in online communities the process of trust making and reputation ascription is more complicated. In this paper, we address a special case of the trust making process where community users need to create bonds with those they have not encountered before. This is a common situation in websites such as amazon.com, ebay.com, epionions.com and many others. The model we propose is able to estimate the possible reputation of a given identity in a any new context by observing his/her behavior in other communities. Our proposed model employs Dempster�Shafer based valuation networks to develop a global reputation structure and performs a belief propagation technique to infer contextual reputation values. The preliminary evaluation of the proposed model on a dataset collected from epinions.com shows promising results.