Catching Lies in the Act: A Framework for Early Misinformation Detection on Social Media
Shreya Ghosh (shreya@psu.edu), The Pennsylvania State University, State College, PA, USA · Prasenjit Mitra (pmitra@psu.edu), The Pennsylvania State University, State College, PA, USA and L3S Research Center, Leibniz University, Hannover, Germany
Published in HT '23: 34th ACM Conference on Hypertext and Social Media · DOI: 10.1145/3603163.3609057 · License: © Copyright held by the owner/author(s). Publication rights licensed to ACM.
Authors: Shreya Ghosh, Prasenjit Mitra
Keywords: Misinformation, social network, discourse analysis.
Session: Social and Intelligent Media: Through the mirror of social media
Conference: HT '23
Abstract
The proliferation of social media has intensified the necessity for automated misinformation detection. Existing methods often struggle with early detection, as key information is not readily available during the initial dissemination stages. In this paper, we introduce a novel model for early misinformation detection on social media by classifying information propagation paths and leveraging linguistic patterns. Our model incorporates a causal user attribute inference model to label users as potential misinformation propagators or believers. Designed for early detection, the model includes two auxiliary tasks: forecasting the scope of misinformation dissemination and clustering similar nodes (users) based on their attributes outperforming the current state-of-the-art benchmarks.
1 INTRODUCTION
Misinformation on social media platforms presents significant challenges to society, as it can sway public opinion [1], intensify polarization [21], and even endanger public health [23]. To lessen the impact of misinformation, early identification of false information is vital, followed by the application of targeted, effective countermeasures [26]. The detection of false information on social media is inherently difficult due to several factors, particularly
Prasenjit Mitra pmitra@psu.edu The Pennsylvania State University State College, PA, USA L3S Research Center, Leibniz University Hannover, Germany
when focusing on early detection. First, misinformation is deliberately crafted to mislead readers, making content-based identification challenging. Secondly, social media data is vast, multi-faceted, predominantly user-created, sometimes anonymous, and often cluttered, which makes detection more complex. Lastly, social media platforms facilitate inexpensive and rapid distribution of news, enabling both accurate and misleading information to proliferate quickly and widely across complex networks. This fast propagation contributes to the challenge of identifying and containing fake news in the early stages. Current methods mainly rely on linguistic patterns [8, 30] and external knowledge bases [18] to identify misinformation, which is insufficient for capturing the intricate interactions and user behaviors driving its spread. Moreover, existing approaches are often constrained by their dependence on singletask learning, which can impede generalizability and robustness across various domains and stages of misinformation dissemination. To address these limitations, our proposed framework (LIEALERT1) combines advanced NLP techniques to discern linguistic differences between false and true information, a graph convolutional network to capture user interactions and propagation features, and attention mechanisms to detect both linguistic patterns and network attributes that characterize misinformation spread. Utilizing a multitask learning approach, our model concurrently addresses the main task of misinformation detection and auxiliary tasks of predicting propagation depth and clustering users based on their reactions to false information. Our contributions (Fig. 1) can be summarized as follows:
• We propose a novel multi-task learning framework that captures both linguistic patterns and network features to effectively detect misinformation, predict its propagation depth, and cluster users based on their reactions to false information. • We introduce a user causal inference model to identify user’s contribution to false information propagation or prevention, and a dynamic attention mechanism that weighs the importance of tokens in the text according to their significance in the misinformation dissemination process, providing a more refined understanding of the linguistic patterns involved in the spread of misinformation. • We demonstrate the efficacy of our framework through extensive experiments on multiple datasets (Anti Vax, Fake News Net, Pheme, Constraint, Russia Ukraine), comparing its performance against baseline models and showcasing
1LIEALERT: Leveraging Information Exchange Ana Lysis for Early Recognition of Misinformation in Tweets"
its generalizability across different domains and stages of misinformation spread. The specific research questions, we are interested to investigate are: RQ: How can we incorporate the user causal model and retweet, follow, and mention network features into a unified framework for early detection of misinformation on social media platforms? How does multi-task learning contribute to enhancing the efficacy of early misinformation detection? How can we leverage user behavioral patterns, such as the frequency and timing of posts, to predict potential misinformation spreaders and better understand their intentions? How can we incorporate real-time, dynamic adaptation mechanisms in misinformation detection models to account for evolving trends and the emergence of new misinformation topics?
Problem Definition In the context of a social network S, denoted by a graph G(V, E) where Vsignifies the set of nodes (users) and Erepresents the set of edges (connections), and given a collection of microblog posts T, let each post t∈Tpossess a timestamp ttand be categorized as either misinformation (M) or factual/accurate/authentic information (F). LIEALERT tackles the subsequent challenges:
(1) Prompt Detection of False Information (PDFI): For every post t∈T, execute a classification task C(t) ∈M, F, reducing the classification duration, tcl, such that tcl≤θ, where θdenotes the maximum acceptable time for prompt identification. Performance is assessed by the F1 score within the bounds of [φ, 1] (i.e., the minimum time needed to achieve at least φF1). This is further enhanced by:
(a) For every post t∈Twith a timestamp ttand labeled as misinformation (M), analyze the temporal dynamics to derive patterns in the frequency, intensity, and origin of misinformation over a specified period. (b) Identify potential sources or originators of misinformation. Using the social network graph, trace back the propagation of misinformation posts to their possible origins.
(2) Categorizing Users (CU): For each user u∈V, allocate a label L(u) ∈Mpropagator, Mcounteractor, Moriginator, Mdoubteraccording to their involvement in the dissemination or counteraction of misinformation. In summary, determine the influence score I(u) for every user u∈Vbased on their connections and frequency of post interactions. Determine how user influence affects the spread of misinformation.
(3) Forecasting Misinformation Spread Depth (FMSD): For every misinformation tweet m∈M, approximate the spread depth d(m,tp) within the social network G(V, E) at a future moment tp, where tp= tt+ Δt, and Δtconstitutes the prediction interval. Subsequently, we group users in Vinto 4 clusters based on their CU labels and attributes. This is further enhanced by:
(a) Assess the social network to determine nodes or clusters more susceptible to misinformation. This could involve analyzing the frequency of misinformation posts, the connectivity of nodes, and the influence scores of users.
(b) Examine how the spread of misinformation affects the temporal dynamics of the social network, such as the creation or removal of connections, user engagement, and overall network health. The multi-objective learning problem addresses the goals (PDFI, CU, FMSD) by considering the linguistic patterns of the posts, the temporal dynamics of social networks, and information propagation.
2 RELATED WORKS
Various studies have examined the use of linguistic features, such as n-grams, syntactic and semantic patterns, and sentiment analysis, combined with machine learning algorithms for detecting misinformation [20, 28]. These approaches typically depend on feature extraction and selection techniques to pinpoint discriminative patterns in textual data. However, they face difficulties in capturing intricate user interactions and the dynamic nature of misinformation spread on social media platforms. Various deep learning models including recurrent neural networks (RNNs), long short-term memory (LSTM) networks, and transformers, have been utilized for misinformation detection [12, 17]. These models can learn highlevel semantic representations from large-scale text data, enhancing detection performance. Nonetheless, they often neglect the significance of network features and user behaviors, which are essential for comprehending the dissemination process of misinformation. Ruchansky et al.[17] introduced a hybrid model called CSI (Capture, Score, and Integrate) for detecting fake news on social media. This study demonstrates that combining textual, publisher, and user interaction features can effectively detect fake news. Ezeakunne et al.[5] concentrated on detecting misinformation by analyzing user behavior on social media. They proposed a deep learning model that uses user behavior patterns, including retweeting, liking, and replying to tweets, to predict the credibility of information. Ma, et al.,[12] devised an attention-based recurrent neural network (RNN) model for detecting fake news on social media. They showed that the attention-based RNN model surpassed other state-of-the-art methods in early misinformation detection. However, these methods have limitations. The CSI model’s[17] performance is constrained by the quality and availability of publisher credibility data and user engagement features, while the BEHIND model [5] depends on user behavior patterns, which may be vulnerable to manipulation by malicious actors or bots. The attention-based RNN model [12] has difficulties detecting misinformation when textual features alone are inadequate or ambiguous. Yang, et al.,[25] employed linguistic cues and user features for early detection of rumors, but the approach might be limited by language-specific attributes and evolving user behaviors. Volkova et al.[22] focused on identifying truthful versus deceptive news headlines using linguistic analysis, but their model has trouble with misleading headlines that are factually correct. Rashkin, et al.,[16] concentrated on identifying truthful versus deceptive news headlines using linguistic analysis, but the model could have issues with misleading headlines that are factually correct. Monti, et al.,[13] suggested a geometric deep learning approach for detecting misinformation, but the model’s performance may be constrained by the structural complexity and scale of real-world social networks.
Liu, et al.,[10] introduced a novel deep neural network that combines crowd response features and user reactions for effective early detection of misinformation. Liang, et al.,[24] proposed a model that incorporates stance information from users to enhance fake news detection, but the model’s performance might be limited by the availability and quality of user-generated stance information. Existing works face limitations in capturing intricate user interactions, dynamic nature of misinformation spread, and reliance on specific features like publisher credibility, user behavior patterns, or language-specific characteristics. Our proposed work addresses these limitations by introducing an innovative multi-task learning framework that combines linguistic patterns, network features, and user causal inference models, enabling a more nuanced understanding of the misinformation dissemination process.
3 LIEALERT: METHODOLOGY
3.1 Network construction
Let G= (V, E) symbolize a directed, weighted multi-layer network, where Vrepresents the set of nodes (users) and Edenotes the set of edges (user interactions) across the three layers. The tri-layered graph is depicted by G3L= (V3L, E3L), where (i) V3L: the amalgamation of VRT, VM, and VF, signifying the set of users engaged in retweets, mentions, or follows. (ii) E3L: the set of edges (u, v,k) traversing the three layers, with k∈RT, M, F, where (u, v, RT) indicates a retweet interaction, (u, v, M) embodies a mention interaction, and (u, v, F) typifies a follow interaction. (iii) w(u, v,k): the weight of edge (u, v,k), illustrating the interaction strength between users u and v in layer k. Subsequently, to incorporate user
credibility within the tri-layered graph, we deploy edge weights based on a personalized Page Rank trust model [2]. The rationale behind this method is that if user B predominantly disseminates untrustworthy content, user A, who interacts with user B, is also likely to share unreliable content. This trust model enables us to capture the significance of channels through which misinformation or accurate information propagates. Here, the personalized Page Rank trust model calculates a trust score for each user contingent on their credibility. Let T(u) represent the trust score of user u. To compute T(u), we employ a personalized Page Rank algorithm [11] with a preference vector that favors users deemed credible by fact-checking organizations or other dependable sources. Utilizing the trust scores T(u) and T(v) for users uand v, we revise the edge weights in the tri-layered graph as follows: w(u, v, RT) = T(u) ∗T(v) ∗NRT(u, v), where NRT(u, v) signifies the number of instances user u retweets content from user v. w(u, v, M) = T(u) ∗T(v) ∗NM(u, v), where NM(u, v) indicates the number of occurrences user u mentions user v. w(u, v, F) = T(u) ∗T(v), as user ueither follows or does not follow user v, and we contemplate the trust scores of both users to allocate the weight. This weight assignment strategy considers the credibility of both users involved in the interaction, rendering the graph more informative for early misinformation detection. Time-Decay Function: One of the inherent attributes of social interactions is their temporal nature. In the context of a multilayered social network, interactions that occurred recently are generally more indicative of current user behavior and the ongoing dissemination of information. To capture this, we introduce a time-decay function that assigns a diminishing weight to interactions as they become older, emphasizing the significance of more recent events.
Figure 1: LIEALERT: building blocks
Let δ(t) represent this time-decay function, where tdenotes the elapsed time since the interaction took place. It is represented as:
δ(t) = e−λt (1)
Here, λis a decay rate constant that determines the rate at which past interactions lose their significance. A higher value of λimplies that interactions lose their relevance faster. Integration with Edge Weights: To embed the time-decay function into the edge weights of our tri-layered graph, we modify the weight computations as follows:
(1) Retweets: Given that t RTis the time since the most recent retweet from user uto user v:
w(u, v, RT) = T(u) ×T(v) × NRT(u, v) × δ(t RT) (2)
(2) Mentions: With t Mbeing the time since the most recent mention from user uto user v:
w(u, v, M) = T(u) ×T(v) × NM(u, v) × δ(t M) (3)
(3) Follows: Since a follow action is usually a singular event, we consider t F, the time since user ubegan following user v:
w(u, v, F) = T(u) ×T(v) × δ(t F) (4)
This time-decayed edge weight assignment is of particular importance when studying the propagation of information. Current interactions can indicate trending or recent information, while older interactions reflect past behaviors and interests. By giving more weight to these recent interactions, we can more accurately capture the current state of information flow within the network. With the inclusion of temporal dynamics, our tri-layered network becomes a more accurate representation of ongoing user interactions and behavior. Especially in the context of early misinformation detection, this enhancement plays a crucial role. Misinformation, by its nature, is often sporadic and can trend quickly; thus, relying solely on historical data might not always be the most effective strategy. By emphasizing recent interactions, the network can better detect and respond to emerging misinformation trends. Moreover, the combination of user trust scores with temporal dynamics means that not only the credibility of the users is taken into account, but also the timeliness of their interactions. This layered approach ensures a more holistic understanding of the information propagation, making the system robust against both historical and emergent threats. Characteristics of Transmitters and Receivers in Misinformation Dissemination Transmitter: Transmitters constitute individuals who initiate and disseminate information on social media, often voicing a supportive stance towards misinformation. We identify the following salient characteristics that play a pivotal role in the propagation of misinformation:
• Response time: The alacrity with which a user shares or propels encountered information.
• Tenacity: Persistence in dispersing information despite challenges or delays in convincing their audience. This persistence can manifest in various intensities, ranging from occasional shares to relentless dissemination by super-spreaders.
• Influence level: A user’s follower count, coupled with their relevance to specific domains (e.g., healthcare or politics),
determines their potential reach and impact in the misinformation ecosystem.
• Emotional Resonance: The degree to which a transmitter emotionally aligns with the misinformation, influencing its tone and vehemence. High emotional investment can result in more compelling, albeit misleading, narratives. Receiver: Receivers are the audience members who consume and may subsequently relay information. LIEALERT delineates various factors that govern the likelihood of receivers proliferating misinformation:
• Disposition: Depending on their pre-existing beliefs and mindset, receivers might quickly adopt a particular stance, require more time and evidence to be persuaded, or remain steadfastly unresponsive.
• Message Frequency: Repetitive exposure to the same misinformation, either from a singular source or multiple users, can boost its perceived authenticity and hence its chances of being disseminated.
• Source credibility: The transmitter’s reputation, evident from their follower count or established expertise in a domain, serves as a primary cue for trustworthiness.
• Discernment: Individual cognitive capabilities influence how misinformation is interpreted. The spectrum spans from immediate acceptance and sharing, neutrality, to skepticism and active counter-propagation.
• Network Topology: The interconnectedness and strength of connections in a receiver’s social graph can modulate the influence of incoming misinformation.
• Confirmation Bias: Receivers are often inclined to accept information that aligns with their existing beliefs, making them susceptible to congruent misinformation. While the transmitter actively engenders the information flow, the receiver, although apparently passive, plays a decisive role in the subsequent dissemination dynamics. LIEALERT discerns these characteristics individually, subsequently amalgamating the insights in the PDFI and FMSD models, thereby devising sophisticated early detection algorithms. This integrated approach equips us with robust strategies to stymie the proliferation of misinformation across digital platforms. We provide two examples in the context of misinformation propagation in Table 12. In the example, the transmitter is responsible for disseminating misinformation, while the receiver, depending on their attitude and susceptibility, may contribute to the further spread of the information.
3.2 User causal model
We propose a Causal User Attribute Inference (CUAI) model, which employs a Graph Attention Network (GAT) to deduce the causal associations between user attributes and their inclination to propagate misinformation. Let G= (V, E) symbolize the social network graph. Each user v∈Vpossesses an associated attribute vector A(v), consisting of features such as response time, tenacity, influence level, discernment, and source credibility. The CUAI model encompasses the following components:
2the actual usernames in the datasets have been changed due to ethical and privacy issues
Dataset Post
3.2.1 Multi-faceted Attribute Embedding Layer: This layer captures both linear and non-linear relationships in the user’s attributes. By modelling these relationships, the system can recognize intricate patterns that might be associated with misinformation spreaders. For instance, a user often spreads accurate news (an attribute), however, occasionally, when they encounter a sensational piece of news, they tend to spread it quickly without verification. A linear relationship might miss this nuance, but the non-linear component in the embedding can catch this inconsistency, flagging the user’s posts for potential verification. We employ a dual-path transformation mechanism that captures both linear and non-linear patterns in the attributes. We represent this as:
h0(v) = W0 ∗A(v) + b0 + σ(W1 ∗A(v) + b1), (5)
where σis a non-linear activation function, W0, W1 are weight matrices, andb0,b1 are bias vectors. This approach provides a richer initial embedding of the user attributes.
3.2.2 User Interaction History Encoder: To further enhance the model, we introduce a User Interaction History Encoder that captures the historical behavior of users in the network. By considering the historical behavior of users, we gain insights into their patterns of information dissemination. Users with a history of sharing misinformation can be scrutinized more closely. A user might have a history of sharing controversial posts every few months. Even if they’ve been dormant or sharing accurate information recently, the model, remembering their historical behavior, would be more vigilant and would prioritize verifying their claims. This component employs a Recurrent Neural Network (RNN) to encode the interaction history of users, providing additional context for the causal inference process:
h H(v) = RNN(h L(v,t) : t∈Tv), (6)
where Tvdenotes the collection of time intervals for user v’s engagements, while h H(v) denotes the historical behavior embedding of user v. By integrating these historical behavior embeddings with the GAT-based causal inference in the subsequent stage, we acquire a more thorough representation of users’ characteristics and their impact on the spread of misinformation.
3.2.3 GAT-based Causal Inference: 3 We introduce a novel Temporal multi-head Graph Attention Network (T-GAT) layer. This approach accounts for the temporal dynamics of user interactions and historical behavior, providing a more comprehensive understanding of users’ misinformation propagation patterns.
3interactions between users and their neighbors in the social network graph to capture the underlying causal structure.
Temporal Graph Attention Network (T-GAT) Layer: This temporal attention mechanism captures the changes in users’ behavior and interactions over time, allowing the model to adapt to evolving misinformation propagation patterns. This component is designed to understand both short-term and long-term interactions between users, adjusting its focus based on the recency of the interactions. Suppose a group of users suddenly start interacting frequently and spreading similar pieces of information. The model, noticing this sudden burst of activity, could interpret this as a coordinated misinformation campaign and flag their shared content for review. T-GAT mechanism integrates a multi-scale temporal convolution operation to capture both short-term and long-term interactions. Formally:
hl+1(v,t) = ||K k=1Temporal Attentionk(hl(v,t), {hl(u,t) :
u∈N(v)}) ∗Convtemporal(hl(u,t)), (7)
where || signifies concatenation, K is the number of attention heads, N(v) denotes the neighbors of user v, and Attentionk() is the kth attention mechanism, Temporal Attentionk() represents the kth temporal attention mechanism, and tdenotes the time step. The convolution operation, Convtemporal(), enables the model to recognize patterns over varying time horizons. The temporal attention mechanism computes the importance of neighboring nodes’ features at different time steps. For the temporal attention mechanism, we introduce a decay factor, ensuring that older interactions have reduced influence:
πk(u, v,t) = softmaxu(Leaky Re LU(WT k[hl(u,t)||hl(v,t)||e−αΔtu,v])), (8) where αis a decay coefficient and eis the exponential function. This addition makes the model’s attention mechanism adaptive to the recency of interactions. Δtu,vsignifies the time difference between user uand user v’s interactions. This temporal attention mechanism enables the CUAI model to consider both static and dynamic features, thus offering a more robust and comprehensive understanding of users’ behavior.
3.2.4 Causal Effect Estimation: We use the learned user embeddings h L(v), where L is the number of GAT layers, to estimate the causal effects of user attributes (transmitter, receiver characteristics described above and network centrality, topic expertise4) on misinformation propagation. For each user v, we define two potential outcomes (i.e., what is their propensity to propagate misinformation): Yv(1) and Yv(0) based on whether the user attribute was set to a specific value (1) or not set (0), respectively. The causal effect for
4Thresholds used in this study for each attribute and discretizing the continuous or categorical values are provided in supplemental document here
Russia-Ukraine War (Transmitter) Propaganda Expert shares a fabricated news story on Twitter, stating, “Breaking: Ukraine launches unprovoked attack on Russian territory! #Ukraine Aggression #Russia Under Attack" (Receiver 1) Patriotic Citizen shares the post with the comment, “Unbelievable! We must defend our country! #Support Russia #No More Ukraine Aggression" (Receiver 2) @Doubtful Reader comments, “Is this true? I haven’t seen any credible news sources reporting this yet. Can anyone confirm? #Fact Check #War News"
Table 1: Illustration of receiver and transmitter in Russia-Ukraine war misinformation scenario.
user vis then defined as the difference between the two potential outcomes:
τ(v) = E[Yv(1) −Yv(0)], (9)
where E[] signifies expectation. To estimateτ(v), we use the learned user embeddings (the historical behavior embeddings, h H(v), and the temporal GAT layer embeddings, h L(v,t)) and fit two separate regressions:
Yv(1) = g1(h H(v),h L(v,t);θ1), Yv(0) = g0(h H(v),h L(v,t);θ0), (10) where g1(;θ1) and g0(;θ0) are regression functions with parameters θ1 and θ0, respectively. We can then estimate the causal effect (how much the user’s attribute contributes to the spread of misinformation) as the difference between the predicted outcomes:
τ(v) ≈g1(h L(v);θ1) −g0(h L(v);θ0). (11)
Our causal effect estimation comprises two levels. At the first level, we gauge the immediate impact of user attributes on their behavior, while at the second level, we assess their aggregated effect on the misinformation ecosystem. For the immediate impact, we stick to:
τ(v) = E[Yv(1) −Yv(0)], (12)
To realize the aggregated effect, we employ a meta-learning approach, where the model observes multiple users and their associated causal effects to derive a meta-impact score:
μ(v) = Meta Learn(τ(vi) : vi∈V), (13)
where Meta Learn() is a meta-learning function that maps individual causal effects to a holistic understanding. By deciphering both immediate and aggregated impacts, we not only detect misinformation sources but also strategize on largescale interventions, ensuring a more resilient social network against misinformation. Individually, a user might seem harmless, with their misinformation having a low immediate impact. However, in the aggregate, if many such users are present, they can form a significant misinformation wave. The model can detect these macrolevel patterns and might, for instance, recommend a platform-wide awareness campaign to counteract the misinformation. Apart from effective misinformation detection, by estimating the causal effects, we can also identify the most influential user attributes that can be targeted for interventions to reduce misinformation propagation.
3.2.5 Misinformation Propensity Prediction: To discern whether a user vwill likely propagate misinformation, we exploit multi-source embeddings: (1) learned user embeddings h L(v,t), (2) historical behavior embeddings h H(v), and (3) contextual embeddings h C(v) that capture the surrounding context of user activities. 1. Contextual Embeddings: To enhance the prediction capability, we incorporate the context in which the user operates. These embeddings capture the overall theme of the user’s posts, their frequent interactions, and sentiment trends. Defined as:
h C(v) = Contextual Encoder(Content(v)), (14)
where Content(v) symbolizes the textual and interaction data associated with user v.
Now, to form a comprehensive representation, we concatenate the embeddings:
htotal(v,t) = [h L(v,t)||h H(v)||h C(v)], (15)
where [||] represents the concatenation operation. 2. Classification with Attention Mechanism: We believe that not all parts of the concatenated embeddings contribute equally to the prediction. Hence, an attention mechanism α() is introduced to weigh different portions of the embeddings:
hatt(v,t) = α(htotal(v,t)) ⊙htotal(v,t), (16)
where ⊙stands for element-wise multiplication. Given the ground truthy(v) ∈{0, 1}, where 1 signifies a misinformation propagator and 0 otherwise, the classification loss function is enhanced with an attention regularization term Ratt():
L(θ) = ∑︁
v∈V Lcls(y(v), f(hatt(v,t);θ)) + λR(θ) + γRatt(α), (17)
where λand γare hyperparameters that manage the trade-offs. The trained classifier predicts the misinformation inclination as:
P(y(v) = 1|hatt(v,t)) = f(hatt(v,t);θ∗). (18)
By amalgamating learned user embeddings, historical behavior, and the context in which they operate, the refined CUAI model offers a superior lens into the users’ propensity to spread misinformation. This paves the way for surgical interventions, ultimately guarding the integrity of information across the social network.
3.3 Temporal Charecteristics
Observation (Anti Vax Dataset) A tweet claiming that the MMR vaccine is linked to autism initially receives retweets and likes from users who agree with the statement. As the tweet spreads, users who express their surprise at such a claim, asking for evidence or research supporting the claim. Further down the line, users begin to question the claim’s validity and ask for reliable sources, engaging in conversations to debunk the misinformation.
3.3.1 Dynamic Attention Value. We propose a novel method that incorporates a dynamic attention value for each post, considering the time interval since the original post, the interaction ratio, and the reachability of the post in the follower-followee network. Let P be a post and t(P) be the time when the post was made. Let Ebe the event corresponding to the initial tweet, and t(E) be the time when the event started. We define the time interval Δt(P) for a post as: Δt(P) = t(P) −t(E). Let N(P) be the set of nodes (users) that have already interacted with post P. Then, we define the interaction ratio R(P) as: R(P) = |N(P) |
V(G) | . Now, let F(G) be the follower-followee network of the users in G, and L1(P) be the set of nodes reachable from Pusing a BFS search algorithm. We define the BFS ratio L(P) as: L(P) = | L1(P) |
V(F(G)) | . The dynamic attention value A(P) for a post P can be calculated as a weighted sum of the time interval, interaction ratio, and BFS ratio: A(P) = α∗Δt(P) + β∗R(P) +γ∗L(P), where α, β,γare weights that can be tuned based on the importance of each factor in the dissemination process. Additionally, we also introduce a dynamic attention mechanism to capture the changing sentiment and interactions over time. |
Sentiment Trend Shifts: Let S(P) represent the sentiment score of post P, computed using sentiment analysis tools. A shift in sentiment ΔS(P) is calculated using a moving average over a predefined window of posts succeeding P. The enhanced dynamic attention value A(P) for a post Pnow becomes:
A(P) = α∗Δt(P) + β∗R(P) + γ∗L(P) + δ∗ΔS(P)
where δis a tunable weight corresponding to the sentiment trend shift.
3.3.2 Post Score. Let S(P) represent the linguistic pattern score for the post P. The overall score for a post P, considering both dynamic attention and linguistic patterns, can be calculated as: Score(P) = γ∗A(P) + (1 −γ) ∗S(P), where γis a weight that balances the influence of dynamic attention and linguistic patterns in the model.
3.3.3 Propagation Path Construction. For a specific news story spreading across social media, we initially establish its dissemination trajectory by pinpointing the users who actively participated in circulating the news.
Influence Score. For each user i, we compute an influence score I(i) based on their followers’ activities related to the news story.
User profiling: We convert user profiles into fixed-length sequences. Let Uibe the fixed-length sequence representing the user profile of user i. For each user i, we create a propagation path Pi that consists of the user profile sequence Uiand the interactions in which user iparticipated. Learning Representations: In layer 1, we apply a Gated Recurrent Unit (GRU) layer to learn the vector representation Vifor each propagation path Pi. Incorporating the influence scores, we adjust the GRU layer to include the influence values: Vi= GRU(Pi, I(i))
In layer 2, we deploy a Graph Convolutional Network (GCN) layer to learn the transformed propagation path Tifor each user profile vector Vi: Ti= GCN(Vi). Concatenation: We concatenate transformed propagation paths by combining all propagation paths Tiinto a single vector C= Concat(T1, T2, . . . , Tn). Prediction: Finally, we deploy a multi-layer feedforward neural network to predict the maximum depth Dfor the corresponding propagation path: D= FNN(C). By incorporating user profiling and propagation paths into the model, we can more effectively capture the propagation patterns observed in the information dissemination process and improve the early detection of misinformation. This approach allows us to account for the impact of individual users on the overall propagation of news stories and to better understand the dynamics of misinformation spread.
3.3.4 Loss Functions. LIEALERT uses a binary cross-entropy loss Lmisinfoas the loss for the primary task of detecting false information. Ldepthis the loss for the auxiliary task of predicting the depth of false information propagation. This is a mean squared error (MSE) loss. Lclusteris the loss for the auxiliary task of clustering users based on their reactions to false information. This is a categorical cross-entropy loss, with clusters represented as one-hot encoded vectors. LIEALERT is enhanced with an additional auxiliary task, Linfluence, capturing the loss for predicting user influence in the
propagation process. This loss uses a hinge loss formulation due to its margin-based nature. The total loss function, Ltotal, becomes:
Ltotal= γ1 ∗Lmisinfo+γ2 ∗Ldepth+γ3 ∗Lcluster+γ4 ∗Linfluence Task-specific weighting factors, γ1,γ2, γ3 and γ4 control the relative importance of the tasks in the joint loss function.
3.3.5 Adaptive Weighting Factors. We also introduce adaptive weighting factors that dynamically adjust the importance of each task during training. We use the inverse training progress as a weight factor: γi(t) = αi
1+βi∗tr. Here, tris the current training step, αiand βiare positive hyperparameters for each task i, and γi(t) is the weighting factor for task iat step t. This formulation ensures that the weighting factors decrease as training progresses, allowing the model to focus on the most relevant tasks at each stage of training.
3.4 Linguistic pattern analysis
LIEALERT utilizes three critical components based on linguistic pattern analysis: (A) Semantic Similarity Analysis (SSA): This method is developed using a pre-trained BERTweet model [14], a multi-layer attention mechanism, and contrastive learning to enhance early misinformation detection. We use the Hugging Face Transformers and spa Cy libraries in Python5. We use the BERTweet model6, denoted as B. A multi-layer attention mechanism A, consisting of Lself-attention layers, each followed by a feed-forward network and layer normalization, is applied. The attention mechanism calculates a weighted representation using the output from the last transformer layer in BERTweet. To enhance the model’s capacity to discern fact from misinformation, we introduced Source Credibility Estimation, where for each reference source rj, a credibility score is assigned based on its historical accuracy. Next, we create a training dataset of paired examples, each pair containing a tweet tiand a reference source rj. For each tweet, we generate positive pairs (ti,r+ j) with verified information sources sharing semantic similarity and negative pairs (ti,r− j) with unrelated or contrasting sources. We fine-tune the Ro BERTa model with the multi-layer attention mechanism on the paired dataset using contrastive learning. The goal is to learn semantic embeddings φ(ti) and φ(rj) that minimize the distance between positive pairs and maximize the distance between negative pairs:
Lcontrastive= ∑︁ i= 1N d(φ(ti),φ(r+ j))−α+max r− j d(φ(ti),φ(r− j)) +,
(19) where d(·, ·) denotes a distance metric (e.g., cosine distance), αis a margin parameter, [·]+ represents the hinge function, and Nis the number of tweets in the dataset. (B) Argument Mining and Logical Fallacy Detection (AMLF): This module is developed to identify argumentative structures and logical fallacies in textual data. Given a dataset Dcontaining text samples ti, our objective is to extract argument components, such as claims Ci, premises Pi, and conclusions Qi. Let Fextdenotes an
5The preprocessing stage involves the spa Cy library with the "encorewebsm" language model to tokenize input text, remove URLs, special characters, and user mentions using regular expressions, and convert text to lowercase. 6https://huggingface.co/vinai/bertweet-base
extraction function, parameterized by a pre-trained BERTweet-base, which is fine-tuned for argument component extraction. We use the dataset [9] for fine-tuning. The extraction process is defined as follows: (Ci, Pi, Qi) = Fext(ti), where tiis a text sample from the dataset D. For each extracted argument component, we aim to identify argumentative relations Rijbetween them, such as support, attack, or neutral. Let Freldenote a relation identification function, parameterized by a pre-trained NLP model fine-tuned on an argument relation dataset. The relation identification process can be defined as: Rij= Frel(Ci, Pj), where Ciand Pjare argument components extracted from the text samples. Next, to detect logical fallacies, we define three fallacy patterns F = f1, f2, f3, such as ad hominem, straw man, or false cause. We aim to recognize and classify these patterns in argumentative structures. Given the argumentative relations Rij and the argument components (Ci, Pj, Qi), the fallacy detection process can be defined as: fk= Ffallacy(Rij, Ci, Pj, Qi), where fk is a fallacy pattern from the set F . The function Ffallacyassigns a fallacy pattern label to each argumentative structure based on the identified relations and components. Some examples of tweets containing logical fallacies from the Anti Vax dataset are provided in Table 27. Specifically, Ad Hominem examples attempt to discredit an individual based on their nationality or potential vested interests, rather than discussing the content or validity of their statements. The posts (Straw Man) misrepresent the views of Ukrainian officials and COVID-19 vaccine advocates by focusing on outliers or exaggerations to create a more easily defeated version of their stance. Finally, examples of False Cause incorrectly assume that a temporal relationship (events following one another) implies causality. (C) Emotion-aware Sentiment Analysis (ESA): The objective is to classify sentiment while recognizing specific emotions. We’ve integrated an Emotional Shift Detector that identifies drastic changes in emotional tone within a text sequence, a potential indicator of misinformation or intentional emotional manipulation. Given a dataset Dcontaining text samples ti, our objective is to classify the sentiment expressed in each text as positive, negative, or neutral, while also considering the specific emotions conveyed in the text. We extract the following features from the emotionaware sentiment analysis: 1. Sentiment Polarity Score: Calculate a sentiment polarity score Pifor each text sample, indicating the degree of positivity or negativity expressed in the text. 2. Emotion Intensity Score: Compute an emotion intensity score Eifor each text sample, representing the strength of various emotions (e.g., anger, joy, fear, or sadness) expressed in the text. 3. Subjectivity Score: Uiis the level of personal opinion, emotion, or judgment in the text i. 4. Emotional Context Vector: Viis a multi-dimensional vector that captures the distribution of various emotions in the text i. 5. Stance Confidence Score: Ciis the model’s certainty in the detected stance towards the target in the text i.Our emotion-aware sentiment analysis model is based on Ro BERTa and is fine-tuned on a dataset of labeled misinformation instances. It utilizes the extracted features (sentiment polarity score, emotion intensity score, subjectivity score, emotional context vector, and stance confidence score) as input to the model. It is observed that for true information,
7More samples from the dataset are provided in supplemental document
the sentiment is positive or neutral, and the stance favors a particular viewpoint or report on factual events. Conversely, false information often exhibits negative sentiment and adopts a stance against specific topics, entities, or claims. We fine-tune the BERTweet8
transformer model on a labeled misinformation dataset with sentiment and emotion annotations, using optimized hyperparameters and stratified 5-fold cross-validation for robust evaluation. We set the initial learning rate to 5 × 10−5 and employ a learning rate scheduler with a linear warm-up period of 0.1 of the total training steps, followed by a linear decay. The batch size is set to 16, and we train the model for 3 epochs. We use the Adam W optimizer with a weight decay of 1 × 10−2 and set the maximum sequence length to 512 tokens. To avoid overfitting, we employ dropout regularization with a rate of 0.1 in the transformer layers and a rate of 0.5 in the classification head.
4 EXPERIMENTAL EVALUATIONS 9 LIEALERT is evaluated on three tasks: (1) early detection of misinformation with an accuracy of ≥85%, (2) forecasting the extent of false information dissemination, and (3) categorizing users based on their attributes. Table 4 presents a comprehensive comparison of the LIEALERT system’s performance with other state-of-the-art misinformation detection models10, including CAMI [26], FNED [10], GRU [12], and [27]. The evaluation is conducted across multiple datasets (See Table 3) and timeframes, including Anti Vax (anti-vaccine), CONSTRAINT (COVID-19-related fake news), Fake News Net, PHEME (rumor detection), and RU War (Russia-Ukraine War). LIEALERT consistently outperforms the other models in terms of F1 score across most datasets and timeframes. The highest F1 score achieved by LIEALERT is 0.942, observed in the RU War dataset within a 24hour timeframe. The LIEALERT system demonstrates competitive performance even in the shortest 30-minute timeframe, where it achieves the best F1 score in all datasets except for PHEME. This highlights LIEALERT’s capability to detect misinformation effectively in near real-time situations. Fig. 2(b) depicts the percentage of users in different clusters as returned by CU model. The Anti Vax dataset has the highest percentage of propagators, while CONSTRAINT has the highest counteractors. Russia-Ukraine War dataset exhibits a balanced distribution between propagators and counteractors. Doubters are the smallest percentage across all datasets. Using a manually labelled 1000-tweets (200 for each dataset), CU model shows 92.6% accuracy in categorizing users. The most important takeaway is counteractors play the most significant role (contributing11 44% of false information detection) in identifying false information, as they actively challenge and debunk misinformation, as well as the interaction between counteractors and propagators (contributing 38% of false information detection). The ablation study (Fig. 2(a)) reveals that incorporating the CU model leads to a noticeable improvement in F1 scores across all
8https://huggingface.co/vinai/bertweet-base 9Our codebase and additional information is available in HERE. We will share the full labelled dataset of Russia-Ukraine war in the camera-ready stage. 10Baselines were chosen as state-of-the-art models for detecting misinformation and early detection of misinformation on social media 11When clustering contribution is considered to 100
Fallacy type Post
Dataset Details
datasets. The CU model helps to categorize users based on their involvement in misinformation dissemination or counteraction. By integrating the Semantic Similarity Analysis (SSA), Argument Mining and Logical Fallacy Detection (AMLF), and Sentiment Analysis modules, we were able to pinpoint crucial linguistic features that facilitate the early identification of misinformation and enhance the accuracy by 6%, 11% and 9% respectively. Adding the FMSD
model also contributes to improving F1 scores, indicating that forecasting the spread of misinformation within the social network is beneficial for detection. Finally, it is observed that our full model (Multi-task learning) consistently achieves the highest F1 scores across all datasets, demonstrating that considering multiple objectives in the learning process results in better misinformation detection and outperforming state-of-the-art methods by a significant margin. Our experiment demonstrates that LIEALERT can
Ad Hominem “Why would we believe a report by XXX? She’s Ukrainian; of course, she’s biased against Russia! #War Propaganda" “Don’t listen to Dr. XXX’s advice on COVID-19 vaccines; he probably has stocks in pharmaceutical companies. #Profit Over Health"
Straw Man “Ukrainian officials claim they want peace, but I saw a single photo of an armed civilian. They obviously want war. #Hidden Agenda" “People who support COVID-19 vaccination think it’s a magical shield that makes you invincible to all diseases. So naive! #Reality Check"
False Cause “The sanctions on Russia started, and the global market crashed the next day. Sanctions clearly ruin the world economy. #Sanction Backfire" “Three celebrities got the COVID-19 vaccine last week, and now they have flu symptoms. Clearly, the vaccine caused it. #Vaccine Risks"
Table 2: Illustration of fallacy patterns (Russia-Ukraine War and COVID-19 Vaccination)
PHEME [4] 330 rumor threads collected from Twitter, with each thread having an average of 100 tweets
Anti Vax [6] Anti-vaccination movement over 1.8 million tweets collected between 2019 and 2021.
CONSTRAINT [15] 17,000 English tweets (COVID-19), annotated as either real or fake, with an equal distribution
Fake News Net [19] Data from two fact-checking websites, Politi Fact and Gossip Cop over 23,000 news articles
RU War [3] Tweets from Feb 22,2022,through Jan 8,2023 collected using hashtags related to RU war
Table 3: Five real-life datasets used for LIEALERT’s performance evaluation
Figure 2: (a) Ablation study (contribution of each modules) of LIEALERT (b) User cluster percentage returned by CU model
Anti Vax
Russia Ukraine War
30 min 12 h 24 h 30 min 12 h 24 h
Dataset / Time F1 Score
LIEALERT CAMI [26] FNED [10] GRU [12] [27]
Anti Vax / 24h 0.952 0.861 0.892 0.814 0.820 Anti Vax / 12h 0.920 0.812 0.842 0.683 0.790 Anti Vax / 30m 0.896 0.752 0.803 0.579 0.748
CONSTRAINT / 24h 0.942 0.843 0.866 0.802 0.801 CONSTRAINT / 12h 0.904 0.808 0.832 0.661 0.772 CONSTRAINT / 30m 0.891 0.736 0.791 0.518 0.721
Fake News Net / 24h 0.931 0.848 0.840 0.849 0.810 Fake News Net / 12h 0.890 0.791 0.831 0.790 0.760 Fake News Net / 30m 0.868 0.736 0.784 0.715 0.691
PHEME / 24h 0.925 0.852 0.848 0.867 0.816 PHEME / 12h 0.863 0.768 0.819 0.828 0.772 PHEME / 30m 0.812 0.701 0.723 0.820 0.607
RU War / 24h 0.951 0.758 0.728 0.810 0.711 RU War / 12h 0.880 0.622 0.548 0.692 0.506 RU War / 30m 0.636 0.573 0.501 0.606 0.481
Table 4: F1 score for detecting misinformation (post): comparison with baseline models and proposed framework (LIEALERT).
Timestep Anti Vax Fake News Net PHEME Constraint RU War
Depth Node Depth Node Depth Node Depth Node Depth Node
T/3 0.832 0.795 0.77 0.78 0.75 0.73 0.837 0.828 0.733 0.681 T/2 0.871 0.83 0.806 0.792 0.781 0.735 0.881 0.86 0.814 0.781 2T/3 0.924 0.861 0.83 0.821 0.837 0.796 0.932 0.928 0.950 0.936
Table 5: LIEALERT’s performance (F1) on predicting maximum depth of false information propagation and predict infected nodes
Dataset Configuration Macro F1 Score AUC
Linguistic Pattern (SSA) 0.721 0.736 0.741 0.780 0.825 0.830 Linguistic Pattern (SSA+AMLF) 0.728 0.744 0.756 0.784 0.831 0.845 Linguistic Pattern (SSA+AMLF+ESA) 0.745 0.750 0.761 0.803 0.858 0.862 Linguistic Pattern (Full) + Temporal (DA) 0.818 0.826 0.874 0.830 0.862 0.899 Linguistic Pattern (Full) + Temporal (DA+PPC) 0.835 0.841 0.906 0.842 0.880 0.921 Linguistic Pattern (Full) + Temporal (Full) + CUAI (UI) 0.878 0.890 0.928 0.882 0.906 0.942 Linguistic Pattern (Full) + Temporal (Full) + CUAI (UI+T-GAT) 0.890 0.911 0.943 0.894 0.930 0.954 Linguistic Pattern (Full) + Temporal (Full) + CUAI (UI+T-GAT+MPP) 0.896 0.920 0.952 0.90 0.946 0.961
Linguistic Pattern (SSA) 0.460 0.485 0.502 0.465 0.490 0.521 Linguistic Pattern (SSA+AMLF) 0.468 0.492 0.528 0.472 0.498 0.535 Linguistic Pattern (SSA+AMLF+ESA) 0.538 0.580 0.614 0.540 0.586 0.630 Linguistic Pattern (Full) + Temporal (DA) 0.590 0.725 0.760 0.602 0.780 0.840 Linguistic Pattern (Full) + Temporal (DA+PPC) 0.621 0.796 0.821 0.624 0.845 0.890 Linguistic Pattern (Full) + Temporal (Full) + CUAI (UI) 0.630 0.842 0.900 0.633 0.860 0.918 Linguistic Pattern (Full) + Temporal (Full) + CUAI (UI+T-GAT) 0.634 0.866 0.912 0.637 0.874 0.928 Linguistic Pattern (Full) + Temporal (Full) + CUAI (UI+T-GAT+MPP) 0.636 0.880 0.951 0.640 0.891 0.957
Table 6: Ablation study (leave one out analysis) of LIEALERT. SSA: Semantic Similarity Analysis. AMLF: Argument Mining and Logical Fallacy Detection. ESA: Emotion-aware Sentiment Analysis. DA: Dynamic Attention Value. PPC: Propagation Path Construction. UI: User Interaction History Encoder. MPP: Misinformation Propensity Prediction.
recognize false information within 3-12 minutes with an F1-score of over 0.85 at the earliest, surpassing state-of-the-art models. Our approach also enables predicting the depth of false information spread, assisting in estimating the potential reach and impact of misinformation (See Table 5). The experiment evaluation is carried out by timesteps, including T/3, T/2, and 2T/3, and compares the performance of LIEALERT in predicting the maximum depth of false information propagation ("Depth") and the number of infected nodes/users ("Node") at each timestep. The highest F1 scores for both depth and node prediction are observed in the 2T/3 timestep across all datasets. The performance of LIEALERT in predicting depth is generally higher than its performance in predicting nodes. This is due to the fact that the depth of information propagation is a more stable and well-defined metric, whereas the number of infected users may be influenced by various factors, such as the structure of the network and the behavior of individual nodes. Across the PHEME, RU War, Anti Vax, Fake News Net, and Constraint datasets, we observed that users who regularly interacted with misinformation often expressed strong opinions or emotions, used divisive language, and displayed biased stances. Conversely, users who served as skeptics or fact-checkers typically adopted a more neutral tone and concentrated on providing evidence or counterarguments to refute false claims12.
LIEALERT’s Ablation Study: Table 6 (ablation study) provides an in-depth analysis of the contribution of various modules to the efficacy of the LIEALERT system when applied to two datasets - Anti Vax and Russia Ukraine War. The primary layer SSA, when used as a standalone module (linguistic pattern), shows a base performance for both datasets. It seems to perform better on the Anti Vax dataset than on the Russia Ukraine War dataset, especially at the 30-min interval. This is due to the nature of the content: Anti Vax discourse possesses repetitive linguistic patterns that are easily captured, while the Russia Ukraine War tweets encompass more diverse semantics due to the dynamic and evolving nature of wartime communication. Introducing AMLF offers a marginal performance improvement for both datasets. Logical fallacies are more prevalent in the Anti Vax dataset, reflecting a higher increment in performance compared to the Russia Ukraine War dataset. ESA shows a more prominent effect on the Russia Ukraine War dataset. Given the emotional charge associated with wartime events and sentiments, ESA likely plays a crucial role in identifying the nuances of misinformation in this context. The addition of DA, especially to the Russia Ukraine War dataset, indicates the importance of temporal patterns in misinformation spread during evolving events. The substantial performance jump, especially between 12h to 24h, suggests that over time, the temporal dynamics of tweets play a pivotal role in distinguishing misinformation. Propagation Path Construction (PPC) component further augments the performance, with the 24h interval showing the most significant rise for both datasets. This underscores the importance of understanding how information disseminates over time and through networks. UI seems vital for both datasets but is particularly significant for the Russia Ukraine War dataset. Given that users’ historical behavior might indicate their inclination towards sharing misinformation,
12Data was obtained from publicly available sources. We took measures to ensure the privacy and anonymity of the individuals whose data was used.
this module’s relevance is underscored in the dynamic landscape of wartime tweets. A further boost in performance is observed with T-GAT’s inclusion, more prominently in the 12h and 24h intervals. This suggests that graph-based attention mechanisms effectively capture user interactions in identifying misinformation. The final configuration, which adds MPP, yields the highest results across the board. Interestingly, for the Russia Ukraine War dataset, there’s a considerable increment in the 24h interval, indicating that predicting misinformation propensity becomes especially effective as more data becomes available over time. The Anti Vax dataset seems more amenable to linguistic patterns, while the Russia Ukraine War dataset leans heavily on temporal and CUAI analyses. This indicates that static misinformation topics might be more predictable linguistically, whereas dynamic events necessitate an understanding of temporal progression and user interactions. The difference in performance at the 30-min interval between the two datasets, especially for the final configuration, suggests that early detection of misinformation in dynamic events like wars is challenging. However, as time progresses, the system becomes more adept, due to accumulating more user interactions and propagation paths that help discern misinformation patterns. In conclusion, the ablation study of LIEALERT highlights the major challenges in early misinformation detection and contributions of each modules. While linguistic patterns lay the groundwork, understanding the temporal dynamics and user interactions become indispensable, especially in rapidly evolving scenarios.
LIEALERT’s Multi-lingual Support: In order to make our framework globally applicable and receptive to diverse linguistic patterns, we incorporate Tw Hin BERT [29], a specialized BERT model trained on 7 billion Tweets from over 100 distinct languages. This integration ensures robustness when dealing with multilingual or codeswitched content, which is commonplace in regions where multiple languages coexist. Given the significant multilingual audience on social media platforms, the ability to handle mixed language data effectively becomes paramount, especially when discerning factual information from falsehoods. Our implementation initializes the pre-trained Tw Hin BERT model13, and we fine-tune it on our training data, mirroring the strategies we employed for our primary models. By leveraging this model, our framework becomes capable of understanding and processing multilingual tweets. Evaluation on Ban Fake News Dataset: The Ban Fake News dataset [7] consists of annotated ≈50K news that can be used for building automated fake news detection systems for a low resource language like Bangla, making it an apt choice for evaluating our framework’s efficacy in a multilingual context. For our evaluation: (1) We preprocess the dataset similarly to our primary data, tokenizing and cleaning the text samples. (2) The fine-tuned Tw Hin BERT model is then applied to the preprocessed dataset, and results are generated based on the three main components of our framework: SSA, AMLF, and Emotion-aware Sentiment Analysis. The results on the Ban Fake News dataset are promising (accuracy: 0.876 (fact) and 0.921 (misinformation)), showcasing our framework’s ability to generalize across different languages. Further details on the evaluation metrics and results are omitted since the major focus of this paper is early misinformation detection.
13https://huggingface.co/Twitter/twhin-bert-base
5 CONCLUSION
In this paper, we introduced a holistic approach to early misinformation detection on social media platforms by incorporating user profiling, linguistic analysis, and network analysis modules. Our proposed framework, LIEALERT, seeks to pinpoint and flag potential misinformation during the initial stages of its dissemination to reduce its spread and societal impact. The promising results achieved by LIEALERT in identifying misinformation emphasize the significance of integrating diverse linguistic features and network attributes for developing a reliable and efficient detection system. In future, we will investigate dynamic assessment of misinformation and temporal intervention techniques to mitigate misinformation propagation. Further, we will also explore explainability of our early misinformation detection system, LIEALERT.
Acknowledgments
This research was partially funded by the Federal Ministry of Education and Research (BMBF), Germany under the project Leibniz KILabor with grant No. 01DD20003.
References
[1] H. Allcott and M. Gentzkow. Social media and fake news in the 2016 election. Journal of economic perspectives, 31(2):211–236, 2017.
[2] Y. Asim, A. K. Malik, B. Raza, and A. R. Shahid. A trust model for analysis of trust, influence and their relationship in social network communities. Telematics and Informatics, 36:94–116, 2019.
[3] E. Chen and E. Ferrara. Tweets in time of conflict: A public dataset tracking the twitter discourse on the war between ukraine and russia. ar Xiv preprint ar Xiv:2203.07488, 2022.
[4] L. Derczynski and K. Bontcheva. Pheme: Veracity in digital social networks. In UMAP workshops, 2014.
[5] U. Ezeakunne, S. M. Ho, and X. Liu. Sentiment and retweet analysis of user response for early fake news detection. In The International Conference on Social Computing, Behavioral-Cultural Modeling, & Prediction and Behavior Representation in Modeling and Simulation (SBP-BRi MS’20), pages 1–10, 2020.
[6] K. Hayawi, S. Shahriar, M. A. Serhani, I. Taleb, and S. S. Mathew. Anti-vax: a novel twitter dataset for covid-19 vaccine misinformation detection. Public health, 203:23–30, 2022.
[7] M. Z. Hossain, M. A. Rahman, M. S. Islam, and S. Kar. Banfakenews: A dataset for detecting fake news in bangla. ar Xiv preprint ar Xiv:2004.08789, 2020.
[8] S. Jiang and C. Wilson. Linguistic signals under misinformation and fact-checking: Evidence from user comments on social media. Proceedings of the ACM on Human Computer Interaction, 2(CSCW):1–23, 2018.
[9] Z. Jin, A. Lalwani, T. Vaidhya, X. Shen, Y. Ding, Z. Lyu, M. Sachan, R. Mihalcea, and B. Schölkopf. Logical fallacy detection. ar Xiv preprint ar Xiv:2202.13758, 2022.
[10] Y. Liu and Y.-F. B. Wu. Fned: a deep network for fake news early detection on social media. ACM Transactions on Information Systems (TOIS), 38(3):1–33, 2020.
[11] P. A. Lofgren, S. Banerjee, A. Goel, and C. Seshadhri. Fast-ppr: Scaling personalized pagerank estimation for large graphs. In Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 1436–1445, 2014.
[12] J. Ma, W. Gao, P. Mitra, S. Kwon, B. J. Jansen, K.-F. Wong, and M. Cha. Detecting rumors from microblogs with recurrent neural networks. 2016.
[13] F. Monti, F. Frasca, D. Eynard, D. Mannion, and M. M. Bronstein. Fake news detection on social media using geometric deep learning. ar Xiv preprint ar Xiv:1902.06673, 2019.
[14] D. Q. Nguyen, T. Vu, and A. T. Nguyen. BERTweet: A pre-trained language model for English Tweets. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 9–14, 2020.
[15] P. Patwa, S. Sharma, S. Pykl, V. Guptha, G. Kumari, M. S. Akhtar, A. Ekbal, A. Das, and T. Chakraborty. Fighting an infodemic: Covid-19 fake news dataset. In Combating Online Hostile Posts in Regional Languages during Emergency Situation: First International Workshop, CONSTRAINT 2021, Collocated with AAAI 2021, Virtual Event, February 8, 2021, Revised Selected Papers 1, pages 21–29. Springer, 2021.
[16] H. Rashkin, E. Choi, J. Y. Jang, S. Volkova, and Y. Choi. Truth of varying shades: Analyzing language in fake news and political fact-checking. In Proceedings of the 2017 conference on empirical methods in natural language processing, pages 2931–2937, 2017.
[17] N. Ruchansky, S. Seo, and Y. Liu. Csi: A hybrid deep model for fake news detection. In Proceedings of the 2017 ACM on Conference on Information and Knowledge Management, pages 797–806, 2017.
[18] N. Seddari, A. Derhab, M. Belaoued, W. Halboob, J. Al-Muhtadi, and A. Bouras. A hybrid linguistic and knowledge-based analysis approach for fake news detection on social media. IEEE Access, 10:62097–62109, 2022.
[19] K. Shu, D. Mahudeswaran, S. Wang, D. Lee, and H. Liu. Fakenewsnet: A data repository with news content, social context, and spatiotemporal information for studying fake news on social media. Big data, 8(3):171–188, 2020.
[20] K. Shu, A. Sliva, S. Wang, J. Tang, and H. Liu. Fake news detection on social media: A data mining perspective. ACM SIGKDD explorations newsletter, 19(1):22–36, 2017.
[21] J. A. Tucker, A. Guess, P. Barberá, C. Vaccari, A. Siegel, S. Sanovich, D. Stukal, and B. Nyhan. Social media, political polarization, and political disinformation: A review of the scientific literature. Political polarization, and political disinformation: a review of the scientific literature (March 19, 2018), 2018.
[22] S. Volkova, K. Shaffer, J. Y. Jang, and N. Hodas. Separating facts from fiction: Linguistic models to classify suspicious and trusted news posts on twitter. In Proceedings of the 55th annual meeting of the association for computational linguistics (volume 2: Short papers), pages 647–653, 2017.
[23] Y. Wang, M. Mc Kee, A. Torbica, and D. Stuckler. Systematic literature review on the spread of health-related misinformation on social media. Social science & medicine, 240:112552, 2019.
[24] L. Wu and H. Liu. Tracing fake-news footprints: Characterizing social media messages by how they propagate. In Proceedings of the eleventh ACM international conference on Web Search and Data Mining, pages 637–645, 2018.
[25] Y. Yang, L. Zheng, J. Zhang, Q. Cui, Z. Li, and P. S. Yu. Ti-cnn: Convolutional neural networks for fake news detection. ar Xiv preprint ar Xiv:1806.00749, 2018.
[26] F. Yu, Q. Liu, S. Wu, L. Wang, T. Tan, et al. A convolutional approach for misinformation identification. In IJCAI, pages 3901–3907, 2017.
[27] Z. Yue, H. Zeng, Z. Kou, L. Shang, and D. Wang. Contrastive domain adaptation for early misinformation detection: A case study on covid-19. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, pages 2423–2433, 2022.
[28] H. Zhang, S. Qian, Q. Fang, and C. Xu. Multimodal disentangled domain adaption for social media event rumor detection. IEEE Transactions on Multimedia, 23:4441– 4454, 2020.
[29] X. Zhang, Y. Malkov, O. Florez, S. Park, B. Mc Williams, J. Han, and A. El-Kishky. Twhin-bert: a socially-enriched pre-trained language model for multilingual tweet representations. ar Xiv preprint ar Xiv:2209.07562, 2022.
[30] C. Zhou, K. Li, and Y. Lu. Linguistic characteristics and the dissemination of misinformation in social media: The moderating effect of information richness. Information Processing & Management, 58(6):102679, 2021.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime