Peer Reviewed

Monitoring misinformation, building discernment: A crowdsourcing study during the 2025 Australian federal election

Article Metrics
CrossRef

0

CrossRef Citations

2

PDF Downloads

6

Page Views

We conducted a design evaluation exploring whether a structured crowdsourcing system using trained volunteers could monitor and categorize election-related misinformation during the 2025 Australian federal election campaign. Over six weeks, 90 active participants surfaced 592 potentially misinforming posts and provided over 3,000 assessments across multiple platforms and recurring narratives. Participants classified content with meaningful accuracy against an independent assessor, with inter-rater agreement varying by category. A follow-up three months later found that participants showed higher misinformation discernment than a separately recruited control group. Structured crowdsourcing may, therefore, serve as a scalable monitoring mechanism and, more tentatively, a potential pathway to building resilience to misinformation.

Image by Geralt on pixabay

Research Questions

  • Can a trained volunteer crowd deployed during a live election identify and categorize election-related misinformation across major social media platforms?
  • What forms of misinformation and themes did participants identify during the 2025 Australian federal election campaign, and how did these forms change over time?
  • How consistently did participants classify the same content, and how well did their judgments match those of an independent assessor?
  • Was participation associated with differences in participants’ ability to distinguish true from false content several months later?

Research note Summary

  • This study deployed a structured crowdsourcing system in which trained volunteers monitored their own social media feeds during the 2025 Australian federal election, submitting and peer-assessing potentially misleading posts using a shared classification framework. 
  • Volunteers surfaced hundreds of instances of election-related misinformation across major platforms. False content, unproven or conspiratorial claims, and out-of-context material dominated submissions. Narratives around electoral integrity grew in prominence towards polling day, while social justice, immigration, and foreign threat narratives persisted throughout.
  • Participants showed meaningful convergence in applying the classification framework and achieved broad agreement with an independent assessor on a sample of submitted posts. 
  • Three months after the campaign, participants distinguished true from false simulated posts more accurately than a separately recruited control group. This difference persisted after adjusting for demographic differences between the groups.
  • The method offers a non-partisan model that may support information literacy: by having participants log, categorize, and rate each other’s contributions, it fosters structured engagement with misinformation without prescribing truth judgements. 

Implications

Implication 1: Crowdsourcing as a scalable citizen-science model for monitoring misinformation

The 2025 Australian federal election, held in May under compulsory voting, returned the center-left Labor government amid widespread concern about election-related misinformation circulating on social media, a concern now common to elections worldwide. One increasingly discussed response is crowdsourcing: distributed users flagging and assessing questionable content, as in X’s Community Notes. Such approaches have emerged as a potential complement to professional fact-checking, which is rigorous but difficult to scale during fast-moving events.

Monitoring misinformation during time-sensitive events requires identifying and classifying large volumes of content at speed. Automated systems currently struggle with contextual, culturally dependent, and multimodal content, particularly where evaluative judgments go beyond simple true/false distinctions (Feng et al., 2025; Kertysova, 2018; Santos, 2023). Professional fact-checking offers careful validation, but it is inherently labor-intensive and difficult to deploy at the speed and scale that live events demand (Zhao & Naaman, 2023).

The primary finding of our design evaluation study is that structured crowdsourcing can provide a practical mechanism for expanding monitoring capacity by distributing elements of this work across a large number of participants. In the system developed for this study, 90 active volunteers identified and submitted 592 potentially misleading posts encountered on their social media feeds and provided over 3,000 assessments using a shared classification framework. Because submissions and assessments were recorded through a structured platform, these distributed judgments were aggregated into a coherent dataset capturing misinformation types, confidence ratings and justifications, enabling analysis of narrative themes and targets, and capturing content directly from participants’ own feeds without reliance on restrictive platform API’s or scraping (Davidson et al., 2023). 

This approach resembles citizen-science models used in fields such as environmental monitoring and astronomy, where distributed volunteers contribute observations that are aggregated into large-scale datasets (Bonney et al., 2009). The underlying model is similarly scalable: similar systems could be expanded to larger participant networks and applied to other domains, such as public health, environmental policy, or international crises, as demonstrated by citizen science applications to real-time weather hazard monitoring (Tan et al., 2022).

The model is also extensible across communities, not only domains. Misinformation monitoring is disproportionately oriented toward the Global North and English-language content, leaving the information environments of some communities under-observed (Mahl et al., 2026). The defining feature of our approach (that participants surface the misinformation they personally encounter) speaks directly to this gap. Where participants are recruited from a specific community and materials translated accordingly, the system could be deployed to surface and categorize the misinformation circulating within and targeting that community: content that centralized, English-language monitoring is poorly positioned to detect. In Australia, this gap is concrete: electoral misinformation targeting migrant and non-English-speaking communities—much of it on platforms such as WeChat and RedNote that sit outside mainstream monitoring and prohibit automated tools—is a recognized and under-addressed concern (Yang et al., 2025). Because these environments resist automated detection, community-embedded human monitoring is one of the few viable means of surfacing what circulates within them. A deployment drawing participants from an affected community could therefore build a clearer picture of the misinformation it actually encounters; information that bears directly on equal and informed democratic participation, yet rarely reaches centralized monitoring. We did not test a community-specific deployment; this is an application the design enables rather than a demonstrated outcome. 

Implication 2: Crowdsourcing participation as a potential discernment intervention

Building resilience to misinformation is a challenge, and experts agree that interventions to improve the public’s ability to identify misinformation should be prioritized (Kruger et al., 2024). Yet, such interventions often attenuate over time without reinforcement. Against this backdrop, we observed that participation in our study was associated with differences in participants’ ability to distinguish true from false simulated social media content three months after the monitoring period concluded (compared with a control group). However, because participants were not randomly assigned, these differences may reflect pre-existing characteristics, such as political interest, motivation or baseline media-literacy, rather than the effects of participation itself. 

If participation did contribute to this difference, there is a plausible mechanism. Sustained, repeated engagement with fact-checking has been shown to durably improve discernment of misinformation beyond the specific claims encountered, effectively functioning as a form of media literacy by familiarizing people with how misinformation works and how it can be evaluated (Bowles et al., 2025). Our system extends this logic from repeated exposure to repeated practice: rather than passively consuming corrections, participants actively classified real-world content, articulated their reasoning, and rated their confidence using shared analytical criteria. This active, effortful evaluation is consistent with research on psychological inoculation, which emphasizes the role of practice and active resistance in building durable resistance to misleading information (Roozenbeek & van der Linden, 2019).

Findings

Finding 1: A distributed crowd provided sustained monitoring of misinformation during the Australian 2025 federal election campaign. 

Over the six-week monitoring period, 90 active participants monitored their social media environment, submitting 592 potentially misinforming posts and over 3,000 assessments. As is typical of crowdsourced projects (Sauermann & Franzoni, 2015; Swain et al., 2015), contributions were unevenly distributed: approximately 10% of participants accounted for 60% of submissions and 70% of assessments. Participants reported 1,228 hours of engagement, with an estimated 614 hours spent actively monitoring. 

Thematic analysis showed that a small number of narratives accounted for a disproportionate share of submissions (Figure 1), while temporal analysis revealed that these narratives shifted over the course of the campaign (Figure 2). In particular, electoral integrity claims were present throughout but increased sharply in the final weeks, becoming the most frequently submitted theme around polling day. Each submission was also linked to a target where one was identifiable. Notably, the incumbent center-left Labor government was the most frequent target (53% of submissions with an identifiable target; see Appendix C).

Figure 1. Top six misinformation themes identified during the 2025 federal election monitoring period. Top themes accounted for 50.7% of all submissions.
Figure 2. Weekly frequency of the six most common misinformation themes identified during the monitoring period.

Finding 2: The volunteer cohort provided broad cross-platform coverage that aligned with narratives independently documented by a major fact-checking organization. 

The misinformation surfaced by participants was heavily concentrated on a small number of major social media platforms. The four platforms in Figure 3 accounted for over 90% of all submissions, with the remainder spread across YouTube, BlueSky, Xiaohongshu, Reddit, and Threads. However, because submissions were drawn from participants’ own social media environments and the participant pool was ideologically skewed, the distribution likely reflects a partial rather than comprehensive view of the ecosystem. It also excludes misinformation in broadcast or print, except where it was later shared on social media.

Figure 3. Concentration of misinformation detections across social media platforms. 

To assess whether our crowd-based method achieved coverage comparable to established fact-checking practice, we compared participant submissions with Australian Associated Press (AAP) fact checks published during the election period (see Appendix E for more detail). Analysis of 34 AAP fact-check articles showed a similar distribution of themes, with electoral integrity, government spending, and cost-of-living claims most prevalent. These correspond closely with the most common themes identified by participants (see Finding 1). Notably, every major election-related narrative documented by AAP was also present in the participant dataset; the crowd surfaced the same narratives that a professional fact-checking organization deemed significant enough to verify. This suggests the method can match an established fact-checking process for the narratives that process judged significant, even if it does not capture the full ecosystem. 

Finding 3: Participants applied a structured classification framework that differentiated distinct forms of misinformation.

In addition to surfacing misinformation across platforms and narratives, participants classified each submission using a shared framework distinguishing among 17 categories, developed for usability by non—expert participants and refined against real posts during development (see Appendix A for more detail). As shown in Figure 4, the most common forms were false content, unproven or conspiratorial claims, and out-of-context material, which together accounted for nearly 60% of all submissions. Other forms (including rage-bait, cherry-picking, political defamation, altered media, and impersonation) appeared less frequently but were consistently identified across the monitoring period.

Figure 4. Distribution of primary misinformation forms assigned by participants. False content, unproven claims, and out-of-context material accounted for the majority of submissions.

Importantly, participants did not indiscriminately label all flagged content as misinformation. Approximately 5.9% of submissions were classified as out-of-scope or not problematic, and 3.2% were coded as true or partly true. This indicates that participants were exercising evaluative judgment, discerning what did not warrant a misinformation label.

Finding 4: Crowd classifications of verifiably false content approached fact-checking accuracy, while inter-rater agreement varied across categories.

To assess accuracy, we compared participant judgments against those of an independent assessor (a postgraduate researcher who documented the basis for each determination). Consistent with established approaches to evaluating crowd-based misinformation detection (e.g., Barbera et al., 2024), this check focused on the verifiable true/false dimension rather than the framework’s more interpretive categories. Posts the crowd classified as false content (n = 24) or true (n = 6) were checked against the assessor’s determination, since these rest on factual claims with a ground truth.

The assessor agreed with the crowd’s classification in 81% of cases. Cohen’s kappa was 0.42, indicating moderate agreement beyond chance (Landis & Koch, 1977). While the sample of posts is small, this level of alignment is broadly consistent with recent findings that crowds can approach professional fact-checking accuracy under structured conditions (Barbera et al., 2024). 

Among posts rated by multiple participants, mean agreement with the modal category was 64.7%, though because participants could see others’ classifications before submitting their own, this may overstate independent agreement.

For posts yielding a unique majority classification (N = 387), agreement was highest for categories with clear structural features (e.g., political scams, unproven or conspiratorial claims, false content) and lower for interpretive ones such as rage-bait and cherry-picking (see Figure 5); these figures reflect agreement among raters, not accuracy against an external standard. 

Confidence ratings of assessments (on a scale 0 – 10) followed a similar pattern. Confidence varied significantly across classification categories, χ²(16) = 144.44, p < .001. Clearer-cut categories such as political scams and false content tended to be rated more confidently than interpretive ones (see Appendix A). 

Figure 5. Majority agreement by category. Agreement is measured as the proportion of raters selecting the modal category for each post. The x-axis displays misinformation categories, and the y-axis shows mean agreement (%). Values above bars indicate the number of posts, and the horizontal line shows the overall mean agreement.

Finding 5: Participation in structured monitoring was associated with higher misinformation discernment.

To examine whether participation in structured monitoring was associated with improved evaluative judgment, we conducted a follow-up assessment approximately three months after the monitoring period concluded (see Appendix F). Participants drawn from the active monitoring cohort (treatment group; n = 56) were compared with a newly recruited control group (n = 132) using a 20-item test comprising simulated social media posts with established ground truth. Participants made binary true/false judgments for each item.

The treatment group achieved a mean accuracy of 73.8% (SD = 0.137), compared with 67.5% (SD = 0.128) in the control group. This difference corresponds to a moderate effect size (Cohen’s d = 0.48, 95% CI [0.16, 0.79]). The groups differed on some demographic measures, but the difference in discernment persisted after adjusting for those on which they differed (age, employment, and education) in a trial-level model (see Appendix F). Because participants were not randomly assigned, however, the difference may still reflect unmeasured pre-existing characteristics. The pattern is best read as associational, consistent with the possibility that active engagement with misinformation is linked to improved discernment, though the design does not permit causal inference.

Methods

Design overview

This study was a field-based design evaluation of a structured crowdsourcing monitoring system during the 2025 Australian federal election. Rather than testing a tightly controlled intervention, the primary aim was to assess whether a trained volunteer cohort could surface and categorize election-related misinformation at scale, apply a shared framework with meaningful consistency, and generate structured outputs suitable for analysis. We also conducted an exploratory follow-up examining whether participation was associated with longer-term differences in misinformation discernment. 

Participants and recruitment

Participants were recruited into two cohorts: a general public cohort (n = 401), recruited via targeted online advertising, and an expert cohort (n = 18), recruited through purposive and snowball sampling for professional or research experience with misinformation (median = ~6 years in the area), drawn from fields including psychology, public health, law, and science communication. Together, these yielded 419 individuals who completed onboarding and consented to participate. Eligibility required participants to be aged 18 or older, active social media users, with a strong interest in Australian politics.

All 419 participants were issued platform accounts. We define active participants as those who made at least one contribution to the database – a submission or an assessment; 90 were active in this sense, with the remainder onboarding but not contributing. The active group comprised both cohorts, 13 of the 18 recruited experts and 77 general public-contributors. The onboarding sample spanned a broad adult age range, skewed toward middle and older age groups, with comparatively high educational attainment and representation across Australian states (see Appendix B).

Volunteer-based detection raises the concern that raters may disproportionately flag content conflicting with their own views (Park et al., 2026). Although we ran targeted recruitment aimed at both left- and right-of-center participants, the active cohort skewed left-of-center, broadly reflecting Australia’s electorate at the 2025 election (Australian Electoral Commission, 2025a, 2025b). Three features mitigated this: training on cognitive bias with instruction to set aside political views; required written justification for each assessment, which has been shown to aid misinformation identification regardless of ideology (Pennycook & Rand, 2019); and a peer-assessment structure that diluted individual skew across multiple judgments. These mechanisms cannot fully eliminate systematic bias, and the dataset may still reflect some ideological filtering in which posts were surfaced and how they were interpreted (see Appendix B).

Procedure

The monitoring phase ran for six weeks during the 2025 Australian federal election campaign (April 2 –May 16, 2025). Training consisted of slides, a participant “hub” where project updates and tips were posted, and a webinar on cognitive biases relevant to misinformation detection. Following onboarding and training, participants were granted access to a custom-built web platform and were instructed to monitor their own feeds as part of regular use and to submit posts they believed constituted election-related misinformation. 

Figure 6. Two-phase study design and participant flow.

For each submission, participants provided the original URL and selected a primary misinformation category (with an optional secondary category), rated their confidence, and gave written reasoning. Screenshots were uploaded to preserve content and submissions could be edited after posting.

Submitted posts were available for peer assessment: other participants could review the content, classification and reasoning before recording their own classification, confidence rating and justification. Participants were encouraged to apply independent judgment using the shared framework. Assessments were visible within the system, and discussion was possible on a separate platform (Slack).

Approximately three months after monitoring concluded, a follow-up compared a treatment group drawn from the active monitoring cohort (n = 56) with a newly recruited control group (n = 132), recruited using mirrored eligibility criteria, on the same 20-item true/false test (see Finding 5; Appendix F).

Measures and analysis

Primary category selections (Appendix A) were used in the reliability and distributional analyses. All submissions were reviewed by both authors using a thematic coding approach; narrative themes and institutional targets were developed from the data and refined iteratively. Independent initial coding surfaced differences in interpretation, which were discussed and resolved through consensus. The coding should be understood as an iterative qualitative analysis rather than a fully independent coding procedure.

Classification consistency was measured as the proportion of raters assigning the same primary category to a post. Accuracy was assessed against an independent assessor on 30 verifiable true/false posts (see Finding 4).

In the follow-up study, discernment was operationalized as accuracy on the 20-item test. We compared treatment and control accuracy using Welch’s t-test (which does not assume equal group sizes or variances) and, using a linear mixed-effects model, tested whether the difference persisted after adjusting for age, employment, and education, the variables on which the groups differed, with random intercepts for participant and item (see Appendix F). Cohen’s d is reported as a descriptive effect size.

Topics
Download PDF
Cite this Essay

Kruger, A., Howe, P., & Saletta, M. (2026). Monitoring misinformation, building discernment: A crowdsourcing study during the 2025 Australian federal election. Harvard Kennedy School (HKS) Misinformation Review. https://doi.org/10.37016/mr-2020-210

Links

Bibliography

Australian Electoral Commission. (2025a). 2025 Federal Election. https://results.aec.gov.au/31496/Website/HouseDefault-31496.htm

Australian Electoral Commission. (2025b). Turnout by state. https://results.aec.gov.au/31496/Website/HouseTurnoutByState-31496.htm

Barbera, D. L., Maddalena, E., Soprano, M., Roitero, K., Demartini, G., Ceolin, D., Spina, D., & Mizzaro, S. (2024). Crowdsourced fact-checking: Does It actually work? Information Processing & Management, 61(5), Article 103792. https://doi.org/10.1016/j.ipm.2024.103792

Bonney, R., Cooper, C. B., Dickinson, J., Kelling, S., Phillips, T., Rosenberg, K. V., & Shirk, J. (2009). Citizen science: A developing tool for expanding science knowledge and scientific literacy. BioScience, 59(11), 977–984. https://doi.org/10.1525/bio.2009.59.11.9

Bounegru, L., Gray, J., Venturini, T., & Mauri, M. (2018). A field guide to “fake news” and other information disorders. Public Data Lab. https://doi.org/10.5281/zenodo.1136272

Bowles, J., Croke, K., Larreguy, H., Liu, S., & Marshall, J. (2025). Sustaining exposure to fact-checks: Misinformation discernment, media consumption, and its political implications. American Political Science Review, 119(4), 1864–1887. https://doi.org/10.1017/S0003055424001394

Davidson, B.I., Wischerath, D., Racek, D. et al. (2023). Platform-controlled social media APIs threaten open science. Nature Human Behaviour, 7, 2054–2057. https://doi.org/10.1038/s41562-023-01750-2

Feng, X., Luo, J., Yang, Y., El Baz, D., & Shi, L. (2025). Health misinformation detection: Approaches, challenges and opportunities. INQUIRY: The Journal of Health Care Organization, Provision, and Financing, 62. https://doi.org/10.1177/00469580251384784

Kertysova, K. (2018). Artificial intelligence and disinformation: How AI changes the way disinformation is produced, disseminated, and can be countered. Security and Human Rights, 29(1–4), 55–81. https://doi.org/10.1163/18750230-02901005

Kruger, A., Saletta, M., Ahmad, A., & Howe, P. (2024). Structured expert elicitation on disinformation, misinformation, and malign influence: Barriers, strategies, and opportunities. Harvard Kennedy School (HKS) Misinformation Review, 5(7). https://doi.org/10.37016/mr-2020-169

Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. https://doi.org/10.2307/2529310

Mahl, D., Zeng, J., Schäfer, M. S., Egert, F. A., & Oliveira, T. (2026). “We follow the disinformation”: Conceptualizing and analyzing fact-checking cultures across countries. The International Journal of Press/Politics, 31(2), 264–290. https://doi.org/10.1177/19401612241270004

Park, S., Lee, J. Y., McGuinness, K., Fisher, C. & Fulton, J. (2026). People rely on their existing political beliefs to identify election misinformation, 7(1). Harvard Kennedy School (HKS) Misinformation Review. https://doi.org/10.37016/mr-2020-195

Pennycook, G., & Rand, D. G. (2019). Fighting misinformation on social media using crowdsourced judgments of news source quality. Proceedings of the National Academy of Sciences, 116(7), 2521–2526. https://doi.org/10.1073/pnas.1806781116

Pew Research Center. (2026, June 10). Political typology quiz: Which of the 9 groups are you?https://www.pewresearch.org/politics/quiz/political-typology/

Roozenbeek, J., & van der Linden, S. (2019). The fake news game: Actively inoculating against the risk of misinformation. Journal of Risk Research, 22(5), 570–580. https://doi.org/10.1080/13669877.2018.1443491

Santos, F. C. C. (2023). Artificial Intelligence in automated detection of disinformation: A thematic analysis. Journalism and Media, 4(2), 679–687. https://doi.org/10.3390/journalmedia4020043

Sauermann, H., & Franzoni, C. (2015). Crowd science user contribution patterns and their implications. Proceedings of the National Academy of Sciences, 112(3), 679–684. https://doi.org/10.1073/pnas.1408907112

Swain R., Berger A., Bongard J., & Hines, P. (2015). Participation and contribution in crowdsourced surveys. PLOS ONE, 10(4), Article e0120521. https://doi.org/10.1371/journal.pone.0120521

Tan, M. L., Hoffmann, D., Ebert, E., Cui, A., & Johnston, D. (2022). Exploring the potential role of citizen science in the warning value chain for high impact weather. Frontiers in Communication, 7, Article 949949. https://doi.org/10.3389/fcomm.2022.949949

Wardle, C., & Derakhshan, H. (2017). Information disorder: Toward an interdisciplinary framework for research and policy making (Report No. DGI(2017)09). Council of Europe. https://rm.coe.int/information-disorder-toward-an-interdisciplinary-framework-for-researc/168076277c

Yang, F., Heemsbergen, L., & Fordyce, R. (2025). Contextualizing critical disinformation during the 2023 Voice referendum on WeChat: Manipulating knowledge gaps and whitewashing Indigenous rights. Harvard Kennedy School (HKS) Misinformation Review, 6(5). https://doi.org/10.37016/mr-2020-185

Zhao, A., & Naaman, M. (2023). Insights from a comparative study on the variety, velocity, veracity, and viability of crowdsourced and professional fact-checking services. Journal of Online Trust and Safety, 2(1). https://doi.org/10.54501/jots.v2i1.118

Funding

Funding for the project was provided by the Susan McKinnon Foundation.

Competing Interests

The authors declare no competing interests.

Ethics

This study was approved by the University of Melbourne Human Research Ethics Committee (Approval ID: 30659). All participants provided informed consent prior to participation in both the monitoring phase and the follow-up discernment study. Demographic categories for gender (male, female, non-binary/third gender, prefer not to say) and age were defined by the investigators based on standard survey conventions. Gender and age data were collected to characterize the sample and assess representativeness, not as analytic variables. Ethnicity data were not collected.

Copyright

This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided that the original author and source are properly credited.

Data Availability

All materials needed to replicate this study are available via the Harvard Dataverse: https://doi.org/10.7910/DVN/UZ42LX. Free-text justification fields and participant account identifiers have been removed to protect anonymity; restricted materials are available on reasonable request, subject to ethics approval.