Research
Understanding Complex Social Systems
02 Political Discourse on Reddit, Mapped
Quantifying all of Reddit's political talk, across its whole history, on 56 dimensions of discourse quality drawn from political philosophy; tracing how communities evolve through that space.
Read project
Trajectories is quantifying political discussion on Reddit in its entirety: every political community, across the platform’s whole history. Using transformer-based language models validated against human annotators, we score discussion on the same set of measures at many points in time, so that every subreddit traces a trajectory through “discourse space.” That opens up a massively general research programme: how does the discourse of communities with different topics, norms, and moderation strategies evolve? What events bend a community’s path, toward antagonism or back toward genuine engagement? Which conditions sustain political talk worth having?
What we measure is not what anyone believes; no position on the political spectrum scores better than any other. We measure how people do discourse. Do you give reasons, or just assert? Do you respond to what your opponent actually said? Can you concede a good point? But also: do you argue with passion? Do you fight your corner, and is your opponent an adversary to beat or an enemy to destroy? Computational research has so far captured only the easiest slice of this at scale: toxicity, incivility, sentiment. To decide which differences actually make a difference, we turn to political philosophy, above all the long-running debate between Jürgen Habermas and Chantal Mouffe. Habermas holds up an ideal of reasoned argument aimed at mutual understanding; Mouffe replies that passionate conflict is the lifeblood of politics, not its failure. The first half of our questions is Habermas’s; the second half is Mouffe’s. From their debate we distill 56 distinct, continuously scored dimensions of discourse quality.
The full dataset and measurement infrastructure will be open-sourced.
- Team
- Tim Booker (lab); Friederike Stock (MPI-HD); Mubashir Sultar (MPI-HD); Fabian Veider (University of Graz)
- Funding
- DeSiRe project
- Responsibility
- Tim Booker
- Duration
- 2025–2027
03 The Geopolitics of Media Bias
How news outlets around the world frame the victims of conflict: 1.36 billion articles, over a hundred languages, ten years of coverage; outlets compared on the same events, so that differences in framing cannot be blamed on differences in the facts.
Read project
When people are killed in a conflict, the news must decide how to tell it. Some victims arrive as people, with names, ages, and families; others arrive as a number. Some perpetrators are named; others vanish into passive constructions, where people are simply “killed in a strike.” Beyond the wording, there is the narrative a death is folded into: a measured response, an unprovoked atrocity, a tragic accident. Whose casualty figures are stated as fact, and whose are “claimed”? Who is quoted, and who is merely described? And which facts, reported by every other outlet covering the story, does an article quietly leave out? We measure these framing choices across the world’s news: roughly 1.36 billion articles in over a hundred languages, from 2016 to 2026, spanning many conflicts at once.
The question is what drives these choices. A benign explanation is severity: bigger events get more names, more detail, more attribution. Our hypothesis is that framing instead tracks the outlet’s geopolitical alignment with the victims’ side; the same death is humanised by one outlet and tallied by another. Testing this fairly means comparing like with like, so the core of the project is an event-matched corpus: articles across languages and outlets are grouped by the real-world event they report, and each is distilled into a structured record of who did what to whom. Holding the event fixed, whatever variation remains in framing belongs to the outlet, not to the facts. It also makes omission measurable for the first time; only by knowing what an event’s other reporters said can you see what one article left out.
The event-matched, fact-extracted corpus and all framing measures will be openly released.
- Team
- Tim Booker (lab); Elisabeth Höldrich (lab); Kirill Solovev (lab); Mathias Angermaier (lab); Ruggero Marino Lazzaroni (lab)
- Funding
- Internal project
- Responsibility
- Tim Booker, Elisabeth Höldrich
- Duration
- 2026–2027
04 Orientation in Conspiration (ORION)
Conspiracy Theories on German-Language Telegram
Read project
For most of the last century, conspiracy theories looked like a harmless curiosity. The pandemic ended that. Online conspiracy groups spread claims that did concrete damage, and what to do about it became a political question nobody has answered well. ORION works on the part of it that is empirically tractable: which factors drive the spread, and how much each contributes. We treat dissemination as something individual, social and technological factors produce together, and try to separate them out.
Our material is the “Schwurbelarchiv”, an archive of German-language public Telegram channels released in 2022, covering roughly 2019 to 2022 and running to over 23 terabytes. Most of it is multimedia. Cleaning, structuring and transcribing that content is a substantial piece of the project in its own right, and it is what makes the archive usable by anyone else.
A second strand asks how far this discourse carries elements of fascist ideologies and narratives. Using Abstract Meaning Representation to extract predicate-argument structures across QAnon and Reichsbürger channels, we measure two things: how far the national collective is spoken about as a living organism rather than a legal or institutional one, and how often the demand is rebirth rather than reform. Both come from Roger Griffin’s definition of the fascist minimum, ultranationalism fused with a myth of rebirth.
- Results
Characterizing the Dynamics of Conspiracy-Related German Telegram Conversations during COVID-19Journal of Quantitative Description: Digital MediaARTICLE The Schwurbelarchiv: a German Language Telegram dataset for the Study of Conspiracy TheoriesEPJ Data Science - Team
- Jana Lasser (PI); Elisabeth Höldrich; Mathias Angermaier
- Funding
- FWF; netidee (Austria's open-source internet funding programme)
- Responsibility
- Jana Lasser, Elisabeth Höldrich
- Duration
- 1 May 2024 – 30 April 2028
05 Computational Affective Science
Using computational methods to understand the emotional life of humans.
Read project
Emotions are a fundamental part of human existence. They inform us about our attitudes towards people, objects, and situations, and guide the actions we take in response. People are not only subject to their emotional experiences, but actively regulate them, downregulating and upregulating both positive and negative emotions depending on the context. This regulation can be directed at one’s own emotions or at the emotions of others. Understanding how humans experience and regulate their emotions means understanding a fundamental mechanism through which people navigate the world.
In this research area, we work on understanding specific emotion regulation strategies and the broader emotion regulation landscape, also with a focus on human-AI interaction. We developed a data-driven, computational method for defining psychological constructs, applying it to map the full space of emotion regulation strategies people use, extending beyond the narrow focus on reappraisal that has dominated prior research. We further zoomed in on reappraisal, mapping different types of reappraisal strategies and testing whether some are more effective than others. We too compared human and AI reappraisal ability, having both humans and GPT-4 reframe negative emotional scenarios, where GPT-4’s reappraisals were judged as more effective than those produced by most humans. Across these projects, we draw on natural language processing methods, including sentence embeddings, clustering, large language models, and classification, to analyze emotional experiences at scale.
Looking ahead, we aim to broaden this work to understand how affect shapes everyday decision-making and how people navigate their world more generally, examining its relationship to cognition and the emotional appraisal processes.
- Team
- Alina Herderich (lab); Amit Goldenberg (Harvard University)
- Responsibility
- Alina Herderich
- Duration
- 2022–today
06 A new ontology of truth
When gut feelings are perceived as honesty, how does that affect the information ecosystem and political functioning?
Read project
Misinformation is increasingly perceived as a major problem for societal cohesion and democracy. Much attention has focused on the role of social media as a vector of misinformation. The role of political leaders has attracted less scrutiny, even though leaders demonstrably influence media coverage and public opinion, and enact political processes crucial for societies to function. Politicians rely on two different rhetorical approaches to express their pursuit of truth: Evidence-based language relies on data and facts to describe elements of external reality. Intuition-based language relies on feelings, instincts and personal values to describe elements of internal experience. A productive democratic discourse strikes a balance between evidence-based and intuition-based conceptions of truth. In this project, we ask whether the language politicians in the US use online and in Congress has changed over time and how it relates to the trustworthiness of information they convey and the functioning of the political system.
We analyse communications by members of the U.S. Congress on Twitter between 2011 and 2022 and show that political speech has fractured into two distinct components related to intuition-based belief-speaking and evidence-based fact-speaking. We also find that conservatives in the US uniquely show a steep rise in the untrustworthy information they spread. We show that belief-speaking is distinctly related to spreading of untrustworthy information for Republicans. Conversely, an increase in fact-speaking language is associated with an increase in the quality of sources cited by both parties. The results support the hypothesis that the current dissemination of misinformation in political discourse is in part driven by an alternative understanding of truth and honesty that emphasizes invocation of subjective belief at the expense of reliance on evidence. Next to the online presence of politicians, we analyse the linguistic traces of evidence- and intuition-based perspectives in speeches in the US Congress from 1879 to 2022. We find that evidence-based language has continued to decline since the mid-1970s, together with a decline in legislative productivity. The decline was accompanied by increasing partisan polarization in Congress and rising income inequality in society. The results highlight the importance of evidence-based language in political decision-making.
- Results
From alternative conceptions of honesty to alternative facts in communications by U.S. politiciansPNAS NexusDOIREPRODUCTION PACKAGE Computational analysis of US Congressional speeches reveals a shift from evidence to intuitionNature Human BehaviourDOI Social media sharing of low quality news sources by political elitesPNAS NexusDOI Republicans are increasingly sharing misinformation, research findsThe Washington Post – Monkey Cage, 2022ARTICLE US politicians tweet far more misinformation than those in the UK and Germany – new researchThe Conversation, 2022ARTICLE Donald Trump's truth: why liars might sometimes be considered honest – new researchThe Conversation, 2023ARTICLE - Team
- Jana Lasser (lab)
- Collaborators
- Stephan Lewandowsky; David Garcia; Segun Taofeek Aroyehun; Almog Simchon; Fabio Carrella
- Funding
- ERC Advanced Grant awarded to Stephan Lewandowsky
- Responsibility
- Stephan Lewandowsky (Stephan.Lewandowsky@bristol.ac.uk)
- Duration
- 1 October 2021 – 30 September 2027
Designing Prosocial Online Spaces
01 Designing Social Media Recommendation Algorithms for Societal Good (DeSiRe)
alternative content recommendation algorithms for social media platforms
Read project
In this project we aim to research alternative content recommendation algorithms for social media platforms. Social media platforms are central for exchange of information in today’s societies. Yet evidence is mounting that they play a causal role in deteriorating social cohesion and civic discourse. User behaviour on social media platforms is governed by content recommendation algorithms that maximise engagement, leading to unintended consequences such as the promotion of outrage or strongly emotionalised content. The EU’s newly enacted Digital Services Act mandates social media platforms to assess and reduce their systemic risks for society – for example by adapting their content recommendation algorithms. The abstract risks, such as the risk to civic discourse however first need to be translated into concrete changes in content recommendation algorithms. In this project we aim to bridge this gap by combining approaches from social science and computer science to incorporate the reduction of risk to civic discourse into content recommendation algorithms of social media platforms.
To this end, we will employ a participation-based approach to develop novel algorithms that consider various aspects of civic discourse, such as information quality and diversity, and the civility of language. To experiment with new algorithms, we will develop Open Source digital twins of social media platforms since experimentation with new algorithms on live social media platforms independent of the permission and influence of platform companies is impossible. Next to the reduction of risk to civic discourse we aim to balance interventions in algorithms with freedom of expression. To this end, we will solicit people’s preferences in different scenarios such as a public health crisis and elections, and develop balanced algorithms. Ultimately, our research will hopefully inform recommendations for the regulation of social media platforms under the Digital Services Act and help transform social media platforms into a technology that is positive for society.
- Results
Alternative recommendation algorithms as an antitrust remedy in digital (democracy) casesJournal of European Competition Law & Practice - Team
- Jana Lasser; Tim Booker; Ruggero Lazzaroni; Nikolaus Pöchhacker
- Funding
- FWF & netidee (Austria’s Open Source internet funding campaign)
- Responsibility
- Research group for Complex Social & Computational Systems
- Duration
- 1 January 2025 – 31 December 2029
10 Countering hate in online spaces
Understanding counter speech as a citizen-based, collective answer to hate on social media platforms.
Read project
Social media platforms are plagued by hate speech from individuals, but also from organized groups that systematically target and discredit politicians and public figures. This coordinated hostility shifts public discourse and distorts the perceived opinion of the majority, as people holding more moderate views are pushed into silence. One common response to online hate is top-down moderation, in which platforms themselves define and remove hate speech. However, platforms are financially motivated rather than democratically legitimized, and their authority to determine what counts as hate speech is neither transparent nor accountable.
An alternative approach is counter speech, in which citizens actively oppose online hate, for example by openly disagreeing with perpetrators. Since online spaces have become a crucial part of public dialogue and are systems naturally resisting top-down control, understanding the dynamics and effectiveness of counter speech is vital for public discourse. In this project, we analyzed a large dataset of over 130,000 conversations on Twitter in Germany during a period when two organized and opposing groups were active: Reconquista Germanica, a hate speech group supporting right-wing political agendas; and Reconquista Internet, a counter speech group founded explicitly to oppose it. We classified the tweets along several discourse dimensions, including argumentative strategies, in- and outgroup utterances, emotions, and hate speech, and found that simply voicing an opinion without resorting to insults is a common strategy that simultaneously shifts discourse in a more civil direction. To this end, we employed a mixed-methods pipeline combining grounded theory, large-scale text classification, and quasi-causal statistical models.
Together with other researchers, we translated these findings into policy recommendations on AI and counter speech. More generally, our work aims to inform and empower anyone wanting to speak up against hate in online spaces.
- Team
- Jana Lasser (lab); Alina Herderich (lab); Mirta Galesic (Complexity Science Hub Vienna); Joshua Garland (Arizona State University); David Garcia (University of Konstanz); Segun Aroyehun (University of Konstanz)
- Funding
- NSF DRMS 1757211
- Responsibility
- Jana Lasser, Alina Herderich
- Duration
- 2022–2026
Curating Open Datasets for Computational Social Science
08 Infini-News: Efficiently Queryable Access to 1.3 Billion Processed Common Crawl News Articles
Enabling reproducible research on online news corpora in hundreds of languages
Read project
Large-scale news corpora are at the core of many a research field, including, but not limited to, computational social science, political science, and natural language processing. Unfortunately, accessing them is non-trivial and is associated with steep resource and monetary costs. The goal of Infini-News is to provide researchers with a way to build their own news corpora in a streamlined manner, reducing the barriers to entry and improving transparency and reproducibility.
Infini-News leverages the entirety of Common Crawl News and applies complex pre-processing and enrichment, as well as infini-gram-mini indexing, to make the entire decade-long corpus searchable in less than a second. To put this in numbers, we provide 1,357,027,742 articles across 331 languages, 17 topics, and 59,885 domains, with a country of origin resolved for 83.4% of articles across 222 countries, distilled from 180 TB of raw web archives, covering August 2016 to April 2026. In particular, Infini-News provides:
- Clean text and metadata: each article is extracted from HTML and contains title, author, publication date, and source.
- Enrichment: we detect the language, the country of origin, and a topic from the standard IPTC news taxonomy.
- Instant search: any word or phrase from across the entire corpus is queryable in under a second, with exact match counts.
- A query API: we host an API allowing users to count, explore, and extract subdatasets with a simple interface (early access — full university deployment in progress).
The corpus (3.4 TB Parquet) is hosted on Hugging Face, with access gated to institutional affiliation and a stated research purpose; the indexes (2.2 TB) and all pipeline code are distributed as free and open source artefacts. Code is MIT-licensed and derived data CC-BY 4.0, while the article text itself remains under the original publishers’ copyright. The dataset and pipeline are described in the accompanying paper (Lazzaroni, Lasser & Solovev, 2026).
- Team
- Ruggero Lazzaroni (lab); Kirill Solovev (lab); Jana Lasser (lab)
- Funding
- European Union's Horizon Europe programme, grant agreement No 101177310. Views and opinions expressed are those of the authors only and do not necessarily reflect those of the European Union.
Developing Methods for Computational Social Science
07 Computational reproducibility and methodological rigour across the sciences
Evidence, tools, and methods for more reliable research
Read project
Research domains have varying modes of establishing validity and reliability of their findings. High-profile cases of research fraud, retractions, and widespread attempts to reproduce previous research findings have contributed to two overlapping discourses that analyse and seek to improve perceived shortcomings: open science and meta-science.
We contribute to this body of work through three research strands:
- Investigating practices that can increase computational reproducibility. For this, we use a range of methods, such as scoping reviews of existing evidence, randomised controlled trials to assess interventions, and computational modelling to expand our understanding of how contextual factors interact with proposed interventions.
- Building tools to assess computational reproducibility. Here we investigate the extent to which we can leverage Large Language Models to automate the computational reproduction of published findings.
- Investigating how causal inference techniques can lead us to reassess existing findings or improve methodological approaches when studying science itself. Using simulations and case studies, we analyse existing measurement approaches in science studies and develop improved computational methods.
- Results
Introduction to structural causal models in science studiesQuantitative Science Studies, 7, 159–178DOI The paradox of competition: How funding models could undermine the uptake of data sharing practicesResearch Policy, 54(10), 105340DOI Open science interventions to improve reproducibility and replicability of research: A scoping reviewRoyal Society Open Science, 12(4), 242057DOI The academic impact of Open Science: A scoping reviewRoyal Society Open Science, 12(3), 241248DOI - Team
- Thomas Klebel (lab)
- Collaborators
- Federico Bianchi; Matthew Cannon; Eva Kormann; Adrian Marangoni; Tony Ross-Hellauer; Eric J Schares; Flaminio Squazzoni; Rebecca Taylor-Grant; Vincent Traag; Olmo van den Akker
- Responsibility
- Thomas Klebel
- Duration
- 2022–today
09 VALPOP: Valuing Public Goods in a Populist World
Extracting signed, typed, multiplex relation networks from news articles at multi-million scale
Read project
VALPOP starts from the notion that populism and the erosion of the rule of law are networked phenomena that could be unearthed by uncovering the relationship networks between political, economic, and media actors. However, to begin this process, we first need to make the relations visible.
To uncover the relationships, we need a data source and machinery to process it at scale. The former are the news articles; they provide a timed, grounded, and multifaceted account of societal happenings. News covers changes in the societal structure both on high and low levels. While some countries might have a more biased coverage, that does not preclude the analysis, as every extracted relation carries the stance of the article that reported it, making the framing an additional source of insight. That is where the latter part of our project comes in: by utilizing an LLM-assisted joint entity and relation extraction pipeline, we are able to process millions of articles in complex ways.
As part of the mission statement at IDea_Lab, we go beyond direct applicability to our projects and try to provide similar tools to a wider research community. For VALPOP this materialises in the extraction pipeline being provided as free and open source, with every part interchangeable for any type of extraction one might want to produce; most importantly, the ontology, prompts, and sources of ground truth can all be replaced in accordance with particular research project needs.
For VALPOP, the pipeline:
- Finds the actors: the people, organizations, and institutions mentioned in the text.
- Resolves who they are: if the entities have an entry in Wikidata or Orbis, they are assigned canonical IDs, allowing us to trace them across countries, languages, and spelling variations.
- Extracts the relationships: a large language model uses our project-specific ontology (110 actor types and 99 relationship types) to extract relations, keeping results comparable across articles, languages, and countries.
- Builds the graph: all the entities and relationships are merged into a single multiplex knowledge graph for further analysis.
Every connection carries a time stamp and traces back to the article that reported it, and the pipeline’s output is validated against a hand-annotated reference standard. The result is a comparable, longitudinal map of societal networks across the 12 VALPOP countries.
- Team
- Kirill Solovev (lab); Jana Lasser (lab)
- Funding
- European Union's Horizon Europe programme, grant agreement No 101177310. Views and opinions expressed are those of the authors only and do not necessarily reflect those of the European Union.