Research

02.1 Active Research
02.1.01

Understanding Complex Social Systems

02

Political Discourse on Reddit, Mapped

Quantifying all of Reddit's political talk, across its whole history, on 56 dimensions of discourse quality drawn from political philosophy; tracing how communities evolve through that space.

Read project

Trajectories is quantifying political discussion on Reddit in its entirety: every political community, across the platform’s whole history. Using transformer-based language models validated against human annotators, we score discussion on the same set of measures at many points in time, so that every subreddit traces a trajectory through “discourse space.” That opens up a massively general research programme: how does the discourse of communities with different topics, norms, and moderation strategies evolve? What events bend a community’s path, toward antagonism or back toward genuine engagement? Which conditions sustain political talk worth having?

What we measure is not what anyone believes; no position on the political spectrum scores better than any other. We measure how people do discourse. Do you give reasons, or just assert? Do you respond to what your opponent actually said? Can you concede a good point? But also: do you argue with passion? Do you fight your corner, and is your opponent an adversary to beat or an enemy to destroy? Computational research has so far captured only the easiest slice of this at scale: toxicity, incivility, sentiment. To decide which differences actually make a difference, we turn to political philosophy, above all the long-running debate between Jürgen Habermas and Chantal Mouffe. Habermas holds up an ideal of reasoned argument aimed at mutual understanding; Mouffe replies that passionate conflict is the lifeblood of politics, not its failure. The first half of our questions is Habermas’s; the second half is Mouffe’s. From their debate we distill 56 distinct, continuously scored dimensions of discourse quality.

The full dataset and measurement infrastructure will be open-sourced.

Team
Tim Booker (lab); Friederike Stock (MPI-HD); Mubashir Sultar (MPI-HD); Fabian Veider (University of Graz)
Funding
DeSiRe project
Responsibility
Tim Booker
Duration
2025–2027
03

The Geopolitics of Media Bias

How news outlets around the world frame the victims of conflict: 1.36 billion articles, over a hundred languages, ten years of coverage; outlets compared on the same events, so that differences in framing cannot be blamed on differences in the facts.

Read project

When people are killed in a conflict, the news must decide how to tell it. Some victims arrive as people, with names, ages, and families; others arrive as a number. Some perpetrators are named; others vanish into passive constructions, where people are simply “killed in a strike.” Beyond the wording, there is the narrative a death is folded into: a measured response, an unprovoked atrocity, a tragic accident. Whose casualty figures are stated as fact, and whose are “claimed”? Who is quoted, and who is merely described? And which facts, reported by every other outlet covering the story, does an article quietly leave out? We measure these framing choices across the world’s news: roughly 1.36 billion articles in over a hundred languages, from 2016 to 2026, spanning many conflicts at once.

The question is what drives these choices. A benign explanation is severity: bigger events get more names, more detail, more attribution. Our hypothesis is that framing instead tracks the outlet’s geopolitical alignment with the victims’ side; the same death is humanised by one outlet and tallied by another. Testing this fairly means comparing like with like, so the core of the project is an event-matched corpus: articles across languages and outlets are grouped by the real-world event they report, and each is distilled into a structured record of who did what to whom. Holding the event fixed, whatever variation remains in framing belongs to the outlet, not to the facts. It also makes omission measurable for the first time; only by knowing what an event’s other reporters said can you see what one article left out.

The event-matched, fact-extracted corpus and all framing measures will be openly released.

Team
Tim Booker (lab); Elisabeth Höldrich (lab); Kirill Solovev (lab); Mathias Angermaier (lab); Ruggero Marino Lazzaroni (lab)
Funding
Internal project
Responsibility
Tim Booker, Elisabeth Höldrich
Duration
2026–2027
04

Orientation in Conspiration (ORION)

Conspiracy Theories on German-Language Telegram

Read project

For most of the last century, conspiracy theories looked like a harmless curiosity. The pandemic ended that. Online conspiracy groups spread claims that did concrete damage, and what to do about it became a political question nobody has answered well. ORION works on the part of it that is empirically tractable: which factors drive the spread, and how much each contributes. We treat dissemination as something individual, social and technological factors produce together, and try to separate them out.

Our material is the “Schwurbelarchiv”, an archive of German-language public Telegram channels released in 2022, covering roughly 2019 to 2022 and running to over 23 terabytes. Most of it is multimedia. Cleaning, structuring and transcribing that content is a substantial piece of the project in its own right, and it is what makes the archive usable by anyone else.

A second strand asks how far this discourse carries elements of fascist ideologies and narratives. Using Abstract Meaning Representation to extract predicate-argument structures across QAnon and Reichsbürger channels, we measure two things: how far the national collective is spoken about as a living organism rather than a legal or institutional one, and how often the demand is rebirth rather than reform. Both come from Roger Griffin’s definition of the fascist minimum, ultranationalism fused with a myth of rebirth.

Results
Characterizing the Dynamics of Conspiracy-Related German Telegram Conversations during COVID-19Journal of Quantitative Description: Digital MediaARTICLE
The Schwurbelarchiv: a German Language Telegram dataset for the Study of Conspiracy TheoriesEPJ Data Science
Team
Jana Lasser (PI); Elisabeth Höldrich; Mathias Angermaier
Funding
FWF; netidee (Austria's open-source internet funding programme)
Responsibility
Jana Lasser, Elisabeth Höldrich
Duration
1 May 2024 – 30 April 2028
05

Computational Affective Science

Using computational methods to understand the emotional life of humans.

Read project

Emotions are a fundamental part of human existence. They inform us about our attitudes towards people, objects, and situations, and guide the actions we take in response. People are not only subject to their emotional experiences, but actively regulate them, downregulating and upregulating both positive and negative emotions depending on the context. This regulation can be directed at one’s own emotions or at the emotions of others. Understanding how humans experience and regulate their emotions means understanding a fundamental mechanism through which people navigate the world.

In this research area, we work on understanding specific emotion regulation strategies and the broader emotion regulation landscape, also with a focus on human-AI interaction. We developed a data-driven, computational method for defining psychological constructs, applying it to map the full space of emotion regulation strategies people use, extending beyond the narrow focus on reappraisal that has dominated prior research. We further zoomed in on reappraisal, mapping different types of reappraisal strategies and testing whether some are more effective than others. We too compared human and AI reappraisal ability, having both humans and GPT-4 reframe negative emotional scenarios, where GPT-4’s reappraisals were judged as more effective than those produced by most humans. Across these projects, we draw on natural language processing methods, including sentence embeddings, clustering, large language models, and classification, to analyze emotional experiences at scale.

Looking ahead, we aim to broaden this work to understand how affect shapes everyday decision-making and how people navigate their world more generally, examining its relationship to cognition and the emotional appraisal processes.

Team
Alina Herderich (lab); Amit Goldenberg (Harvard University)
Responsibility
Alina Herderich
Duration
2022–today
06

A new ontology of truth

When gut feelings are perceived as honesty, how does that affect the information ecosystem and political functioning?

Read project

Misinformation is increasingly perceived as a major problem for societal cohesion and democracy. Much attention has focused on the role of social media as a vector of misinformation. The role of political leaders has attracted less scrutiny, even though leaders demonstrably influence media coverage and public opinion, and enact political processes crucial for societies to function. Politicians rely on two different rhetorical approaches to express their pursuit of truth: Evidence-based language relies on data and facts to describe elements of external reality. Intuition-based language relies on feelings, instincts and personal values to describe elements of internal experience. A productive democratic discourse strikes a balance between evidence-based and intuition-based conceptions of truth. In this project, we ask whether the language politicians in the US use online and in Congress has changed over time and how it relates to the trustworthiness of information they convey and the functioning of the political system.

We analyse communications by members of the U.S. Congress on Twitter between 2011 and 2022 and show that political speech has fractured into two distinct components related to intuition-based belief-speaking and evidence-based fact-speaking. We also find that conservatives in the US uniquely show a steep rise in the untrustworthy information they spread. We show that belief-speaking is distinctly related to spreading of untrustworthy information for Republicans. Conversely, an increase in fact-speaking language is associated with an increase in the quality of sources cited by both parties. The results support the hypothesis that the current dissemination of misinformation in political discourse is in part driven by an alternative understanding of truth and honesty that emphasizes invocation of subjective belief at the expense of reliance on evidence. Next to the online presence of politicians, we analyse the linguistic traces of evidence- and intuition-based perspectives in speeches in the US Congress from 1879 to 2022. We find that evidence-based language has continued to decline since the mid-1970s, together with a decline in legislative productivity. The decline was accompanied by increasing partisan polarization in Congress and rising income inequality in society. The results highlight the importance of evidence-based language in political decision-making.

Results
From alternative conceptions of honesty to alternative facts in communications by U.S. politiciansPNAS NexusDOIREPRODUCTION PACKAGE
Computational analysis of US Congressional speeches reveals a shift from evidence to intuitionNature Human BehaviourDOI
Social media sharing of low quality news sources by political elitesPNAS NexusDOI
Republicans are increasingly sharing misinformation, research findsThe Washington Post – Monkey Cage, 2022ARTICLE
US politicians tweet far more misinformation than those in the UK and Germany – new researchThe Conversation, 2022ARTICLE
Donald Trump's truth: why liars might sometimes be considered honest – new researchThe Conversation, 2023ARTICLE
Team
Jana Lasser (lab)
Collaborators
Stephan Lewandowsky; David Garcia; Segun Taofeek Aroyehun; Almog Simchon; Fabio Carrella
Funding
ERC Advanced Grant awarded to Stephan Lewandowsky
Responsibility
Stephan Lewandowsky (Stephan.Lewandowsky@bristol.ac.uk)
Duration
1 October 2021 – 30 September 2027
02.1.02

Designing Prosocial Online Spaces

01

Designing Social Media Recommendation Algorithms for Societal Good (DeSiRe)

alternative content recommendation algorithms for social media platforms

Read project

In this project we aim to research alternative content recommendation algorithms for social media platforms. Social media platforms are central for exchange of information in today’s societies. Yet evidence is mounting that they play a causal role in deteriorating social cohesion and civic discourse. User behaviour on social media platforms is governed by content recommendation algorithms that maximise engagement, leading to unintended consequences such as the promotion of outrage or strongly emotionalised content. The EU’s newly enacted Digital Services Act mandates social media platforms to assess and reduce their systemic risks for society – for example by adapting their content recommendation algorithms. The abstract risks, such as the risk to civic discourse however first need to be translated into concrete changes in content recommendation algorithms. In this project we aim to bridge this gap by combining approaches from social science and computer science to incorporate the reduction of risk to civic discourse into content recommendation algorithms of social media platforms.

To this end, we will employ a participation-based approach to develop novel algorithms that consider various aspects of civic discourse, such as information quality and diversity, and the civility of language. To experiment with new algorithms, we will develop Open Source digital twins of social media platforms since experimentation with new algorithms on live social media platforms independent of the permission and influence of platform companies is impossible. Next to the reduction of risk to civic discourse we aim to balance interventions in algorithms with freedom of expression. To this end, we will solicit people’s preferences in different scenarios such as a public health crisis and elections, and develop balanced algorithms. Ultimately, our research will hopefully inform recommendations for the regulation of social media platforms under the Digital Services Act and help transform social media platforms into a technology that is positive for society.

Results
Alternative recommendation algorithms as an antitrust remedy in digital (democracy) casesJournal of European Competition Law & Practice
Team
Jana Lasser; Tim Booker; Ruggero Lazzaroni; Nikolaus Pöchhacker
Funding
FWF & netidee (Austria’s Open Source internet funding campaign)
Responsibility
Research group for Complex Social & Computational Systems
Duration
1 January 2025 – 31 December 2029
10

Countering hate in online spaces

Understanding counter speech as a citizen-based, collective answer to hate on social media platforms.

Read project

Social media platforms are plagued by hate speech from individuals, but also from organized groups that systematically target and discredit politicians and public figures. This coordinated hostility shifts public discourse and distorts the perceived opinion of the majority, as people holding more moderate views are pushed into silence. One common response to online hate is top-down moderation, in which platforms themselves define and remove hate speech. However, platforms are financially motivated rather than democratically legitimized, and their authority to determine what counts as hate speech is neither transparent nor accountable.

An alternative approach is counter speech, in which citizens actively oppose online hate, for example by openly disagreeing with perpetrators. Since online spaces have become a crucial part of public dialogue and are systems naturally resisting top-down control, understanding the dynamics and effectiveness of counter speech is vital for public discourse. In this project, we analyzed a large dataset of over 130,000 conversations on Twitter in Germany during a period when two organized and opposing groups were active: Reconquista Germanica, a hate speech group supporting right-wing political agendas; and Reconquista Internet, a counter speech group founded explicitly to oppose it. We classified the tweets along several discourse dimensions, including argumentative strategies, in- and outgroup utterances, emotions, and hate speech, and found that simply voicing an opinion without resorting to insults is a common strategy that simultaneously shifts discourse in a more civil direction. To this end, we employed a mixed-methods pipeline combining grounded theory, large-scale text classification, and quasi-causal statistical models.

Together with other researchers, we translated these findings into policy recommendations on AI and counter speech. More generally, our work aims to inform and empower anyone wanting to speak up against hate in online spaces.

Team
Jana Lasser (lab); Alina Herderich (lab); Mirta Galesic (Complexity Science Hub Vienna); Joshua Garland (Arizona State University); David Garcia (University of Konstanz); Segun Aroyehun (University of Konstanz)
Funding
NSF DRMS 1757211
Responsibility
Jana Lasser, Alina Herderich
Duration
2022–2026
02.1.03

Curating Open Datasets for Computational Social Science

08

Infini-News: Efficiently Queryable Access to 1.3 Billion Processed Common Crawl News Articles

Enabling reproducible research on online news corpora in hundreds of languages

Read project

Large-scale news corpora are at the core of many a research field, including, but not limited to, computational social science, political science, and natural language processing. Unfortunately, accessing them is non-trivial and is associated with steep resource and monetary costs. The goal of Infini-News is to provide researchers with a way to build their own news corpora in a streamlined manner, reducing the barriers to entry and improving transparency and reproducibility.

Infini-News leverages the entirety of Common Crawl News and applies complex pre-processing and enrichment, as well as infini-gram-mini indexing, to make the entire decade-long corpus searchable in less than a second. To put this in numbers, we provide 1,357,027,742 articles across 331 languages, 17 topics, and 59,885 domains, with a country of origin resolved for 83.4% of articles across 222 countries, distilled from 180 TB of raw web archives, covering August 2016 to April 2026. In particular, Infini-News provides:

  • Clean text and metadata: each article is extracted from HTML and contains title, author, publication date, and source.
  • Enrichment: we detect the language, the country of origin, and a topic from the standard IPTC news taxonomy.
  • Instant search: any word or phrase from across the entire corpus is queryable in under a second, with exact match counts.
  • A query API: we host an API allowing users to count, explore, and extract subdatasets with a simple interface (early access — full university deployment in progress).

The corpus (3.4 TB Parquet) is hosted on Hugging Face, with access gated to institutional affiliation and a stated research purpose; the indexes (2.2 TB) and all pipeline code are distributed as free and open source artefacts. Code is MIT-licensed and derived data CC-BY 4.0, while the article text itself remains under the original publishers’ copyright. The dataset and pipeline are described in the accompanying paper (Lazzaroni, Lasser & Solovev, 2026).

Team
Ruggero Lazzaroni (lab); Kirill Solovev (lab); Jana Lasser (lab)
Funding
European Union's Horizon Europe programme, grant agreement No 101177310. Views and opinions expressed are those of the authors only and do not necessarily reflect those of the European Union.
02.1.04

Developing Methods for Computational Social Science

07

Computational reproducibility and methodological rigour across the sciences

Evidence, tools, and methods for more reliable research

Read project

Research domains have varying modes of establishing validity and reliability of their findings. High-profile cases of research fraud, retractions, and widespread attempts to reproduce previous research findings have contributed to two overlapping discourses that analyse and seek to improve perceived shortcomings: open science and meta-science.

We contribute to this body of work through three research strands:

  1. Investigating practices that can increase computational reproducibility. For this, we use a range of methods, such as scoping reviews of existing evidence, randomised controlled trials to assess interventions, and computational modelling to expand our understanding of how contextual factors interact with proposed interventions.
  2. Building tools to assess computational reproducibility. Here we investigate the extent to which we can leverage Large Language Models to automate the computational reproduction of published findings.
  3. Investigating how causal inference techniques can lead us to reassess existing findings or improve methodological approaches when studying science itself. Using simulations and case studies, we analyse existing measurement approaches in science studies and develop improved computational methods.
Results
Introduction to structural causal models in science studiesQuantitative Science Studies, 7, 159–178DOI
The paradox of competition: How funding models could undermine the uptake of data sharing practicesResearch Policy, 54(10), 105340DOI
Open science interventions to improve reproducibility and replicability of research: A scoping reviewRoyal Society Open Science, 12(4), 242057DOI
The academic impact of Open Science: A scoping reviewRoyal Society Open Science, 12(3), 241248DOI
Team
Thomas Klebel (lab)
Collaborators
Federico Bianchi; Matthew Cannon; Eva Kormann; Adrian Marangoni; Tony Ross-Hellauer; Eric J Schares; Flaminio Squazzoni; Rebecca Taylor-Grant; Vincent Traag; Olmo van den Akker
Responsibility
Thomas Klebel
Duration
2022–today
09

VALPOP: Valuing Public Goods in a Populist World

Extracting signed, typed, multiplex relation networks from news articles at multi-million scale

Read project

VALPOP starts from the notion that populism and the erosion of the rule of law are networked phenomena that could be unearthed by uncovering the relationship networks between political, economic, and media actors. However, to begin this process, we first need to make the relations visible.

To uncover the relationships, we need a data source and machinery to process it at scale. The former are the news articles; they provide a timed, grounded, and multifaceted account of societal happenings. News covers changes in the societal structure both on high and low levels. While some countries might have a more biased coverage, that does not preclude the analysis, as every extracted relation carries the stance of the article that reported it, making the framing an additional source of insight. That is where the latter part of our project comes in: by utilizing an LLM-assisted joint entity and relation extraction pipeline, we are able to process millions of articles in complex ways.

As part of the mission statement at IDea_Lab, we go beyond direct applicability to our projects and try to provide similar tools to a wider research community. For VALPOP this materialises in the extraction pipeline being provided as free and open source, with every part interchangeable for any type of extraction one might want to produce; most importantly, the ontology, prompts, and sources of ground truth can all be replaced in accordance with particular research project needs.

For VALPOP, the pipeline:

  • Finds the actors: the people, organizations, and institutions mentioned in the text.
  • Resolves who they are: if the entities have an entry in Wikidata or Orbis, they are assigned canonical IDs, allowing us to trace them across countries, languages, and spelling variations.
  • Extracts the relationships: a large language model uses our project-specific ontology (110 actor types and 99 relationship types) to extract relations, keeping results comparable across articles, languages, and countries.
  • Builds the graph: all the entities and relationships are merged into a single multiplex knowledge graph for further analysis.

Every connection carries a time stamp and traces back to the article that reported it, and the pipeline’s output is validated against a hand-annotated reference standard. The result is a comparable, longitudinal map of societal networks across the 12 VALPOP countries.

Team
Kirill Solovev (lab); Jana Lasser (lab)
Funding
European Union's Horizon Europe programme, grant agreement No 101177310. Views and opinions expressed are those of the authors only and do not necessarily reflect those of the European Union.