All publications

Can AI debias the news? LLM interventions improve cross-partisan receptivity but LLMs overestimate their own effectiveness

Can a language model rewrite partisan news so that the other side will listen? In two preregistered experiments with Faisal Feroz, swapping emotive words for milder ones did nothing, while reframing the headline made conservatives rate liberal headlines as more trustworthy — without putting liberals off. But when AI models stood in for the readers, they predicted effects that people did not show.

Read the paper ↗

Computers in Human Behavior Reports · 2026

The paper, in plain language

Two ways of rewriting a headline

We took ten opinion headlines with their teasers from MSNBC — the ten that conservatives in a pretest trusted least — and had a language model, OpenAI's o3-mini, rewrite them. In the first experiment the model was only allowed to replace emotionally charged words with more moderate synonyms: an 'attack' became a 'criticism'. In the second it was allowed to rephrase the headline, lowering the emotional intensity and adopting a conservative communication style while keeping the facts and the viewpoint. Participants in the United States, recruited on Prolific, saw each headline in either its original or its rewritten form and rated how trustworthy and complete it was, and how open they were to the perspective behind it.

What readers did

Word swaps changed nothing: among 176 Republicans, not a single outcome moved. Reframing did. In the second experiment, with 348 participants, conservatives rated reframed headlines as less biased, more trustworthy and more complete, and were more willing to consider the perspective. Liberals did not rate the reframed headlines less favourably on any outcome, so there was no backfire among the outlet's own audience. The effects are modest — fractions of a point on a seven-point scale. And rewriting has a price: independent raters felt that the factual claim had shifted in roughly a quarter of their judgments of the word swaps and in more than a third for the reframed headlines.

What the models predicted

We then let six OpenAI models play the participants, each simulated reader given the profile of one real person. In the first experiment every model predicted clear effects where people showed none. In the second the models pointed in the right direction but predicted significantly larger effects on some outcomes, and three of them expected a backfire among liberals that did not occur. All six models suggested that reframing would work better among people who identify strongly with their political group. No such pattern appeared among people, where — tentatively — those with low trust in the media responded most.

What it does not show

This is a proof of concept on headlines and teasers, not on articles: one outlet, one political direction, one country, self-reported judgments rather than behaviour. The headlines were not selected for being true, and the same technique applied to false content would make misinformation easier to accept — reframing would have to be coupled with verification. The numbers on this page come from the accepted manuscript; the conclusions follow the published article. Language models can generate candidate interventions, but they cannot yet judge whether their own interventions work. That still takes people.