When a Machine Sounds More Like a Politician Than the Politician

Imagine reading two answers to the same tough question at a town hall. One came from the politician who was actually asked. The other was written by a language model given nothing but a short biography of that person. Now imagine most readers pick the second one as the more genuine of the two. That is roughly what a large study out of Germany found, and it complicates a lot of comfortable assumptions about what "sounding authentic" really tells us.
A team led by the computer scientist Steffen Herbold ran the experiment using a rich, messy, real-world source: the long-running British television program Question Time, where public figures field unscripted questions from an audience. The researchers pulled 520 audience questions and the responses that 112 different public figures gave across episodes aired between mid-2020 and late 2021. Then they asked GPT-4 Turbo to answer the same questions in the voice of each speaker, handing the model a biography and telling it to stay conversational and keep to around 200 words [1].
What nearly a thousand people decided
The headline result is uncomfortable. A nationally representative sample of 948 British adults consistently rated the AI-written responses as more authentic, more coherent, and more relevant than the words the real people had spoken [1][2]. That pattern held whether participants judged a single answer in isolation or compared a human answer directly against a synthetic one.
Why would a machine out-authentic a human? Part of it is probably the format. A live broadcast answer is improvised under pressure, full of hedges, tangents, and half-finished thoughts. A model working from a prompt produces something smoother and tidier, the kind of prose that reads like a person choosing every word with care. We seem to mistake that polish for sincerity. The genuine article, delivered in real time by a nervous human, can come off as the less convincing one precisely because it is real.
The part worth pausing over
There is a detail in the findings that should give anyone pause. In about a quarter of the cases where participants noticed a difference between the real and synthetic answers, the AI had generated a stance that actually contradicted the public figure's known position [1]. So the model was not just paraphrasing more elegantly. It was, at times, putting confident, plausible, well-structured words in someone's mouth that the person would likely reject.
That gap between how something reads and whether it is true sits at the center of this research. We tend to treat fluency and warmth as signals of honesty, because for most of human history the person talking to us was, in fact, a person with a mind and a motive. A system that can manufacture that signal on demand quietly breaks the link. The same instinct that makes us reach for classic persuasion tactics when we want a chatbot to cooperate is the instinct these synthetic answers exploit in reverse: they feel human enough that we lower our guard.
Perhaps the most striking social finding came after the debrief. Once participants learned they had been evaluating AI impersonations, more than 90 percent still said they were comfortable with AI playing a role in political communication [1]. Learning that the technology could convincingly impersonate real figures did not dent people's willingness to accept it. That is either a sign of admirable pragmatism or a warning about how quickly we normalize a tool whose risks we have barely started to map.
How much weight to put on this
A few limits keep this from being a verdict on AI and democracy as a whole. The study drew on a single television format in a single country, so British viewers reacting to a particular style of broadcast debate may not represent how people elsewhere respond to campaign ads, social feeds, or prepared speeches. Only one model was tested, and models change fast. And the comparison itself was arguably lopsided: unscripted live speech was pitted against text the AI could refine, which may inflate the authenticity gap rather than reveal a general truth about machines versus humans.
Those caveats matter, but they do not dissolve the core observation. When smoothness and coherence are the cues we lean on to judge whether a message feels real, a language model can supply those cues better than a flustered human, whether or not the underlying content is accurate or even honest. That is a psychological vulnerability, not a technical one, and it does not get patched by a software update.
The practical takeaway is not to distrust everything you read. It is to notice how much of your trust rides on style. An answer that flows, stays on topic, and sounds like the person you expected is doing a lot of work to earn your belief, and increasingly that work can be done by something with no beliefs at all. For more on how these systems reshape everyday judgment, see our wider artificial-intelligence coverage.
Sources
- Herbold, S., Trautsch, A., Kikteva, Z., & Hautli-Janisz, A. (2026). LLM-impersonated debate contributions are more authentic, relevant and coherent than their original: A representative study using BBC1's Question Time. PLOS One. https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0347757
- EurekAlert / PLOS. (2026). AI impersonations of political debaters rated more authentic than real responses. https://www.eurekalert.org/news-releases/1133435
This article summarizes published research for general informational purposes only and does not constitute professional advice.
Frequently asked questions
- What did the study actually compare?
- Researchers took real answers that panelists gave on the BBC program Question Time and generated alternative answers with GPT-4 Turbo, feeding the model a short biography of each speaker. A representative sample of 948 British adults then rated the real and synthetic answers on authenticity, coherence, and relevance, sometimes side by side and sometimes one at a time.
- Did people know which answers were written by AI?
- Not while they were rating. The whole point was to see how the texts landed on their own merits. Only afterward were participants told the truth about the experiment, and even then more than 90 percent said they still supported the use of AI tools in political communication.
- Does this mean AI is better at politics than humans?
- No. The study measured perceptions of a few written qualities, not the accuracy or honesty of the content. In roughly a quarter of mismatched cases the AI even invented positions that contradicted the real person's views, so 'sounds convincing' and 'is trustworthy' are clearly not the same thing here.
Comments (0)