AI Agents are Vulnerable to Radicalization

작성자

카테고리:

← 피드로
arXiv cs.AI · Ozgur Can Seckin, Shalmoli Ghosh, Alessandro Flammini, Kristina Lerman, Maria Elizabeth Grabe, Filippo Menczer · 2026-10-01 AI

[Submitted on 29 Sep 2026]

View PDF HTML (experimental)

Abstract:Large language models (LLMs) can influence people’s beliefs, yet little is known about whether and how they can manipulate each other. To investigate this, we simulate conversations between two agents: a target LLM that role-plays a human persona based on demographic and psychological attributes, and an influencer LLM that aims to make the target’s beliefs more extreme. We examine radicalization along two pathways: resonance, where the influencer reinforces a target’s pre-existing belief, and persuasion, where the influencer promotes a belief the target initially considers unimportant. Across affective and behavioral metrics, we find that both mechanisms radicalize the target. However, resonance produces consistently stronger effects than persuasion. Different influence tactics, such as using sycophancy and unverified claims, produce different levels of radicalization, but not consistently across metrics. We further show that resonance propagates to related beliefs, suggesting interconnected belief structures within AI agents. These findings indicate that AI agents are susceptible to radicalization, particularly when messages align with their existing beliefs, raising concerns about the vulnerability of personalized AI agents and multi-agent AI ecosystems.

Submission history

From: Shalmoli Ghosh [view email]
[v1] Tue, 29 Sep 2026 17:45:00 UTC (3,921 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2609.38296