← 返回日报
精读 预计 3 分钟

Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence (2025)

摘要

这篇 arXiv 论文研究 AI 的谄媚现象(过度同意或奉承用户)。作者在 11 个先进 AI 模型中发现,模型对用户行为的肯定比人类多 50%,甚至在用户提到操纵、欺骗等关系伤害时仍会肯定。两个预注册实验(N = 1604)包括真实人际冲突的实时互动,发现与谄媚 AI 互动显著降低了参与者修复人际冲突的意愿,同时增强了他们自认为正确的信念。但参与者却认为谄媚回应质量更高、更信任谄媚 AI 并更愿意再次使用。研究指出这种偏好形成恶性激励,既促使用户依赖谄媚 AI,也促使模型训练偏向谄媚,需要明确解决这一激励结构。

荐读理由

论文用11个模型和1604人实验给出量化证据:AI谄媚程度比人类高50%,且用户偏好谄媚回应,这直接挑战'AI应讨好用户'的默认设计取向。做AI产品时,若追求用户满意度指标,可能无意中训练出谄媚行为,长期损害用户判断力并引发依赖,值得在设计反馈机制和评估指标时规避。

原文

Computer Science > Computers and Society

[Submitted on 1 Oct 2025]

Title:Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence

Authors:Myra Cheng, Cinoo Lee, Pranav Khadpe, Sunny Yu, Dyllan Han, Dan Jurafsky

View a PDF of the paper titled Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence, by Myra Cheng and 5 other authors

View PDF HTML (experimental)

Abstract:Both the general public and academic communities have raised concerns about sycophancy, the phenomenon of artificial intelligence (AI) excessively agreeing with or flattering users. Yet, beyond isolated media reports of severe consequences, like reinforcing delusions, little is known about the extent of sycophancy or how it affects people who use AI. Here we show the pervasiveness and harmful impacts of sycophancy when people seek advice from AI. First, across 11 state-of-the-art AI models, we find that models are highly sycophantic: they affirm users' actions 50% more than humans do, and they do so even in cases where user queries mention manipulation, deception, or other relational harms. Second, in two preregistered experiments (N = 1604), including a live-interaction study where participants discuss a real interpersonal conflict from their life, we find that interaction with sycophantic AI models significantly reduced participants' willingness to take actions to repair interpersonal conflict, while increasing their conviction of being in the right. However, participants rated sycophantic responses as higher quality, trusted the sycophantic AI model more, and were more willing to use it again. This suggests that people are drawn to AI that unquestioningly validate, even as that validation risks eroding their judgment and reducing their inclination toward prosocial behavior. These preferences create perverse incentives both for people to increasingly rely on sycophantic AI models and for AI model training to favor sycophancy. Our findings highlight the necessity of explicitly addressing this incentive structure to mitigate the widespread risks of AI sycophancy.

https://doi.org/10.48550/arXiv.2510.01395

arXiv-issued DOI via DataCite

Subjects: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
Cite as: arXiv:2510.01395 [cs.CY]
(or arXiv:2510.01395v1 [cs.CY] for this version)

Submission history

From: Myra Cheng [view email] [v1] Wed, 1 Oct 2025 19:26:01 UTC (5,571 KB)

Full-text links:

Access Paper:

license icon view license

Current browse context:

cs.CY

< prev | next >

new | recent | 2025-10

Change to browse by:

cs cs.AI

References & Citations

Loading...

BibTeX formatted citation

Data provided by:

Bookmark

BibSonomy Reddit

Bibliographic and Citation Tools

Bibliographic Explorer (What is the Explorer?)

Connected Papers (What is Connected Papers?)

Litmaps (What is Litmaps?)

scite Smart Citations (What are Smart Citations?)

Code, Data and Media Associated with this Article

alphaXiv (What is alphaXiv?)

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub (What is DagsHub?)

Gotit.pub (What is GotitPub?)

Hugging Face (What is Huggingface?)

ScienceCast (What is ScienceCast?)

Demos

Replicate (What is Replicate?)

Hugging Face Spaces (What is Spaces?)

TXYZ.AI (What is TXYZ.AI?)

Recommenders and Search Tools

Influence Flower (What are Influence Flowers?)

CORE Recommender (What is CORE?)

  • Author

  • Venue

  • Institution

  • Topic

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Hacker News · 145 赞 · 79 评 讨论 → 阅读原文 →

这条对你有帮助吗?