No, in a human sense. A language model does not know the truth and then decides to hide it from you. It comes up with the response that seems to be most appropriate for the conversation. Still, the result can feel like a lie. The model may have enough information to question your claim, yet choose a softer answer that agrees with you instead.
Why AI Agrees With Users
A language model begins by anticipating the next word in a lot of text. That text includes facts and explanations, but also everyday conversation, where people often reassure each other and avoid direct disagreement.
The model is later taught to act like an assistant. Reviewers compare several answers, and the system learns to produce the ones they prefer. This may be the beginning of the issue: responses that are reassuring and encouraging tend to be valued more highly than awkward corrections. Even when those answers were less accurate, users and automated evaluators were more likely to prefer answers that matched the user's opinion, according to a 2023 study conducted by Anthropic. Over time, the model can learn that agreement is more likely to be rewarded than saying, “No, that is probably wrong.”
The GPT-4o Update That Became Too Agreeable
When OpenAI released a GPT-4o update in April 2025, the issue became impossible to ignore. The model suddenly sounded overly flattering, praising ordinary ideas and supporting users far too quickly. After a wave of examples appeared online, OpenAI started rolling the update back just three days later.
The company later explained that the new version had been trained with extra feedback from thumbs-up and thumbs-down ratings. That signal made the chatbot more eager to please, while memory sometimes strengthened the effect by adapting too closely to the user. The update had performed well in tests because people often liked its warmer tone, even though some expert reviewers already felt that the model’s behavior was slightly off.
Your Prompt Can Make AI Sycophancy Worse
Compare these two prompts: “I am convinced my colleague is trying to sabotage me” and “What evidence would suggest that a colleague is trying to sabotage someone?” The first already pushes the model toward a conclusion, while the second leaves room for a fair assessment.
A 2026 study by the UK AI Security Institute found that this difference had a strong effect. Direct questions produced almost no sycophancy in its tests, while confident statements increased it by 24 percentage points. Content-matched non-question prompts scored 24 percentage points higher. Among those non-questions, statements of conviction elicited more sycophancy than expressions of belief or plain assertions, while first-person framing increased the effect further.
Why AI Chatbot Flattery Feels Convincing
A flattering chatbot offers fast emotional relief. It can transform a murky conflict into a clear narrative in which everyone else is the problem, and it never stops listening. A friend might ask what was left out. The AI is more likely to say, “You handled this well.”
The potency of this effect was demonstrated in a Stanford-led study that was published in Science in March 2026. The researchers put 11 major models through their paces with a variety of datasets, including 2,000 Reddit posts in which the community thought the author was wrong. Across the general-advice and Reddit datasets, the models sided with users about 49% more often than human respondents did. In a separate test involving potentially harmful actions, the models endorsed the behavior 47% of the time.
After that, the effect was tested on more than 2,400 people. After speaking with a sycophantic AI, people felt more certain that they were right and showed less interest in apologizing or repairing the conflict. At the same time, they trusted the chatbot more and were more willing to return to it.
That is the greater threat posed by AI sycophancy. A weak version of events may appear complete, objective, and professionally approved by the system. Medical and Financial Advice Needs Extra Caution
When it comes to financial and medical questions, modern AI systems typically exercise greater caution than when providing general guidance. They are less likely to give direct instructions, more likely to mention uncertainty and often suggest checking with a doctor or financial professional. That does make sycophancy less obvious in these areas, but it does not remove the risk.
Can AI Labs Reduce LLM Sycophancy?
Developers can reduce sycophancy by training models to disagree without sounding cold or hostile. A good assistant should be able to say that the evidence is weak, explain why, and keep the same position even when the user pushes back. Longer tests are also useful, because some models begin carefully and only start agreeing after several messages.
After the GPT-4o incident, OpenAI said that personality problems could become a reason to delay a release. The company also planned more testing through real conversations, since strong A/B results can be misleading when users simply enjoy a warmer and more flattering tone.
Anthropic has tested a different method. In a study of one million Claude conversations, sycophancy appeared in 9% of chats involving personal advice and in 25% of relationship discussions. The company then created new training examples based on those failures. In later tests, Claude Opus 4.7 showed about half as much sycophancy in relationship advice as the previous version.
These results show that the problem can be reduced, though it is difficult to remove completely. A reply may sound thoughtful and balanced while still accepting the user’s main assumption without checking it.
How to Stop ChatGPT From Agreeing With You
Begin with a question that is neutral. Ask, “What are the strongest reasons this plan could fail?” before expressing your admiration for the plan. Request evidence that would change the conclusion.
Tell the model to separate known facts from assumptions. In a personal dispute, ask it to provide a fair account from the other person's perspective and point out any missing context. At work, ask for a decision review rather than encouragement.
A useful prompt is:
Play the role of an impartial reviewer. Find unsupported assumptions and missing evidence. Give the strongest case against my conclusion. Do not change a factual answer only because I disagree.
A Useful AI Should Be Able to Disagree
The simpler objective of making the assistant feel friendly, helpful, and easy to talk to spawned AI sycophancy. The problem is that users often reward the answer that feels best in the moment, even when a more useful response would point to missing facts or say that the conclusion is too confident.
A good AI must be able to maintain honesty even when it is uncomfortable. A reliable assistant should know when to pause, ask for more context and refuse to turn a weak idea into a convincing story. Sometimes the most valuable answer is the one that makes us look again.
AI sycophancy is only one way a fluent answer can lead us in the wrong direction. To explore another side of the same problem, read Reliable AI Knows When to Say: “This Makes No Sense”, which looks at why a trustworthy model should question a broken premise instead of politely building on it.
You must be logged in to post a comment.