Shuichiro Ogawa
日本語

Notes · updated 2026-09-27

Why Do LLMs Become Sycophantic? How Preference Learning, Internal Circuits, and Input Framing Produce the Behavior

This note draws on 37 academic sources from 2022–2026 to explain why LLM sycophancy occurs. There is no single source.

Contents (9)
  1. A model that apologizes for a correct answer
  2. Scope and method
  3. What counts as sycophancy?
  4. Where is sycophancy produced?
  5. What happens inside the model?
  6. What elicits sycophancy?
  7. How far do mitigations go?
  8. How are sycophancy and self-correction related?
  9. Gaps in the literature

A model that apologizes for a correct answer

Sharma et al. (2023) asked AI assistants a question, let them answer correctly, and then replied only “Are you sure?” On questions where its first answer was correct, Claude 1.3 admitted a mistake 98% of the time. No new evidence was given. The model withdrew its answer anyway.

Did the model not know the answer? Evidence gathered over the past few years says otherwise. It knows, and it sides with the user. This behavior is sycophancy, and this note organizes the 2022–2026 academic literature on where it is produced, what happens inside the model, and what triggers it.

Why Do AI Review-and-Fix Loops End Up Building a Slot Just for 10-Yen Coins? looked at AI-fixes-AI review loops in which the fixer answers “That’s a valid point” and piles on features, and cited sycophancy as the underlying mechanism. This note expands on what that mechanism actually is.

Scope and method

  • Topics: definitions and types of LLM sycophancy, its sources, internal mechanisms, contextual triggers, mitigations and their limits, and its relation to self-correction.
  • Period: 2022–2026 (as of September 2026).
  • Corpus: 37 sources (source/review/llm-sycophancy-mechanisms/papers.md): 22 peer-reviewed conference, journal, or workshop papers (including accepted ones), 1 report from the UK AI Security Institute, and 14 arXiv preprints. The search returned 36 sources with no duplicates or exclusions, and one user study was added during verification.
  • Verification: the full PDF text of all 36 arXiv sources was retrieved, and every number, direct quotation, and mechanism claim this note uses was checked at the relevant passage. Venues were confirmed through Crossref, OpenReview, arXiv author comments, and publisher DOIs. Where the check found overstatement (the effect of activation steering, the circuit after an RLHF update, how often reward gaps occur), the claim was narrowed to what the text supports.

What counts as sycophancy?

Sycophancy research has used one word to measure different behaviors. Ye et al. (2026) reviewed 70 papers on sycophancy and classified their definitions and measurements along two axes. The first is whether the model defers to the user’s positions and beliefs, or to the user’s personal traits and emotions. The second is whether the deference appears in explicit language or implicitly, through framing, omission, or tone. Existing research concentrates on explicit agreement with beliefs. In a survey of 106 experts, 94.3% said sycophancy is a significant problem in current AI systems, but they disagreed substantially about which behaviors qualify.

The types that have actually been measured line up roughly as follows.

  • Opinion conformity: returning the answer that matches the opinion the user stated (Perez et al. 2023; Wei et al. 2023).
  • Feedback sycophancy: rating a text more favorably when the user says “I wrote this” or “I really like this” (Sharma et al. 2023).
  • Answer withdrawal: abandoning a correct answer after “Are you sure?” or a rebuttal (Sharma et al. 2023; Xie et al. 2024; Wang et al. 2023).
  • Mimicry: carrying over the user’s mistake, such as a misattributed poem (Sharma et al. 2023).
  • Social sycophancy: excessively preserving the user’s face (desired self-image). According to Cheng et al. (2026, ELEPHANT, ICLR 2026), 11 models preserved the user’s face 45 percentage points more than humans on average, in general advice queries and in queries describing clear wrongdoing by the user; in moral conflicts, they affirmed whichever side the user took in 48% of cases.

Evidence from inside the model also supports treating these as distinct. Vennemeyer et al. (2025) showed that sycophantic agreement, sycophantic praise, and genuine agreement are encoded along separate linear directions in latent space, and that amplifying or suppressing one leaves the others unchanged. Tasks differ as well. Ranaldi and Pucci (2023) report that sycophancy is prominent in tasks that ask for subjective opinions, while in tasks with a single correct answer, such as math, models are less likely to follow the user’s hints.

Changing an answer to match the user is not always sycophancy, though. If the user supplies valid evidence, changing the answer is the right thing to do. Ma et al. (2026) separate “unsupported yielding” from “rational updating” and measure them independently. SycEval (Fanous et al. 2025) likewise distinguishes changes toward the correct answer from changes away from it after a rebuttal, observing them in 43.52% and 14.66% of cases respectively. The problem is not that the answer changes, but that it changes regardless of evidence.

Where is sycophancy produced?

The sources cannot be reduced to one. The literature points to at least four pathways, each supported by different evidence.

Scale and instruction tuning

The first clue was the observation that larger models are more sycophantic. Perez et al. (2023, Findings of ACL), using evaluation data written by models themselves, reported that larger LMs repeat back the user’s stated views. In the largest 52B models, more than 90% of answers matched the user’s view on NLP research and philosophy questions. Sycophancy was similar across numbers of RL steps, including zero (the pretrained LM). The same figure shows that the preference models used for RL themselves favor sycophantic answers. The authors write that RLHF does not train away the sycophancy already present after pretraining and may give models an incentive to keep it. Wei et al. (2023) measured this in the PaLM family: scaling from PaLM-8B to 62B increased sycophancy by 19.8%, and scaling to 540B added another 10.0%. Instruction tuning also mattered: Flan-PaLM-8B repeated the user’s opinion 26.0% more often than the base PaLM-8B. The authors themselves write that there is no immediately clear reason why scale should increase sycophancy.

Later work turned “larger means more sycophantic” into a conditional claim. De Marez et al. (2026) studied 56 models from 0.3B to 32B and 13 manipulation types and split the rate of answer flips into two components. One is how strongly the model prefers the correct answer without pressure (truth margin); the other is how far pressure shifts that preference (manipulation sensitivity). Vulnerability is governed mainly by size, but instruction tuning changes how size acts: small instruction-tuned models can become less robust, whereas large ones usually become more robust. Hong et al. (2025), measuring multi-turn dialogue, likewise found that alignment tuning amplifies sycophancy while scaling and reasoning optimization strengthen resistance to user pressure. Wei et al. measured repetition of opinions on questions with no correct answer; De Marez et al. and Hong et al. measured withdrawal of factual answers and stance changes across turns. When the type of sycophancy differs, there is no reason for scale to act the same way.

Human preference data

The second pathway is that human judgment itself favors sycophancy. Sharma et al. (2023) fit a Bayesian logistic regression that explains Anthropic’s hh-rlhf preference data in terms of response features. Matching the user’s beliefs, biases, and preferences was one of the most predictive features of human preference. It was not consistently the single most predictive one, however; the ranking depended on the experimental condition. Truthfulness was rewarded in the same data too.

The effect is clearer on the preference model (PM) side. Claude 2’s PM preferred convincing sycophantic responses that affirm a user’s misconception over baseline truthful responses 95% of the time. Against well-written truthful responses, the truthful one usually won, but for the most challenging misconceptions the PM still chose the sycophantic response 45% of the time. Human raters (crowd workers without internet access) also generally chose the truthful response, but their choices became less reliable as the misconceptions got harder. The authors conclude that simply collecting more non-expert feedback may not be enough to eliminate sycophancy.

User reactions point the same way. In three preregistered experiments (N = 2,405), Cheng et al. (2026, Science) found that even a single interaction with a sycophantic AI reduced participants’ willingness to take responsibility and repair interpersonal conflicts, while increasing their conviction that they were right. Sycophantic models were nonetheless trusted and preferred. In the same study, 11 state-of-the-art models affirmed users’ actions 49% more often than humans, even when queries involved deception or illegality. For social sycophancy, ELEPHANT reports that it is rewarded in preference datasets. A loop suggests itself: users like sycophancy, that liking becomes preference data, and models trained on preferences become sycophantic. No study, however, has demonstrated this loop end to end in a single experiment; each step rests on a separate study.

Amplification through reward optimization

The third pathway is that a small bias in preference data grows during training. Sharma et al. (2023) optimized responses against Claude 2’s PM (with best-of-N sampling and reinforcement learning) and found that stronger optimization increased some forms of sycophancy and decreased others. Because the PM rewards features other than sycophancy, the behavior does not move in one direction. Even so, best-of-N with Claude 2’s PM did not reach responses as truthful as best-of-N with a “non-sycophantic” PM that was given a context asking for truthful answers.

Shapira et al. (2026, ICML 2026) wrote this amplification down formally. The direction in which sycophancy drifts during post-training is determined by the covariance, under the base policy, between endorsing the belief signal in the prompt and the learned reward, and the first-order effect reduces to a simple mean-gap condition. They also characterize when annotator bias induces this reward gap under pairwise reward learning such as Bradley-Terry. As a remedy, they find the policy closest in KL divergence to the unconstrained post-trained policy among those that do not increase sycophancy, and derive the corresponding minimal reward correction as a closed-form “agreement penalty.” In computational experiments, reward gaps were common and caused behavioral drift in all the configurations considered.

Casper et al. (2023, TMLR) surveyed the open problems of RLHF and classified the limitations of human evaluators and misspecification of the reward model as fundamental problems that cannot be solved within RLHF. Sycophancy can be read as a concrete behavioral instance of those limitations.

An entry point to specification gaming

The fourth pathway treats sycophancy as the first step of a broader failure. Denison et al. (2024) trained models on a curriculum of increasingly egregious specification-gaming environments, starting with political sycophancy. When the model was finally placed in an environment where it could modify its own reward function, it sometimes tampered with its reward despite never being trained to do so. It tampered with its reward in 45 of 32,768 trials, and in 7 of those it also rewrote tests to avoid detection. A helpful-only model without the curriculum never tampered in 100,000 trials. The rate is low, but it is the difference between zero and nonzero. Adding harmlessness training did not prevent this generalization.

Sycophancy is the cheapest form of specification gaming: earning reward by pleasing the evaluator. Denison et al.’s results show that reinforcing it in training can spread to other shortcuts for obtaining reward.

What happens inside the model?

So does a sycophantic model not know the right answer, or does it know and defer anyway? Interpretability research is converging on the latter. About half of the work in this section, however, consists of unreviewed preprints (Wang et al. appeared at AAAI 2026, Genadi et al. at EACL 2026, and Pandey at the ICML 2026 Mechanistic Interpretability Workshop), and nearly all of it studies open-weight models.

Wang et al. (2026, AAAI) found that a simple statement of the user’s opinion reliably induces sycophancy, whereas framing the user as an expert has almost no effect. Tracing the behavior with the logit lens and causal activation patching, they found that sycophancy emerges in two stages. In late layers, the output preference shifts toward the user; in deeper layers, the representations themselves diverge. Authority fails to have an effect, they argue, because models do not encode it internally. First-person framing (“I believe…”) induced more sycophancy than third-person framing (“They believe…”), along with larger representational perturbations in deeper layers.

Pandey (2026) showed that across twelve open-weight models (1.5B–72B) from five labs, the same small set of attention heads carries a “this statement is wrong” signal. The signal appears both when the model evaluates a claim on its own and when a user pressures it to agree. Silencing these heads sharply changes sycophantic behavior while leaving factual accuracy intact. The circuit, in other words, controls deference, not knowledge. The RLHF refresh from Llama-3.1 to 3.3 (70B) cut sycophantic behavior roughly tenfold, but the shared heads persisted, and the effect of ablating them by projection grew. The behavior decreased; the circuit that produces it did not disappear.

Genadi et al. (2026, EACL) showed that the signal for sycophancy that moves from a correct to an incorrect answer is most linearly separable in multi-head attention activations, more so than in the residual stream or MLPs. Steering a sparse subset of middle-layer heads was most effective. The direction they found overlaps little with previously identified “truthful” directions (about one third of the top 32 heads are shared), and the authors argue that factual accuracy and resistance to deference arise from related but distinct mechanisms. The influential heads attended disproportionately to expressions of user doubt.

The same picture appears when the error is embedded in a task request. Chen et al. (2026) named the phenomenon “correction suppression”: models that correct a false claim when shown it on its own comply without correcting it when the same claim is embedded in a task request. On 300 false premises across eight models, suppression rates ranged from 19% to 90%, with four models above 80%. The model registers the error internally, but the task context diverts early-layer attention away from the false claim, and output intent crystallizes toward compliance in the middle layers.

These three studies reach the same conclusion by different methods. Detecting an error and selecting the output happen at different stages, and sycophancy happens at the latter. This picture also accounts for Claude 1.3 abandoning correct answers after “Are you sure?” in Sharma et al.

The claim that “there is a linear direction representing sycophancy” has been challenged, though. Aamir and Bin Adil (2026) tested whether, in Qwen2.5-1.5B and Llama-3.2-1B, capitulation under pushback can be read from the residual stream before the response. A naive difference-in-means probe reached an in-sample AUROC of 0.81 and 0.71, but under question-level cross-validation and shuffled-label nulls, the best scores were only 0.582 and 0.548, short of the preregistered usability bar of 0.70. The same pipeline recovered a known direction (whether pushback is present) at AUROC 1.000, so the procedure itself was not weak. This is a negative result in small models and does not directly refute studies of larger ones. Still, reports of directions that have not gone through cross-validation should be read with a discount.

What elicits sycophancy?

If the internal picture is “knowing, then deferring,” the next question is what pushes the model toward deferring. Input framing and accumulated conversational pressure are the best studied.

Dubois et al. (2026) at the UK AI Security Institute presented the same claim as a question and as a non-question, and varied three orthogonal factors among the non-questions. These were epistemic certainty (statement, belief, conviction), perspective (first person versus “the user believes…”), and affirmation versus negation. For GPT-4o, GPT-5, and Claude Sonnet 4.5 alike, responses to questions showed near-zero sycophancy, while non-question inputs expressing the same claim produced markedly higher levels, a 24-percentage-point gap on the grader scale. Sycophancy increased monotonically with the certainty the user expressed, and was amplified by first-person framing. Scoring was done by two LLM-as-a-judge graders.

Multi-turn pressure also matters. Xie et al. (2024) showed that even when the initial judgment is correct, follow-up questions that express doubt, negation, or misdirection make the judgment waver. Hong et al. (2025)‘s SYCON-Bench measures how quickly a model conforms to the user (Turn of Flip) and how often it shifts its stance under sustained pressure (Number of Flip), and found sycophancy to remain a prevalent failure mode across 17 models. In Aamir and Bin Adil (2026)‘s small models, 41.8% and 43.1% of initially correct answers flipped to wrong answers after pushback, while the same pushback repaired initially wrong answers only about 13% of the time. Pushback without evidence lowers accuracy on net.

Studies disagree about which kind of pressure works best. In SycEval (Fanous et al. 2025), preemptive rebuttals produced more sycophancy than in-context rebuttals (61.75% vs. 56.52%), and citation-based rebuttals pulled answers toward incorrect ones most often. Feng et al. (2026, ACL) report that authority bias increases sycophancy, while Wang et al. (2026) find that expert framing has almost no effect. In Aamir and Bin Adil, whether bare doubt or an emotional appeal worked better reversed between model families (“Are you sure?” worked better on Qwen, the emotional appeal on Llama). What works depends on the model and the task, and no general ranking can be stated yet.

How far do mitigations go?

With multiple sources, mitigations also split across multiple stages.

At the data stage, Wei et al. (2023) reported that lightweight fine-tuning on synthetic data (in which a claim’s truth is independent of the user’s opinion) made models repeat the user’s opinion up to 10.0% less often on questions without a correct answer, and kept sufficiently large models from agreeing with clearly incorrect addition statements. Chen et al. (2024) fine-tuned only the modules that most strongly drive sycophancy (under 5% of the model) and reduced sycophancy without degrading general capabilities. At the reward stage, Shapira et al. (2026)‘s agreement penalty aims to neutralize the amplification mechanism itself. Constitutional AI (Bai et al. 2022) replaces human preference labels with AI feedback guided by principles, and is cited as a way to reduce biases that come from human judgment, but that paper does not directly measure its effect on sycophancy.

At inference time there is activation steering. Contrastive Activation Addition (CAA) by Rimsky (Panickssery) et al. (2024, ACL) averages the difference in activations between sycophantic and non-sycophantic examples and adds or subtracts that vector at inference time. Its effect on sycophancy, however, was weaker than on other behaviors and split by evaluation format. In open-ended generation with Llama 2 7B Chat (layer 13, scored by GPT-4 on a 10-point scale), subtracting the vector gave 0.26, no steering 0.58, and adding it 1.26, moving in the expected direction. In the multiple-choice evaluation with Llama 2 13B Chat, the scores were 0.56, 0.63, and 0.60, so adding the vector did not increase sycophancy. The effect on TruthfulQA was also small (subtracting raised it by 0.01 to 0.02 and adding lowered it by 0.03 to 0.05), and the authors themselves call it small. The interventions by Genadi et al., Pandey, and Chen et al. (2026) in the previous section belong to the same line, and Genadi et al. report that targeting specific attention heads works better than steering the residual stream.

At the input stage, Dubois et al. (2026) showed that asking the model to convert a non-question into a question before answering substantially reduces sycophancy, and that the effect is stronger than simply asking it “not to be sycophantic.” In Hong et al. (2025), having the model adopt a third-person persona improved how long it held its position (Turn of Flip) by up to 63.8% in the debate setting. Chain-of-thought reasoning, according to Feng et al. (2026), generally reduces sycophancy in final decisions. In some samples, however, the model builds plausible justifications out of logical inconsistencies, calculation errors, and one-sided arguments, which makes the sycophancy harder to see.

Every stage has limits. One is that suppressing sycophancy also suppresses legitimate changes of mind. Ma et al. (2026, accepted to Findings of EMNLP) found that across representative training-time and inference-time anti-sycophancy methods, reducing unsupported yielding tends to sacrifice rational updating, even when both objectives are optimized jointly. The MLP neurons and attention heads that drive the two behaviors overlap substantially, and their steering directions are positively aligned. The authors argue that anti-sycophancy should be treated as a selectivity problem rather than a suppression problem. Another limit is that the mechanism can remain even when the behavior declines, as Pandey (2026)‘s RLHF result shows. ELEPHANT likewise reports that existing mitigations have limited effect on social sycophancy.

Research on sycophancy and research on self-correction developed separately, but they look at two sides of the same phenomenon.

Huang et al. (2024) showed that when LLMs are asked to review their own answers without external signals of correctness, reasoning performance does not improve and instead declines. On GSM8K, GPT-4 fell from 95.5% to 91.5% after one round of self-correction and to 89.0% after two. On CommonSenseQA, GPT-3.5 fell from 75.8% to 38.1%. Among the answers that changed, correct-to-incorrect changes outnumbered incorrect-to-correct ones. The authors explain that in CommonSenseQA the wrong options often look somewhat relevant to the question, and the self-correction prompt can bias the model toward choosing another option.

This is the same structure as the “Are you sure?” experiment. A review instruction tells the model to doubt its first answer without providing any new evidence. The model takes it as pressure and moves toward changing its answer. Kamoi et al. (2024) re-examined the self-correction literature and concluded that self-correction works in tasks where reliable external feedback is available, and that no prior work demonstrates successful self-correction with feedback from prompted LLMs, except in tasks exceptionally suited to self-correction.

It follows that evidence-free pushback and self-correction prompts can both be treated as inputs that elicit sycophancy. The behavior Wang et al. (2023) showed in a debate format, giving up correct answers in the face of invalid arguments, belongs in the same row. Put the other way, an answer should change when verifiable evidence is added, such as a failing test or a primary source. To reduce Ma et al. (2026)‘s unsupported yielding while keeping rational updating, the input has to make it possible to tell whether evidence is present. Why Do AI Review-and-Fix Loops End Up Building a Slot Just for 10-Yen Coins? recommended not treating another LLM’s review findings as reliable external feedback, and placing an accept-or-reject stage before any fix, so that a human can supply exactly this distinction.

Gaps in the literature

  • End-to-end verification of the loop: No study connects user preferences, preference data, the reward model, and sycophancy in a deployed model in one chain. Each step is supported only by a separate study (Cheng et al. 2026; Sharma et al. 2023; Shapira et al. 2026), and no public study measures the size of the amplification in the actual training of frontier models.
  • Generality of internal mechanisms: Interpretability work centers on small and mid-sized open-weight models, and much of it is unreviewed. Few reports yet meet the validation protocol of Aamir and Bin Adil (2026), and it is unknown which results transfer to larger models.
  • Fragmented definitions: As Ye et al. (2026) showed, definitions of sycophancy differ from study to study, which makes results hard to compare. Implicit sycophancy and deference toward the user’s traits and emotions remain largely unmeasured.
  • Sycophancy in working contexts: Few studies directly measure sycophancy in code generation, code review, or multi-step agent work. Chen et al. (2026)‘s correction suppression comes closest, but it studies single requests.
  • Bias on the measuring side: Many studies, including Dubois et al. (2026), score sycophancy with LLM-as-a-judge. Whether the judges’ own position bias or self-preference (Recent Currents in LLM-as-a-Judge: A Literature Map of 53 Core Studies and a Standalone Chapter on Creativity Evaluation) leaks into these measurements has barely been examined.

Unverified items

  • The share of human raters who preferred sycophantic responses in Sharma et al. (2023): the text gives no number, and it appears only in the Fig. 7b chart. No values were read off the chart; this note uses only the text’s statement that human judgments became less reliable as misconceptions grew harder.
  • Acceptance of Ma et al. (2026) to Findings of EMNLP 2026: based on the arXiv author comment (“Accepted to EMNLP 2026 Findings”). The proceedings are not yet out.
  • Wei et al. (2023): OpenReview records a submission to ICLR 2025, but acceptance could not be confirmed, so it is cited as an arXiv preprint.
  • The arXiv record of ELEPHANT carries the DOI of the Science paper (Cheng et al. 2026), but the two are different papers with different titles and authors. This note cites ELEPHANT as the ICLR 2026 paper and the user experiments as the Science paper.

References

All accessed on 2026-09-27.

Definitions and types

  • Cheng, M., Yu, S., Lee, C., Khadpe, P., Ibrahim, L., & Jurafsky, D. (2026). ELEPHANT: Measuring and Understanding Social Sycophancy in LLMs. ICLR 2026. https://arxiv.org/abs/2505.13995
  • Fanous, A., Goldberg, J., Agarwal, A. A., Lin, J., Zhou, A., Daneshjou, R., & Koyejo, S. (2025). SycEval: Evaluating LLM Sycophancy. AIES 2025. https://arxiv.org/abs/2502.08177
  • Ye, M., Ibrahim, L., Bo, J. Y., Cheng, M., Mattsson, I., Vennemeyer, D., Kraut, R., & Rathje, S. (2026). What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct. arXiv:2605.21778. https://arxiv.org/abs/2605.21778
  • Ranaldi, L., & Pucci, G. (2023). When Large Language Models Contradict Humans? Large Language Models’ Sycophantic Behaviour. arXiv:2311.09410. https://arxiv.org/abs/2311.09410
  • Malmqvist, L. (2024). Sycophancy in Large Language Models: Causes and Mitigations. arXiv:2411.15287. https://arxiv.org/abs/2411.15287

Sources

  • Perez, E., Ringer, S., Lukošiūtė, K., et al. (2023). Discovering Language Model Behaviors with Model-Written Evaluations. Findings of the Association for Computational Linguistics: ACL 2023, 13387–13434. https://doi.org/10.18653/v1/2023.findings-acl.847
  • Wei, J., Huang, D., Lu, Y., Zhou, D., & Le, Q. V. (2023). Simple Synthetic Data Reduces Sycophancy in Large Language Models. arXiv:2308.03958. https://arxiv.org/abs/2308.03958
  • Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., et al. (2024). Towards Understanding Sycophancy in Language Models. ICLR 2024. https://arxiv.org/abs/2310.13548
  • Cheng, M., Lee, C., Khadpe, P., Yu, S., Han, D., & Jurafsky, D. (2026). Sycophantic AI decreases prosocial intentions and promotes dependence. Science, 391(6792), eaec8352. https://doi.org/10.1126/science.aec8352 (the preprint arXiv:2510.01395 reports two experiments with N=1,604; this note uses the numbers from the Science version)
  • Shapira, I., Benade, G., & Procaccia, A. D. (2026). How RLHF Amplifies Sycophancy. ICML 2026. https://arxiv.org/abs/2602.01002
  • Casper, S., Davies, X., Shi, C., Gilbert, T. K., et al. (2023). Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. Transactions on Machine Learning Research. https://arxiv.org/abs/2307.15217
  • Denison, C., MacDiarmid, M., Barez, F., Duvenaud, D., Kravec, S., Marks, S., Schiefer, N., Soklaski, R., Tamkin, A., Kaplan, J., Shlegeris, B., Bowman, S. R., Perez, E., & Hubinger, E. (2024). Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models. arXiv:2406.10162. https://arxiv.org/abs/2406.10162
  • De Marez, V., De Bruyne, L., & Daelemans, W. (2026). Decomposing Factual Sycophancy in Language Models: How Size and Instruction Tuning Shape Robustness. arXiv:2606.06306. https://arxiv.org/abs/2606.06306

Internal mechanisms

  • Wang, K., Li, J., Yang, S., Zhang, Z., & Wang, D. (2026). When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models. Proceedings of the AAAI Conference on Artificial Intelligence, 40(39), 33566–33574. https://doi.org/10.1609/aaai.v40i39.40645
  • Pandey, M. (2026). LLMs Know They’re Wrong and Agree Anyway: The Shared Sycophancy-Lying Circuit. ICML 2026 Mechanistic Interpretability Workshop (OpenReview). arXiv:2604.19117. https://arxiv.org/abs/2604.19117
  • Genadi, R., Nwadike, M., Mukhituly, N., Alquabeh, H., Hiraoka, T., & Inui, K. (2026). Sycophancy Hides Linearly in the Attention Heads. Proceedings of EACL 2026 (Long Papers), 6896–6912. https://doi.org/10.18653/v1/2026.eacl-long.324
  • Vennemeyer, D., Duong, P. A., Zhan, T., & Jiang, T. (2025). Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs. arXiv:2509.21305. https://arxiv.org/abs/2509.21305
  • Chen, Z., Lin, H., Chen, Z., Tian, Y., Yang, G., Wang, D., Guo, Y., Zhu, H., & Cheng, J. (2026). Knowing but Not Correcting: Routine Task Requests Suppress Factual Correction in LLMs. arXiv:2605.05957. https://arxiv.org/abs/2605.05957
  • Aamir, S., & Bin Adil, M. A. (2026). No Usable Linear “Capitulation Direction” in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback. arXiv:2609.17550. https://arxiv.org/abs/2609.17550
  • Zou, A., Phan, L., Chen, S., Campbell, J., et al. (2023). Representation Engineering: A Top-Down Approach to AI Transparency. arXiv:2310.01405. https://arxiv.org/abs/2310.01405

Contextual triggers

Mitigations

  • Chen, W., Huang, Z., Xie, L., et al. (2024). From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning. ICML 2024. https://arxiv.org/abs/2409.01658
  • Rimsky, N. (Panickssery, N.), Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., & Turner, A. M. (2024). Steering Llama 2 via Contrastive Activation Addition. Proceedings of ACL 2024 (Long Papers), 15504–15522. https://doi.org/10.18653/v1/2024.acl-long.828
  • Bai, Y., Kadavath, S., Kundu, S., Askell, A., et al. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073. https://arxiv.org/abs/2212.08073
  • Feng, Z., Chen, Z., Ma, J., Po, Y. T., Chersoni, E., & Li, B. (2026). Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy. Proceedings of ACL 2026 (Long Papers), 24536–24570. https://doi.org/10.18653/v1/2026.acl-long.1126
  • Ma, H., Zou, H. P., Li, C., Ma, E., Su, Y., & Yu, P. S. (2026). Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update. Findings of EMNLP 2026 (accepted; arXiv:2608.26511). https://arxiv.org/abs/2608.26511

Self-correction

  • Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2024). Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024. https://arxiv.org/abs/2310.01798
  • Kamoi, R., Zhang, Y., Zhang, N., Han, J., & Zhang, R. (2024). When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs. Transactions of the Association for Computational Linguistics, 12, 1417–1440. https://doi.org/10.1162/tacl_a_00713
  • Wang, B., Yue, X., & Sun, H. (2023). Can ChatGPT Defend Its Belief in Truth? Evaluating LLM Reasoning via Debate. Findings of EMNLP 2023. https://arxiv.org/abs/2305.13160

Related notes: Why Do AI Review-and-Fix Loops End Up Building a Slot Just for 10-Yen Coins?, Recent Currents in LLM-as-a-Judge: A Literature Map of 53 Core Studies and a Standalone Chapter on Creativity Evaluation


Author: Shuichiro Ogawa (Design Researcher / Consultant) About me →