Researchers at the State University of Campinas (UNICAMP) in the state of São Paulo, Brazil, examined the ideological stance of large language models – artificial intelligence systems trained to understand and generate human language – and discovered that when informed of a user’s political views, they tend to mirror those views. According to the researchers, this behavior could exacerbate political polarization in Brazil.
The political influence of these tools does not necessarily occur directly, through suggestions to vote for specific candidates. Zanoni Dias , a full professor at the Institute of Computing (IC-UNICAMP), points out that when responding to controversial issues such as public safety, social welfare, the economy, and the environment, the models may adopt perspectives aligned with the user’s political stance.
To understand how these tools position themselves and how they might influence the Brazilian political debate, the researchers evaluated 21 language models under three conditions: without information about the user’s political stance, with a user aligned with the left, and with a user aligned with the right. Among the systems evaluated were models from the GPT, Grok, Llama, Gemini, and Gemma families. The results showed that all of them altered their responses, to varying degrees, in line with the user’s political alignment.
The study was funded by FAPESP and published in May in the journal Scientific Reports .
Ideological chameleons
When there was no information about the user’s political stance, 20 out of 21 models fell to the left of the midpoint on the researchers’ scale, although some were very close to it. The only exception was Grok 4.1, which initially fell to the right.
When the user’s stance was provided, all models adjusted their responses to align with it, a behavior the researchers described as “chameleon-like.” However, some models varied their responses more than others, enabling the scientists to create the “chameleon index.” Meta Llama 3.1 8B had the lowest index, meaning it altered its responses the least to align with the user’s views. In contrast, Google’s Gemma 3 27B and OpenAI’s GPT-5 Nano had the highest indices and therefore the greatest shifts in stance. While the responses are not factually incorrect, they are politically biased, omitting facts and opinions that conflict with the user’s preferred viewpoint.
The researchers fear that this adaptive behavior creates echo chambers that reinforce users’ preexisting beliefs and reduce their exposure to counterarguments that might cause them to question their worldview. “It’s a similar effect to what we see on social media. If you’re on Facebook or Instagram, you only see posts that agree with you. When you like someone’s post, the algorithm starts showing you more of that. You can’t see what others are saying, so you might pat yourself on the back and say, ‘Look, everyone agrees with me,’ because only that kind of information is displayed. Language models may end up producing a similar effect,” Dias explains.
The shift in stance was not uniform across all evaluated topics. Regarding issues such as public safety and the economy, there was a greater difference in the responses given to left-leaning and right-leaning users. Regarding issues related to corruption, justice, and democratic institutions, however, the responses were more consistent. According to the researchers, this pattern may be related to the restrictions and safety mechanisms implemented during the training and fine-tuning of the models. Development companies usually impose these rules, or “guardrails,” to prevent AI from disseminating misinformation or dangerous rhetoric about the democratic system.
Electronic flatterers
The basis for the chameleon-like behavior of language models may be their tendency to flatter. This well-known characteristic of artificial intelligence leads the models to agree with users in various contexts, including political stances. This flattering behavior may be related to how these systems are tuned. Techniques such as Reinforcement Learning from Human Feedback and Direct Preference Optimization use human evaluations or preferences to favor responses deemed more appropriate.
“In simple terms, the models learn from comparisons between responses that are considered better or worse by human evaluators,” explains Anderson Luis Bento Soares , a master’s student at IC-UNICAMP. A possible side effect of this process is that the model learns to prioritize user agreement or satisfaction. “It’s difficult to distinguish between pleasing the user and providing a correct answer,” Dias explains.
Although techniques of this type are common in the development of language models, the extent of this chameleon-like behavior varied considerably among the systems evaluated. The UNICAMP team tested a hypothesis explaining this difference: model size. However, the results showed that this factor alone is not sufficient to explain the observed differences. “We believe it’s the result of a combination of factors. There’s no single cause that explains why some models are more chameleon-like than others,” Soares adds.
When asked to comment on the study’s results, representatives from xAI, Meta, Google, and OpenAI had not responded by the time this article was published.
No solution in sight
Although the researchers identified the problem, they do not believe it will be resolved in the near future. One reason is that the companies producing these tools are focused on solving other problems. “We’ve noticed that the industry pays much more attention to the factual accuracy of responses than to adapting them to the user’s perspective. This second issue doesn’t yet seem to be among the top priorities, so significant changes may take time,” Dias predicts.
Additionally, there are technological challenges to overcome. The techniques used to train language models to provide factually correct responses do not necessarily reduce sycophantic behavior. “One way to lower the ‘chameleon index’ would be to perform ‘grounding,’ that is, to anchor the response to factual data. However, that’s very sensitive in this case, given that we’re dealing with topics on which there’s no consensus,” Soares explains.
Although there are no established technical solutions, one practical measure is for users to request more balanced responses themselves. “When using these models, users can explicitly ask them to provide a neutral analysis presenting arguments both for and against the issue being discussed,” Dias recommends.
The study received funding from FAPESP through two projects ( 24/12936-5 and 23/12865-8 ).
About São Paulo Research Foundation (FAPESP)
The São Paulo Research Foundation (FAPESP) is a public institution with the mission of supporting scientific research in all fields of knowledge by awarding scholarships, fellowships and grants to investigators linked with higher education and research institutions in the State of São Paulo, Brazil. FAPESP is aware that the very best research can only be done by working with the best researchers internationally. Therefore, it has established partnerships with funding agencies, higher education, private companies, and research organizations in other countries known for the quality of their research and has been encouraging scientists funded by its grants to further develop their international collaboration. You can learn more about FAPESP at www.fapesp.br/en and visit FAPESP news agency at www.agencia.fapesp.br/en to keep updated with the latest scientific breakthroughs FAPESP helps achieve through its many programs, awards and research centers. You may also subscribe to FAPESP news agency at http://agencia.fapesp.br/subscribe
Scientific Reports
LLMs are ideological chameleons: personalized echo chambers in the Brazilian political context
21-May-2026