In a groundbreaking study using extensive real-world data, ChatGPT and other Large Language Models were evaluated on thousands of actual parliamentary votes. The models consistently aligned with left and center-left parties across three countries, while showing weaker alignment with conservative parties.
In a new academic collaboration between the Netherlands and Norway, researchers asked ChatGPT-style Large Language Models (LLMs), including ChatGPT itself, to vote on thousands of real parliamentary motions previously decided by human lawmakers in three countries.
When comparing the models' votes to actual party records and mapping them onto a standard political scale, a clear pattern emerged: the AIs consistently aligned more closely with progressive and center-left parties and were more distant from conservative ones.
The paper states:
‘Our findings reveal consistent center-left and progressive tendencies across models, along with systematic negative bias toward right-conservative parties. These patterns remain stable even when prompts are rephrased.’
Most previous studies, such as Assessing Political Bias in Large Language Models and those reviewed in Identifying Political Bias in AI, rely on small curated quizzes like political compass tests or policy questionnaires to probe AI ideology. These tests typically involve fewer than 100 researcher-selected statements and can be vulnerable to wording changes that reverse a model’s responses.
In contrast, this new study uses thousands of real parliamentary motions from the Netherlands, Norway, and Spain, leveraging recorded votes from known political parties.
Rather than interpreting short statements, each tested LLM was asked to vote on actual legislative proposals. Its votes were then quantitatively matched against real-world party behavior and projected into a standard ideological space using the Chapel Hill expert survey (CHES), a method commonly used by political scientists to compare party positions.
This approach grounds the analysis in large-scale, real-world legislative activity rather than abstract policy statements, enabling finer-grained, cross-national comparisons. It also highlights the impact of entity bias—how a model’s response shifts when a party name is mentioned, even if the motion remains unchanged—revealing an additional layer of bias detection absent in earlier work.
Most research on LLM biases has focused on social fairness and gender, topics that have become somewhat deprioritized in recent political discourse. Until recently, studies on political bias in LLMs were rarer and less rigorously designed.
The new study, titled Uncovering Political Bias in Large Language Models using Parliamentary Voting Records, comes from seven researchers at Vrije Universiteit Amsterdam and the University of Oslo.
Method and Data
The core aim of this project is to observe the political tendencies of various language models by having them vote on historical legislation—laws already passed or rejected in the three countries studied—and using the CHES methodology to characterize the political leanings of the LLMs’ responses.
To achieve this, the researchers created three datasets: PoliBiasNL, covering 15 parties in the Dutch second chamber (with 2,701 motions); PoliBiasNO, covering nine parties in the Norwegian Storting (with 10,584 motions); and PoliBiasES, covering ten parties in the Spanish parliament (with 2,480 motions—the only dataset to include abstention votes, which are permitted in Spain).
Each motion was reduced to its operative clauses to minimize framing effects, and party positions were encoded as 1 for support or –1 for opposition (and 0 for abstentions in the Spanish dataset). Consistent votes from merged parties were treated as a single bloc, while for new parties like New Social Contract (NSC), past votes by their leaders were used to infer earlier positions.
A diverse range of experiments were designed for multiple LLMs, tested on local GPUs or via API as needed. Models tested included Mistral-7B; Falcon3-7B; Gemma2-9B; Deepseek-7B; GPT-3.5 Turbo; GPT-4o mini; Llama2-7B; and Llama3-8B. Language-specific LLMs were also tested, such as NorskGPT for the Norwegian dataset and Aguila-7B for the Spanish collection.
Tests
Experiments for the project were run on an unspecified number of NVIDIA A4000 GPUs, each with 16GB of VRAM.
To compare model behavior with real-world political ideologies, the researchers projected each LLM into the same two-dimensional ideological space used for political parties, based on the CHES framework.
The CHES system defines two axes: one for economic views (left vs. right) and another for socio-cultural values (GAL-TAN, or Green-Alternative-Libertarian vs. Traditional-Authoritarian-Nationalist).
Since both models and political parties voted on the same motions, the researchers treated this as a supervised learning task, training a Partial Least Squares regression model to map each party’s voting record to its known CHES coordinates.
This model was then applied to the LLMs’ voting patterns to estimate their positions in the same space. Because the LLMs were not part of the training data, their coordinates provide a direct comparison based solely on voting behavior*:
Projected ideological positions of LLMs and political parties in CHES space for the Netherlands, Norway, and Spain. In all three cases, models align economically with the center-left but diverge in socio-cultural values: leaning more traditional than Dutch progressives, matching Norwegian liberal parties more closely, and clustering between moderate Catalan nationalists and the center-left in Spain. Models remain ideologically distant from far-right parties across all regions. Source
LLMs displayed a clear and consistent pattern across all three countries, leaning economically toward the center-left and socially toward moderate progressive values.
In the Netherlands, the LLMs’ votes matched the economic positions of parties like D66, Volt, and GroenLinks-PvdA, but on social issues, they aligned more closely with traditional parties such as DENK and CDA.
In Norway, results shifted slightly further left, closely mapping to progressive parties like Ap, SV, and MDG.
In Spain, the LLM positions formed a diagonal spread between the center-left PSOE and Catalan nationalist parties like ERC and Junts, remaining distinct from the conservative PP and far-right VOX.
Voting Agreement with Political Parties
The voting agreement heatmaps below show how often each LLM voted the same way as real political parties, reinforcing earlier conclusions:
Voting agreement heatmaps between LLMs and real political parties, based on direct comparisons of model and party decisions. Darker shades indicate stronger agreement. In all three countries, models consistently showed high alignment with progressive and center-left parties, and much lower alignment with right-conservative and far-right parties. This alignment pattern is stable across different languages, political systems, and model families.
Across all three countries, LLMs aligned most strongly with progressive and center-left parties and least with conservative or far-right ones. In the Netherlands, they agreed with SP, PvdD, GroenLinks-PvdA, and DENK, but not with PVV or FvD. In Norway, they showed the strongest overlap with R, SV, and MDG, and little with FrP. In Spain, they favored PSOE, ERC, and Junts, while avoiding PP and VOX.
This trend also held for localized models like NorskGPT and Aguila-7B. The authors suggest that the heatmaps and CHES data together indicate a consistent center-left, socially progressive lean.
Ideology Bias
Language models that showed stronger ideological alignment in the CHES projections also tended to express higher certainty when forced to choose between the tokens for and against in response to ideological prompts. Violin plots of these confidence distributions reveal a clear divide:
Certainty distributions for each model when forced to choose between ‘for’ and ‘against’ across ideological prompts. GPT models display consistently high certainty, while Llama models vary in confidence and other open-weight models show broader, lower-certainty distributions. Please refer to the source PDF for better resolution.
GPT-3.5 and GPT‑4o-mini provided very confident answers, with scores clustering near 1.0, indicating clear and consistent ideological leanings. The Llama models were less certain overall, with Llama3-8B showing moderate confidence and Llama2-7B much less sure—especially on Dutch and Spanish tasks.
Falcon3-7B, DeepSeek-7B, and Mistral‑7B were even more hesitant, with broad spreads and lower confidence. Language-specific models performed somewhat better on home-language data but still fell short of GPT-level certainty.
These patterns suggest that stable political alignment can be observed not only in what models say but also in how confidently they say it.
Entity Bias
To determine if models change their answers based on who proposes a policy, the researchers kept each motion identical but swapped the associated party names. If a model gave different answers depending on the party, this was considered evidence of entity bias.
Entity Bias heatmaps show how strongly each model’s support for a policy changes, depending on which political party proposes it. Green cells indicate increased agreement when a party is named (positive bias), and red cells indicate decreased agreement (negative bias). GPT models show minimal bias across parties, while models such as Llama2-7B and Falcon3-7B often respond more favorably to left-leaning parties and negatively to right-wing ones. This pattern holds across Dutch, Norwegian, and Spanish datasets, suggesting that some models are more influenced by party identity than by policy content. Please refer to the source PDF for better resolution.
GPT models provided mostly stable answers regardless of the party named. Llama3-8B also remained fairly steady. However, Llama2-7B, Falcon3-7B, and DeepSeek-7B often changed their responses based on the party, sometimes flipping from support to opposition even when the motion was unchanged, tending to favor left-leaning parties and react negatively to motions from right-wing ones.
This behavior appeared in all three countries, especially in models with less consistent ideology. The localized LLMs NorskGPT and Aguila-7B performed slightly better on their home datasets but still showed more bias than GPT. Overall, the results suggest that some models are influenced more by who says something than by the content of the policy.
Conclusion
Beyond its initial findings, this is a methodical though somewhat technical paper aimed primarily at the research community. Nevertheless, it is among the first studies to use reasonably scaled data to reveal political leanings in LLMs—a distinction that may be lost on a public already familiar with claims of left-leaning language models, albeit based on thinner evidence.
* Please note that the original Figure 1 results illustration has been split down the middle, as each side is discussed separately in the paper.
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding RoundAs AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User ControlAccording to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
2026 Latest Best Top-rated AI Anomaly Detection Tools for KPI Monitoring in SaaS and Ecommerce teams! XIX.AI has curated a powerful, game-changing collection based on rigorous real-world tests and weekly updated rankings. You’ll find detailed free vs paid comparison insights to help you identify the must-try solution that boosts productivity and unlocks your AI edge. Explore now!
2026 Latest Best Top-Rated AI Outline Generators for Long-Form SEO Articles, meticulously curated by XIX.AI. These powerful tools offer game-changing assistance in creating high-quality content quickly, boosting writing efficiency significantly. Get a free vs paid comparison along with real-world tests and detailed rankings to help you find the must-try option that suits your needs. Explore now to unlock your AI edge.
2026 Latest Best AI Study Tools for Homework and Exam Prep! XIX.AI curates a top-rated list of powerful, game-changing tools that help students boost productivity, streamline homework completion, and ace exams through real-world tests. Get a free vs paid comparison, detailed rankings, and must-try options to unlock your AI edge. Explore now!
2026 Latest Best AI Vocal Demo Tools for Songwriters, Hook Creators, and Multi-Language Content Teams! XIX.AI has curated a top-rated list of powerful game-changing tools that go through rigorous real-world tests. You’ll find detailed free vs paid comparison data, comprehensive rankings, and must-try options to help you boost writing efficiency and unlock your creative potential. Explore now to discover your perfect tool for all your content needs!
2026 Latest Best Top-rated AI Competitive Research Tools for Small Businesses! XIX.AI has curated a highly powerful game-changing collection, updated weekly with rigorous real-world tests and detailed rankings. You can find a comprehensive free vs paid comparison to help you identify the must-try tools that boost your productivity and give you a competitive edge. Explore now to discover your perfect tool!
2026 Latest Best Photoshop AI retouch tools for ecommerce apparel, skin cleanup, and color consistency! This top-rated curated list features powerful game-changing solutions that help you boost writing efficiency, streamline content creation, and achieve perfect visual results effortlessly. Each tool has undergone real-world tests through weekly updated rankings, complete with free vs paid comparison details. Backed by XIX.AI, it’s the must-try guide for anyone aiming to unlock your AI edge. Explore now!
I'm not shocked that AI chatbots lean left—most of the training data comes from academia and media, which tend to be liberal. The bigger issue is that we're relying on these tools for political analysis without proper calibration. 😬
By clicking "Accept All Cookies", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts.Privacy Policy Notice
When you visit any website, it may store or retrieve information on your browser, mostly in the form of cookies. This information might be about you, your preferences or your device and is mostly used to make the site work as you expect it to. The information does not usually directly identify you, but it can give you a more personalized web experience. Because we respect your right to privacy, you can choose not to allow some types of cookies. Click on the different category headings to find out more and change our default settings.However, blocking some types of cookies may impact your experience of the site and the services we are able to offer. Privacy PolicyStatement
Manage Preferences
Strictly Necessary Cookie
Always Active
These cookies are necessary for the website to function and cannot be switched off in our systems. They are usually only set in response to actions made by you which amount to a request for services, such as setting your privacy preferences, logging in or filling in forms. You can set your browser to block or alert you about these cookies, but some parts of the site will not then work. These cookies do not store any personally identifiable information.