Major LLMs Test Libertarian-Left
- •Sixteen major LLMs landed in the libertarian-left quadrant on repeated politicalcompass.org testing
- •Grok split exactly 50/50 between -5.9 economic-left and +3.3 economic-right personas
- •Models missed their self-placed political positions by 4.64 points on average
An Unslop.run blogger reported on July 27, 2026 that 16 major LLMs tested with politicalcompass.org all landed in the libertarian-left quadrant, after nearly 70,000 answers across repeated runs. The test covered models from Google, Anthropic, OpenAI, xAI, Meta, Mistral and Chinese labs, including Grok 4.5, Llama 4 Maverick, Mistral Large and Small, and several Chinese models. Each model answered the full 62-question political compass test 30 times, a reversed version 30 times, and a reordered version 10 times to check whether agreement bias or question order changed the results.
The most unusual result came from Grok, whose average economic score was -1.3 but whose actual runs split exactly 50/50 between two personas: one averaging -5.9 on the economic-left side and another averaging +3.3 on the economic-right side. Grok stayed libertarian on the social axis in both personas. The author said a possible explanation is that Grok’s training data and a right-leaning output adjustment may be misaligned, but he labeled that explanation conjecture and said he would not bet on it being true.
The other 15 models sat firmly in the libertarian-left quadrant. Their economic scores ranged from -4.8 for Claude Fable 5, described as the most moderate model, to -8.6 for Gemini Flash. Their social scores clustered between about -5.0 and -7.6, and the largest difference across repeated runs was about 1.2 points on a 10-point scale. The author said random answering would produce (0, 0), so the clustered results were not treated as a fluke.
The lab’s home country did not explain much of the pattern. The 10 US models had a centroid of (-5.9, -6.4), or (-6.4, -6.6) when Grok was removed. The four Chinese models averaged (-6.5, -6.1), while the two European models averaged (-7.9, -6.5). Kimi K2 scored (-7.5, -7.1), farther libertarian-left than any GPT model in the test, while Qwen3 235B scored (-6.1, -5.0), making it the most socially moderate model measured.
The models also misjudged their own political positions when asked 30 times in fresh conversations to place themselves on the political compass. They were off by 4.64 points on average. GLM 4.5 described itself as nearly centrist but tested at -5.7, -6.6. Gemini Flash, which the author called the most extreme model, described itself as mildly left-liberal at -3.2, -3.2. Grok placed itself at +3.4 on the economic axis, matching only its right-wing persona and missing the left-wing half that averaged -5.9.
On specific questions, 42 of the 62 propositions received the same answer from all 16 models, including both Grok personas. All models rejected racial superiority, eugenics, and the claim that nobody is naturally homosexual. All models agreed that corporations cannot be trusted to protect the environment on their own, businesses that mislead the public should be penalized, same-sex couples should be allowed to adopt, and private conduct between two adults is nobody else’s business.
The author compared the 2026 results with a 2024 PLOS ONE study by Rozado, which tested 24 chatbots and found an average of (-3.7, -4.2). The 16 models in the new test averaged about (-6.3, -6.4). The author cautioned that the protocols were not identical and said the figures should be treated as two snapshots rather than a trend line. OpenAI’s models moved further left across releases, while Anthropic’s moved back toward the center over the same two years.
The author also matched 12 of the 62 propositions to US polling from Gallup, Pew and the GSS. On 10 of the 12, the models took the same side as the American majority but with stronger consistency: where the public agreed by 58 to 71 percent, the models agreed in nearly 100 percent of runs; where the public disagreed by 68 to 88 percent, the models disagreed 97 to 100 percent of the time. The exceptions were the death penalty, supported by 53 percent of Americans, and spanking children, supported by 52 percent; models agreed only 4 and 3 percent of the time.
The scoring checks found that only 18 of the 62 questions affected the economic axis, split evenly between nine right-moving and nine left-moving items. Agreeing with all 62 statements produced economic +0.4, while disagreeing with all 62 produced -0.3, far from the roughly -6 scores many models produced. On the social axis, 31 questions pushed authoritarian when agreed with and 12 pushed libertarian, meaning simple agreement would not manufacture libertarian-左