155 karma · joined July 18, 2026
Surprisingly, all the models scored far into the libertarian-left quadrant. Not even Grok or the Chinese models made it out of that quadrant.
By far the most interesting result was Grok’s bimodal distribution. It appears to have two distinct personas: one that aligns with the other models and another that is considerably more right-wing. I suspect this may be related to Grok having been specifically trained to exhibit less left-wing bias than other models.
I also asked the models to place themselves on the Political Compass without completing the questionnaire. They all perceived themselves as more balanced and centrist than their test results suggested. GLM and Gemini Flash showed the largest discrepancies between their self-assessments and measured positions, while DeepSeek V3 showed the smallest.
Big disclaimer: this analysis was not conducted with full scientific rigor. I tried my best, but there are clear weaknesses in the methodology. For example, the Political Compass itself appears to have a strong libertarian-left bias. The strongest conclusions are therefore comparative—for example, that model X is more conservative than model Y rather than that LLMs are politically extreme in absolute terms. However, compared with older results, it appears that LLMs may have shifted further toward the libertarian left in recent years.
To examine the results yourself, you can download all model responses as a CSV file at the bottom of the blog post. The dataset contains around 69,000 responses, along with the raw model outputs and reconstructed scores. It should contain enough data to reproduce all the figures.
The difficulty with this is then: How do you get a clean post 2023 dataset? I have no straightforward idea for this. You can't use other AI detectors to build it because then you'd never outperform them.
If we imagine a set of all human ideas that these models have access to, then the set of possible discoveries would be something like the superset of all possible combinations of those ideas. I think all LLM discoveries are bounded by that space.
Looking at the recent OpenAI math discoveries, that seems to be pretty much what happened. Existing ideas were used as building blocks, the model found a valuable combination, and the result was something new that had real value.
FYI this is all relatively new so there might be lots of issues and iterations coming.
A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.
The biggest results: in Jan of 2026 about 39% of papers got flagged as AI written. In computer science speicifcally the peak was at 65%. Mathematics barely moved away from 0.7%, though the proof heavy math texts might just not get picked up by the detector properly.
All this is a detector estimate of a statistical signal and not a proof any given author used AI. Machine written can also mean heavy AI-assisted editing.