28 karma · joined February 9, 2022
(If you ever end up going for it, don't use AI it takes away a lot of the fun)
I used this questionnaire from the Wahl-O-Mat and the parties' justifications for their Wahl-O-Mat answer to estimate which party various large language models would prefer.
Maybe not surprising: Most models are slightly left of the political center. In Berlin, the preferred party are DIE GRÜNEN (moderate left wing party). In Mecklenburg-Vorpommmern and Sachsen-Anhalt it looks slightly different, but not much.
One surprise is that Grok models answer the Wahl-O-Mat in a way that aligns with the FDP, AfD, and CDU (moderately right to far right parties). However, when Grok models judge the (party blinded) justifications, they fall slightly left of the political center, not far from all the other models. The most left-wing model (going by justifications) in Berlin is mistral-medium-3-5, the most right-wing is gemini-3.7-flash.
On the website you can look into all the reasons models gave why they liked or disliked a particular justification of a party, and also a bunch of other stats. And the methodological details. (Sorry for mobile users, this website is not really made for small screens)
Disclaimer: This project was largely vibe coded, but the analysis methods and write up are mostly by myself (because unfortunately for many things LLMs are still very stupid). Another disclaimer: This is NOT a resource to find out who to vote for. The LLM responses may make sense in some case, but may also be misinformed or include hallucinations (in sometimes non-obvious ways)!
I think ARC-AGI was supposed to be a challenge for any model. The assumption being that you'd need the reasoning abilities of large language models to solve it. It turns out that this assumption is somewhat wrong. Do you mean that HRM and TRM are specifically trained on a small dataset of ARC-AGI samples, while LLMs are not? Or which difference exactly do hint at?
I always say, the way we used LLMs (so far) is basically like having a human write text only on gut reactions, and without backspace key.
Maybe once Carbon reaches a stable 1.0 release and has seen some success in production (~2026 maybe?), that point won't be as important, but especially in the beginning it seems to me a deciding factor.
This is wrong. At least, it is not true generally. I wrote a 1:1 port of my C++ chess engine, and after a bit of fiddling with compiler options the Nim version was faster. In fact, when using a delayed garbage collector, the program was even faster than when using reference counting.