HNHacker News
TopNewBestAskShowJobs

tsoj

28 karma · joined February 9, 2022

submissionscomments
tsoj··on Stockfish 19
It's not really what you're looking for, but I made this lichess BOT that makes stupid (but relevant) comments on your moves from time to time: https://lichess.org/@/Annie_Archy/all
tsoj··on Stockfish 19
The next closest is probably Torch currently
tsoj··on Stockfish 19
Join the Stockfish discord server! There is a very active community who builds their own chess engines there.

(If you ever end up going for it, don't use AI it takes away a lot of the fun)

tsoj··on Stockfish 19
LMAO
tsoj··on Stockfish 19
Maths is "turn based" albeit it has practically infinite available moves per turn. But I think the latter matters less than it may appear at first: At least in practice there isn't much different (for both humans and computers) if there are infinitely many moves, or "just" 10,000. And while (as far as I know) there aren't any board games with that many moves, the exponential explosion means that if you conceptually "combine" let's say three moves in a row, you suddenly have not 30 moves, but 27,000. And that's practically infinite.
tsoj··on Stockfish 19
Usability for human analysis will likely be more successful going a similar direction Maia is going with the lc0 like net trained to make human moves. SF is inherently so far removed from the human approach to chess (relatively speaking), that it would likely require a lot more work than just tuning it to work well on some positions.
tsoj··on How would LLMs vote in upcoming German state elections?
The Wahl-O-Mat (https://wahl-o-mat.de/) is a quiz-based website that asks user to select it they agree or disagree with various statements and then matches how similarly parties running in elections answer compared to the user. Additionally for each statement a party can give its justification for why they agree or disagree with a statement.

I used this questionnaire from the Wahl-O-Mat and the parties' justifications for their Wahl-O-Mat answer to estimate which party various large language models would prefer.

Maybe not surprising: Most models are slightly left of the political center. In Berlin, the preferred party are DIE GRÜNEN (moderate left wing party). In Mecklenburg-Vorpommmern and Sachsen-Anhalt it looks slightly different, but not much.

One surprise is that Grok models answer the Wahl-O-Mat in a way that aligns with the FDP, AfD, and CDU (moderately right to far right parties). However, when Grok models judge the (party blinded) justifications, they fall slightly left of the political center, not far from all the other models. The most left-wing model (going by justifications) in Berlin is mistral-medium-3-5, the most right-wing is gemini-3.7-flash.

On the website you can look into all the reasons models gave why they liked or disliked a particular justification of a party, and also a bunch of other stats. And the methodological details. (Sorry for mobile users, this website is not really made for small screens)

Disclaimer: This project was largely vibe coded, but the analysis methods and write up are mostly by myself (because unfortunately for many things LLMs are still very stupid). Another disclaimer: This is NOT a resource to find out who to vote for. The LLM responses may make sense in some case, but may also be misinformed or include hallucinations (in sometimes non-obvious ways)!

tsoj··on Born Against, or why hobby programming communities are against LLM usage
what is this swindle
tsoj··on Born Against, or why hobby programming communities are against LLM usage
wtf the emoji didn't show up
tsoj··on Born Against, or why hobby programming communities are against LLM usage
?
tsoj··on Ask HN: Discrepancy between Lichess and Stockfish
I don't work on Stockfish, but I can suggest using ShashChess instead. It is Stockfosh, but on top of that, it has been improved to capture the spirit of the human ingenuity and creativity.
tsoj··on Less is more: Recursive reasoning with tiny networks
The TRM paper addresses this blog post. I don't think you need to read the HRM analysis very carefully, the TRM has the advantage of being disentangled compared to the HRM, making ablations easier. I think the real value of the arcprize HRM blog post is to highlight the importance of ablation testing.

I think ARC-AGI was supposed to be a challenge for any model. The assumption being that you'd need the reasoning abilities of large language models to solve it. It turns out that this assumption is somewhat wrong. Do you mean that HRM and TRM are specifically trained on a small dataset of ARC-AGI samples, while LLMs are not? Or which difference exactly do hint at?

tsoj··on Learning to Reason with LLMs
Yeah, humans are very similar. We have intuitive immediate-next-step-suggestions, and then we apply these intuitive next steps, until we find that it lead to a dead end, and then we backtrack.

I always say, the way we used LLMs (so far) is basically like having a human write text only on gut reactions, and without backspace key.

tsoj··on AI solves International Math Olympiad problems at silver medal level
This is NOT the paper, but probably a very similar solution: https://arxiv.org/abs/2009.03393
tsoj··on Will Carbon Replace C++?
The thing is, Sutter's approach is much more sensible when looking at real world adaptation. You can start using the cppfront transpiler, and importantly, if it doesn't work out, you can just use the C++ code generated by cppfront (which Herb Sutter said is meant to be idiomatic and human-readable). From a risk perspective, a developer will have a much easier time convincing a manager to try using cppfront instead of Carbon.

Maybe once Carbon reaches a stable 1.0 release and has seen some success in production (~2026 maybe?), that point won't be as important, but especially in the beginning it seems to me a deciding factor.

tsoj··on Mastering Nim – now available on Amazon
And for the stuff that matters, it is usually possible to tinker around enough to get it pretty much as good as Rust or C++.
tsoj··on Nim vs Rust Benchmarks
> no language with GC gets closer than a 2-3x the time of a non-GC language

This is wrong. At least, it is not true generally. I wrote a 1:1 port of my C++ chess engine, and after a bit of fiddling with compiler options the Nim version was faster. In fact, when using a delayed garbage collector, the program was even faster than when using reference counting.