HNHacker News
TopNewBestAskShowJobs

jasisz

4 karma · joined July 19, 2015

submissionscomments
jasisz··on Show HN: Aver – a language designed for AI to write and humans to review
Tooling tax is a real argument. But “just use macros + linting” usually gives you policy-flavored Rust, not a genuinely different language model. That works if enforcement is the goal; it works much less well if the artifact itself is meant to read differently.
jasisz··on Show HN: Aver – a language designed for AI to write and humans to review
Author here.

Aver is an experimental statically typed language for AI-written, human-reviewed code.

What’s different is that intent (`?`), explicit effects (`!`), design decisions (`decision`), and behavior checks (`verify`) are part of the source itself rather than split across code, comments, docs, and tests.

If you want to evaluate it quickly, I’d suggest these places:

- medium-sized example: https://github.com/jasisz/aver/tree/main/projects/workflow_e...

- Lean proof export for the pure subset: https://github.com/jasisz/aver/tree/main/docs/lean.md

- examples of effectful programs / replay: https://github.com/jasisz/aver/tree/main/examples/services

The main question I’m testing is whether this deserves to be a language, or whether the same idea should just be tooling and conventions on top of an existing language.

jasisz··on Introduction to Monte Carlo Tree Search
From the "An analysis of UCT in multi-player games", Nathan Sturtevant, 2008: "Multi-player UCT is nearly identical to regular UCT. At the highest level of the algorithm, the tree is repeatedly sampled until it is time to make an action. The sampling process is illustrated in Figure 2. The only difference between this code and a two-player implementation is that in line 5 the average score for player p is used instead of a single average payoff for the state."

I think this is kind of a clear statement that original paper (and after it a lot of writing on the topic) may be lacking. Of course people used this simple generalization before and it is pretty straightforward, but it is not that obvious at a first glance. And I've seen quite a lot of code examples, images explaining UCT for games and articles that were just not saying a word on this. Or even worse - just doing it wrong for multiplayer games.

Choice of action is a different topic, as I remember correctly there was also a paper proving that win rate and most robust branch are in the end performing the same ;)

Hope you will continue this series, because it is really good and code examples are really nice!

jasisz··on Introduction to Monte Carlo Tree Search
In your example script yes, but basic UCT does not do that - simply because UCT was not meant only for multi-player games in the beginning. And this is some assumption we make about our opponents (actually that they want to maximize their payout, not to e.g. win or minimize our score). Of course this is a very straightforward "application" of UCT or MCTS to the multiplayer games that was done in works of Sturtevant and Cazenave in 2008.

But it is not so easy to know about this, e.g. this is a very recent change to wikipedia page on the topic https://en.wikipedia.org/w/index.php?title=Monte_Carlo_tree_... and very often people writing on the topic are not pointing this out at all, which I find very strange and misleading.

Also in the classic MCTS you should select move which has most visits, not the one with the highest percentage of wins.

jasisz··on Introduction to Monte Carlo Tree Search
This is exactly the problem I've written about - AFAIK basic UCT algorithm does not model your opponent in the selection phase (only in simulation when some heavier than random logic is applied).
jasisz··on Introduction to Monte Carlo Tree Search
But this requires storing and back propagating this info for the other players - something I really haven't seen in any examples (nor in this article). We cannot also assume that game is always zero-sum game and this information is not needed.
jasisz··on Introduction to Monte Carlo Tree Search
This algorithm (in it's regular form used often in games and examples) has one interesting "downside" I was exploring some time ago - selection is performed using the UCB formula. So basically it tries to maximize the player payout. But in the most games this is in fact impractical assumption, because we end up tending to expand branches that will be most likely not chosen by our opponent. As in the example (I assume gray moves are "our" moves) - we will much more likely choose to expand 5/6 branch instead of the 2/4, that will be in fact more likely chosen by our opponent.