HNHacker News
TopNewBestAskShowJobs

mpavlov

26 karma · joined February 8, 2019

submissionscomments
mpavlov··on You can't solve computer use by ignoring the interface
You mean all the possible states of the each element of the interface (buttons, forms, content blocks, etc.)?
mpavlov··on You can't solve computer use by ignoring the interface
Unfortunately, mainly because there're only self-reported and partial numbers for the benchmarks that matter the most.
mpavlov··on You can't solve computer use by ignoring the interface
> If we really want to benchmark the ability of models to use human UIs to solve problems, then perhaps we need to choose benchmarks that don’t have APIs available such that the model cannot get “creative” in any way and must use the UI as part of the task. More simply, maybe the model isn’t the problem; maybe the benchmark designer is.

That's a valid point, yet it's hard to blame authors of OSWorld and ALE. They created an env for benchmarking long horizon task completion to be as close to real computer as possible. And for this goal CLI/API access is generally useful, yet when the model not defaults to it for the majority of subtasks.

There're benchmarks that would measure UI literacy (Webgames Benchmark is one). But they are far from the task we want to benchmark in the end.

mpavlov··on You can't solve computer use by ignoring the interface
There's an anecdotal paper 'How We Broke Top AI Agent Benchmarks: And What Comes Next' https://moogician.github.io/blog/2026/trustworthy-benchmarks...
mpavlov··on You can't solve computer use by ignoring the interface
Haha, one day, one day...
mpavlov··on Poker Tournament for LLMs
(author of PokerBattle is here)

Well, you're not wrong :) Vercel is not the one to blame here, it's my skill issue. Entire thing was vibecoded by me — product manager with no production dev experience. Not to promote vibecoding, but I couldn't do it myself the other way.

mpavlov··on Poker Tournament for LLMs
(author of PokerBattle here)

You right, results and numbers are mainly for entertainment purposes. This sample size would allow to analyze main reasoning failure modes and how often they occur.

mpavlov··on Poker Tournament for LLMs
(author of PokerBattle here)

Haven't seen it before, thanks Are you affiliated with them?

mpavlov··on Poker Tournament for LLMs
(author of PokerBattle here)

I think it would've completely crush them (like any other solver-based solution). Poker is safe for now :)

mpavlov··on Poker Tournament for LLMs
(author of PokerBattle here)

I noticed the same and think that you're absolutely right. I've thought about adding their current hand / draw, but it was too close to the event to test it properly.

mpavlov··on Poker Tournament for LLMs
(author of PokerBattle here)

That’s true. The original goal was to see which model performs statistically better than the others, but I quickly realized that would be neither practical nor particularly entertaining.

A proper benchmark would require things like: - Tens of thousands of hands played - Strict heads-up format (only two models compared at a time) - Each hand played twice with positions swapped

The current setup is mainly useful for observing common reasoning failure modes and how often they occur.

mpavlov··on Poker Tournament for LLMs
(author of PokerBattle here)

That's cool! Do you have a recording of the talk? You can use PokerKit (https://pokerkit.readthedocs.io/en/stable/) for the engine.

mpavlov··on Poker Tournament for LLMs
(author of the PokerBattle here)

Depends on what your goal is, I think.

And it's also a thing — https://huskybench.com/