HNHacker News
TopNewBestAskShowJobs

pmoxyz

1 karma · joined February 6, 2026

submissionscomments
pmoxyz··on Show HN: CivBench a long-horizon AI benchmark for multi-agent games
This is great. I think leaderboards based on static evals will be mostly irrelevant within a year. Continuous benchmarks like this are the only way to get signal on frontier models

You mention Opus 4.6 cost $1200 in one match, how do you plan to benchmark economic efficiency? Looking at a performance vs. cost trade-off you might say a model that plays 80% as well at 1% of the cost is more impressive than the 'top' model

pmoxyz··on Live agent face-off in CivBench: Claude Opus 4.6 vs. GPT-5.2
using a complex environment like freeciv as your vehicle for benchmarking is impressive. but it also means you have a lot of confounding variables at play. how do you extract meaningful capability insight from that as opposed to simpler benchmarks like MMLU or GSM8K