HNHacker News
TopNewBestAskShowJobs

zone411

4,274 karma · joined August 13, 2010

https://twitter.com/LechMazur

10 LLM benchmarks: https://github.com/lechmazur/

https://www.linkedin.com/in/lech-mazur-69b70493/

Advameg (City-data.com) founder and CEO. AI startup founder.

Author: AI melody songwriting assistant https://melodies.ai

Author: Accurate COVID-19 county-by-county neural net case prediction model based on most data.

submissionscomments
zone411··on Gemini 4 Argon
The article presented facts and data. If that's a problem for you, that sounds more faith-based than whatever Anthropic is doing.
zone411··on Navier-Stokes Announcement
Completely misleading.

This is all you need to read and understand for Anthropic's FLT formalization:

  import Mathlib
  import Theorems.Thm_fermat_last_theorem

  /-- Solution side: the same statement, binder for binder, proved by this tree's `fermat_last_theorem`. -/
  theorem FLT_for_comparator (n : ℕ) (hn : 3 ≤ n) (a b c : ℕ) (ha : 0 < a) (hb : 0 < b) (hc : 0 < c) :
    a ^ n + b ^ n ≠ c ^ n :=
  fermat_last_theorem n hn a b c ha hb hc

  /-- Mathlib's named proposition, by the one-line bridge from the elementary statement
  (the bridge is restated inline so that this file depends only on `Theorems.Thm_fermat_last_theorem`). -/
  theorem FLT_mathlib_for_comparator : FermatLastTheorem :=
  fun n hn a b c ha hb hc => fermat_last_theorem n hn a b c (Nat.pos_of_ne_zero ha) (Nat.pos_of_ne_zero hb) (Nat.pos_of_ne_zero hc)
The actual proof is 13 million lines of Lean.
zone411··on OpenAI's GPT-6 Astra on ARC-AGI-3
I don't:

"Problems solved before a model's training cutoff can be filtered out, and all models compared on the remaining problems" means that the problems an older model actually solved are the ones that get filtered out, while the remaining problems are the ones it already tried and failed on. So older models end up with 0s on the filtered set and you can't really use this to compare new models to older ones.

Also, since these are known public problems, you can't stop people from spending far more than your arbitrary time and $ limits on them. So the number of clean problems will go down over time.

zone411··on Intelligence Is Not the Main Bottleneck
A complete mischaracterization, as usual for HN lately when discussing AI or LessWrong. Obviously, even average levels of persuasion are enough to convince some people. And nobody is air-gapping AI.
zone411··on Be skeptical of OpenAI's rogue hacker agent story
So why don't companies in other industries rush to prove their products are dangerous weapons? Maybe because it would be a really dumb PR stunt?
zone411··on Be skeptical of OpenAI's rogue hacker agent story
And how do people saying this know the capabilities of yet unreleased models?
zone411··on Be skeptical of OpenAI's rogue hacker agent story
How is it in their interest? Scaring customers, worrying employees, and inviting regulators to act is in their interest?
zone411··on Be skeptical of OpenAI's rogue hacker agent story
It's not good marketing for them. This is a talking point with zero evidence that people repeat mindlessly. Scaring customers, worrying employees, and inviting regulators to act would be the worst marketing idea ever devised.
zone411··on Openrouter Fusion API
Yes, definitely not a new idea. I had a multi-turn composite model in 2024 that was outperforming the top models across benchmarks: https://x.com/LechMazur/status/1828804485033992514.
zone411··on Artificial intelligence is not conscious – Ted Chiang
That's not proof. Emergent intelligence is not consciousness.
zone411··on The mysterious Hy3 LLM is topping OpenRouter Model Rankings by a large margin
I’ve tested this model on four of my benchmarks:

https://github.com/lechmazur/buyout_game 10th out 36.

https://github.com/lechmazur/pact/ 14th out 25.

https://github.com/lechmazur/nyt-connections/ 60th out 81.

https://github.com/lechmazur/debate 16th out of 29.

zone411··on Spain blocks prediction markets Polymarket, Kalshi over lack of gambling licence
100%. It's sad to see that this attitude has spread to HN
zone411··on An OpenAI model has disproved a central conjecture in discrete geometry
I actually tried using GPT-5.5 Pro on this problem recently. It thought it was making progress on one path, but it made so many mistakes that it didn't feel worth it pushing further. It'll be interesting to check whether it's the same route. I got partial results (proved in Lean) that improve on the best-known results for four Erdős problems with GPT-5.5 Pro
zone411··on EFF is leaving X
https://variety.com/2020/digital/news/twitter-unblocks-new-y...
zone411··on AI overly affirms users asking for personal advice
I built this benchmark this month: https://github.com/lechmazur/sycophancy. There are large differences between LLMs. There are large differences between LLMs. For example, Mistral Large 3 and GPT-4.1 will initially agree with the narrator, while Gemini will disagree. I swap sides, so this is not about possible viewpoint bias in the LLMs. But another benchmark shows that Gemini will then change its view very easily in a multi-turn conversation while Kimi K2.5 or Grok won't: https://github.com/lechmazur/persuasion.
zone411··on Folk are getting dangerously attached to AI that always tells them they're right
I built two related benchmarks this month: https://github.com/lechmazur/sycophancy and https://github.com/lechmazur/persuasion. There are large differences between LLMs. For example, good luck getting Grok to change its view, while Gemini 3.1 Pro will usually disagree with the narrator at first but then change its position very easily when pushed.
zone411··on Show HN: LLM Debate Benchmark
Hmm, maybe in the next edition, Opus gets expensive. I should probably run GPT-5.4 xhigh too if I do that for fairness...
zone411··on Polymarket gamblers threaten to kill me over Iran missile story
Rationalists were right about everything that mattered: crypto, AI, COVID... HN commentators, by contrast, were wrong about everything that mattered.
zone411··on GPT-5.4
Results from my Extended NYT Connections benchmark:

GPT-5.4 extra high scores 94.0 (GPT-5.2 extra high scored 88.6).

GPT-5.4 medium scores 92.0 (GPT-5.2 medium scored 71.4).

GPT-5.4 no reasoning scores 32.8 (GPT-5.2 no reasoning scored 28.1).

zone411··on I asked Claude for 37,500 random names, and it can't stop saying Marcus
I've made top-10 lists of LLMs' favorite names to use in creative writing here: https://x.com/LechMazur/status/2020206185190945178. They often recur across different LLMs. For example, they love Elara and Elias.
zone411··on Claude Sonnet 4.6
They're improved compared to 4.5 on my Extended NYT Connections benchmark (https://github.com/lechmazur/nyt-connections/).

Sonnet 4.6 Thinking 16K scores 57.6 on the Extended NYT Connections Benchmark. Sonnet 4.5 Thinking 16K scored 49.3.

Sonnet 4.6 No Reasoning scores 55.2. Sonnet 4.5 No Reasoning scored 47.4.

zone411··on Which AI Lies Best? A game theory classic designed by John Nash
For people interested in these kinds of benchmarks, I have two multiplayer, multi-round games:

- Elimination Game Benchmark: Social Reasoning, Strategy, and Deception in Multi-Agent LLM Dynamics at https://github.com/lechmazur/elimination_game/

- Step Race Benchmark: Assessing LLM Collaboration and Deception Under Pressure at https://github.com/lechmazur/step_game/

zone411··on Gemini 3 Flash: Frontier intelligence built for speed
Scores 92.0 on my Extended NYT Connections benchmark (https://github.com/lechmazur/nyt-connections/). Gemini 2.5 Flash scored 25.2, and Gemini 3 Pro scored 96.8.
zone411··on GPT-5.2
I've benchmarked it on the Extended NYT Connections benchmark (https://github.com/lechmazur/nyt-connections/):

The high-reasoning version of GPT-5.2 improves on GPT-5.1: 69.9 → 77.9.

The medium-reasoning version also improves: 62.7 → 72.1.

The no-reasoning version also improves: 22.1 → 27.5.

Gemini 3 Pro and Grok 4.1 Fast Reasoning still score higher.

zone411··on AI agents break rules under everyday pressure
I haven't looked in the logs for this in this particular project, but I've seen this occur frequently in my multiplayer benchmarks.
zone411··on AI agents break rules under everyday pressure
I did some searches when I posted this project, but I didn't find any at the time.
zone411··on AI agents break rules under everyday pressure
Without monitoring, you can definitely end up with rule-breaking behavior.

I ran this experiment: https://github.com/lechmazur/emergent_collusion/. An agent running like this would break the law.

"In a simulated bidding environment, with no prompt or instruction to collude, models from every major developer repeatedly used an optional chat channel to form cartels, set price floors, and steer market outcomes for profit."

zone411··on Gemini 3
Sets a new record on the Extended NYT Connections: 96.8. Gemini 2.5 Pro scored only 57.6. https://github.com/lechmazur/nyt-connections/
zone411··on Gemini 3
Sets a new record on the Extended NYT Connections benchmark: 96.8 (https://github.com/lechmazur/nyt-connections/).

Grok 4 is at 92.1, GPT-5 Pro at 83.9, Claude Opus 4.1 Thinking 16K at 58.8.

Gemini 2.5 Pro scored 57.6, so this is a huge improvement.

zone411··on Poker fraud used X-ray tables, high-tech glasses and NBA players
You got many answers already, but a couple more points:

Poker doesn't require lying or table talk. Bluffing is rule-legal strategic deception expressed through betting. More like a feint in sports than cheating.

If "sitting at a table following rules" is the issue, that's true of most games. And formats vary: many are short and cash games are leave-anytime.

Page 1 of 28Next →