HNHacker News
TopNewBestAskShowJobs

z7

906 karma · joined April 15, 2022

submissionscomments
z7··on Bend 2 and the Vibe-Coding Trap
> The field in question is formal verification. It’s notable that those two words appear nowhere on Bend’s webpage or in its codebase. The developer has built an entire language around a field seemingly without realising that said field exists.

I checked the developer's X account, they have written numerous posts about formal verification, so this specific claim ("without realising that said field exists") seems to be false.

z7··on GPT-6 Astra
François Chollet wrote in February that he expected ARC-3 to be saturated in "about one year".

"Frontier models today perform very poorly with a minimal harness. However if big labs start directly targeting the benchmark like they did for ARC-2, numbers will go up fast."

https://x.com/fchollet/status/2022054537293705260

z7··on GPT-6 Astra
Chollet writes he expects AGI now sooner than 2030, "given progress is happening faster than I expected."

https://x.com/fchollet/status/2095607046129463577

z7··on Ten advances in mathematics and theoretical computer science
> The cost of generating the proofs for all 10 of these breakthroughs combined was under $2,000 at Sol API prices.

https://x.com/polynoamial/status/2083470822258467194

z7··on Claude Fable produced a counterexample to the Jacobian Conjecture
> The incredible Yitan Zhang (https://newyorker.com/magazine/2015/02/02/pursuit-beauty) worked on proving this conjecture for 7 years. Moh, his advisor, wrote that Zhang "failed miserably" in proving the Jacobian conjecture, "never published any paper on algebraic geometry" after leaving Purdue, and "wasted seven years of his own life and my time".

https://x.com/aminkarbasi/status/2079129649830137989

https://en.wikipedia.org/wiki/Yitang_Zhang

z7··on GPT‑Live
"Hey Leibniz, how do you live with yourself knowing that your binary system helped eventually replace human conversations?"
z7··on Midjourney Medical
"I just tested my hand in a mini version of this scanner. Images that are higher quality than MRI, whole body captured in <1 minute, virtually free to run. This is going to change medicine."

https://x.com/SebastianCaliri/status/2067452733356122303

z7··on Child's Play: Tech's new generation and the end of thinking
"As Alexander predicted in 'AI 2027,' OpenAI did release a major new model in 2025; unlike in his forecast, it’s been a damp squib. Advances seem to be plateauing; the conversation in tech circles is now less about superintelligence and more about the possibility of an AI bubble."

I'm not sure how many AI researchers would find this accurate. It seems to me that under conditions of ambiguity people often default to describing their preferred version of reality.

z7··on Tesla’s autonomous vehicles are crashing at a rate much higher tha human drivers
The comparison isn't really like-for-like. NHTSA SGO AV reports can include very minor, low-speed contact events that would often never show up as police-reported crashes for human drivers, meaning the Tesla crash count may be drawing from a broader category than the human baseline it's being compared to.

There's also a denominator problem. The mileage figure appears to be cumulative miles "as of November," while the crashes are drawn from a specific July-November window in Austin. It's not clear that those miles line up with the same geography and time period.

The sample size is tiny (nine crashes), uncertainty is huge, and the analysis doesn't distinguish between at-fault and not-at-fault incidents, or between preventable and non-preventable ones.

Also, the comparison to Waymo is stated without harmonizing crash definitions and reporting practices.

z7··on Over 36,500 killed in Iran's deadliest massacre, documents reveal
> The West is not complicit in the actions of the Iranian regime

What about the 1953 CIA/MI6 coup that overthrew Iran's elected prime minister?

z7··on Self-hosting a NAT Gateway
"You only live once."

Why state this as absolute fact? Seems a bit lacking in epistemic humility.

z7··on Hi, it's me, Wikipedia, and I am ready for your apology
Here's the Grokipedia submission (currently censored / flagged):

https://news.ycombinator.com/item?id=45726459

z7··on It's insulting to read AI-generated blog posts
Hypothetically, what if the AI-generated blog post were better than what the human author of the blog would have written?
z7··on The dawn of the post-literate society – and the end of civilisation
List of dates predicted for apocalyptic events:

https://en.wikipedia.org/wiki/List_of_dates_predicted_for_ap...

z7··on DeepMind and OpenAI win gold at ICPC
Current cope collection:

- It's not a fair match, these models have more compute and memory than humans

- Contestants weren't really elite, they're just college level programmers, not the world's best

- This doesn't matter for the real world, competitive programming is very different from regular software engineering

- It's marketing, they're just cranking up the compute to unrealistic levels to gain PR points

- It's brute force, not intelligence

z7··on An LLM is a lossy encyclopedia
An encyclopaedia is a lossy representation of reality.
z7··on His psychosis was a mystery–until doctors learned about ChatGPT's health advice
Meanwhile this new paper claims that GPT-5 surpasses medical professionals in medical reasoning:

"On MedXpertQA MM, GPT-5 improves reasoning and understanding scores by +29.62% and +36.18% over GPT-4o, respectively, and surpasses pre-licensed human experts by +24.23% in reasoning and +29.40% in understanding."

https://arxiv.org/abs/2508.08224

z7··on GPT-5
Yes, but the jump in performance from o3 is well beyond marginal while also fitting an exponential trend, which undermines the parent's claim on two counts.
z7··on GPT-5
>The actual benchmark improvements are marginal at best

GPT-5 demonstrates exponential growth in task completion times:

https://metr.org/blog/2025-03-19-measuring-ai-ability-to-com...

z7··on GPT-5
GPT-5 is #1 on WebDev Arena with +75 pts over Gemini 2.5 Pro and +100 pts over Claude Opus 4:

https://lmarena.ai/leaderboard

z7··on OpenAI claims gold-medal performance at IMO 2025
Some previous predictions:

In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025.

He thought there was an 8% chance of this happening.

Eliezer Yudkowsky said "at least 16%".

Source:

https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...

z7··on Grok 4 Launch [video]
How do you explain Grok 4 achieving new SOTA on ARC-AGI-2, nearly doubling the previous commercial SOTA?

https://x.com/arcprize/status/1943168950763950555

z7··on Grok 4 Launch [video]
"Grok 4 (Thinking) achieves new SOTA on ARC-AGI-2 with 15.9%."

"This nearly doubles the previous commercial SOTA and tops the current Kaggle competition SOTA."

https://x.com/arcprize/status/1943168950763950555

z7··on O3 beats a master-level GeoGuessr player, even with fake EXIF data
Quoting Chollet:

>I have repeatedly said that "can LLM reason?" was the wrong question to ask. Instead the right question is, "can they adapt to novelty?".

https://x.com/fchollet/status/1866348355204595826

z7··on AI 2027
It's just predicting tokens:

https://old.reddit.com/r/singularity/comments/1jl5qfs/its_ju...

z7··on Marine Le Pen banned from running in 2027 and given four-year sentence
Why are you hallucinating feelings? Also, appeal to authority. ("Why are your feelings relevant to the wizarding laws of Hogwarts?")
z7··on 4o Image Generation
>For starters, this completely blocks generation of anything remotely related to copy-protected IPs

It did Dragon Ball Z here:

https://old.reddit.com/r/ChatGPT/comments/1jjtcn9/the_new_im...

Rick and Morty:

https://old.reddit.com/r/ChatGPT/comments/1jjtcn9/the_new_im...

South Park:

https://old.reddit.com/r/ChatGPT/comments/1jjyn5q/openais_ne...

z7··on Blocklist for AI Music on YouTube
The beginning of a new kind of discrimination - call it 'synthetic racism.' AI-generated music is being dismissed outright even before listening to it, not based on quality or enjoyment but purely on its artificial origin. Just as past prejudices dismissed art based on heritage rather than merit, we're now seeing a new bias against anything not 'human-made.'
z7··on US and UK refuse to sign AI safety declaration at summit
Waymo's driverless taxis are currently operating in San Francisco, Los Angeles and Phoenix.
z7··on Musk-led group makes $97B bid for control of OpenAI
I don't understand it either. Why is Elon Musk a "terrorist"? And why is this the most upvoted post? Maybe being European limits my ability to comprehend American political rhetoric.
Page 1 of 8Next →