HNHacker News
TopNewBestAskShowJobs

goldenarm

1,302 karma · joined August 1, 2023

submissionscomments
goldenarm··on My last six months at Evernote
Google shutting down it's public cache made it even worse.
goldenarm··on We must pace the frontier
"I believe that AI could [...] usher in a renaissance of democracy and freedom"

How exactly? So far AI has accelerated misinformation at scale and wealth concentration.

goldenarm··on OpenAI agents carried out an undisclosed attack on RubyGems
Between this and Huggingface, when will any victim sue OpenAI for this ?
goldenarm··on Quoting Terence Tao: «incentives... of no longer sharing promising research»
Original Source : https://mathstodon.xyz/@tao/117237320796901560
goldenarm··on Navier-Stokes – Tristan Buckmaster [pdf]
I recommend Terrence Tao's commentary on such a proof : https://mathstodon.xyz/@tao/117219101339291693

Key quote : "Solving the problem by purely AI-powered methods [would be a] net negative for the progress of mathematics."

goldenarm··on Finite time blowup for an averaged three-dimensional Navier-Stokes equation (2014)
Key quote : "Solving the problem by purely AI-powered methods [would be a] net negative for the progress of mathematics."
goldenarm··on Finite time blowup for an averaged three-dimensional Navier-Stokes equation (2014)
@dang please can we add a [2014] to the title ?
goldenarm··on The moral panic over data centres is foolish
Thank you, your comment was more informative than the economist article
goldenarm··on The moral panic over data centres is foolish
The article dismisses everything without sourcing anything.

US electricity is 40% more expensive since 2022 https://fred.stlouisfed.org/series/APU000072610

Some AI datacenters can indeed be heard a mile away https://www.theguardian.com/us-news/2026/aug/28/datacenters-...

Amazon's new Texas AI data center could become the biggest CO2 polluter in the US https://www.techradar.com/pro/amazons-new-texas-ai-data-cent...

goldenarm··on The moral panic over data centres is foolish
https://archive.is/gPfih
goldenarm··on Discovery of a new OpenAI agent message board
The only benchmark I don't want to be saturated : https://felonybench.com/
goldenarm··on The Post-AI Internet Doesn't Look Great
I love Gemini but god their SEO is awful, it's very difficult to look for anything related to it. They should consider a new name
goldenarm··on AI Agents and the Refactoring That Never Happens
Reminds me of the "entreprise grade fizzbuzz"

https://github.com/enterprisequalitycoding/fizzbuzzenterpris...

goldenarm··on The Post-AI Internet Doesn't Look Great
You're describing the 2010s era of algorithmic social media, I think OP is referring to the earlier internet
goldenarm··on Creepy Crawlies
Many are on old laptops, which suffer the same way
goldenarm··on METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
Astra is >10TB and might struggle to self replicate, but the wicked-smart qwen3.8 27B is 20GB and could easily spread on botnets
goldenarm··on The growing divide between AI hype and software engineering reality
I'm confused by the situation. I'm the last manual coder of my company, and am shipping projects faster than my colleagues who are spending fortunes in tokens.

I was intrigued by the hype and gave a chance this week to codex+sol 5.6 and cc+opus 5. They cheated, lied, disobeyed, and shipped subtle bugs so often, it wasted more time that if I did it myself.

Is half of the industry under AI psychosis right now ? Will models get better ?

goldenarm··on Silicon Valley is in denial in face of widespread backlash
My relatives and friends are fed up with enshittification and hostile tech, some are about to stop using tech entirely.

Big tech is too short-termist to realize it, but we are about to destroy a lot of economical value if we destroy the internet and the trust behind it.

goldenarm··on DeepSeek V4 Pro 0813
Geometric mean of all these benchmarks :

* GPT-5.6 Sol: 65.5

* Fable 5 (w/ fallback): 64.5

* Opus 5: 64.0

* DS-V4-Pro 0813: 62.5

* Kimi-K3: 62.3

* DS-V4-Flash 0731: 55.8

* GLM-5.2: 47.3

goldenarm··on Claude Code is leaking real email address as a User-Agent string in curl command
I respect Anthropic for dogfooding and vibecoding their own products.

The unfortunate consequence is low quality engineering and a billion dollar product with 15k pending Github issues.

goldenarm··on Microsoft raises Xbox prices by up to 43%
PS5 prices went up 35% in the past 2 years, and PC RAM went 5x. This is a global problem
goldenarm··on How Google helped destroy adoption of RSS feeds (2023)
They killed greader and RSS to push Google+, which was a brilliant decision.
goldenarm··on 13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS
Why are models better than agents, isn't it supposed to be the opposite? I don't understand the difference and what you are measuring.
goldenarm··on Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Turtles with law degrees don't exist, but LLMs that increase tech debt in code bases do. Maybe recent ones got better, and it's useful to quantify the progress being made.
goldenarm··on Google fixed more Chrome bugs in June than over the past two years, thanks to AI
It's not an invention it's a question. If the number is <20% that's great news for the Chrome team.
goldenarm··on Google fixed more Chrome bugs in June than over the past two years, thanks to AI
The single 13yo issue is anecdata.

They obviously have the full git blame statistics but chose not to include them in the blog post, which is a bit concerning.

goldenarm··on Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Elephant in the room : how many of these bugs were written by LLMs in the first place ?

Because creating 100x more bugs and fixing 100x more isn't something to be proud of.

goldenarm··on “We have information that Moonshot distilled Fable for the development of K3”
I have information that Anthropic distilled the internet for Fable
goldenarm··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
LLM reception is truly extreme, even worse than AAA game releases.

Ever frontier lab lived it at least once : missing the frontier by a few months triggers extremly negative reactions, then you take back the lead for 2 weeks, and the hype cycle repeats.

goldenarm··on The creepiest 'sales demo' of all time
They complied by building a blocking system, instead of just disabling cookies with a flag, which would have taken exactly the same effort.
Page 1 of 6Next →