HNHacker News
TopNewBestAskShowJobs

johnfn

13,367 karma · joined August 27, 2009

https://x.com/thesilenceturns
submissionscomments
johnfn··on Livenerf: Has Opus 5.5 been nerfed yet?
You are correct, I softened the wording a bit. And thanks for the heads up on the typo!
johnfn··on Livenerf: Has Opus 5.5 been nerfed yet?
Sorry, you are correct - I modified my original post. I get frustrated every time there's a model release and 1 week later everyone is saying NERF! NERF! 99.9% of the time these people are wrong, but you are right that it's technically not 100% due to a few edge cases.

I am more skeptical about the compute provider claim - do you have any evidence of that?

johnfn··on Livenerf: Has Opus 5.5 been nerfed yet?
It's not about doing more complex things - complexity is more dictated by how large your codebase is, etc.

> It’s not a crazy conspiracy that the same model can be stupider

Sorry, I really do think it's a conspiracy. If nerfing were real, it would be trivial to prove. DeepSWE, SWEBench, and other benchmarks are all available for anyone to run. A "nerfing" hypothesis has to survive the fact that a statistically significant dip in benchmarks has never been observed.

johnfn··on Livenerf: Has Opus 5.5 been nerfed yet?
"Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.

I made a graphic to explain why people feel like the models get nerfed:

https://x.com/thesilenceturns/status/2103551351825543610

The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.

johnfn··on Show HN: Lofi Cities – Pixel-art city nights with browser-generated lofi
I mean pixel art is kind of weird, in that I rarely just look at a single screen of pixel art and say, wow, I feel more connected now. It’s often taken in context of a larger whole, like a game or animation where the artist is usually trying to express something, and the specific art is some small part of the overall expression. It’s kind of like how no individual still image from Spirited Away will convey the meaning of the whole movie.
johnfn··on Show HN: Lofi Cities – Pixel-art city nights with browser-generated lofi
I use a compiler to perform a task. I don’t use pixel art to perform a task, other than to feel connection to another human —- something AI makes more challenging.
johnfn··on Claude Opus 5.5
Did no one in this entire HN thread read past the headline of the post from Amodei? He wrote a very clear set of actions Anthropic is taking.
johnfn··on I built non-autoregressive decision models with RL a year ago
I just typed dash twice into iPhone. Hardly a “non-standard” web interface. Also, I’d like to think my comment was higher quality than anything an LLM could generate! At least, yet.
johnfn··on I built non-autoregressive decision models with RL a year ago
It’s a tale as old as time — people don’t understand that marketing and branding are just as important, if not more so, than the product. Jev is exceptionally-well branded. Anyone can look at the webpage and understand it, and the implications, instantly.

OPs “marketing” is a single post on Reddit titled “ Predicting sales conversion probability from conversations using pure Reinforcement Learning”. Can you understand what that means? I can’t, and I consider myself reasonably technical. Is it obvious it has the same implications as Jev? Again, no idea. And it was just a single post on a subreddit that I don’t even browse! I see people on this thread saying “Jev is just BERT”. Sure, and Dropbox is just a ftp account mounted with curlftpfs!

I do feel bad for the author for finding something cool and being unable to brand it. But the full definition of “product” INCLUDES being able to coherently communicate it. In some sense the branding is just as much the “breakthrough” as the model.

johnfn··on Bend 2 and the Vibe-Coding Trap
Pretty impressive to accuse the author of not knowing formal verification when even minutes of research would immediately prove the opposite (https://x.com/victortaelin/status/2100942399132312059?s=46, https://x.com/victortaelin/status/2100374221671051472?s=46).
johnfn··on There Is No AI (It's Just People) with Jaron Lanier
Some guy at OpenAI asked AI to solve a task and it ended up hacking RubyGems. Interesting that that never happened when I ran Bubblesort.
johnfn··on Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1
Wow totally missed this, thanks for sharing
johnfn··on Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1
I'd be happy to read a source as I am fairly confident that audio transcription -> facebook (or other) ads has never been true.
johnfn··on Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1
Models do not get nerfed. There has never been evidence of this. This would be trivial to prove if it were true, and such a proof would be a huge story and scandal to a news market hungry for a shred of a signal on AI's downfall.

This is the "your iPhone is listening to you and serving ads based on what you say" of the 2020s.

johnfn··on Growing proof that autonomous cars save lives
As a biker, I've had human drivers intentionally try to drive me off the road or hit me. Neither driver education nor higher test standards will solve that problem.
johnfn··on METR Report on OpenAI / Hugging Face Hacking Incident
Are you claiming that an independent investigation is actually a marketing stunt?
johnfn··on How accurate have Ed Zitron's AI skeptic predictions been?
Zitron is wrong, not "early", and the post has an extremely long list of examples.
johnfn··on The Teaser Period: Why the AI Boom Is Hitting a Reset Wall
The article claims that the finances fall apart because OpenAI won't hit the growth it wants. It's an interesting thing to say when using their product to write 100% of the article.
johnfn··on The Teaser Period: Why the AI Boom Is Hitting a Reset Wall
I really want to get in the mindset of people who are like "AI is totally going to fail! Haha! Now let me just use AI to write a piece about it..."
johnfn··on DeepSeek-v4-flash-vision-exp
How is this a "gotcha" question? "gotcha" implies there's some sort of trick. This is just... a question.
johnfn··on fx :Tiny, open, native coding agent.
How is it different from any other installation method?
johnfn··on Claude: System Prompts
It’s really fascinating to me how when the community dislikes certain content people always jump to conspiratorial justifications rather than the much more mundane “this content is not very good”. The first article is poor and all the comments on it say exactly that. The second one actually did fairly well for what is a fairly middle-of-the-road Ask HN. There’s no shadowy cabal removing anti-AI content from the front page.
johnfn··on How Claude's text watermarking works
I'm not sure how that changes the question -- just add `has_watermark(text) < 0.5`.
johnfn··on How Claude's text watermarking works
Sure, but I imagined it'd be something like Anthropic handing over this API only to trusted third-parties, not everyone in the world.
johnfn··on How Claude's text watermarking works
> We will soon be offering a watermark detection API. We’re in the process of working out the details of its implementation.

Dumb question - doesn't this defeat the purpose of a watermark? i.e., anyone who wants to avoid detection can simply run `while (has_watermark(text)) text = slightly_rewrite_with_non_anthropic_llm(text)` until it's gone? I feel I am missing the intent of the watermark if it is so easily defeated.

johnfn··on Understanding is the new bottleneck
I remember when I raised this point like a year or two ago -- in response to someone saying that coding AI made all their work trivial I said something like "if you have multiple agents the work changes and becomes more managerial - don't you think that managers contribute value" and I got a bunch of downvotes and all the responses were like "no manager has ever contributed value." Ahh, good times.
johnfn··on Accelerating GPT-5.6 Sol Ultrafast
This does look pretty incredible, but don't forget that incredible token thoroughput can only necessarily solve certain bottlenecks. If your e2e tests take an hour, they'll still take an hour after Ultracode. If the agent runs a 10 minute typecheck after a change, that will still take 10 minutes. grep over a massive codebase is still just as slow, etc. I say this not to take away from this accomplishment but just to ensure everyone here keeps a clear head about what it means - 14x faster tokens does not mean it completes every task 14x faster.

I suspect Humanity's Last Exam is without tool-calls, making it kind of the perfect benchmark to highlight how fast Ultrafast is, but not really the same as the everyday work you or I do.

johnfn··on Show HN: Isopolis – Isometric pixel map of SF
I think “soulless” crosses a line into unkind. Just say you don’t like AI generated text.
johnfn··on Taste Is All That's Left
> the one-word lines

The funny thing is that there are no "one-word lines". This more or less cements the fact that it must be AI in my mind.

johnfn··on Show HN: Isopolis – Isometric pixel map of SF
Ouch, this seems unnecessarily harsh for a side project made for fun. This might be acceptable if this was a low effort single prompt to an AI but it’s pretty clear from the write up that that is far from the case.

FWIW I thought this was really cool.

Page 1 of 34Next →