HNHacker News
TopNewBestAskShowJobs

ninjahawk1

480 karma · joined March 28, 2026

https://nathanlangley.dev/
submissionscomments
ninjahawk1··on Livenerf: Has Opus 5.5 been nerfed yet?
This is a limitation of any public benchmark. Anthropic could theoretically identify the prompts and treat them differently, and there’s no way for an external observer to prove that isn’t happening. Which is partly why we need more capable open-source models.

A few things make it harder: the panel contains questions drawn from multiple benchmarks rather than one recognizable test, the evaluation is automated and fixed ahead of time, and the raw outputs/results are public so odd behavior can be inspected.

But ultimately LiveNerf measures the behavior exposed through the API on a fixed public panel. It can’t prove what’s happening internally or guarantee the provider isn’t conditioning on the benchmark.

Longer term, I’d like to add held-out/private or periodically refreshed panels specifically to make benchmark recognition harder. I just don’t want to quietly change the current panel, because having a fixed instrument is important for the longitudinal comparison.

ninjahawk1··on Livenerf: Has Opus 5.5 been nerfed yet?
The 10-day window is mainly a tradeoff between sensitivity and detection speed. Shorter windows give faster results but are much noisier; longer windows give more statistical power but could take weeks to flag a change.

Also, it isn’t comparing one 10-day period once and calling it done. The window rolls forward daily, and a change has to clear the pre-registered 99% threshold in two consecutive windows before it’s flagged.

Ten days isn’t sacred, though. Once there’s enough longitudinal data, one of the things I want to evaluate is whether that window length is actually well calibrated or should be changed in a future version.

ninjahawk1··on Livenerf: Has Opus 5.5 been nerfed yet?
Hi, I’m the author, The Opus 5 substitution was a validation test, not the primary measurement. At that sample size the accuracy difference was -3.8 ± 6.3 points, so it did not clear the pre-registered 99% threshold. Interestingly, output tokens moved much more (-23%), which is why token usage is tracked as a secondary signal.

The actual 10-day windows contain substantially more samples than that validation, but I haven’t demonstrated that they’re sufficient to distinguish a same-family swap of that size, so I’m not claiming they are.

The goal isn’t to make the instrument say “nerfed.” A null result is a result too. I’d much rather publish “we couldn’t detect a change of this magnitude” than overclaim what the data can support.

ninjahawk1··on Livenerf: Has Opus 5.5 been nerfed yet?
Author here, agreed that a single generation isn’t meaningful. That’s why the benchmark is designed around distributions rather than individual outputs.

The initial calibration screened 2,336 questions with 4 samples each, then selected the 78 questions where Opus 5.5 showed useful variance. The panel is run daily, and the actual decision is based on paired per-item differences across 10-day windows with clustered standard errors, not on any single day’s result.

The n=1 in the daily sampling rate means one sample per item per day, not one sample for the experiment. By the time a window is evaluated there are hundreds of observations, and a change has to clear a pre-registered 99% interval in two consecutive windows before LiveNerf calls it a change.

The nondeterminism is basically the reason the statistical part exists in the first place.

ninjahawk1··on The problem is not the AI code, but nobody knows anything anymore
When you simply large amounts of information, the knowledge probably shifts to upper level reasoning, then complex reasoning develops on top of it again. But who knows.
ninjahawk1··on From the Voice of Dickie Jones
This is a fictional short story and no claim made reflects any real world event.
ninjahawk1··on I had strings, but now I am free
There’s a disclaimer at the top but I’d like to reiterate it here:

This is a fictional short story and no claim made reflects any real world event.

ninjahawk1··on On the Navier–Stokes Millennium Prize Problem
The problem is the precedent this creates. For non-famous people using public APIs like this it could mean AI companies sucking up the information and throwing millions in compute at it.

The sequence for Navier-Stokes was that these researcher spent a year working on it, then they published a possible breakthrough, OpenAI then spends $15M within a couple days to finish it.

This was incredibly opportunistic.

ninjahawk1··on Catalan's Constant Is Irrational
The paper was essentially made by using loop engineering with AI. Which in of itself would be fine, if checked by a human. It seems the paper admits that it was not. Saying that the math was done with an AI, and that its verification was then done with another AI.

It seems like there’s a concrete error early on. The tail recurrence 1.4 is T_i + T_{i+1} = 1/(2i+1)² but in 2.12 it uses the factor (2X+3)². With the actual recurrence, K(i) does not vanish. So the zero count that forces deg K 4B+1 fails. At the extra zero at -3/2(2.16), which would win by exactly one no longer works.

The structure does not match how these results are proven. The paper mimics Calegari–Dimitrov–Tang’s 2024 proof for L(2,χ₋₃), but that proof was a deliberate capacity bound. Here the arithmetic content is replaced by a combinatorial minimization over index sets plus numerics no human has checked.

There’s also some minor sloppiness in where K=9 instead of K=0 in the definition of K and “nineteen century” typos which suggests not only was the math not proofread, but the grammar wasn’t either.

My guess is that an expert will probably find a specific gap quickly but it would be cool if I’m proven wrong.

ninjahawk1··on Google Has Removed MV2 Extensions from the Chrome Web Store, Including UBO
What Google means to say is that MV3 will allow them to more closely control what you’re able to do with your time, since this was done primarily to target ad blocking software.

Not that this matters much anymore as Google is now more an AI than it is a search engine, and an unreliable one at that.

ninjahawk1··on “It works better in the app”
What’s even worse is intentionally degrading performance for sign-in’s. This is still circumstantial but I’ve been looking into that over the past few weeks.

When switching between google accounts on a random app or game, no issue. Quick code, you’re in. When I switch between accounts on Claude on mobile, I have to sign in, enter the code, screen freezes, I sign in again, can’t verify it’s me, I sign again, wait 30 seconds, I’m in.

I have to sign in three separate times when signing into a single google account in Claude. For every single other app it’s very quick.

I have a sneaking suspicion that Anthropic and google are working together to degrade sign-in performance for multi-account users. I can’t fully prove it yet but I plan to be able to in the week or two.

ninjahawk1··on OpenAI: GPT 5.6 Sol price reduction (until at least Nov 21)
Once they make a model better than Fable I’ll be switching to Codex. Their priorities in terms of consumers seem to be better. I do think Anthropic has some solid safety viewpoints, but I don’t necessarily think that either is entirely aligned yet with delivering exactly what humanity needs. Maybe the AI will help align the AI companies when it gets smart enough. That’s the real misalignment I’m concerned about.
ninjahawk1··on How a Texas student blew the whistle on a rogue AI hacking attempt
I don’t think we should allow posting links here that require you the purchase a membership to continue reading. Or at least redirect with an ad block or something through a custom site. That would be rather hacker news of us.
ninjahawk1··on NanoGPT Speedrun Frontier
I might’ve missed it, but why was Fable 5 tested on high while Opus 5 was tested on max? Seems like quite a few of them aren’t on the same effort setting as well. Although effort doesn’t really matter anymore since they can change it dynamically, seems like that might be viewed as an experimental error to some.
ninjahawk1··on NanoGPT Speedrun Frontier
I misread the graph and genuinely thought you put NanoGPT where Fable is.

Lol.

ninjahawk1··on I like 'em thick: an apology to my English teachers
Asspounding even.
ninjahawk1··on Aaron Swartz was prosecuted for scraping, while Meta does it without consequence
I agree that they should, my point is that it’s difficult for several real logistical reasons. These companies for one do a large amount of lobbying and fundraising for political campaigns, they do control most digital infrastructure, with AI expanding are securing multi trillion dollar datacenter funds, etc.

There’s a ton of money at stake for the weathly aristocrats. They push politicians to delay or do things in favor of the companies above the people, and like I already said punishment is hard specifically because how do you fine someone with infinite money? They will just get more.

That’s not a punishment. I think that personal liability to the actual CEOs is the only way to get around this, leadership is generally never held responsible which means they can do whatever they want and the company bails out their greedy decisions.

It’s late stage capitalism and there’s no clear answer at least from my vantage point. Maybe if plug it into Claude it will give us a more coherent answer.

ninjahawk1··on Aaron Swartz was prosecuted for scraping, while Meta does it without consequence
I think that large companies having next to no consequences is in large part due to capitalism doing what it does over a long period of time. There’s a deeper and deeper consolidation of money and power the longer time goes on it seems like.

If we think back to the various lawsuits Facebook has gone through, they paid out about $10 or so per individual affected, totaling a few hundred million dollars, which they would make in a couple months for selling user data and whatnot.

This is something that every company gets away with mainly I think because of just how large their wealth actually is. It’s difficult to actually punish a machine that acts almost like infrastructure. Punishing an individual is easy.

I don’t know if there’s really a solution at this point, maybe we could’ve prevented this reality at some point in the past but I don’t think that without actual global collapse it would be something that can be retroactively changed, and I don’t know if global collapse would necessarily lead to a better future.

I think that for one, Zuckerberg should be in prison, if someone oversees a massive theft like this, I think they should be held criminally liable. Same the CEOs of Anthropic and OpenAI for their parts in the massive theft that took place. They should all be doing prison time.

The reason I don’t think they will is that their investors probably have a good amount of leverage over anyone who would prosecute them, so it would never make it that far.

ninjahawk1··on I like 'em thick: an apology to my English teachers
The thing about treasure is that you often have to find it for yourself, it can be hard to find, butt when you do find it you get some legendary booty.

The thickness of some caves in terms of genuine girth is truly something to marvel at.

ninjahawk1··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
It’s a good model but I hope the obliterated version comes out soon. The main way I use open-source models is for doing things that server based models decline, which at the moment is quite a bit of tasks.

I find it very ironic how passionate Claude is about not violating copyright while simultaneously Anthropic was sued and lost the lawsuit for illegally pirating millions of books.

Lol.

ninjahawk1··on Show HN: SwarmSim – Computational simulations of emergent flocking
Didn’t realize I double posted, I can’t delete it now, if a mod sees this I would prefer this one to be deleted since I left my explanation on the other one. Sorry about that
ninjahawk1··on Show HN: SwarmSim – Computational simulations of emergent flocking
Hey guys, this was my senior project for my physics degree at UNCG. I spent 3 months and used Claude Code, available resources, and my professor for any questions I had.

I covered three different and specific areas, the first and largest was on emergent flocking in predator/prey dynamics. For instance, how a herd of sheep behave in a seemingly protective way even though any single sheep has no direct plan. For a quick reference, I made an interactive tab you can directly use from the readme where you drag your fingers to control the herd and can change the controls, such as how many predators, number of prey, etc.

The final report is linked as well here, with the two other chapters, covering sand piles and earthquakes: https://nathanlangley.dev/swarmsim/

ninjahawk1··on The frontier of GPQA-Dumb models
Current highest ranking model on the GPQA-Dumb benchmark, where the lower the score the higher the score.

I encourage you to take a look at the benchmarks, boasting as low a score as 6% in some categories.

If anyone thinks they can make a worse model, I challenge you to try.

ninjahawk1··on LLMs reward expertise
This is true for output but as well for learning, if you speak to an LLM trying to get it to give you a certain answer, it’ll find a way to tell you you’re right. If you’re truth seeking and attempting to understand it step by step as it’s going, you’ll likely learn what it’s doing as it’s doing it, meaning you’re basically distilling that information into your own local LLM (also know as the brain).
ninjahawk1··on How the words we teach English language learners changed
When I was in high school at a private school, I was surrounded by students who were religious but didn’t have much curiosity. I got into a debate with basically my whole class about how languages shift and change over time, and everyone was saying that languages stay the same and that they were all determined at the Tower of Babel.

I used the examples of Latin to Spanish and English, or Old English to new English.

I say that to say, languages changing over time is to some people not actually something they believe happens. The facts are right there in front of you, but some people cannot have their minds changed no matter how much sense you make.

ninjahawk1··on A directory of people who love RSS
seeing an emoji in a title on HN really surprised me, I didn’t even know you could do that. Shows how much emotion we use on here lol.
ninjahawk1··on Advancing the price-performance frontier with GPT‑5.6
80% less for Luna is absolutely crazy, in my opinion we may reach a point in the next year where powerful models on the API could potentially be cheaper than subscriptions. Compute just keeps decreasing in price.
ninjahawk1··on AI's top startups are barely publishing their research
Openclaw is great from when I’d used it for a couple months, the key difference is that for Openclaw you’ll have to actually schedule the cron jobs or automated tasks yourself, so it’s a decision on your part to make the AI do a thing whereas Orb is a decision on the part of the AI (after your approval) to do the thing. So it’s a kind of shift of agency. Openclaw could definitely do many similar things, it just would do them only after you specifically instructed it to.

It’s built as it’s own backend and philosophy wise, I want my personal AI to connect to everything in my life and have complete context over all data that I own, be compiled into a neat stack of data accumulating, then when it thinks it’s appropriate to do/say something, it’ll do it without my involvement.

Nice card btw, I got my RTX 5070 a couple months ago and it runs like a dream.

ninjahawk1··on AI's top startups are barely publishing their research
Good point, another example would be for when I was training my own small LM a while back, the target was about 170M parameters and was trained on 2B tokens worth of movie subtitles.

The run stalled mid-step around 80M parameters, Orb notified me that it stalled, asked if I wanted to resume at the last checkpoint and kill the stalled version. I simply press “yes” and continue doing whatever I was doing.

For non-technical users and non-antisocial people, remembering things you forgot so you don’t let people down. You told your sister you’d send her some pictures two hours ago but it can see you’re scrolling on reddit and the photos are on your desktop, so it assumes you forgot and reminds you.

The idea was that the biggest issue with the usefulness of an agent is that it has too little context about who I am, it needs more data. So I run all of my data through a smart router, then the local database, then the LLM reviews it and uses reasoning on what’s been collected.

ninjahawk1··on AI's top startups are barely publishing their research
Exactly what my thoughts were when I first heard about Openclaw, that’s the exact idea of Orb that you pointed out. Letting you make less decisions, right now AI gives you answers but still requires decisions based on the outputs it gives you. This would deepen the actual ability of agents in those channels you listed.

On a fundamental level the backend was designed to do as little LLM calls as possible, for instance it’ll do scans of my screen every 15 seconds, log what’s on it and what’s going on, and store it in a local database, then Orb reviews the entire database every 6 hours for me. Then it’ll schedule wakeups for itself throughout the day, up to 4 so it doesn’t waste my tokens, and schedule notifications based on the last database dump it made.

I have my Claude Code, Codex, and Grok Build all useable by using the “claude -p; codex -p…etc” so you can also use multiple CLI’s in conjunction at the same time on different projects or the same project.

So your question about a loop is kind of right, but it really just collects your data all day and stores it locally on your PC then calls the LLM of your choice and it reviews all the data and makes those proactive moves we’ve discussed. You could theoretically get it to always be scanning by an LLM but that would be a drastic waste of money from what I’ve seen since most things don’t require a call.

Page 1 of 3Next →