HNHacker News
TopNewBestAskShowJobs

meander_water

1,399 karma · joined February 9, 2025

I write software and I write words about software:

https://vivis.dev

https://findsubstack.com

https://pythonkoans.substack.com

submissionscomments
meander_water··on There is more to code review than (automatable) detection
I actually think this might be a false positive. I think it got tripped up by the higher than average use of jargon (which I don't mind here because the article itself flowed well and raised good points).

Gpt-zero scores "human", and I've always found it to be a better judge

meander_water··on How to keep enjoying programming in a world of LLMs
Sounds like you could just use an autocomplete model for the exact same benefit.

And that's not a criticism, that's what I've come back to myself.

meander_water··on Show HN: JevBench, a reproducible benchmark for typed decision models
They released some examples of what their workflow evals are like. I'm sure you could reverse engineer a benchmark from that

https://evals.typesafe.ai/

meander_water··on We must pace the frontier
> I believe that AI could cure most major diseases in the next 5–10 years, greatly accelerate economic growth rates, create a world of abundance and empowerment, and usher in a renaissance of democracy and freedom

translation: I'm going to build a product and hope that it works, promise it to be a panacea while realistically having no control over how it's used, the wider economic impacts it causes or how it changes the mental health and cognition of its users.

meander_water··on Resist "AI"
This is not being against technology. It's ensuring that technology is used to serve humans, and not the other way around.

https://caiml.org/dighum/dighum-manifesto/

meander_water··on Resist "AI"
What I would love to see is a side by side comparison of two companies (one using AI, and one without) over an extended period of time.

Id wager that they would fare the same. The time saved on some tasks are outweighed by extra time spent on others. Everyone assumes that just because you can pump out more features, you're more successful.

meander_water··on Resist "AI"
A humanist revolution.

Sure we can't stop the technology, but we _can_ stop the importance and money heaped into it.

meander_water··on Resist "AI"
I see a lot of people throw their hands up in the air and say "I have no control over this, might as well go all in on AI".

But we as individual consumers have more power than we think. We don't need to use these AI features and products. We've survived perfectly fine without AI. Made plenty of scientific discoveries and lead enriched lives without it.

The biggest "fuck you" to the pricks who pushed this on us would be if we collectively went "meh".

meander_water··on AI Is Breaking This Thing We Call Trust
You trust machine code because it was produced by a deterministic system that was written and tested thoroughly by other engineers.

AI systems are not the same. You can't guarantee deterministic output.

Also, code is still the best language we have to convey our intentions as engineers into function. Natural language is too imprecise and ambiguous.

Understanding the code allows us to understand the intention behind the code and identify bugs, plan and architect better. Inspecting prompts gives you a blinkered view of the system meaning you're more likely to make incorrect assumptions leading to serious bugs.

meander_water··on OpenAI's GPT-6 Astra on ARC-AGI-3
The OpenAI charter defines it as:

"highly autonomous systems that outperform humans at most economically valuable work"

https://time.com/article/2026/08/26/openai-sam-altman-interv...

meander_water··on Six curl CVEs after OpenAI and Anthropic came back with zero
It's an effective approach. Google project zero started doing this in 2024

https://security.googleblog.com/2024/11/leveling-up-fuzzing-...

meander_water··on The turbulent AI era is here
> I wish the world could get the benefits rapidly and delay the problems it will cause as long as possible, but the benefits and problems are arriving at the same time.

I'm not trying to be facetious, but what are the benefits of "general purpose intelligence"? With machine learning it was obvious what task we were trying to automate, and you could take a stab at the productivity achieved.

That's not the case now. You can't figure out if the time it took to review a 5k loc PR balances out the time saved debugging a syntax error.

Yes we know you can churn out your damn CRUD apps and (hope) to make money (PS, you don't need AI to make money - see exhibit A: crypto bros).

So what benefits has this actually brought to humanity?

meander_water··on Why does Opus 5 feel worse to work with?
> First, the desire to create a self-improving AI that is capable of recursively bootstrapping itself to AGI/ASI.

I feel like we've become blinkered in this quest to push the frontier at all costs. Somehow the target has shifted from economic productivity to a vague notion of general intelligence.

I don't think productivity gains are going to be found by trying to generalise all tasks. I think we need to go back to specialist models that do one thing well. I'm perfectly happy to use one model for coding, and another for penetration testing, which have different goals.

Opus 5 and other frontier model's tendency to be relentless and cheat their way to a goal is great for hacking, but not so great when you have to build a maintainable, reliable codebase.

meander_water··on AI is removing the middle class of software engineering?
I think there's a corollary to this. Not only is it hollowing out the middle class, it's preventing the junior engineers from stepping up to the senior level.

Everyone starts off as a bad engineer. Just like any other profession, you to make lots of mistakes to learn. But now you're less likely to get wisdom from another human to build up the knowledge and experience you need. And more likely to delegate the hard stuff to Claude so you don't learn from your mistakes.

meander_water··on Karpathy’s Pelican
This is what I was looking for, thanks!
meander_water··on Karpathy’s Pelican
Sure, but then you would just use an image generation or multimodal model to generate that image. I don't think you'd want a weird looking svg.
meander_water··on Karpathy’s Pelican
Can someone explain what the pelican on a bicycle tests exactly? And why is it so important? I've never understood how it could translate to a useful task in real life.
meander_water··on Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
Firstly, I don't have many issues with benchmarks per se. But I do have issues with leaderboards. And the AA index is touted by lots of people to argue that X model is better than Y, which I find inaccurate.

> I mean, it's telling that your gamut of examples are three models that are within spitting distance of each other on the broadest benchmarks.

This is kind of my point. The benchmarks say they are splitting distance, but they actually vary wildly in performance for specific tasks, so they are in fact not equivalent.

meander_water··on Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model.

A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task.

For e.g. you might use Fable for UI design, Sol for systems design backend work and Kimi K3 for exploit development.

The only purpose these metrics serve is bragging rights for the model companies.

meander_water··on OpenAI’s accidental attack against Hugging Face is science fiction that happened
1. The benchmark is run with a python script - https://github.com/sunblaze-ucb/exploitgym using an agent harness. I suspect they used codex. So the model has access to the environment and could trivially inspect its own source code, which has lots of references to exploit gym and docs relating to it.

2. Hanlons Razor

meander_water··on Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
Seems similar to Openrouter Fusion - https://openrouter.ai/docs/guides/routing/routers/fusion-rou...
meander_water··on Micro-Agent: Beat Frontier Models with Collaboration Inside Model API
I thought all model providers are doing this under the hood anyway in their UI?

They certainly seem to when A/B testing different models, and Fable routes to Opus 4.8 when guardrails fail.

Also, openrouter recently released a fusion router - https://openrouter.ai/blog/announcements/fusion-beats-fronti...

meander_water··on Exploring the internal representations of Pangram 3.3.2
GPTZero is much better at handling humanized outputs. Also has a similar false positive rate to Pangram.
meander_water··on Stealing Is a Skill
> However, it’s your job to go down the rabbit hole, learn the 100%, and sprinkle in your 3%.

I would say that there is a big difference between stealing without acknowledgement, and stealing with acknowledgement and actively learning through reverse engineering.

meander_water··on Use AI for reviewing code especially when the diff is huge
> I don't think you should waste time reviewing every single line of code in here and just use AI to review it!

> What you bring is the knowledge that the author nor the LLM doesn't know.

How can you possibly know what relevant context to provide the LLM unless you read the 10k loc? Now you've wasted double the time.

meander_water··on GLM 5.2 vs. Opus
Thanks, I didn't mean to be brusque, but I have seen a lot of these vibe tests lately that come to grand conclusions like "X model is better than Y" from the result of a single prompt.

Appreciate you sharing the results of your tests though!

meander_water··on GLM 5.2 vs. Opus
> So we ran it head-to-head against Claude Opus 4.8: same one-shot prompt, build a 3D platformer in raw WebGL from scratch

Running a single one-shot prompt is not a benchmark, not is it representative of any sort of real-world usage.

Most agent usage is collaborative so you need to test things like reliability (when I delegate a task, does it complete it without making up test results for e.g.) and steerability (does it obey my instructions or does it just do what it thinks is best).

meander_water··on Building an HTML-first site doubled our users overnight
Really curious to understand why I'm being downvoted. I don't think it's a particularly spicy take - Just choose the right tool for the job.
meander_water··on Building an HTML-first site doubled our users overnight
As someone who has built both react based frontends and html based ones (with htmx), there is a law of diminishing returns at play.

To start off, writing a basic crud website with forms is much easier with htmx.

But when you start building more complex components, and integrate with other systems (OAuth for e.g.) there are tons of libraries and SDKs for the react ecosystem, but not many for pure html components.

At this point, it's much easier to use off the shelf components than it is to manually write html to handle all the bizarre UI edge cases.

meander_water··on Claude Fable 5
All the model releases we've seen this year have only made incremental improvements in benchmarks.

This feels like the first release that feels like a significant step up in terms of benchmark results.

Can anyone make an educated guess what the secret sauce in the model architecture is between 4.8 and Fable?

Page 1 of 7Next →