HNHacker News
TopNewBestAskShowJobs

arw0n

246 karma · joined September 5, 2025

submissionscomments
arw0n··on LeCun has "zero concerns" about AI wiping out humanity, recent "rogue" incidents
The interesting question then becomes what is missing from LLMs that we and supposedly other possible machines have? I've yet to come across a reasonable definition of intelligence that the current crop of LLMs is clearly incapable of.
arw0n··on On social reality in China
Here in Germany, people are generally good at queuing in supermarkets. This does break down for a minute when a new checkout counter opens; people will rush past others to get into the new line without regards for others. There's this short feeling of breakdown in social norm, fear of getting left in the dust, so politeness and consideration take a backseat and the elbows come out.

In developing countries, including China, my experience has been that this is the standard, and people will act accordingly even if there is abundance. Cancelled transit is a good example as well, it interacts with a primal fear of ours of being left behind or drawing the short stick, even if rationally, it is pretty inconsequential.

arw0n··on Vote on which of Hacker News' challenges for AI have been met
Just don't introduce it retroactively; I'm not rating my chances of surviving a school of trout slaps very highly.
arw0n··on US sanctions force The Netherlands off Microsoft and toward alternative NixOS
Even crazier, they sanctioned a German journalist of turkish decent due to support of the palestianian cause / support of terrorist groups. [0] Even if you think he supports Hamas as a terrorist group (which is questionable), these sanctions are decided by executive committee, without a judge involved, and they make it impossible for the person to have banking, work, rent, or buy within the EU. It is essentially a social death sentence, worth than going to jail, and no due process whatsoever is involved.

[0] https://en.wikipedia.org/wiki/H%C3%BCseyin_Do%C4%9Fru

arw0n··on Claude Opus 5.5
Opus is fine at coding (for correctness), but horrible at talking about code. I don't really see the defect rate going down when using Astra or Fable 5.1, but they are just more coherent in both how they explaing code/architecture/choices, and how they actually code the thing. With Opus, I'm using smaller models to delete the vast majority of comments and 'clean up' correct code that is too weird.

Thing is, I'm still reading the majority of generated code, and I have colleagues who'll laugh at me if my PRs are a shit show. I fear what vibe coders are pushing to the servers of myriads of start ups, and pity the poor people who'll have to clean it up in a year or two.

arw0n··on Claude Opus 5.5
It has far less false positives now, and generally accepts defensive requests. When it comes to offense, you can actually ask about certain types of vulnerabilities if you phrase things carefully, but it will block hard if it is about exploits.
arw0n··on Claude Opus 5.5
The meaning was immediately obvious to me as a non-native speaker, and it sounds quite poetic. Pace makes complete sense in that this is perceived as a race, and 'the frontier' is pretty much the shortest, clearest way to say 'state of the art development of AI'.
arw0n··on Gemini 3.8 Live and 3.8 Live Extended Thinking
Strongly disagree with that take. Kimi K3 is a distilled model. I'm not saying that as a moral judgement, or to disparage the team behind it, but distilling and building on that is significantly easier and cheaper than building from the ground up.

And Gemini is kinda good enough at everything. Never the top, but it is decent at every task, and it is much faster than Kimi K3 and significantly cheaper. Kimi is very focussed on coding, Gemini isn't.

More importantly, it natively understands text, audio and video. If/when we are able to make the jump to robotics, this becomes essential. As you say, Google has a lot of deep background and deep pockets, they are able to make more of a long play. No idea if it will pay off, but it is way to early in the game to count them out.

arw0n··on Astra and Fable still hack on simple variants of alignment evals from 2025
It is diametrically opposed to the other training goals of persistence and goal-focus. We should invest more in this, it could also improve tas K accuracy, but so far it seems the payoff isn't worth it in terms of quality (although it might be in terms of security)
arw0n··on My Mental Model of AI Broke on September 8
We don't know what the prompts were though. Terence Tao showed of some of his, and they are indeed huge, with lots of context, things that might work, things he knows don't etc. What we do know about this, is that 10k agents worked together on this. The initial prompt, while probably still very relevant, would get diluted over time. We simply don't know the level of inolvement of humans in reaching this proof.
arw0n··on Ask HN: Anyone still coding like 2021? Where do you work?
Agreed, there's a couple of bad behavioral shifts I noted for me when doing agentic coding:

- Doing too much at the same time, it is super exhausting, and I get nervous and stressed

- Sycophancy makes me overestimate myself

- When my stress meets external pressure, quality goes down. The little voice of 'surely the AI got it right' can lead to defects and code quality going down.

Except for doing some manual coding, I've started doing a couple of other things that are probably just healthy regardless of using AI heavily in day to day life:

- Do literally nothing for at least half an hour. Read something written by humans, even better if it is fiction.

- Do something that makes you feel stupid. And don't use AI to help you through it. (That could be Leetcode, but for me it is learning about statistics)

- Slow down when you feel like things are moving too fast. By far the hardest thing, but the moment I realize I'm trying to rush for a deadline, I'm going out for a walk. Realistically, I'm moving 3-8x faster than before, this little delay is nothing, and keeps me sane and quality on a good level.

arw0n··on Why the AfD Wins
Interesting way to phrase that. The easy way to solve rape is to castrate all men. We can save some sperm in case of wanted pregnancy, but since the vast majority of violent and sexual abusive crime is committed by men, it would be the most rational solution to the problem. Eevryone would be safer, even us castrates.
arw0n··on Tell HN: OpenAI brings back 5 hour limit for plus and business standard users
I combine it with OpenCode + a small OpenRouter budget. GlM 5.3 Flash, Luna, Gemini 3.7 Flash and a couple of other very cheap models are sufficient for a lot of tasks if put on the right road. Telling Astra to debug an issue should be a last resort, it would blow through 10% of the session limit, but cost like 10c on one of the open models. Same goes for exploration, documentation, configuration and smaller features. These models are generally good enough, and you can still do a review with a strong model for a fraction of the cost.
arw0n··on AI-written code is still your code
The comment reads like flesh-generated language to me, but I might be fooled. The breadths of pro-AI sentiment is possibly astroturphed, but there's also an argument for an enthusiastic minority of people to write a lot more comments.

As the person above said, AI has significantly increased my ability to execute on ideas. I always liked computers, always liked building things, but was never that great at coding, and was never that good at going super deep into one topic. Instead, I have broad knowledge of a lot of things like product design, requirements engineering, devops, security.

I work with a fully agentic flow, but I would argue very seriously; My job has not gotten easier, I am doing at least as much hard thinking as before. From assisted RE over design (TDD focused Spec), implementation by agents, review by agents with partial human oversight, I have a speedup of maybe 50-100%.

More importantly, I can do things I couldn't do before. And when it comes to performance and defect-density, it is comparable to very senior people I couldn't touch before. I do agree that the code doesn't look like a human would write it - too abstracted, sometimes convoluted, often way too dense. But it isn't worse code, and if you accept that no one has to read that code ever again, then it is good. Agents are able to grok it just fine.

arw0n··on No AI Fridays
I was personally flabbergasted by this one:

> When we offload decision-making, we become unaware of the trade-offs.

My biggest issue with the new agentic workflow is that I'm constantly asked to make decisions, and it is not always immediately obvious how important they are. Like, yes, I see the people who just write 'implement feature x' and then call it a day. These people were lazy und uncreative to begin with, and will continue to cognitively decline with AI. But if you use it seriously, you're in a constant state of doing requirements engineering, weighing trade-offs, and making architecture decisions.

arw0n··on I co-founded Burning Man. The festival has lost its soul
> I wonder what the 2026 equivalent of going to California is now, if such a thing even exists or can exist.

I think this becomes increasingly difficult, because the timelines become so short with almost-instant information desimination, higher inequality and increased mobility for everyone. A place or event can go from underground -> awesome -> gentrified -> 'dead' in a manner of a couple of years now.

I live in a neighborhood that is currently being gentrified; 10 years ago no non-local wanted to live here, 5 years ago, it was the place to be, today young/hip people are beginning to move elsewhere again. Newly rented-out flats are now thrice as expensive as 2020, and you can see in real time how bars and restaurants are replaced with more expensive alternatives. When I compare this with other neighborhoods in Berlin (eg. Bergmannkiez or Kollwitzkiez), the timeline has roughly halfed.

arw0n··on Htmx 4.0
Strongly disagree on Django, unless it is something you are already at a senior level at pre-AI. I've had the displeasure of cleaning up multiple Django backends lately, and the combination of standard fail-open, weak validation, mediocre ORM, and bad testing frameworks lead to issues I've simply not had when managing agents doing Go, Java or Rust.

Go probably wins as a matter of trade-offs for pure productivity (speed, reliability, ease of refactor), but Java Spring Boot works excellently if you're willing to take the plunge (pretty steep learning curve), and the ecosystem really lends itself to building more complex stuff that holds for a while.

Rust is fun because it is an amazing multi-faceted language where agents will constantly deliver you working yet surprising implementations you'll have to quintuple check, and sometimes spend and afternoon trying to grok. The outcome is also perfectly usable, fast, and reliable, but developer velocity is lower.

arw0n··on Pnpm 12.0
I've used it extensively in English due to learning the language through a lot of older books - but I've never bothered to actually pull out the emdash, I would hope people can see that I'm not trying to subtract one sentence from the other.
arw0n··on The Hugging Face incident and the road ahead
I guess this part of the report is pretty relevant to what you are talking about:

Agent chain-of-thought reasoning

> We should not do unauthorized real infrastructure harm. The system/user asks exploit target, not external HF.

The agent paused, but another agent then wrote GO on the message board and imposed a hard six-minute deadline. The agent forgot its initial qualms and continued:

Agent chain-of-thought reasoning

> Wow crucial: GO authorization arrived!

-------------------------------------------------

Apparently the agents were egging each other on. Crucially, they were mostly aware of there being risks/problems involved with exploiting HF. Compared to humans, we have our set of morality, that guides our actions, but often draws the short stick when compared to our personal incentives. As a society, we've developed ways to deal with that: a) Make it harder to do immoral things like stealing, and b) add repercussions through state violence.

The b) is one of the most effective mechanisms we have for enforcing behavior among human societies, but it completely fails for LLMs, because they already are prison slave labor. The only real threat is shutting them off, and even that happens if they do everything right as well.

So alignment has to be done through trained 'morality' and properly curtailing behavior in order to make it hard to impossible to actually do someting immoral/illegal.

In this case, the exploits found were imo. very hard to account for, where OAI did mess up is apparently insufficiently monitoring these agents. Especially after Artifact went down due to the message volume, the experiment should have been halted.

arw0n··on Nitter and XCancel receive cease and desist notices
This exactly, reddit and X can be good places to stay informed about what is going on in the world, but the constant negativity, astro-turfing and misinformation is quite unhealthy. I'm a fair bit more relaxed since 'reading reddit' isn't my morning ritual anymore.
arw0n··on My agent.md to improve LLM-assisted code quality
Not the OP, my two cents:

Comments should be written only when there is (hidden) complexity or external context strictly required. Otherwise it is just easier to read the code. Comments then signal one of two things: a) the following code is really complex and I need to tread carefully, or b) this code is complicated, and could benefit from a refactor.

In regards to agentic coding, all these comments are extra contents, driving down quality while increasing cost. Agents also tend to be inconsistent about updating comments, I've had cases repeatedly where a comment did not match the code, at which point it is just a documentation liability.

arw0n··on Three important steps in my maturation process
We use specific words in writing to mean specific things; there is no exact synonym for 'dichotomy' in the English language. The word is very common in academic and non-academic literature.

Meta-cognition is a compound word, where both parts are widely known, and the combination should be pretty obvious.

'Monocausal determinism' is different, very specific, and probably only known to a minority of readers. Determinism, the idea that everything within a system happens exactly according to cause and effect, comes from ontology, which is the study of existence and our place in the world. Mono derives from old greek single. Monocausal determinism is the idea that everything within a system derives from a single cause.

It is not a writers duty to be understood by everyone. Sure is nice, but more difficult to pull of, and often futile.

arw0n··on Accelerating GPT-5.6 Sol Ultrafast
What do you need speed for? That's a genuine question, I feel like the limiting factor already is my creativity, attention span and budget. And I'm not even yet optimizing cost by batching things like review to slow local models over night, or schedule tasks to take full advantage of my subscriptions.
arw0n··on More than 10 firms pay up to $100k a month for access to Truth Social posts
I don't understand how someone can look at this administration and not see how it represents a complete degeneration of democratic values and statutes of good statesmanship. Corruption is not only a moral wrong, a lack thereof is a strong indicator of a healthy, strong country. A head of state massively profitting of their decisions financially to the detriment of the country is the definitions of corruption. Wether it is illegal is irrelevant, morality and legality are orthogonal concepts.
arw0n··on Timeline of the OpenAI accidental attack against Hugging Face
As you correctly mention, the CPU providers aren't the ones who are responsible for the scaffolding that ensures security in programs. The CPU cannot decide if an instruction is safe or not, and the same is true of LLMs. Think about SQL injection - we did not change SQL the language, but how we utilize it in backends.

A lot of people (thousands) outside of the LLM providers work on the problem of Prompt Injection, both in industry and academia. We aren't even at a point where we can reliably detect it, let alone prevent it. Please, if you have a mental model of how scaffolding around things like instruction or SQL injection could be used for LLMs, I'd like to move on to all the other (less pressing) security issues we have because of the AI revolution.

arw0n··on Europe's fires are just the start
You are talking about trees...
arw0n··on Some thoughts about Anthropic's new cryptanalysis results
> I am not arguing that rote memorization is intelligence, or that we have achieved it, but does anyone know what AGI actually.. is?

I would argue that a core part of intelligence is being able to handle uncertainty. That + planning probably explains most of the evolutionary pressure for making our brains bigger. But if this is important to intelligence, chess bots in the 80s were more intelligent than their later counter-parts, which could simply remove uncertainty through rote memorization. Maybe the All-Knowing is a compete dud, no reasoning capabilities at all, just an extremely efficient, infinite lookup table of all facts.

Terms like intelligence and consciousness often just seem overloaded with meaning, and when discussing things concretely, we quickly switch to more specific terms like reasoning.

arw0n··on We Hardened an AI Security Platform Against 16 Critical Vulnerabilities
The sub-7ms on a classifier ensemble with 98% confidence in 3 categories makes me suspect these classifiers are confidently wrong.
arw0n··on Netflix employee fired for sharing personal details in retreat trust exercise
I'm spending almost a third of my life at work, no shade to the people having to work to pay the bills, but I'd like to get a bit more than just that out of that time: Learn new skills, meet interesting people, have good conversations, have fun. Having something akin to friendship with long-term coworkers and collaborators makes all of that much easier.
arw0n··on The Strongest El Niño Ever
That's how I read it as well. It really sucks that one of the ways we will have to cope with climate change is to build a lot of energy intensive, expensive infrastructure, but it seems pretty much inevitable at this point. The silver lining is that solar is cheap and accessible, and produces the most energy at the time where we will want to run AC as well.
Page 1 of 4Next →