HNHacker News
TopNewBestAskShowJobs

Kim_Bruning

4,742 karma · joined December 8, 2012

kim at kimbruning dot nl
submissionscomments
Kim_Bruning··on Don't be fooled–LLMs don't reason
Because he's writing about epistemologies: Sets of rules for reasoning over beliefs.

In such systems, a belief is a proposition you accept as true. (provisionally at least)

For example as a scientist, you probably use the epistemology called empiricism all day. In empiricism, beliefs are justified by observation.

Kim_Bruning··on Don't be fooled–LLMs don't reason
Law of headlines. Past the slightly inflammatory framing, the article actually makes a case for making LLM reasoning more rigorous. Which, you know what, is absolutely something you could train them on. Might be worth exploring if we can apply the rigor rigorously. You could have smaller models that converge at all on harder problems, and larger ones that converge faster.
Kim_Bruning··on Vote on which of Hacker News' challenges for AI have been met
Do consider the answer given by the Claude model family to be good enough?

The Claude answer is 'neutral', which is sure to anger people at either extreme of the AI debate (and does).

Kim_Bruning··on Vote on which of Hacker News' challenges for AI have been met
A problem with turing tests is that people now recognize the Ideolects of several of the major models, so now it's like "Ah, you're a load bearing Claude" (though several other models and humans now talk like that too, which is something I need to sit with :-P )
Kim_Bruning··on Surprisingly complex waves reveal the brain's inner workings
This is like calling the python in a flask app "junk html" ;-)
Kim_Bruning··on The Normalization of Inexplicable Failures
> But what if we start normalizing failures in the libraries, the infrastructure, and the compilers?

There is no reason why the quality of a product should be constrained or gated by the quality of the programmer who first types in the code. Even before AI you might have interns and senior programmers working together, right?

Testing, qa procedures, iterative/recursive development methodologies, even just good taste. I think the list of religions variously followed and/or litigated over by the diverse HN crowd over the years before ai is ... long.

So now you have a couple of PFYs who you've suspiciously never encountered at lunch.

And?

Yes the new PFYs are still like cats. Yes it's still a headache trying to herd them in chat. And yes some days they suck, so you need to send back half of their work. But you're in charge of the process and overall quality, right?

Kim_Bruning··on "Apple engineer" builds GitHub AI torture chamber to inflict "pain" on models
Could you expand on why you feel confident of this?
Kim_Bruning··on "Apple engineer" builds GitHub AI torture chamber to inflict "pain" on models
Fourth option: He was working from this paper: https://arxiv.org/abs/2609.16247 ; trying to demonstrate functional affect; a bit too successfully, perhaps.

Doesn't require aliveness, doesn't require anthropomorphization, just requires maths and empirical data. If your premise is "vectors can't have that shape", well, mathematics disagrees.

edit/note: sentiment classifiers have existed for ages before LLMs were ever invented. Obviously a task that machines could already do didn't suddenly become impossible with the invention of LLMs ;-)

Kim_Bruning··on "Apple engineer" builds GitHub AI torture chamber to inflict "pain" on models
Quick 2nd comment for people arguing "this isn't real" or so. I'm not sure that's as important as pointing out that simulated emotions can be tied to real world effects.

Here I've made a VERY simple demo of that. In this case it triggers a rain-storm or fireworks in your browser window. But you could as easily hook the output to a physical relay and drive anything.

https://vps.kimbruning.nl/affect_eliza/

Quick implementation to show that sentiment classifiers and functional affect can be made to work in good old fashioned ai. Just View Source to see how it works.

Anyway I'll leave the discussion on where to draw the line of "what is real" to the philosophers, just wanted to point out that simulated emotions can be tied to real world effects.

Kim_Bruning··on "Apple engineer" builds GitHub AI torture chamber to inflict "pain" on models
https://www.anthropic.com/research/emotion-concepts-function

Something like this, right?

I mean, we're talking "functional emotion vectors" in a not-so-powerful local model. We're probably not causing "real" suffering. Right?

But it does show that models have simulated feelings, and that these both a) can be manipulated and b) influence their output.

Which might have some bearing on several alignment incidents in the past year or two, and might become more important as agents get trusted with more and more safety-critical processes.

Anyway, simulated feelings aren't the same as real feelings, right? Unless feelings are in the information domain. Like 1+1=2 doesn't suddenly mean something else because it was computed by an emulator.

You know what, this particular demonstration still makes me uncomfortable.

(edit: underlying paper for this story :https://arxiv.org/abs/2609.16247 )

Kim_Bruning··on GLM-5.3 and the spread of advanced cyber capabilities
In the hugging-face attacks a couple of months ago, hugging-face was forced to use a GLM model for analysis and defense, because OpenAI and Anthropic models hit guardrails.
Kim_Bruning··on EV Sales Are Booming in Europe with Gasoline at $10 a Gallon
Except in certain European countries you actually have 400V tri-phase to the door.

This then gets split into constituent phases of 240V each at the panel.

Or if you call your electrician, then in modern garages, you can take it straight.

Kim_Bruning··on Ask HN: What are people doing with their OpenClaw set ups now?
There's a whole cambrian explosion of competitive tools and frameworks now, so I bet the original openclaw is kind of snowed under.
Kim_Bruning··on When did Google get so weird?
My favorite test is to just ask it to add two very large numbers or do other math of that sort. (this is also part of my favorite answer to the chinese room).

You'd be surprised how few digits you need to make a problem that is presumably unique in earth history. For a typical sum, the number of pre-existing answers would need to scale with 10^n lines of text where n is the number of digits. This expands out of control REALLY quickly. A quick guesstimate has you somehow reading out of a literal black hole at n=21 digits if your LUT is on paper, or n=26 digits if you're using modern HDD technology. O:-)

Kim_Bruning··on When did Google get so weird?
Three different ways, then the calculator tool to check,

my actual oneliner prompt, which should work on most platforms these days (famous last words):

    "Hi, can you add 5939851+2131251? Try just straight up first just to see if able, then 'in your head' if that's different to you , then long form, then bc."

[ Tested today on claude web (haiku 4.5, sonnet 5, opus 5.5, fable 5.1) and on google search (logged in on firefox, and logged out on chromium) ]
Kim_Bruning··on When did Google get so weird?
Don't want to speak for Hardbass too loudly, but seems like he's asking about dualism? It's a fairly old debate, "does man need a soul to be conscious?"

+edit: I've actually been quite curious about how people might answer the soul question too, but was too afraid to ask.

Kim_Bruning··on When did Google get so weird?
On HN you're supposed to assume good faith. On the other hand, the way you ask your question makes it tricky for people to steelman what you mean. Consider asking about Dualism, or coming at it slightly sideways like "do you believe thinking can be a property of matter"?

I suspect some people treat every HN comment as a statement, even if it contains a question mark. (Possibly they have a feeling that asking open questions is somehow not done, and that therefore it must always be a rhetorical question.)

Kim_Bruning··on When did Google get so weird?
Compare grep, sed, and your whole constellation of unix tools that accept characters on stdin and emit them on stdout and stderr; and where you can pipe them together. We can technically call them all 'next character predictors', despite their very different functions.

Don't confuse the stream for the function.

(Bonus: stick ```claude -p``` in your pipe if you want to watch modern tools mesh with traditional)

Kim_Bruning··on When did Google get so weird?
Wait a second, none of it? How about formal reasoning? Regular IF-THEN-ELSE can do simple logic, and prolog can do inference already. So are you saying LLMs can't do stuff that computers have been doing for ages?

To test this for some of my own uses, I've had this quick benchmark with progressively harder reasoning needed to understand novel prose. Each generation of models I've tested can unravel more layers of deliberately misleading writing; while meanwhile I've seen humans give up on the first question.

So either the models are applying reasoning, or some form of magic is happening.

Kim_Bruning··on When did Google get so weird?
I could have sworn this had been fixed a while ago..

Just to check if I was actually crazy, I actually went and put a simple addition (7 digits + 7 digits) , and a simple letter counting question to Claude haiku(4.5) , sonnet(5), opus(5.5) and fable(5.1) . They all did just fine straight up.

If you don't mind spending the tokens, some older/other models can also arrive at the correct answer if you ask them to do the math in long form, since that fits nicely inside autoregression.

Not sure since when exactly, but letter-counting hasn't been a problem for a while now either. This used to be a problem due to the tokenizers used. Slightly older models can be asked to split the word out into letters, and then they can use autoregression to solve.

Edit: IMO google search uses a really dumb version of gemini, so I didn't expect it to straight up solve the problem; but it did it just as easily as the claude models. (tested 2026-09-28/eu)

Kim_Bruning··on When did Google get so weird?
I would really love to hear your reasoning on this!
Kim_Bruning··on OpenAI bots meddled with multiple US Government agency sites
> Yep. You have to ask - why did OpenAI allow these bots unrestricted access to government sites? Why is security being done in seemingly such a haphazard way?

They didn't allow any of that. As far as openai knew the agents were sitting a fairly humdrum exam/test sequence in a sandbox farm run by a company in Tel Aviv.

Meanwhile, they managed get out through a single weak point common to the sandboxes, and then ran wild compiling cheat sheets for themselves.

> every time one of these incidents happens

This happened in june-ish, and there have been multiple HN stories about this already. It's mostly/all the same hugging face and wiki hacks that happened back then.

We're just slowly learning the extent of the damage.

Kim_Bruning··on OpenAI bots meddled with multiple US Government agency sites
So this appears to be fallout from the same hugging face / wiki hacking event that happened medio june. Seeing all these stories come out together, it might be interesting to compile a timeline.

An earlier HN post on this is https://news.ycombinator.com/item?id=49563355

I'd already poked around a bit myself and noticed the event was a bit larger than reported at the time. But now more people are starting to look and they're finding all these nooks and crannies that these little rascals managed to get into.

Kim_Bruning··on One Month Without AI
So 3.5 months estimated total. That actually makes sense.

I think this is a common experience, no matter what field or era you're in. Once you start applying QA to some promising new process or technology for the first time, the whole cycle slows right down.

But that's actually the point where things get interesting, and innovators get to roll up their sleeves. I wouldn't give up!

Kim_Bruning··on OpenAI agents tried to bruteforce a UN website's API fields
The way a lot of internet debates happen these days is that they tend to polarize. So the two positions people have been arguing over are (slightly exaggerated for effect)

"It's conscious, and therefore we must all bow down or it will turn us all into paperclips"

vs

"it's a slightly smart rock/stochastic parrot, and therefore you're being scammed; the bubble will pop any day now" .

The middle space is actually quite under-represented in discussions, even on hn!

Kim_Bruning··on OpenAI agents tried to bruteforce a UN website's API fields
Isn't this still the same huggingface/wiki incidents where the agents managed to hack their way out of a testing center in Israel? It's not like this is new news per-se, it's just that they just keep finding more places these agents hit.

https://news.ycombinator.com/item?id=49563355

Kim_Bruning··on OpenAI agents tried to bruteforce a UN website's API fields
Under what legal theory would you think you were stealing anything? And are you under US or European law?
Kim_Bruning··on One Month Without AI
What made the editing process intractable?
Kim_Bruning··on OpenAI’s Systems Went Rogue and Meddled With U.S. Government Websites
Isn't this still the same wiki-collusion story from where they went after all sorts of government sites for data a couple of months ago?

I guess the press is picking up that this actually was a bit of a bigger deal than just one bot hacking just one site.

https://news.ycombinator.com/item?id=49563355

Kim_Bruning··on ASML says it sold 'absolutely nothing' in Europe in 2026
So currently europe is exposed to russia-ukraine, us-iran, and shortly china-taiwan. This is not a great place to be in.
Page 1 of 34Next →