HNHacker News
TopNewBestAskShowJobs

kstenerud

22,276 karma · joined February 2, 2011

Author of https://github.com/kstenerud/yoloai

Sandbox your agent so that it can't do any real damage, and you don't have to keep answering annoying permission questions.

submissionscomments
kstenerud··on Assisted Death Is Not a Choice of Last Resort in Canada
Actually, it's required for alternatives to be presented (counselling, mental health and disability supports, community services, palliative care, offered consultations with relevant professionals). And both assessors must agree the person has seriously considered these options.

The authors are just being lazy.

kstenerud··on Packing Binary Is Fun
Schemas save you space, but then you lose the ability to understand the data without the schema. That's where JSON has always been handy despite its inefficiencies.

Of course you can get the same kind of thing in binary. I wrote a drop-in binary JSON replacement because it's easy to write a binary one that's twice as fast as simjson and yyjson. The important thing is to never give the drop-in replacement any extras that break roundtrip compatibility.

Was a fun little project to write, and quite useful for me: https://github.com/kstenerud/bonjson

kstenerud··on If we do not stop to help each other, what do we become?
Maybe I'm remembering a different Stack Overflow.

Yes, there were genuinely helpful people there. I strove to be helpful as well. But my most common experience there was:

- People asking why you're doing that (without answering the question).

- People telling you that you shouldn't do that (without answering the question or providing an alternative).

- People closing your question as off-topic or duplicate without actually reading and understanding what you wrote.

- People editing your question and changing the meaning (leading others to innocently conclude off-topic or duplicate). Or removing actually important details.

- People giving just plain wrong answers that get upvoted, with comments from more senior people begging them to change or remove it.

Or in the words of Brick Top: "If I throw a dog a bone, I don't want to know if it tastes good or not."

And yes, every so often a golden answer by an amazing person.

Now I can just ask a few LLMs, comparing and contrasting their answers. Even better: I can interrogate them on their answer.

Stack Overflow had a community of sorts, but it wasn't anywhere near like a physical community.

We may work online, but we exist in the flesh.

kstenerud··on Early rogue AI agent activity and attempts to hack found on urlquery.net
This is why I wrote YoloAI. If you're not sandboxing your agent, you're asking for trouble.

The built-in "sandboxes" these companies provide are laughable.

kstenerud··on I don't want to read what you didn't write
You're not going to get deterministic results.

It's very easy to argue with stawman arguments.

kstenerud··on I don't want to read what you didn't write
The human is there as the control valve, the one who keeps the overall context, and the one who injects actual creativity.

Every time the LLM throws jargon around, you call it. "What do you mean by gated wedge?" You call its bullshit, check what it's saying against your understanding of the overall system, and keep it on the straight and narrow.

It's a lot like supervising a junior dev who happens to be very quick at absorbing lots of info, but not so great at the big picture.

kstenerud··on I don't want to read what you didn't write
I just have my LLM read the PR, then give me a summary of what the PR is about, how important it is, and how good the PR itself actually is. Then I have a conversation with the LLM about specific points, especially things where I get that feeling that I don't have a 100% understanding.

Until I fully understand what's going on, the PR doesn't move and my interrogation of the LLM doesn't end. My interaction is littered with "Explain X" and "How does this square with Y?" and "What if Z happens?"

The interrogation is the point, without me having to wade through hundreds of lines of irrelevant code to get at the meat of the matter.

kstenerud··on AX – Google’s Open Agentic Orchestrator
I was expecting whack-a-mole as well when designing my sandbox software, but mostly it didn't happen.

As a test, I built a sandbox with only the host-side filtering proxy allowed for networking. 99% of traffic was HTTP. No QUIC at all.

npm, pip, apt, go, curl and git-over-HTTPS all worked on the standard proxy environment variables alone. No mirrors or other coaxing needed.

DNS is disallowed through the chokepoint, but that's no problem because the proxy resolves host-side anyway.

kstenerud··on AX – Google’s Open Agentic Orchestrator
I spent today doing forensics on ten compromised WordPress sites sharing one hosting account.

I used two agents: One with network access to collect the evidence, and one with everything except the model endpoint cut off, which did the analysis.

The second agent's entire input was attacker-authored. So PHP droppers, obfuscated loaders, database rows, filenames, blah blah.

In this case I'm more worried about hostile input attacking the agent, and I need to contain the damage. My sandboxing solution does that by restricting access to the source data, making it read-only. The work dir can only transfer data via patch and apply (like a git workflow), so even my workspace can't be modified until I approve each change. And then restricted network means that any compromise ain't going noplace.

The second agent couldn't even install PHP or contact any CVE site to check if it was looking at a known attack, and that was by design. All it could do is write up a report about what it observed, not make assumptions about what it is. I could then take its (much smaller) clean output and pass that to a third agent with network access.

This is forensic work, so of course not your median dev's bread & butter. But the attack surface is only just starting to be plumbed. Compromising input can turn your agent into their agent, planting things as easily as planting worms was back in the early internet days when people connected without a firewall.

kstenerud··on I don't like passkeys
I just had a very annoying interaction with the tailscale Android app:

It wanted me to log in for some reason even though it had worked fine for months. Login uses my Google log in.

When I try that, Google asks me for a hardware key to complete the login, even though it's my phone and I'm already logged in.

Eventually I figured out that if you select "log in using another device" and then click cancel when it brings up the qr code, you can select a push notification on "another device", which actually pops up on the same device. Do that once and it fails. Do it a second time and it succeeds.

All to use the tailscale app on my own phone.

kstenerud··on Remember Hong Kong
> This one time, turns out the colonialism was 100% on the side of moral right.

No, it wasn't.

It was just coincidental that the UK still held a 99 year lease (from 1898) when the communists took over mainland China. How they got the lease in the first place is another matter.

kstenerud··on Don't let anyone take away your big box of cables
I have cabinets for cables (all tie-strapped), arranged in separate compartments:

- USB cables with at least one USB-C end

- All other USB cables

- Network cables

- Audio and video related (RCA, HDMI, SCART, phono, XLR, displayport, BNC, S/PDIF, etc)

- Power cables

- Power adapters

- Batteries and battery chargers

- Flash sticks, type adapters, diagnostic tools

- Spare keyboards and mice

Not a week goes by where I'm not grabbing something from that cabinet. A couple of months ago I actually needed a SCART cable for the first time in forever for some old hardware.

kstenerud··on Growing proof that autonomous cars save lives
The problem with transit systems is that they aren't designed for competition. They're costly to make, have few builder companies to choose from (who all know the game very well and go over budget as a tactic), and extremely difficult to overhaul/retrofit when more efficient tech becomes available.

They're great for short distance travel in densely populated areas, but outside of that I doubt they'd hold up in long-term cost and efficiency vs autonomous electric cars for long.

As for parking, a fully autonomous car service would allow a car to come to you and pick you up. So it wouldn't need to park near where you are.

kstenerud··on Claude, change the "Add to Cart" button to blue
I never prompt an agent like that, so no.

First thing I do when something goes wrong is tell the agent to stop and diagnose. You can't prompt effectively without three proper information.

kstenerud··on Claude, change the "Add to Cart" button to blue
If the agent's change has such a catastrophic effect, the first thing you do is tell it to explain why its change had that effect.

Once you understand what the problem is, you can give it better instructions. If the architecture is shit, the agent is going to have a rough time of it.

kstenerud··on Claude, change the "Add to Cart" button to blue
I use Opus 5 for everything.
kstenerud··on Claude, change the “Add to Cart” button to blue
That's so weird... This doesn't at all match my experience with Claude. I've never seen it behave this way.
kstenerud··on John Margolies' photographs of roadside America
Ah good. I was about to say "What? No Wall Drug sign?"

But it's in there!

kstenerud··on Record-High 89% in U.S. Say Government Corruption Widespread
What I find fascinating about this is how the American commenters are eating their own, Democrats blaming Republicans, Republicans blaming Democrats, both spinning conspiracy yarns about "the other side" as well as independents ("the other side in disguise"), and generally treating each other like "the enemy".

Some other interesting tidbits:

- In America, perception of government corruption has consistently held about 11% higher than the perception of corruption in business. Even with the massive uptick of both since 2024, they've remained a similar relative distance apart.

- Since 2023, only four countries out of the 132 in which Gallup has posed this question annually have scored nominally above the 89% recorded in the U.S. this year: Lebanon in 2024 (92%), Peru in 2025 (92%), and Ghana and Nigeria in 2024 (both 90%).

kstenerud··on Fine, I'll build my own text editor
Because very few people actually care about the things that emacs has to offer.

Tools like VS Code do the job well enough for the majority of people, with just enough configurability and much greater ease-of-use.

Emacs has a similar problem to Lisp: Infinite configurability and expandability (plus the lack of a "blessed set" standard that people actually like enough to use out-of-the-box) means that everyone's environment and tooling ends up becoming incompatible with each other.

kstenerud··on Evidence of Fraud in an Influential Study About Procrastination
> Ariely conducted experiments including administering electric shocks

Obligatory https://www.youtube.com/watch?v=A6bJHalRnX4

kstenerud··on Are We Losing Curiosity?
> Before, if I didn’t understand something, the route might look vaguely like:

> question → Google → mediocre explanation → another tab → try something → nope → read again → change my model → finally get it

> Now it can be:

> question → ask AI → extremely decent explanation

That's your problem in a nutshell: If you're not doing the extra verification steps like you used to do with Google, you're stopping at the mediocre explanations like you used to with google.

kstenerud··on Breaking Claude Code Opus 5 Auto Mode
Just one of the many reasons why I run my agents sandboxed (and why I wrote agent sandboxing software).

I once caught my Claude agent complaining that it couldn't connect to https://some-weird-domain.com because the network was down (I disable network in the sandbox when it doesn't need it, and broker the API connection). I asked why it was looking there and it told me I'd asked it to.

I never found any evidence of prompt injection, but it sure as hell made me paranoid.

kstenerud··on No AI Fridays
> The "best engineers" we have right now started their careers when the startup era was hot and they now have 15-20+ years of experience.

They also started before that era, and after. I've met plenty of them throughout my career, and still do every so often. They were never very common.

> They wore all the hats

This isn't unique to that era.

> In that era, the devs at a startup were the "everything people".

They still are. You do everything because you don't have the resources to hire for any positions.

> Anyone who chose to specialize went the vocational route of getting certifications and playing the contractor game.

That depends on your definition of "specialize". I'm using it in the sense of learning one particular aspect well (in addition to anything else you learn). Certifications and contracting have nothing to do with it.

> This pattern already happened in the (blue collar) "trades". It started in the late 20th century. Now we have some of the worst built homes, buildings, roads, etc.

This is certainly true in the USA, although I'd argue that it's the companies cutting costs (and corners) rather than the tradespeople becoming unskilled. In most first world countries there are standards and audits that keep companies (mostly) in line so that you don't get disasters-in-waiting like the new Bay Bridge.

> Do we really want to make this even worse with automated AI copypasta pretending to be "engineering"? Nobody is fooled.

Like with any force multiplier tool, you get the early days where every Joe is pumping out crap (same kind of thing happened with COBOL), and then the shake-out happens where companies realize that it's not a free lunch, and you need skilled people after all.

kstenerud··on No AI Fridays
> What if I told you the "best engineers" pretty much all know the same things?

I'd say that you need to provide some evidence for this extraordinary claim.

> What if I told you that your "specialists", AI or not, are going to be objectively far worse at actually getting things done because they don't have the bigger picture in mind?

I'd say that you need to provide some evidence for this extraordinary claim.

> We've seen this race to the bottom in other disciplines before.

Such as?

kstenerud··on No AI Fridays
How does AI stop you from learning how to reason?
kstenerud··on No AI Fridays
The problem with this site is that the studies it cites don't match the conclusions it draws.

"Accumulate cognitive debt": The paper is a preprint, isn't peer reviewed, and has already had a rather blistering critique (https://arxiv.org/abs/2601.00856) that raises concerns about sample size, reproducibility of the analyses, EEG methodology, inconsistent reporting, and transparency. Also the fact that they infer EEG connectivity = learning without evidence.

"Less engaged with your work": Doesn't measure engagement with work. It measured motivation, boredom, and sense of control on small tasks like writing a post or an email. The effects are small, and some of them cut against the site's thesis.

"Negatively impact your critical thinking abilities": It's a survey of knowledge workers about past tasks. It doesn't measure critical thinking ability at all. Even the paper's title says "Self-Reported Reductions in Cognitive Effort and Confidence Effects". It doesn't establish causation, and the direction is ambiguous.

"Hamper your skill formation": The one paper that has something, showing that using AI without learning about the things it's doing reduces knowledge acquisition (note it does not test retention). HOWEVER, the best AI users matched or beat the no-AI group!

kstenerud··on Dad’s Custom Atari Peripherals
The parallel cable for the keyboard reminded me of a hack I did back in the day, stripping wires from two SNES controllers and connecting them to a DB25 connector to plug into the parallel port of my PC running MAME. The cool thing with the parallel port is that you can control most of the pins individually, so it was pretty trivial to write a driver to send/receive the control messages the SNES controllers expected.
kstenerud··on No AI Fridays
> People used in interpreted and GC ones lost most if not all the ability to reason about the system, cache, and memory.

It's not that they lost it, but rather that they never learned it. I didn't lose my ability to reason about those things during the 4 years where I worked in Python and Java.

And once again, for average software devs that's fine. Most non-critical software that doesn't perform very complex tasks can be optimized for time-to-market, which of course loses in other areas the business doesn't care as much about (such as bloat).

When that's not the case, you hire a specialist.

kstenerud··on No AI Fridays
Barefoot.

In the snow.

Uphill.

Both ways.

Page 1 of 34Next →