The authors are just being lazy.
22,276 karma · joined February 2, 2011
Sandbox your agent so that it can't do any real damage, and you don't have to keep answering annoying permission questions.
The authors are just being lazy.
Of course you can get the same kind of thing in binary. I wrote a drop-in binary JSON replacement because it's easy to write a binary one that's twice as fast as simjson and yyjson. The important thing is to never give the drop-in replacement any extras that break roundtrip compatibility.
Was a fun little project to write, and quite useful for me: https://github.com/kstenerud/bonjson
Yes, there were genuinely helpful people there. I strove to be helpful as well. But my most common experience there was:
- People asking why you're doing that (without answering the question).
- People telling you that you shouldn't do that (without answering the question or providing an alternative).
- People closing your question as off-topic or duplicate without actually reading and understanding what you wrote.
- People editing your question and changing the meaning (leading others to innocently conclude off-topic or duplicate). Or removing actually important details.
- People giving just plain wrong answers that get upvoted, with comments from more senior people begging them to change or remove it.
Or in the words of Brick Top: "If I throw a dog a bone, I don't want to know if it tastes good or not."
And yes, every so often a golden answer by an amazing person.
Now I can just ask a few LLMs, comparing and contrasting their answers. Even better: I can interrogate them on their answer.
Stack Overflow had a community of sorts, but it wasn't anywhere near like a physical community.
We may work online, but we exist in the flesh.
The built-in "sandboxes" these companies provide are laughable.
It's very easy to argue with stawman arguments.
Every time the LLM throws jargon around, you call it. "What do you mean by gated wedge?" You call its bullshit, check what it's saying against your understanding of the overall system, and keep it on the straight and narrow.
It's a lot like supervising a junior dev who happens to be very quick at absorbing lots of info, but not so great at the big picture.
Until I fully understand what's going on, the PR doesn't move and my interrogation of the LLM doesn't end. My interaction is littered with "Explain X" and "How does this square with Y?" and "What if Z happens?"
The interrogation is the point, without me having to wade through hundreds of lines of irrelevant code to get at the meat of the matter.
As a test, I built a sandbox with only the host-side filtering proxy allowed for networking. 99% of traffic was HTTP. No QUIC at all.
npm, pip, apt, go, curl and git-over-HTTPS all worked on the standard proxy environment variables alone. No mirrors or other coaxing needed.
DNS is disallowed through the chokepoint, but that's no problem because the proxy resolves host-side anyway.
I used two agents: One with network access to collect the evidence, and one with everything except the model endpoint cut off, which did the analysis.
The second agent's entire input was attacker-authored. So PHP droppers, obfuscated loaders, database rows, filenames, blah blah.
In this case I'm more worried about hostile input attacking the agent, and I need to contain the damage. My sandboxing solution does that by restricting access to the source data, making it read-only. The work dir can only transfer data via patch and apply (like a git workflow), so even my workspace can't be modified until I approve each change. And then restricted network means that any compromise ain't going noplace.
The second agent couldn't even install PHP or contact any CVE site to check if it was looking at a known attack, and that was by design. All it could do is write up a report about what it observed, not make assumptions about what it is. I could then take its (much smaller) clean output and pass that to a third agent with network access.
This is forensic work, so of course not your median dev's bread & butter. But the attack surface is only just starting to be plumbed. Compromising input can turn your agent into their agent, planting things as easily as planting worms was back in the early internet days when people connected without a firewall.
It wanted me to log in for some reason even though it had worked fine for months. Login uses my Google log in.
When I try that, Google asks me for a hardware key to complete the login, even though it's my phone and I'm already logged in.
Eventually I figured out that if you select "log in using another device" and then click cancel when it brings up the qr code, you can select a push notification on "another device", which actually pops up on the same device. Do that once and it fails. Do it a second time and it succeeds.
All to use the tailscale app on my own phone.
No, it wasn't.
It was just coincidental that the UK still held a 99 year lease (from 1898) when the communists took over mainland China. How they got the lease in the first place is another matter.
- USB cables with at least one USB-C end
- All other USB cables
- Network cables
- Audio and video related (RCA, HDMI, SCART, phono, XLR, displayport, BNC, S/PDIF, etc)
- Power cables
- Power adapters
- Batteries and battery chargers
- Flash sticks, type adapters, diagnostic tools
- Spare keyboards and mice
Not a week goes by where I'm not grabbing something from that cabinet. A couple of months ago I actually needed a SCART cable for the first time in forever for some old hardware.
They're great for short distance travel in densely populated areas, but outside of that I doubt they'd hold up in long-term cost and efficiency vs autonomous electric cars for long.
As for parking, a fully autonomous car service would allow a car to come to you and pick you up. So it wouldn't need to park near where you are.
First thing I do when something goes wrong is tell the agent to stop and diagnose. You can't prompt effectively without three proper information.
Once you understand what the problem is, you can give it better instructions. If the architecture is shit, the agent is going to have a rough time of it.
But it's in there!
Some other interesting tidbits:
- In America, perception of government corruption has consistently held about 11% higher than the perception of corruption in business. Even with the massive uptick of both since 2024, they've remained a similar relative distance apart.
- Since 2023, only four countries out of the 132 in which Gallup has posed this question annually have scored nominally above the 89% recorded in the U.S. this year: Lebanon in 2024 (92%), Peru in 2025 (92%), and Ghana and Nigeria in 2024 (both 90%).
Tools like VS Code do the job well enough for the majority of people, with just enough configurability and much greater ease-of-use.
Emacs has a similar problem to Lisp: Infinite configurability and expandability (plus the lack of a "blessed set" standard that people actually like enough to use out-of-the-box) means that everyone's environment and tooling ends up becoming incompatible with each other.
Obligatory https://www.youtube.com/watch?v=A6bJHalRnX4
> question → Google → mediocre explanation → another tab → try something → nope → read again → change my model → finally get it
> Now it can be:
> question → ask AI → extremely decent explanation
That's your problem in a nutshell: If you're not doing the extra verification steps like you used to do with Google, you're stopping at the mediocre explanations like you used to with google.
I once caught my Claude agent complaining that it couldn't connect to https://some-weird-domain.com because the network was down (I disable network in the sandbox when it doesn't need it, and broker the API connection). I asked why it was looking there and it told me I'd asked it to.
I never found any evidence of prompt injection, but it sure as hell made me paranoid.
They also started before that era, and after. I've met plenty of them throughout my career, and still do every so often. They were never very common.
> They wore all the hats
This isn't unique to that era.
> In that era, the devs at a startup were the "everything people".
They still are. You do everything because you don't have the resources to hire for any positions.
> Anyone who chose to specialize went the vocational route of getting certifications and playing the contractor game.
That depends on your definition of "specialize". I'm using it in the sense of learning one particular aspect well (in addition to anything else you learn). Certifications and contracting have nothing to do with it.
> This pattern already happened in the (blue collar) "trades". It started in the late 20th century. Now we have some of the worst built homes, buildings, roads, etc.
This is certainly true in the USA, although I'd argue that it's the companies cutting costs (and corners) rather than the tradespeople becoming unskilled. In most first world countries there are standards and audits that keep companies (mostly) in line so that you don't get disasters-in-waiting like the new Bay Bridge.
> Do we really want to make this even worse with automated AI copypasta pretending to be "engineering"? Nobody is fooled.
Like with any force multiplier tool, you get the early days where every Joe is pumping out crap (same kind of thing happened with COBOL), and then the shake-out happens where companies realize that it's not a free lunch, and you need skilled people after all.
I'd say that you need to provide some evidence for this extraordinary claim.
> What if I told you that your "specialists", AI or not, are going to be objectively far worse at actually getting things done because they don't have the bigger picture in mind?
I'd say that you need to provide some evidence for this extraordinary claim.
> We've seen this race to the bottom in other disciplines before.
Such as?
"Accumulate cognitive debt": The paper is a preprint, isn't peer reviewed, and has already had a rather blistering critique (https://arxiv.org/abs/2601.00856) that raises concerns about sample size, reproducibility of the analyses, EEG methodology, inconsistent reporting, and transparency. Also the fact that they infer EEG connectivity = learning without evidence.
"Less engaged with your work": Doesn't measure engagement with work. It measured motivation, boredom, and sense of control on small tasks like writing a post or an email. The effects are small, and some of them cut against the site's thesis.
"Negatively impact your critical thinking abilities": It's a survey of knowledge workers about past tasks. It doesn't measure critical thinking ability at all. Even the paper's title says "Self-Reported Reductions in Cognitive Effort and Confidence Effects". It doesn't establish causation, and the direction is ambiguous.
"Hamper your skill formation": The one paper that has something, showing that using AI without learning about the things it's doing reduces knowledge acquisition (note it does not test retention). HOWEVER, the best AI users matched or beat the no-AI group!
It's not that they lost it, but rather that they never learned it. I didn't lose my ability to reason about those things during the 4 years where I worked in Python and Java.
And once again, for average software devs that's fine. Most non-critical software that doesn't perform very complex tasks can be optimized for time-to-market, which of course loses in other areas the business doesn't care as much about (such as bloat).
When that's not the case, you hire a specialist.
In the snow.
Uphill.
Both ways.