HNHacker News
TopNewBestAskShowJobs

everforward

2,890 karma · joined December 10, 2021

submissionscomments
everforward··on You said no MCP
I’m not sure; the integrated GUI seems like a major differentiator for them.

Pi’s agent is supposed to be simple, and a simple ACP agent is like a couple hundred lines of code. Making a system that allows UI plugins is way harder.

Also not sure if you’ve seen but you can get ACP from Pi with https://github.com/svkozak/pi-acp It bridges Pi’s RPC mode to ACP, works okay but not amazingly. My thinking level selector in Zed has never worked with it but everything else I use has worked (I’m sure other things don’t but I must not use them).

everforward··on You Said No MCP
I’m still not entirely sold. I do hear you about team workflows, I’m just not sure MCP does that dramatically better (as of today).

I can see the snap appeal. One protocol, we can chuck an auth reverse proxy in front of all the MCPs, compliance has their integration point, etc.

I don’t think MCP is structured enough to give a huge edge over bash there. Looking at MCP messages, they aren’t immediately more legible than a bash command, output and exit code. You also don’t own a lot of the MCP servers you use, so backtracking for audits will require knowing what MCP commands did what back then.

I do suspect something more like MCP than bash will be the winner. MCP just feels very open source rather than enterprise. Eg I don’t think I’ve seen any sort of privilege escalation and logging scheme. The enterprise will want some sort of “request admin privileges” scheme. Likewise they’ll probably want more context on ACP requests; who is calling this MCP, using what agent, and for what project?

everforward··on Dots: Always-on agents
I'm working on a project that solves some of this [1]. It runs agents in Docker and proxies the ACP connection via websocket, with WASM-based plugins in that proxy so you can get in between your client and the agent if you want.

It currently tears down the container after the session, but it wouldn't take much to leave it running post-connection and make a mode that continually re-uses the same container.

It's also possible to intercept ACP read/write file and shell commands via one of the WASM-based plugins if you wanted to execute them in a separate VM/container. I'd have to double check on FS permissions for the Docker socket; I think plugins have no file access currently because I haven't figured out a permission system for it yet. The whole plugin system is new and I'm still working out some of the edges.

Feedback and feature requests welcome!

[1] https://github.com/SethCurry/abyss

everforward··on America.gov
They can take it as authoritative, but that won't change the consequences if they're wrong. This is already a thing between two people; "my cop friend told me it was okay" isn't a legal defense. Changing the 2nd person to an AI doesn't change much, both are still agents of the government.

Doesn't really have any bearing on a similar system for HR. The government is sort of unique in a bunch of ways, but not really being responsible for what their low level agents say is one of them.

I'm not sure how to fix that, to be honest. I'm also not sure how much worse it is than trying to Google it? Google is full of stuff that's either wrong, or won't apply for a reason that takes some reading comprehension to grasp (e.g. state-specific rules/programs/etc).

everforward··on 500k facial scans at UK stations yield no arrests, 1 false positive
That’s not hard to value, it just doesn’t make Flock look particularly good. If these actually work well, you would expect the average catch rate to be nearly 100%. We had decent rates just having humans look, Flock should be clearing those like crazy if it works.

It doesn’t. I had a car stolen, called it in as they pulled away. I live a block from a Flock camera, 2 others apparently captured it. Car wasn’t intercepted, car was never recovered, no arrests were ever made. They could have put up scarecrows and it would have done as much good for a lot cheaper.

> the system will come off pretty badly in terms of ROI.

It’s a bad evaluation function. A lot of the cost of crime isn’t directly the crime, it’s the economic deadweight of trying to deter the crime. Every car needs anti-theft parts, lots of people buy home security systems, armored cars to do bank drops, etc, all this stuff we pay for that exists solely to try to stop crime.

Flock would look great if it worked so well it made armored cars and car anti-theft obsolete. It won’t, but if it did the ROI would be crazy.

everforward··on The problem is not AI code, but not knowing about system architecture or intent
> But today you can do even better, with AI can do true waterfall by rewriting from scratch many times.

This does nothing other than ensure you end up in the "joy" of running a v0.0.1 product but for years on end instead of for a few months.

I hear this sort of thing a lot, and I can't help but internally translate it to "I've never had to take oncall for a product directly after a rewrite". It will be broken; not even because the AI is "wrong", but because the rewrite has new edge cases. No one rewrites a project to have the exact same edge cases. Those edge cases will become outages. No one will learn anything, because a month from now it will be rewritten and those edge cases will get swapped for something else; you can pick which edge of the CAP theorem you want to live on, but you can't pick "none of them".

everforward··on The problem is not AI code, but not knowing about system architecture or intent
I agree with you, but it's a moot point for software engineering.

> Your concerns have moved up the stack to managing requirements, context, and verification processes.

This has always been the concern. "Add oauth to this app, there are no requirements beyond oauth working" has been an intern level task for ages. What makes software engineering hard is when the requirements start adding "well it has to use this oauth backend that isn't technically spec compliant, and the user is going to send some kind of random token you need to translate to oauth, and...". The problem isn't in writing code that will do the thing, it's figuring out exactly how that backend isn't oauth compliant and what chain of API calls I have to make to convert their random token into an oauth one, and etc.

Producing software that complies with a test suite isn't really novel. You've been able to outsource that forever. This falls apart in the same places outsourcing does; I'm sure India/Phillipines/etc/ is more than capable of iterating on code until it passes a test suite.

everforward··on Flock Wants the Most Detailed Map of Its Surveillance Cameras Taken Offline
I think you misunderstand the balance of power. If you're talking about power over you specifically, your local clerk almost certainly has more than your senator. If they both want to make your life suck, the clerk is the one with the relationship to the local PD, the ability to make your filings get lost in a pile for a couple days so they were technically filed late, etc. Your senator or representative is likely in 0 of the processes you are likely to engage in unless you're notable enough that you specifically get targeted by the state legislature.

> Also, there are private individuals who knowingly exercise more power than that clerk.

This is partially a pet peeve of mine, but again, almost certainly not. I struggle to think of someone with more ability to kill a project than a clerk that refuses to do their job.

Kim Davis was a clerk that refused a literal Supreme Court order. As best I can tell, she did 5 days in jail and launched a million dollar speaking career out of it. Bezos or Musk can be absolutely sidelined by a clerk who is pissed off enough to accept eventually getting fired as a result.

everforward··on Ember-1
I vaguely recall a project from a while back that did something similar without LLMs.

I’m really pushing my recall, but I want to say it was written in Ruby and stored pre-configured commands that it just did traditional search over.

I vaguely recall it working okay because 99.99% of the questions people asked were the same (“tar command to gzip a directory and strip the prefix” is something I google like once a month).

everforward··on On caring for user data: NeoVim caused Vim undo files to be deleted
At one point vim lacked asynchronous plugins. If a plugin was running a builder or linter it locked the editor up (from what I recall).

That fell apart when people wanted vim to do some more modern IDE kind of things like all the “… on save” stuff (build on save, test, lint, etc). I think LSP support is native in neovim as well.

I believe vim merged asynchronous plugin support a while back though, so I’m not sure how different they really are anymore.

everforward··on Turning GLM-5.3-Flash into a Jev-like decision model
I think the truth is sort of halfway between yall.

Hallucinations in tool call results _do still exist_, but basically everyone asks for a JSON schema for the tool call and uses that to validate and re-prompt the LLM until it emits something with a valid schema.

That all goes out the window when a string field has a “hidden” schema in that only particular strings are valid, but that restriction isn’t in the JSON schema. I have had failures when I want a field to be specifically formatted Markdown or something.

We’ll probably see something that handles this better in the future like jsonnet or Cue or dhall that has some execution capabilities so you can write a custom validator beyond what JSON schema supports.

everforward··on U.S. appeals court upholds designation of Anthropic as supply chain risk
Bombing civilian infrastructure is a war crime because the end state is the same as bombing civilians. Killing power shuts down hospitals and emergency responders (generators run out eventually) and desalination plants. The military is largely unphased, they’re the first ones to get gas for generators.

Also, this is carpet bombing. Carpet bombing is the targeting of civilian infrastructure with effectively willful ignorance of the collateral damage.

We started a war knowing the only way to win was either boots on the ground or war-criming our adversary into submission. We don’t get to pretend our hands are tied and we have to send out the bombers. We don’t have to do this because Iraq forces us to, we have to do this because we elected a man who can barely read Post It notes and ignored half a century of military intelligence. We are at fault for every dollar of damage caused to infrastructure and every life lost.

everforward··on Plan mode is dead
I haven’t tested with Claude specifically in a while, but I see this a lot on larger features.

It tends to be small decisions way down the stack that bubble up, or an incoherent data model that can’t handle what you’re asking for cleanly.

Eg I was messing with a state tracker the other day. The state tracker assumes a container is either currently running, or fully removed from disk.

The LLM chose to remove the state file when the container is stopped and then to remove it after, which leaks container storage.

The LLM is kind of stuck though, because every option other than “rewrite the data model” has negative outcomes and it probably violates user expectations to launch a massive rewrite there.

everforward··on U.S. appeals court upholds designation of Anthropic as supply chain risk
Militaries have a lot of experience, and an exceptionally poor track record.

I don’t remember the last war that didn’t have credible evidence of war crimes occurring. Were bombing civilian infrastructure in Iran, Iraq had Abu Ghraib among all the Collateral Damage stuff, the Highway of Death in the Gulf was probably a war crime, Vietnam had My Lai, WWII was the advent of carpet bombing civilian infrastructure. I can’t think of any for Korea, but I also know very little so that doesn’t say much.

We still haven’t charged anyone for the second strike on that fishing boat in South America, and I haven’t heard a single rationale for why that’s not a war crime other than “fog of war”.

The US doesn’t even really have a meaningful system for finding and prosecuting these, because we aren’t signatories for the ICC and have a bill saying we’ll invade if they charge one of our service members with a war crime. We aren’t basically the furthest thing from having any experience prosecuting war crimes. I can probably count on my fingers the number of cases we’ve tried. We rarely charge our own service members, and we usually kill foreign combatants rather than capture and charge them.

everforward··on U.S. appeals court upholds designation of Anthropic as supply chain risk
He didn’t say they were _the_ corporate party, just that they’re dominated by corporations. It’s notable because it’s an accusation that they pitch at Republicans often (not that R is doing any better there, but it’s part of their platform more or less).

The dems haven’t done anything notably anti-corporate in ages. There are rumbles about doing it (anti trust, supporting unions, climate change), but it never _quite_ seems to actually materialize into anything.

Cynically, that’s why both sides lean so hard into the culture wars. The men in suits would rather we argue about how to interpret history than whether a wealth tax makes sense.

everforward··on Oracle cites 'force majeure' to shield itself on controversial data center
The data tier is one of the lowest tiers in the stack (for bigcorps with centralized DBs), and changes bubble up the stack. Change an API and a few upstreams have to change. Change the core DB and _everything_ has to change.

That makes the schema calcify, and the DB becomes the most stable format of the data. I've worked more than one place where "upgrade the core DB between major versions" was a multi-quarter effort.

everforward··on Linux support is coming to Snapdragon X2 series
I don’t hear people suggesting to replace an RPi with a similar device (OrangePi or what not).

What I hear a _lot_ is that the thing could have been an esp32 (or Arduino rarely).

That’s a way wider price gap. I can order Costco-sized lots of esp32’s for the price of a single RPi.

RPis pricing has really, really narrowed the space where their products make sense. Low power devices can be esp32, high power can be x86 NUC things (or interconnected esp32s if you need tons of pins).

I don’t encounter a ton of things in “too big for an esp32 but I’m positive I don’t even want the option of a beefier x86 CPU”.

No hate if it works for you. I don’t even dislike RPi, they’re just in a narrower band for me these days.

everforward··on Strands Harness
I use Pi and mostly open weight models. I pay for the $20/month Ollama plan and use Deepseek and GLM through that. I’ve never hit the limits on it, but I tend to ask for targeted things rather than “implement a whole feature in one prompt”.

I do keep an OpenRouter account topped up for things that Ollama doesn’t have. 99% of my usage there is embeddings, the other 1% is wanting to test some new model Ollama doesn’t have.

everforward··on Claude Code reads AGENTS.md only when telemetry is on [fixed]
I did this at one point with Jinja templates.

I wrote an agent launcher sort of bash script. Pass in the command to start the agent, the script checks if there’s a Jinja file in a special directory matching that name, and builds it to AGENTS.md. Then it launches the agent.

I was trying to use it as a sort of janky RAG. I had a bunch of snippets (one for DB architecture, one for how load balancing works, etc), and my Jinja files were mostly a list of snippets to pull in. Voila, a bunch of agents that share little pieces of info but have a single source of truth.

I never got a ton of value tbh, it was very good at just grepping the snippets.

everforward··on Apple has added persistent 'ads' to iOS, and it's driving users crazy
Search on the app store sucks too. I haven't been able to find either a way to search for only apps that charge for downloads (I don't want to watch an ad for features, or nickel and dimed on IAP), or some kind of "minimum price" slider I can use.

That feels like a bog-standard search feature on a marketplace. I can only assume the omission is intentional.

everforward··on New bill to protect American citizens access to AI – Please read and share
I’m not a lawyer either, but this feels a bit like a shotgun blast of reasoning for the bill that will surely confuse the courts when they have to interpret it.

Eg I haven’t finished reading it, but it cites Heller on the first page and enshrines that AI is/can be used for self defense. That feels like a huge can of worms for the court, because now every 2A case in our history is relevant to any AI bill.

I also don’t know that it will accomplish what the author intends. Even if we hypothetically establish that AI use is covered by the first and second amendments, the government can still regulate it if they can pass strict scrutiny, ie if they have a worrisome enough complaint.

What the author wants, at the level of certainty they want it, would require a constitutional amendment and dismantling the strict scrutiny system. No idea what the impact of that would be, but my wild guess is “probably bad”.

everforward··on Show HN: Drop – a rootless Linux sandbox with gVisor support
Small world, we’re all working on the same thing!

I went with Docker because history has taught me that new isolation strategies _will_ have escape bugs at some point, and I’m distrustful that the LLM can’t find one if it wants to.

> Probably the biggest (and least hardened) feature that I spent time on in mine was trying to figure out how to allow arbitrary GUI apps so that I could run agents in it via Zed.

I actually have code for this if you want to use/fork/borrow it. TLDR, mine pretends to be an ACP agent so you run Zed on the host, but under the hood that binary is just creating a Docker container with your image and agent, copying files, etc, and then proxying ACP messages via web socket back and forth. Except for the built-in ACP read/write file and shell endpoints. Those get executed inside the container by the proxy by default, though there’s a config option to pass either or both through to the host.

Repo is https://github.com/SethCurry/abyss and the code you’d want would be in ‘internal/websockets/wsacp’.

It does wrap them in protobuf and there’s some router-like stuff so you can add your own non-ACP messages between the two ends of the proxy.

main also has experimental wasm plugin support in the proxies so you can block prompts/tool calls/whatever in an agent-independent way, or add RAG that works for every agent in the world, or whatever.

everforward··on MCP was always a bad idea?
I've taken to sandboxing my entire agent in a Docker container. I wrote a tool that pretends to be an ACP client but is actually making Docker containers, copying files I specified in, bind-mounting, etc, and then proxying ACP via websocket to an agent in the container (except the ACP terminal/FS commands, those happen in the container).

It works well, though there is some leakiness around paths. I opted to make it place/mount files at the same path as on the host so paths are the same (as opposed to manipulating the ACP messages to modify paths on the fly, that felt messy and buggy).

Configurable networking is on my list for the future, but I haven't decided whether to start with IP-level firewalls or if it's better to start with a proxy and firewall rules to force traffic to it. IP firewalls suck for APIs that might have semi-dynamic IPs.

[1] https://github.com/SethCurry/abyss

everforward··on Porsche puts wireless EV charging into production
> A car capable of driving itself onto this pad could as easily drive itself to a spot to plug in. Extendo-arm from the car could easily slot into an outlet in the floor.

Those two things really aren't comparably difficult. Eg in the dumbest case for pads, you could make the pads slightly smaller than axle diameter (to prevent dead shorts) and have the car randomly jostle around the space until it detects current flow. I've seen some amateur systems that work vaguely like this. Some signal that charging is nearby, and then a semi-random search to find exposed contacts.

That's not going to cut it on a recessed outlet. Your odds of hitting it randomly are tiny, and it will probably the damage the port with enough repetitions of ramming it into the floor.

> Planes refueling mid flight is hardly an apt metaphor given the charging station isn’t moving at all and the car can move as slowly as needed.

Change your reference point to the plane carrying the fuel and they get an awful lot closer. The planes are moving a lot relative to the ground, they're moving very little relative to each other. I'm not sure if you've watched them doing it, but relative to each other it's actually a very slow process.

everforward··on Porsche puts wireless EV charging into production
Robo vacuum charging is almost inherently safe. They probably charge at something like 12V/2A. Voltage too low to penetrate skin, wattage too low to cause serious injury unless you’re really trying.

You can kind of just leave them exposed and live.

None of that applies to a 220V, 50A (or whatever amperage) charging. That can absolutely kill somebody; jet fuel may not melt steel beams, but 11,000W sure will.

You’d end up needing some kind of recess probably, and then getting cars to connect to it like planes refueling mid-flight is a nightmare. You could try some sort of USB-C like signaling before power is sent, but people are never going to enjoy “exposed 220V contacts that may or may not be live depending on if things malfunctioned”.

everforward··on You can run Git on object storage if you re-make packfiles
Doesn’t Git store data in sparse files? My impression is that git’s on-disk layout is exactly the kind of thing you’d avoid on object storage. Interactions involve a ton of small reads, resulting in poor performance and enormous bills.

It only really makes sense to tarball repos into cold storage on S3 or Glacier, but short of GitLab cloud I suspect there aren’t enough repos cold enough to be worth the dev costs.

S3 that doesn’t have wild costs for git objects sounds like it’s just NFS

everforward··on What Zig felt like, coming from Rust
I don’t think this delineation is that clear, unless by “application” you mean the app tier of a 3 tier app.

Postgres and Nginx make sense in system programming languages; they’re extremely performance sensitive and that granular level of control offers them features. Interpreters are sort of the same, they interact with the OS a ton, it makes sense to work in the same language as the OS.

I do generally agree for the app tier of a web app. I wouldn’t build a CMS in Rust, but I also wouldn’t build a reverse proxy in Python.

everforward··on Wax motor
Some of them are a sort of hydraulic one that uses a spring and some fluid to make the spring release slowly. No idea which you would have lol.

The slow closing prevents water hammer, at least in water systems. In air systems it might just prevent a blast of back pressure blowing all the dust out of your vents onto your carpet.

everforward··on Replacing Pull Requests with Delta
I _think_ they fixed the file sync issue; I haven’t had desynced files in a few weeks.

Still a bad bug that lingered entirely too long.

> when a file is changed asynchronously, which for me, most happens due to a pull or LLM edits!

If you’re using Pi (I am), that’s partially because Pi doesn’t have “real” ACP and violates the spec as a result. File read/writes are supposed to happen via ACP RPC calls from the agent to the client, so Zed would see them and know to refresh. Pi wasn’t built for ACP so it embeds its own read/write tools and ignores the ACP ones, and Zed never gets the “refresh this file” signal it expects.

everforward··on An empirical study of harness design for coding agents
I think it’s sort of self-defined. If a model is able to use bash well enough to not need specific tools.

The research seems to agree with you, though. The paper calls out that for “bash capable” models, adding tools to do things bash can already do doesn’t improve performance.

Vaguely the same result as RAG. Unless you’re in specific domains, you won’t beat handing the agent a shell and grep.

Page 1 of 34Next →