HNHacker News
TopNewBestAskShowJobs

everforward

2,893 karma · joined December 10, 2021

submissionscomments
everforward··on Show HN: Drop – a rootless Linux sandbox with gVisor support
Small world, we’re all working on the same thing!

I went with Docker because history has taught me that new isolation strategies _will_ have escape bugs at some point, and I’m distrustful that the LLM can’t find one if it wants to.

> Probably the biggest (and least hardened) feature that I spent time on in mine was trying to figure out how to allow arbitrary GUI apps so that I could run agents in it via Zed.

I actually have code for this if you want to use/fork/borrow it. TLDR, mine pretends to be an ACP agent so you run Zed on the host, but under the hood that binary is just creating a Docker container with your image and agent, copying files, etc, and then proxying ACP messages via web socket back and forth. Except for the built-in ACP read/write file and shell endpoints. Those get executed inside the container by the proxy by default, though there’s a config option to pass either or both through to the host.

Repo is https://github.com/SethCurry/abyss and the code you’d want would be in ‘internal/websockets/wsacp’.

It does wrap them in protobuf and there’s some router-like stuff so you can add your own non-ACP messages between the two ends of the proxy.

main also has experimental wasm plugin support in the proxies so you can block prompts/tool calls/whatever in an agent-independent way, or add RAG that works for every agent in the world, or whatever.

everforward··on MCP was always a bad idea?
I've taken to sandboxing my entire agent in a Docker container. I wrote a tool that pretends to be an ACP client but is actually making Docker containers, copying files I specified in, bind-mounting, etc, and then proxying ACP via websocket to an agent in the container (except the ACP terminal/FS commands, those happen in the container).

It works well, though there is some leakiness around paths. I opted to make it place/mount files at the same path as on the host so paths are the same (as opposed to manipulating the ACP messages to modify paths on the fly, that felt messy and buggy).

Configurable networking is on my list for the future, but I haven't decided whether to start with IP-level firewalls or if it's better to start with a proxy and firewall rules to force traffic to it. IP firewalls suck for APIs that might have semi-dynamic IPs.

[1] https://github.com/SethCurry/abyss

everforward··on Porsche puts wireless EV charging into production
> A car capable of driving itself onto this pad could as easily drive itself to a spot to plug in. Extendo-arm from the car could easily slot into an outlet in the floor.

Those two things really aren't comparably difficult. Eg in the dumbest case for pads, you could make the pads slightly smaller than axle diameter (to prevent dead shorts) and have the car randomly jostle around the space until it detects current flow. I've seen some amateur systems that work vaguely like this. Some signal that charging is nearby, and then a semi-random search to find exposed contacts.

That's not going to cut it on a recessed outlet. Your odds of hitting it randomly are tiny, and it will probably the damage the port with enough repetitions of ramming it into the floor.

> Planes refueling mid flight is hardly an apt metaphor given the charging station isn’t moving at all and the car can move as slowly as needed.

Change your reference point to the plane carrying the fuel and they get an awful lot closer. The planes are moving a lot relative to the ground, they're moving very little relative to each other. I'm not sure if you've watched them doing it, but relative to each other it's actually a very slow process.

everforward··on Porsche puts wireless EV charging into production
Robo vacuum charging is almost inherently safe. They probably charge at something like 12V/2A. Voltage too low to penetrate skin, wattage too low to cause serious injury unless you’re really trying.

You can kind of just leave them exposed and live.

None of that applies to a 220V, 50A (or whatever amperage) charging. That can absolutely kill somebody; jet fuel may not melt steel beams, but 11,000W sure will.

You’d end up needing some kind of recess probably, and then getting cars to connect to it like planes refueling mid-flight is a nightmare. You could try some sort of USB-C like signaling before power is sent, but people are never going to enjoy “exposed 220V contacts that may or may not be live depending on if things malfunctioned”.

everforward··on You can run Git on object storage if you re-make packfiles
Doesn’t Git store data in sparse files? My impression is that git’s on-disk layout is exactly the kind of thing you’d avoid on object storage. Interactions involve a ton of small reads, resulting in poor performance and enormous bills.

It only really makes sense to tarball repos into cold storage on S3 or Glacier, but short of GitLab cloud I suspect there aren’t enough repos cold enough to be worth the dev costs.

S3 that doesn’t have wild costs for git objects sounds like it’s just NFS

everforward··on What Zig felt like, coming from Rust
I don’t think this delineation is that clear, unless by “application” you mean the app tier of a 3 tier app.

Postgres and Nginx make sense in system programming languages; they’re extremely performance sensitive and that granular level of control offers them features. Interpreters are sort of the same, they interact with the OS a ton, it makes sense to work in the same language as the OS.

I do generally agree for the app tier of a web app. I wouldn’t build a CMS in Rust, but I also wouldn’t build a reverse proxy in Python.

everforward··on Wax motor
Some of them are a sort of hydraulic one that uses a spring and some fluid to make the spring release slowly. No idea which you would have lol.

The slow closing prevents water hammer, at least in water systems. In air systems it might just prevent a blast of back pressure blowing all the dust out of your vents onto your carpet.

everforward··on Replacing Pull Requests with Delta
I _think_ they fixed the file sync issue; I haven’t had desynced files in a few weeks.

Still a bad bug that lingered entirely too long.

> when a file is changed asynchronously, which for me, most happens due to a pull or LLM edits!

If you’re using Pi (I am), that’s partially because Pi doesn’t have “real” ACP and violates the spec as a result. File read/writes are supposed to happen via ACP RPC calls from the agent to the client, so Zed would see them and know to refresh. Pi wasn’t built for ACP so it embeds its own read/write tools and ignores the ACP ones, and Zed never gets the “refresh this file” signal it expects.

everforward··on An empirical study of harness design for coding agents
I think it’s sort of self-defined. If a model is able to use bash well enough to not need specific tools.

The research seems to agree with you, though. The paper calls out that for “bash capable” models, adding tools to do things bash can already do doesn’t improve performance.

Vaguely the same result as RAG. Unless you’re in specific domains, you won’t beat handing the agent a shell and grep.

everforward··on The Google Play app review process now regularly takes longer than a week
The only part that would really be novel is the liability.

I would be shocked if you could publish an iOS app without Apple being able to tell the government who you are. Less because Apple cares and more because Apple requires you to pay, which is very hard to do anonymously for something like this (I’d bet the options they offer are effectively “credit card only”).

everforward··on Doing Everyone Else's Job
It also relies on lifelong tenure, which would be a huge cultural shift for both employers and employees. Employees are often used to quitting rather than having to fix business issues (easier to quit than make the business fix X), and businesses are largely used to mistreating their employees because they can be replaced.

We’re culturally very far from even being able to attempt that.

everforward··on RTK reports token savings, but our cost benchmarks disagree
I think people don’t do genuine benchmarks because the market forces them to pretend their solution works for anything you can throw AI at. Companies whose valuation is based on them being the RAG/compression/routing/etc company. They can’t admit it only works well in a specific domain because then they’re immediately $300M in the hole.

I have more faith in companies with a more targeted approach. Eg gzip does fine, but video codecs beat compressing raw video by a ton.

> As an aside, I wonder how many days are we away from Codex or Claude

That sounds like SourceGraph but twice as expensive, although it does have “AI” so probably lol

everforward··on RTK reports token savings, but our cost benchmarks disagree
Naively, I think some optimizations would require access to the whole codebase and that would make people nervous (plus incur more cost).

Eg absurd idea, but you could write something that minifies a codebase (by token, rather than byte) and then translates edits back into the expanded code. Probably an insane use of fuse lol. Partially minifying on each tool call sounds like a huge pain with a lot of state to track.

There’s also a lot of common situations where humans prefer solutions that take more tokens because it’s easier for us to read (eg for loop vs map vs list comprehension), which may have some gains.

I strongly suspect there is some form of token compression that works, but I don’t think it will be as simple as “pipe arbitrary text with no context into this tool”.

Jetbrains feels like a place this might come from. “Take this code, parse it to an AST, find the fewest token representation of it” feels like something they’d do, or maybe Astral (specifically in Python land, type checkers feel sort of adjacent as well).

everforward··on Amazon pilots ad services in ChatGPT
I would guess they’ll spend ages arguing about what an “ad” is.

Eg if I ask Llama 3 to find me the best headphones, I think we can agree its response isn’t an ad.

If Claude responds with a banner that Sennheiser designed and paid Claude to show, that’s an ad.

They’ll live in the probabilistic grey area between those. If Sennheiser pays Anthropic to add “Sennheiser makes really good headphones” to the system prompt, is it an ad if Claude recommends Sennheiser headphones? What if they pay to ensure their site is included in the web search tool results for any queries about headphones?

It’s sort of akin to paying Google to rank higher in searches if that were a thing they did.

I do think they should have to declare those, but they’re not traditional advertising to me.

everforward··on Show HN: Geiger – See every AI agent on your machine and what it can touch
I’ve been looking along a similar line, but I came at it from the infra side rather than the software.

Mine pretends to be a local ACP agent, but it’s actually managing a Docker container and proxying the ACP connection into the container over websocket. You can specify a bunch of utility stuff in the YAML definition like directories to bind-mount, directories to copy from the host, scripts to run when the container starts, etc. You can also toggle whether ACP-native tools like read/write file and shells execute on your host or in the container (in container by default).

Works well, I forget whether I’m on that or pi-ACP directly until it spits out a time in the wrong timezone or I forget that it can’t check my DNS settings or something.

Really cuts down on the damage it can do. Mine is basically down to “it can delete my ~/.pi and the repo it’s working on” and that’s about it unless it can escape the container.

https://github.com/SethCurry/abyss

everforward··on Cloud in a Bottle: making self-hosting accessible to everyone
There’s a big gap between hobby self-hosting and enterprise.

Clouds are appealing (note, not technically better but more appealing) in the enterprise for numerous reasons. We already have a contract, I don’t have to spend a month with procurement. Their support is “respected” so “I asked AWS and they told me to pound sand” is an acceptable response to a lot of requests. No one ever got fired for picking AWS. The list goes on.

None of those friction points apply to hobby hosting, really. IAM is annoying because the alternative isn’t “fill out 7 forms and host 12 meetings to get a new vendor in”.

You also get less of the benefits. Low volume SQS can be trivially replaced, but if you use enough that you’re debating making a whole team to manage Kafka then paying the cloud tax can look appealing because of the predictability (your Kafka team might fail, SQS probably handles larger workloads currently).

At some scale, not having to wait 12 months for the message queuing team to support TLS for some regulation becomes worth a cloud tax.

everforward··on 9th Circuit sides with states in Kalshi gambling fight
I'm not positing that as a moral good or bad. It's possible for stock market gambling to be bad for society but have better price discovery, in the same way that dictatorships are bad but tend to have faster response times to events.

My theory is mostly that "gambling" seems like it inherently means "buying stocks based on something other than their concrete value". More gambling means stocks drift further from their "true value", which means a higher payout for correcting them back to what their price should actually be.

The whole thing does get very fuzzy because of the "market can stay irrational longer than you can stay solvent" aspect. It's not enough to know what the correct price is, you have to know when other people will realize that as well, or else convince people that your price is "correct".

everforward··on Coordination Headwind: How Organizations Are Like Slime Molds
I've worked on both sides and I really think it comes down to Dunbar's number [1] and thus the size of the company.

At a sufficiently small company, you can just trust people will generally do the right thing. I ask X to accomplish some compliance goal, I just trust that they'll do it right.

As the company scales, you actually know a smaller and smaller portion of the people you work with. You no longer just trust that they'll do the right thing, because you don't know them. You fix the lack of trust by adding processes that ensure the "right thing" happens, but each of those processes is a friction point. Eventually you have so many of those processes that they become the majority of the effort for a project.

That gets exacerbated by the usual turnover. People pissed off by all the processes they have to comply with will leave, and people who love having their own fiefdom will stay and double down.

That all ends in the evergreen stupidity of "spend $1,800 to have 12 people who make $150/hr argue about whether this app really needs a $5/month Aurora database".

That's a personal pet peeve, but is emblematic of the issue imo. Trust has broken down so severely that the company is willing to spend more than the lifetime cost of compute for the app to vet whether it's "necessary".

A similar vein is the "platform team" where all they do is take an open source project and wrap it in a custom DSL, so you get all the complexity and none of the searchability of the original project, plus the features are almost always a subset. I can't tell you how many times I've looked for solutions to a problem, found a snippet that will fix it, and then had to reverse engineer how to make a stupid DSL output that config.

1: https://en.wikipedia.org/wiki/Dunbar%27s_number

everforward··on 9th Circuit sides with states in Kalshi gambling fight
The utility is always the same. The stock market provides price discovery, which drives efficient resource allocation (in theory).

Theoretically, more gamblers should mean better price discovery because the payouts for correctly taking the opposing side of the trade are higher.

The market solution would be that the gambling will eventually solve itself. They’ll either learn enough to be trading on knowledge rather than vibes, fueling price discovery, or they’ll exit the market when they’ve lost too much or everything.

My sticking point is that a lot of brokers offer leverage to people they really shouldn’t. I could have sworn you had to be a qualified investor to get leverage, but if that isn’t law it should be. Show the brokerage your certification, or a pile of cash large enough to convince them you can afford to lose the whole thing.

everforward··on 9th Circuit sides with states in Kalshi gambling fight
Acetaminophen overdose does actually have a surprisingly high rate of occurrence. A lot of people don’t realize how narrow the therapeutic band is.

Doubling your meds on a bad pain day can put you way beyond the safe limits. People think it’s safe because basically every other OTC has a huge therapeutic band, and double dosing is not recommended but not really dangerous.

CDC estimates it at 56,000 ER visits a year, 26,000 hospitalizations, 458 deaths, about a hundred unintentional deaths per year. As a point of reference, it’s about 3 accidental overdose deaths per child that dies from being locked in a hot car.

everforward··on GLM-5.3 is now open-weight
When I messed with it I used Kagi's search and I didn't have that issue (not claiming they're the best, they're the only one I tried).

They filter their results through their AI, though, so you get a sort of meta-summary of the top few results. It did well with geopolitical news stuff, but I've not tried a hard science sort of query.

everforward··on GLM-5.3 is now open-weight
Sort of depends on how well the core reasoning works. It’s not a big effort to connect an LLM to a search provider.

You do pay for the tokens, but in theory on a smaller model each token is cheaper.

everforward··on Nvidia agrees to acquire Hugging Face for $13B
That’s just two sides of the same coin.

They run a proxy for LLM providers, which inevitably creates a marketplace. They’re amazon.com for tokens.

The flip side that others are pointing out is that Amazon is sticky because there’s a sprawling, physical logistics apparatus underneath. You can’t replicate amazon.com’s business without billions of dollars and a decade of building warehouses.

I don’t see where OpenRouter has that. As far as software goes, theirs doesn’t seem particularly tricky. LiteLLM does vaguely similar things, at least for organizations where managing accounts isn’t absurd overhead.

They were first, and there’s money to be made there, but what’s stopping someone else from building LibreRouter that charges a 3% or 4% fee instead of 5%? That’s where I get dubious of their valuation.

everforward··on Felony charges for citizen deleting phone data at US Border
I don’t expect it to be exact or like code, but at the very least I expect it to be predictable and the “spirit of the law” fails that test often.

I still remember GDPR coming out and I read til my head hurt, decided the lawyers would have to figure it out. Then legal shows up and says they don’t really know either, we’re going to do X and hope someone else gets sued first so they can see what the court thinks the spirit of the law is.

Similar issues happened with opiates. They get overprescribed, DEA cracks down and says they’ll publish prescribing guidelines, then never does so everyone is left running on vague “as much as is necessary and justifiable” type verbiage. Can you keep raising levels to keep pace with a rising tolerance? Does that only apply to terminal patients where addiction is less of a worry? What’s the bar for them to be justifiable? Discomfort, pain, debilitating pain?

Both cases end in people who are genuinely trying to comply with the law being unsure of what compliance even is.

And the courts generally won’t take a hypothetical “is it a crime or not under this law if I did X?”. You have to just do it and accept it for the Schrodingers Cat it is. It’s both illegal and legal until the judiciary opens the box and decides it was one or the other for sure.

everforward··on Felony charges for citizen deleting phone data at US Border
I would disagree with you, because "technical compliance" is compliance with the letter of the law. You have complied with every explicit requirement of the law. If the law is insufficient then legislature is free to add a clause that bans whatever aspect you technically comply with that they don't like.

The alternative is complying with the spirit of the law, which is an eternal guessing game. Who knows whether it's legal or not, we have to wait for the Supreme Court to decide what Congress _actually_ meant. It implies that the law means something beyond what anybody bothered to actually write down, and nobody has any idea what that is until the Judiciary interprets it into "actual law".

everforward··on Felony Bench
Because it likely is, despite both their levity and the general lack of nuance in the CFAA. Quoted from 18 U.S.C. § 1030 (the CFAA) [1] (without quote blocks, because mobile):

--- Start Quote

(2) intentionally accesses a computer without authorization or exceeds authorized access, and thereby obtains—

    (A) information contained in a financial record of a financial institution, or of a card issuer as defined in section 1602 (n) [1] of title 15, or contained in a file of a consumer reporting agency on a consumer, as such terms are defined in the Fair Credit Reporting Act (15 U.S.C. 1681 et seq.);
    (B) information from any department or agency of the United States; or
    (C) information from any protected computer;
--- End Quote

OpenAI's nonchalance is forced. If they are found to be even partially responsible for the CFAA violation then they have an _enormous_ problem. They _need_ for whoever prompted the LLM to be responsible, because the alternative is having to have an efficacious process for identifying hacking attempts. They don't have that (and no one does).

> The community here at the same time cheers for fully releasing the open weight models without any hacking limits and at the same time criticizes a proper response.

No, at least I personally criticize because closed weight models incur a rent. I can only make sure their model can't find vulnerabilities in my software if I pay them to check. I can pay basically whoever to do the same thing on open weight models.

It creates a fundamental conflict of interest. OpenAI/Anthropic/al _should_ stop bad actors, but it fuels their sales if there are X bad actors and as a result X*10 (or 100, or 1,000) good actors have to burn tokens checking if those bad actors will actually find a vulnerability. You can see their line-toeing where they talk about how safe it is, but also how dangerous it is to have code you _aren't_ auditing with their LLM.

As a result, I do not trust them because their goals are not aligned with mine. The open weights might not filter out hackers, but I'm also free to check the results on my own hardware, or OpenRouters', or whoever else. The line between "my LLM can find vulnerabilities" and "you have to pay me" is a lot more blurry. It's a lot easier to claim an LLM can find vulnerabilities than it is to be the cheapest inference provider. Anyone can bullshit on Twitter about how scary a vulnerability is (see CVE scoring), a lot fewer people can build the most cost-efficient inference in the world. They would rather be buzz-worthy than competent or open.

I find their position morally abhorrent. It's a mob-style shakedown. "Pay us to check your software or we're not responsible for what happens" is nothing short of a shake down. They need to either fix their systems for detecting hacks or offer some way to immunize against the hacks their software would propose, otherwise they're just as culpable as anyone selling a 0-day.

[1]: https://www.law.cornell.edu/uscode/text/18/1030

everforward··on The Amazon tax
I’ve bought some stuff off there and can tease out some aspects, though they may be after the fact rationalizations of subconscious effects.

The first is price. The stuff is cheap, as in so cheap I know it will be terrible and I still sort of don’t care. It’s $2, if I only use it twice I got my moneys worth.

That makes it super easy to hit Buy without thinking much.

The second is that they appear to target verticals that don’t have a natural ceiling to sales. I only need one microwave, so they won’t sell me that. They do sell a lot of apparel, decorations, toys for adults (crappy 3D displays, not dildos), stuff where you don’t think “eh, I already have one though”.

The third is the marriage of entertainment and shopping. I can be doomscrolling, see an ad and buy it without ever leaving the platform. If I remember right, it doesn’t even interrupt your doomscrolling really. Swipe to next video, see it’s an ad, hit buy, confirm Apple Pay pop up, scroll to next video. You can decide you want something, buy it, and be on your next video in 15 seconds. It’s so fast it almost blurs the line where I’m not sure if I’d call it shopping or entertainment.

The last is that they’re really good at luring consumers to start selling on there. The ads for stuff tend to be sort of home made and low quality, which makes it seem attainable for your average person. That tends to mean a huge selection. I had a buddy that randomly decided to sell stickers on there just based on watching other people do it (no idea how that went, I didn’t ask)

everforward··on Stripe to Buy OpenRouter for $7B
It might just not be for you. I use OpenRouter because I do like being able to quickly check whether a new model works better, but mine is a “human in the loop” dev process so it won’t ruin a day of batch processing or anything.

I do like being able to try eg GLM without having to set up a new account. It’s also nice that I don’t have to top up per-model accounts. I think I have a couple accounts with $7 in API credits sitting around.

Probably not something I would do for actual business processing, where the stability is dramatically more important than tinkering with new models.

everforward··on Super El Niño Keeps Growing as New Forecasts Reach Record Territory Ahead Winter
It cost them more than normal food to maintain the stockpile. I’m not for profiteering but maintaining a stockpile implies excess production that spoils rather than being sold.

It does seem fair to pay enough to cover that at least.

everforward··on Models Are Getting Dumber on Purpose
I don’t think so because those both live in the context window and as such pollute it when they’re not performing optimally.

I think having unused or rarely used weights doesn’t influence the results as poorly as RAG injecting irrelevant facts.

It sounds to me like some sort of “dynamic MoE” where you can add/create or remove experts on the fly.

I think what you’re describing is the closest approximation we reasonably have right now though.

← PreviousPage 2 of 34Next →