HNHacker News
TopNewBestAskShowJobs

throw10920

4,183 karma · joined October 15, 2021

put your email in your HN profile!

quiet.can8525@fastmail.com

Marketers: you do NOT have permission to add this email to any databases or use for any advertising and solicitation whatsoever - personal correspondence only

submissionscomments
throw10920··on Show HN: I spent 2 years designing a mechanical Magic Keyboard
> To everyone who is about to pile on regarding the lack of a fingerprint sensor

This is a cheap way to try to subvert legitimate criticism by trying to shame others into not voicing their criticisms, and is exactly the opposite of what HN is intended for.

throw10920··on Pi's Minimalism Is Its Advantage
> This is getting silly, I can’t tell whether you are kidding, never launched a Node-based app, or never checked htop.

You're very much operating based on emotion rather than reason. That's fine in general, but it's not OK for HN, which is not a place to blast out your emotions and try to assert that your opinion is factually true by getting maximally angry about it.

> That indeed seems like a simple and elegant setup. I can’t recall a single instance of using a harness when I didn’t want the tool calls to be local, though. Why would I be SSHing into a SBC just to run the agent back on my laptop?

You don't understand how harnesses work. Let me explain: in a harness, the tools that the harness provides are the only way for the agent to interact with its environment, and as such, if the harness is run on one host, and the tools on another, the agent's perspective is that it is running on the second host. This is effectively true modulo some small set of scenarios that are mostly irrelevant in the case of "I'm SSHing into an SBC" (e.g. the harness logic and LLM API calls would both be happening on the SBC, which you definitely don't want when you're resource-constrained).

> Sounds even more elegant! Presumably, they communicate via HTTP?

No. Why would you think that? Again, I don't think you understand how harnesses work.

> I don’t need to, but I can. That’s the entire point, in case you’re still missing it.

That's a useless point, if so, akin to saying "I want to have Vim running at address 0x00000001beef0000 in physical memory" - you're conflating use-case (needing a large number of agents running in parallel so you can work on things in parallel) with implementation detail (having a large number of separate processes running on your computer).

> How beautifully ironic.

I'm meeting every single one of your technical arguments and factual inaccuracies (e.g. your misunderstanding of the word "contrarian"), and you're continuing to get very angry at Javascript. I don't think there's any irony here.

throw10920··on Pi's Minimalism Is Its Advantage
> For the record, I measure acceptable latency in milliseconds. I think it has something to do with my brain.

That's a personal preference thing. You're not really being inconvenienced.

> They’re very clearly stating that they’re expressing their opinion because it is underrepresented.

That's literally the dictionary definition of being contrarian:

"A person who likes or tends to express a contradicting viewpoint, especially from one held by a majority of people, usually because of nonconformity or spite."

https://en.wiktionary.org/wiki/contrarian

https://dictionary.cambridge.org/dictionary/english/contrari...

https://www.merriam-webster.com/dictionary/contrarian

I suggest consulting the dictionary before using terms you're not familiar with.

> No.

If you're going to copy-paste random pieces of text without making a point, you probably shouldn't be on HN. HN is for intellectual curiosity and debate, not Reddit-esque expressions of randomness conflated wuth humor.

throw10920··on Pi's Minimalism Is Its Advantage
The claims being made here are fallacious in three extremely basic ways.

First, "I regularly have 10+ instances of my harness running or idling on my laptop" implicitly asserts that you need that many instances running at once.

You don't. You can have a single instance running that handles your sessions in parallel.

Second, "I sometimes run my harness on machines that are low-power, low-RAM single-board computers." presumes the whole agent has to run on the SBC itself. It doesn't. There are agents (including Pi, which makes this even sillier) that can run their tool calls in a separate process/host, meaning that the agent can live on your development machine while the tiny tool-call daemon runs on the SBC.

Finally, "If my harness was written in JS, this would be both annoying in terms of responsiveness and limiting in terms of how much work I can do in parallel." could only be written out of a complete lack of understanding about how harnesses are actually implemented.

Agent harnesses don't do very much. The limiting factor is not the harness itself, but in the filesystem (reads, writes, file searches, greps), the computational tool calls (compilers, linters, evaluating the program), and the LLM API itself.

Your arguments against Javascript appear to be based on fallacies, lack of understanding of technology, and strong emotions against JS rather than actual fact.

throw10920··on Pi's Minimalism Is Its Advantage
>> Being contrarian for the sake of it is not good form.

> Why would you say that?

Because it's true.

> Is it good form to post that someone is being contrarian for the sake of it?

It's orthogonal to good form - it's preserving the value of HN, which is meant for intellectual curiosity and not knee-jerk emotional reactions, like your comment.

> OP’s comment looks like constructive criticism with plenty of good points.

You must have missed the very first line in their comment where they said:

> Lots of praise for Pi in this thread, so I'll offer up a diverging opinion.

...which is literally admitting to being reflexively contrarian.

Additionally, their criticisms were extremely shallow and/or misleading, and I addressed every one.

> So what?

Nobody is inconvenienced by a 3.5s startup time of an application meant to run for hours or days. That's a trivial complaint akin to arguing about the color of paint on a bikeshed.

> Don’t be silly, JS is a joke. I was going to check Pi out, but now that I know it’s based on a shitty stack, I’m staying the fuck away from it. I don’t have the RAM to spare.

You're just operating based on emotion and not a state of reason or intellectual curiosity. And I highly doubt that you "were going to check Pi out" - this wording follows the pattern of fabricating falsehoods out rage due to an opinion expressed that you didn't like.

>> If these are your complaints, then this is one of the strongest endorsements of Pi that I've ever seen. I think I'm bookmarking this comment.

>> good form

Did you forget to make a point here?

throw10920··on Pi's Minimalism Is Its Advantage
This is factually false.

A contrarian is "A person who likes or tends to express a contradicting viewpoint, especially from one held by a majority of people, usually because of nonconformity or spite."

The post I replied to is literally being contrarian:

> Lots of praise for Pi in this thread, so I'll offer up a diverging opinion.

Meanwhile, my post was not being contrarian in the slightest. I didn't offer a contradictory viewpoint - I provided factual information and context to dispel false claims/misinformation in the post I was responding to (e.g. the claim "the standard C-p and C-n bindings don't work" is highly misleading/false because you can trivially rebind them), and the opinions I provided were completely orthogonal to the opinion expressed ("I suspect that if the harness was written in Lua [...] you would have far more issues with it."

I expressed no opinion either for or against "Given all the hype, I was a bit underwhelmed by Pi. It definitely has some good ideas around customization, but it annoyed me in many little ways."

It always surprises me when someone so confidently makes an objectively incorrect statement when they could have spent 30 seconds looking up the definition of a word and realized that they were in the wrong.

https://en.wiktionary.org/wiki/contrarian

https://dictionary.cambridge.org/dictionary/english/contrari...

https://www.merriam-webster.com/dictionary/contrarian

throw10920··on Goodhart's Law Comes for Every Benchmark You Trust
The obvious solution is to have non-public benchmarks.

It is exceedingly difficult to train on a proprietary benchmark administered by someone with half a brain (i.e. don't sign up for a ChatGPT account with your benchmark@artificialanalysis.ai email) - you have to find a tiny needle in a vast haystack.

In fact, it can be difficult enough that it's simply not economically viable - that is, that it's cheaper to make the model better than it is to try to find the account running the benchmark.

In the limit case, the benchmark is indistinguishable from...normal problems that need to be solved.

throw10920··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
This is such a ridiculous and shallow cop-out.

And it also is completely irrelevant to my challenge to show how proprietary benchmarking can be gamed, because it presumes (absolutely insane and divorced from reality) circumstances that have nothing to do with benchmarking as a concept or process.

throw10920··on Pi's Minimalism Is Its Advantage
> Lots of praise for Pi in this thread, so I'll offer up a diverging opinion.

Being reflexively contarian is not what HN is for.

https://news.ycombinator.com/item?id=45530593

> For a program that's minimal it sure takes a long time to start up

It took the same amount of time (3 seconds) as cursor-agent and codex on my machine, which isn't particularly high-spec.

> the standard C-p and C-n bindings don't work

It has programmable keybindings and you can ask the agent to remap them in five minutes, if not 30 seconds.

> it doesn't follow the XDG Base Directory Specification and just pollutes my $HOME directory

You can make it put .pi/agent anywhere with $PI_CODING_AGENT_DIR in 10 seconds.

> scriptable in a simple (aka non-JS) scripting language

This doesn't make any sense. The very point of Pi is that instead of implementing all of your features directly in the harness, you make the harness minimal and then implement the functionality you want as extensions, which necessarily means that you have some powerful and expressive extension language, ideally the one the harness was written in.

I suspect that if the harness was written in Lua (which I love and is probably the least bad choice of "scripting" language), you would have far more issues with it.

If these are your complaints, then this is one of the strongest endorsements of Pi that I've ever seen. I think I'm bookmarking this comment.

throw10920··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Please imagine that the benchmark runners are not making mistakes of the fourth grade level - which they won't be. If you assume this level of incompetence, then literally everything is possible.
throw10920··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Artificial Analysis was an clearly meant to be an example. I was obviously talking about the ideal scenario of proprietary benchmarking, not how it might be being screwed up in practice.
throw10920··on Pi's Minimalism Is Its Advantage
> Are subagents basic? I found them only useful in very few situations.

I've found them to be extraordinarily helpful, because they allow me to much more carefully control context and reduce token spend by using a smart model for the parent agent and cheap models for the subagents. Do you just have a big token budget?

throw10920··on DeepSeek V4 Flash on a Single AMD MI300X
> e.g. sentiment analysis when the user begins cursing at the agent, or checking whether the user continued another session with the generated code, or started a new session with the same starting point as before, i.e. they git-stashed.

Thank you for elaborating, that's already useful. Anywhere I can learn more about this? I'm very interested in it!

throw10920··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
> Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs).

Please explain how, if I'm OpenAI and I'm making ChatGPT 5.7, and I release it, and Artificial Analysis goes off and runs one of their proprietary benchmarks on it from a random account, how I can optimize for that benchmark.

throw10920··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
> Building a new benchmark won’t solve the problem, either.

It will if the benchmark is proprietary. If you can't train on it, then it's extremely difficult to game, and if it's hard enough, then it's economically more efficient to just...make the model smarter

throw10920··on Show HN: cMCP, deny an AI agent's tool call and get a signed receipt
It's AI-generated, which is against the guidelines:

> Don't post generated text or AI-edited text. HN is for conversation between humans.

throw10920··on DeepSeek V4 Flash on a Single AMD MI300X
> should not discount that DeepSeek also gets paid in data, which is probably more valuable to them

That's agentic feedback loops for training, right? Any more detail on this, such as how they actually tell whether that data is good or not? That seems like a very hard problem, and like the value of that data is low compared to just building their own, controlled RL gyms.

throw10920··on Running Kimi K3 on MI355X at Better Performance per Dollar Than B300
This doesn't have anything to do with Google or OpenAI. I'm not Ilya Sutskever and it's not 2017, either.

I'm not "whining" about anything. You made the claim "The Chinese have actually been very open about training" and I showed that that was false. That's all that there is to it.

throw10920··on Running Kimi K3 on MI355X at Better Performance per Dollar Than B300
> You can replicate the architectural innovations, and try them for yourself with your own dataset.

That's not related to my comment. My comment was pointing out that you can't verify something that wasn't published. You have no idea what fraction of their techniques they're not publishing, and how much they contribute to their model performance, because you cannot replicate the models, because they don't publish their training data.

> and the reason the Chinese are not sharing data are no more nefarious than why the American companies are not sharing

This is moving the goalposts. Your claim was that "The Chinese have actually been very open about training", which is false, as discussed. Nobody ever claimed that the American labs were open.

throw10920··on Kimi Work
That's generally true, but consider this:

If the consolidated (stolen) IP of millions of books remains under US LLM company control, then at the very least there's a chance of justice to be served to the creators in the future - a large lawsuit victory and change to the legal system that results in the ill-gotten gains being transferred to the original creators.

If those proprietary models are distilled by the PRC, then that will never happen, because the PRC does not care about any sort of law (including their own) and will simply never return a single cent to the original US authors.

throw10920··on Running Kimi K3 on MI355X at Better Performance per Dollar Than B300
> The Chinese models are mostly very well documented in terms of architecture and training processes/flows, with what is missing to recreate them being the training data.

...and because that training data is missing, they can't be replicated. Which means that you cannot assert that the Chinese are being open in their LLM development, because there's no way to verify that the techniques they describe are actually the ones being used.

The reason that the training data is missing is that they're trained on a large amount of American copyrighted data and distilled on American models, which is where a lot of their performance comes from.

throw10920··on Kimi K3-256k
The Reddit comments are both older.
throw10920··on Handbook.md shows that long policy documents do not reliably govern agents
> In-context, your reply reads like you're asserting that anthropic's offerings give you all the same control as a local model

It absolutely does not.

throw10920··on Handbook.md shows that long policy documents do not reliably govern agents
> "Almost like magic" is cliché hyperbole. It also means "not magic" in the same way that "almost like a dog" means "not a dog." It was immediately followed by the explicit clarification that was surgically omitted above, in absolute and conscious bad faith.

Incorrect. The omitted text doesn't change the relevant meaning of the quotation in the slightest. Why are you taking such offense?

> Want it to go away, almost like magic? Local inference. When its under your control, and no longer being forced to hold it wrong, all of the common LLM defects will go away.

> "Want it to go away, almost like magic? Local inference. [...] all of the common LLM defects will go away.

There's no difference between these two statements in the relevant context of the article, which is specifically "long policy documents do not reliably govern agents", or the comments above, which are specifically addressing how that issue is not solved by local agents.

throw10920··on Handbook.md shows that long policy documents do not reliably govern agents
I'd rather just use my actual human brain to compute the answer at that point. I don't see the value at throughput that is this low.
throw10920··on Superlogical
> Sessions can be accessed through the web and native macOS/iOS applications

...only native macOS/iOS applications. Not a good look.

throw10920··on Discovering Cryptographic Weaknesses with Claude
That might be true, but it's irrelevant to the current discussion, which is arguing about whether or not Chinese models can do this now. One of the commentators above is arguing that it's not only possible but significantly cheaper (i.e. Chinese models are not only at par with frontier but several months ahead).
throw10920··on Netflix employee fired for sharing personal details in retreat trust exercise
Huh. Reading your comment and the parent above - maybe retreats aren't "therapeutic", but "exacerbatory" - as opposed to making interpersonal relationships uniformly better (bad -> good and good -> better), they make good relationships better but bad relationships worse?
throw10920··on Nvidia, Microsoft, Meta warn against overregulating open-weight models
I've had repeated conversations with Opus about cybersecurity and never gotten a refusal.
throw10920··on Private healthcare makes industries less innovative. It's time for change
> Capitalism created the computers whose Soviet copies were Tetris' native platform, it made the languages Pajitnov used and the machines he ported it to. It incentivized people to bring Tetris to the rest of the world, and who knows, it might have spread on its own eventually, just for free, but maybe not.

Yeah, GP's claim here was kind of insane. It's trivial to go look up lists of the 100 biggest innovations of the past 100 years (or whatever) and 90-99% of them were capitalism.

← PreviousPage 3 of 34Next →