HNHacker News
TopNewBestAskShowJobs

Systemerror7A69

203 karma · joined June 30, 2026

submissionscomments
Systemerror7A69··on Owners mourn spoiled food after firmware update bricks Samsung smart fridges
Not the person you asked but I can answer that from my perspective:

I am always on the lookout to automatate as much boring, everyday tasks. A Smart Fridge like that would allow you to generate shopping lists and reduce food waste by alerting you to possible soon-to-be-spoiled food.

A bit like what roombas are doing for cleaning. You still need to do it yourself a bit, but they help.

Not necessary, not even a gamechanger but another neat home automation. Would be nice, sad to hear they don't work well.

Systemerror7A69··on AI Is Antithetical to Learning
I'm not sure I fully agree with the article either, but I think that comparison is kinda nonsensical, over generalized and just handwaves a bunch of points and arguments about the article.

I don't think this brings your point across at all. If you disagree with the article, just saying "no, it's a tool, like coffee" does not meaningfully engage at all with any of the points made.

Systemerror7A69··on How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip
This is just speculation on my part, but LLMs work best when they get immediate, verifiable feedback on their task, and the kind of physical optimizations they mean might not give that to LLMs.
Systemerror7A69··on An empirical study of harness design for coding agents
I'd really love to see more studies about effectiveness of AI in general. As in, what works best and how to use it and such.

Because I feel that the technology and space is - so - hyped and fast moving that a lot of cultish feeling rituals seem to pop up, none of which are backed by evidence. Anthropic openly recommend giving the agents.md file an architectural overview of the code, and the one time this was studied they found the opposite - that the agents.md file is best for concrete commands about how to build stuff and such, and - not - huge overviews. This was, and still is, the official recommendation from Anthropic as far as I can tell.

And then there are the benchmarks, how feel vague and not concrete, and everyone kind of knows they're not the best cuz you can't just assign these tools one fixed number ( for multiple reasons ), but everyone still looks at them and compares them.

People share skills and superpowers and plugins and mcps and very, very few of them have and kind of proof they do much at all.

It all feels a bit weird to me, and I've been on the lookout for exactly these kinds of studies more lately, because I think having this research, even if not done on the exact newest models or not the exact, newest thing, are still - vastly - superior to the alternative.

Systemerror7A69··on An empirical study of harness design for coding agents
I have to say, I am starting to hate this line of reasoning. Yes, LLMs move extremely fast and a lot of improvements are done in a short amount of time.

And there might be a point to these arguments, vaguely. However:

There never seems to be - any - kind of counter example or reasoning behind the rationale. You have an in depth and empirical study, done by researchers who, frankly, now their shit (most of the time)

And on the other hand a random internet comment saying "nope" because...the models aren't the latest.

If the latest models really would make a difference, you should at least provide some kind of evidence towards that. As it stands though, every time these comments come up this is missing.

There seems to just be a vaguely defined understanding that "everything changes all the time, and nothing you ever research is transferable to state-of-the-art models"

Which brings me to my second point about these kinds of arguments:

LLM models often - aren't - fundamentally different. Yes, they are vastly more capable. And yes, there are emergent properties. But at their core, they function very much similarly. And for quite a while now, there have not been any of these drastic changes we saw when LLMs first become "good enough" for agentic coding.

I am tired of dismissing empirical evidence and studies every. single. time for reasons without evidence and seemingly a vague sense of "no, but my model is different"

Systemerror7A69··on Shapelearn Qwen 3.8 27B (13.1 GB VRAM)
AMD 7900 XTX with Vulkan here as well, wasn't faster on my test either. Might be much different on Nvidia though.

I assume the limit for me is memory bandwith, as the 7900 XTX has the same bandwith as the 3090 from what I can gather and I already reached ~60 t/s with Unsloth. Those would fit with the numbers Byteshape has for their cards.

4090 and 5090 have much higher bandwith apparently, so on those cards you can probably get much more out of the kinds of performance improvements they are doing.

Systemerror7A69··on Qwen 3.8 Omni Flash
My approach to this problem is to just...not try them all.

As long as the model you're using solves the problems you have to your satisfaction, there is no need to try any other models, except for financial reasons maybe.

So I start with a relatively cheap model (GLM 5.3 flash for me) and as long as it accomplishes the task (it did so far) I don't have to change. And even if it can't do something, the first thing I change is see if I can give it more tools or better context (useful even if I switch models later) or trying a different approach to the problem.

If google/gemma-4-31b works, you don't need to overthink it.

Systemerror7A69··on Our framework for reporting model misalignment
So, alignment does need to be taking seriously, you're right.

But keep in mind this is a report from OpenAI about OpenAI, who have a financial incentive to present this in a certain light. Take these things with a grain of salt.

This does not mean that models are now self-aware.

Systemerror7A69··on Nine coding harnesses vs. your laptop
It's absolutely amazing to see this harness. I've been on the lookout for something like this for a while now. I've used both Pi and Maki in the past but was unhappy with certain aspects for both. Pi is not respecting XDG and the author refuses to change and maki you curl an install script into bash.

So the points about it being a well behaved unix tool, installing it via brew and it not being react are points I - love - to see.

Thank you for making it, I will definitely try it out.

Systemerror7A69··on Mistral raises €3B
And my wallet is thought to contain a billion dollars, so long as I don't open it.

Seriously "thought to be" is such a baseless statement. Thought to be by whom? And on what basis?

Systemerror7A69··on “Next-token predictor” is the wrong mental model for LLMs
To be honest, I believe I get the point the article is trying to make, and to an extent I agree, but I also think the point is not really made very well.

The core of the argument as I understood it is that LLMs aren't just using existing data is training but also new ones. That's fine and good, and you can't simply assume an LLM is simply mashing together all it's data to give you an average of all that got fed into it - but at least I would still call it a "next token predictor"

It's not using just training data, but what it's doing is predicting the next token to get to the solution. As far as my amateur knowledge goes, LLMs still roughly go token by token, deciding which one fits best given the context.

It's just not predicting based on it's training data, but predicting based on RLVR & more, trying to get to the optimal solution ( as much as the solutions CAN be optimal)

And I honestly think keeping this very much in mind is helpful in understanding and dealing with LLMs.

Systemerror7A69··on AI Can Make You Suck Faster Too
Even as someone using AI on the regular I'm starting to hate the "You didn't actually use this exact most expensive model so your point is invalid" argument.

This is fair to say if someones last experience with AI was copy-pasting code into GPT3 chat windows years ago, but Deepseek is a more than capabale model and enough for someone to get an informed opinion about the technology.

If people have actual counter argument, use those. And if some of those counter argument are "What you say isn't possible, the neweste model can do and here are examples of that", that is fine.

But a blanket "Nuh-uh, it wasn't Model X" is not only a poor argument but also automatically invalidates any criticism when a new, better model comes out - and that can't be the basis of a good argument.

Systemerror7A69··on GLM-5.3-Flash
I think the assumption here that might not hold is simply that increases in efficiency and smaller size will be achieved by linearly just training smaller models better.

You are absolutely right that there is a physical limit about these things, but very often I find that the solution is a clever way to work around the problem. Maybe the problem with knowledge of the models will be improved by them looking the information up in a better way - so smaller models will not have to have the knowledge trained in but will default to checking. Maybe Models will, I dunno, focus on training in assembler and start to only ever check the compiled output so they only ever need to learn assembler and will then compile the solution to reason about the assembler code.

Obviously that last part is a ridiculous example because I'm not gonna be able to come up with a solution myself - I'm not nearly smart enough for that. But I h ope you get what I mean. Not going the direct route but instead finding solutions people didn't think of before.

Systemerror7A69··on Headlong: A microharness for persistent agents
One of the big problems is that no one actually does read these scripts. You could say "Oh but it's their own fault, duh" but theres a very legitimate argument to be made users going the path of least resistance and that you shouldn't offload this responsibility on your users.

Regarding appImage or rpm, attackers need to build and package these to inject these, while this curl | bash pipe opens up the possiblity of payloads simply by taking over the domain. And this isn't really that far fetched, just think about the Notepad++ update payload recently. The regular package was unaffected while the domain used for the update was taken over.

Then theres also the argument about normalization. Just like he said, this isn't just something he said, this is a very real argument. You don't want to teach users bad habits. Even if / you / inspect the code you get, not everyone will. And ultimately, we should strive to make the Internet a safer place, if only to get less botnets.

Systemerror7A69··on Feature Request: Support AGENTS.md
Yes, but since we are specifically talking about claude.md, Anthropic themselves claim on their website these files "serve[s] multiple purposes: providing architectural context, ..." (https://claude.com/blog/using-claude-md-files)

They also recommend starting with an /init command, which is also something the study very specifically called out as "having a marginal negative effect"

And from personal experience, I can only confirm that many people seem to see this as the main purpose of agents/claude md files - a persistent architectural overview of your project.

Quick Edit: My point simply being that I think it's understandable if some people don't understand the big deal about these files because they've had a drastically different experience than other people - anthropic themselves recommend apparently totally ineffective practices on their website, and the starter tool present in many harnesses seems to even have a (marginal) negative effect.

Systemerror7A69··on Feature Request: Support AGENTS.md
Theres a study from earlier this year which suggests the opposite: https://arxiv.org/abs/2602.11988

Most of the agents.md and what people use it for / write into is does, in fact, not make a difference.

Now, sure, this study is a bit old for LLM standards - as everything beyond the current month is - but

a) I haven't seen any tangible evidence to the contrary and

b) Since the basic inner workings of LLMs haven't changed I'd be sceptical of this not still applying.

I think one major side effect of LLMs moving so fast is that best practices and how to use this tool is very much not catching up as fast.

No one knows what is best and what actually makes a difference, doubly so because LLMs are / very / hard to quantify - even benchmarks themselves are very rough estimations.

People do, in fact, use stuff which makes no difference all the times.

Systemerror7A69··on Unsloth Dynamic 3.0 GGUFs
Since it seems like this not only improved sizes but also performance I can't wait for some benchmarks and comparisons. If you don't have a separate GPU for inference, every single GB matters so a comparison between specific Q4 Quants is really interesting to me.

Currently I very much can't decide between going for a bit of a lower Q4 Quant to squeeze out a bit of buffer and ctx or wondering if a slightly higher (IQ4_XS vs Q4_K_M/XL) is worth it

Systemerror7A69··on Pi coding agent: config folder is out of place on Linux
From what I understood in the ticket ( I haven't confirmed independently) it's the later. You can set an environment variable to move the .pi directory ( into your XDG Config dir) but it's both the cache and config combined.

Quick Edit: Since github is back up, heres the comment I was referencing: https://github.com/earendil-works/pi/issues/534#issuecomment...

Systemerror7A69··on Pi coding agent: config folder is out of place on Linux
Oh sorry - basically just a github ticket to an explanation about the .config folder. Pi doesn't respect XDG_BASE_DIR specification, and even if you use the env var it combines cache and config.

The ticket is closed and this will not be changed, period.

Systemerror7A69··on Pi coding agent: config folder is out of place on Linux
Sure but I'm really not that invested in one single harness. I tried out pi because people were recommending it so much. It turns out I personally have some things which annoy me, so I try out others now as well.

If I don't find anything which fits me I might fork it, but even that little effort is not really worth it if theres something which fits me better.

I already found maki which...seems to do the exact same job pi did for me, and I wanna check out crush as well.

Systemerror7A69··on Pi coding agent: config folder is out of place on Linux
This has been frustrating me for a while and is part of why I explore other coding agents.

As many advantages as pi has in some areas, there are definitely areas where I believe the hype to be a bit overstated. While the config folder is, ultimately, not relevant for performance, how this was and is continued to handled is a bit frustrating for me.

They have made it abundantly clear it's not going to change however so I'm looking at how other coding agents perform currently.

Systemerror7A69··on Choosing an AI model: one prompt, 11 models, different results
Am I wrong or are these evaluations, while interesting, not really meaningful for anyone doing serious development work?

I'm asking because I personally only use AI with specific and detailed instructions, building my projects piece-by-piece.

I mostly don't look at the low level code and some of it I don't understand as much as I'd like, but I very much give much more technical instructions than a simple, two sentence prompt.

So, it seems to me this "oneshot from a simple prompt" eval is fairly meaningless when it comes to model evaluation itself, as it is in no way representative of real world application.

This would be more something for "vibe coders", people with little to no programming background wanting a website?

Systemerror7A69··on The Human Is the Loop
I'm in a similar situation is you are, having ADHD, and my observations have been similar. I've been more hestiant to use AI at all for a while at the beginning but even now I only ever use it for one thing at a time. It helps me tremendously with the things I want to have done but not do - find out who clutters my home directory, build a small DMS plugin for local models, stuff like that. But even if I work on a bigger task and idea, it's mostly one thing I use it for at a time.

Maybe you're right in the fact that the coping mechanism help with that. Maybe using LLMs is similar in that it is easy to lose focus for neurotypicals so having lived with that we know how to handle it. Or maybe it's just us, who knows.

Systemerror7A69··on How I use LLMs to learn complex topics
It's also relying on the assumption that the checking LLM only ever corrects wrong statements and never incorrectly "corrects" an already correct statement, which might not always be the case as well.
Systemerror7A69··on AMD acquires Taalas to boost inference performance by etching models in silicon
It's not thinking. Not in the way she probably meant. It can "think" that fast the same way a calculator can "think" that fast (kind of).

Because it's not human and not "thinking", it's a mathematical algorithm

Systemerror7A69··on Qwen3.8 Max now ranked as the best overall model by agentic index
If they already require your constant supervision the reason is money.
Systemerror7A69··on Pi's Minimalism Is Its Advantage
I just wish it wouldn't ask me to pipe an install script directly to shell to install.

Yes, I can probably inspect that but I do think installing through package managers is the best practice.

It looks better than pi with XDG and not being JS but that is it's own red flag for me.

Systemerror7A69··on Pi's Minimalism Is Its Advantage
I've absolutely had the same experience. Pi has been praised a lot, people speaking so highly of it's code quality.

And then I installed it and found all the problems you mention. Not only that but I read Github Issues about the XDG problem and was a bit taken aback by the reaction of the developer.

It's one of the best agents I have used so far but I'm still looking for a very lightweight, token efficient agent NOT written in a JS framework and which respects XDG

Systemerror7A69··on $100 AI Music Video: Claude Fable 5 vs. GPT-5.6 Sol
I don't think that example applies at all here. The quote you quoted itself said it - "subjects for high art". Theres a difference between treating the banalities of life as SUBJECTS for your art and making human, non-mass produced art from that ( like the paining you linked) vs just treating the soup cans themselves as art.

At least that's my interpretation of it

Systemerror7A69··on Inkling: Our Open-Weights Model
No, Open weights US models would not break as well - this isn't related to China or USA, it's about Open Weights and the fact that you can download the models.
Page 1 of 2Next →