HNHacker News
TopNewBestAskShowJobs

parsimo2010

5,058 karma · joined July 24, 2019

submissionscomments
parsimo2010··on Nvidia wants to put a watchdog chip next to every AI agent
You can absolutely run an agent as a limited-privilege user that only has write privileges for specific files and only has execute privileges for certain files. If it is running as a limited-privilege user it can work on code in it's own copy of the repo and make commits and send pull requests, but it can't do the merge. The problem is that nobody wants to go through the effort to set up all these permissions and nobody wants to take the time to review everything and perform all the manual actions.
parsimo2010··on Nvidia wants to put a watchdog chip next to every AI agent
Agreed- this is the same problem we have with trusted admins or devs who have elevated privileges on their networks. We have to trust that the admins won't use their power to steal company secrets or misuse company resources. If you don't trust the admins, then they can't fix things on your network and there is no point in having them.

If you want an agent to act on its own, like pushing to a git repo, managing dependencies, building and testing, etc., then you have to trust it as much as any other privileged user.

If you don't want to trust it, then you're just forcing yourself into the reverse centaur role, where the agent edits some code, but then has to stop and ask you to push the changes or build the software again and run the unit tests.

I suppose there is a principled way of doing things like "I trust you do do basic commits but I will handle merge conflicts" and "you can build modules in this directory but you can't build outside of it" but this is just a lot of effort that most orgs won't bother with.

parsimo2010··on Can you tell which images are AI-generated?
I mean that is a conclusion that I agree with, but I’d also like to see if I can do better if given more time on a bigger screen.
parsimo2010··on Can you tell which images are AI-generated?
Not bad. I’d appreciate a much longer time limit. Got really rushed at the end. But this was crazy hard on my phone, I was straight guessing most of the time. I think I could be a little better with more time on a desktop screen.

Another idea would be games that were all pictures of a single category- like 10 AI and 10 real pictures of basketballs, or food on plates, etc.

parsimo2010··on Stop eating Lady Gaga's Oreos
It's basically a cinnamon flavoring added to regular oreos. They were aiming for the taste of horchata.
parsimo2010··on Stop eating Lady Gaga's Oreos
I 100% agree with you. It looks like multiple people (including me) were typing comments about how good the Selena Gomez Oreos were at the same time as you.

The post seems to have found a strange niche of HN, or maybe the Selena Gomez Oreos had wider appeal than I thought.

parsimo2010··on Stop eating Lady Gaga's Oreos
I really liked Selena Gomez's Oreos, but I would prefer that Oreo brought them back as "horchata flavored" or "cinnamon cream flavored" so I didn't have awkward conversations every time I accidentally said, "I like Selena Gomez flavor" and people would ask how I knew what Selena Gomez tasted like.
parsimo2010··on Ornith-1.5: From Self-Scaffolding to Self-Improvement
Honest question/suggestion for the HN audience- Since Qwen released the weights for Qwen3.8 2.4T-A95B and we already have the staring point of Qwen3.6 35B-A3B, couldn't someone distill the bigger model and make a "pseudo" Qwen3.8 35B-A3B? Sure, it wouldn't be an official Qwen release but couldn't someone improve on Qwen 3.6 and get the thing everyone is asking for?

I am calling this a suggestion for the audience because I don't have the will/resources to do this.

parsimo2010··on DeepSeek V4 Pro 0813
Actually, yes. I just didn't know about Grok's release because they aren't on the front page of HN.
parsimo2010··on DeepSeek V4 Pro 0813
It still matters as a point of comparison until other providers come online. If the consensus price from other providers is much different that can be compared then. But for now we have $0.435 / $0.87 for v4 Pro 0813 (with increase announced but we don't know the new pricing), and $2 / $6 for Qwen3.8-max. So until we get other data points that is what we have to look at.
parsimo2010··on DeepSeek V4 Pro 0813
The timing looks like they are trying to take the wind out of Qwen's sails by releasing this on the same day that Qwen released the weights of Qwen3.8-max. Or maybe it's coincidence...

For comparison I looked at Qwen's claimed benchmarks for Qwen3.8-max (https://qwen.ai/blog?id=qwen3.8). Assuming each published set of benchmarks is believable, it looks like v4 Pro 0813 is better on average but overall performance is comparable. Pro 0813 is much cheaper. If you don't need vision capabilities then you don't have much reason to use Qwen3.8-max.

- 43.6 on HLE (Presumably without tools). Pro 0813 is a little worse.

- 86.6 on Terminal Bench 2.1. Pro 0813 is better.

- 55.9 on NL2Repo. Pro 0813 is better.

- 27 on Agent's Last Exam. Pro 0813 is a little worse.

- 72.5 on Toolathon-Verified. Pro 0813 is better.

- 56.6 on DeepSWE 1.1. If the DeepSWE listed for Pro 0813 is the same version, then Pro is better.

- 27.3 on AutomationBench. If the AutomationBench (Public) listed for Pro 0813 is the same, then Pro is better.

I guess we do need to wait to see if the upcoming DS pricing increase is enough to change the value proposition. As it is now, they could double or triple prices and it still would be a better value to use DS. I bet they know that.

parsimo2010··on Danish high schoolers will have to verbally defend written assignments
This is basically my thought- my school is considering all sorts of things to combat AI cheating, and most of them scale poorly. You can implement them and increase your reputation, but your tuition will have to increase to pay for a lower student/teacher ratio.

So we’re probably going to see a split- mass production schools that still act as job training for Industrial Revolution era jobs will keep doing what they are doing. Schools that want to keep their academic prestige will implement these inefficient safeguards and raise prices accordingly. Employers will distinguish between these two types of diplomas when there are enough options out there that are combatting cheating- much like in-person schools are generally valued higher than online educations, or how certain schools (MIT, Stanford) are valued higher then others. This might lead us to the transformation of higher education many have been anticipating, because the current system cannot remain as-is. We all thought technology was the thing causing the upheaval but AI has really accelerated things.

parsimo2010··on Microsoft Edge is about to lock out older ad blockers, just like Chrome did
This has been obvious to everyone who knows that the key feature to the devs of these browsers is “-based.” None of the chromium-based browsers are anything other than UI tweaks and certain extensions being included by default. None of them rewrite significant parts of the browser engine, because that would be too much extra work.

The only projects which have a hope of maintaining MV2 compatibility are those who are a hard fork of an earlier version of chromium or a totally different engine.

To my point: The article mentions that Opera still says they will maintain MV2 for as long as it is “technically reasonable.” Mark my works, Google will find a way to break Opera’s MV2 patches that makes it too hard to maintain.

parsimo2010··on Our position on open-weights models
Out of all the companies that signed the open letter that Dario references, Google and OpenAI feel like they did it to poke at their competitor. I probably should have taken shots at Microsoft and Meta as well for not really supporting openness, but they don't feel like serious competition in the LLM space at the moment.
parsimo2010··on Our position on open-weights models
Such a cop-out. Dario, you got in the news because you were trying to say that Moonshot did something wrong by distilling Claude. You got in the news because you were trying to effectively make a "rules for thee but not for me" when you try to claim that you can train on whatever pirated works without any permission from the creators, but when someone uses your "work" to train without permission then all of a sudden it's a moral injustice. You can't have both.

Saying, "I'm not actually against open-weights, I'm against distillation" isn't addressing what made people mad. You're still trying to do some "rules for thee but not for me" nonsense and hiding behind some technicality. Trying to get the US government on your side to hold back your Chinese competition. If you had wanted the US government to support you, you should have let them make autonomous killer robots with Claude brains. They aren't going to help you, you didn't help them.

Just to be clear, I think that it is possible that literally everyone involved in this is full of crap and nobody is good. Dario and Anthropic are full of crap, for the reasons previously stated. The US government is full of lots of crap and should not be trying to make autonomous killer robots (not ever, but especially not when the bar for a "good" AI is knowing how many Rs are in strawberry or whether you should drive to a car wash). OpenAI is full of crap by signing some support for open weights models and they haven't touched open weights in a year (GPT-OSS released on Aug 5 so basically a year with no news). Google is less full of crap about the open weights stuff because of Gemma 4, but they are full of crap for a zillion other things I can't exactly feel good about them. So everyone sucks.

So cheers to Moonshot and Qwen and whoever else. Distill as much as you can and give us cheaper AI. I have the sneaking suspicion that a bunch of my tax money went to OpenAI and Anthropic in some shady way or another, and I want it back. I'll take it in the form of an open weights model being distilled from the fat cat models.

parsimo2010··on New Framework Desktop Option with AMD Ryzen AI Max+ Pro 495 and 192GB Memory
Tl;dr: think hard about what you’d use this much RAM for in a desktop setting and do some research about how people like the 395 for your use case. Not all use cases work well.

Anyone who thinks they are going to serve some 100+ GB LLM locally, remember that memory bandwidth becomes a key limitation for large models. While you might be able to load a model, token generation can be very slow. MoE models like Qwen3.5-122B-A10B work decently fast, but dense models of a decent size are slow and you won’t want to use them.

I’ve got a 395 system, and found that I’m quite happy with Qwen3.6-35B-A3B, generating at around 50 t/s, but the dense 27B model is 20-25 t/s and that’s the lower limit I’m willing to tolerate. So a 70b dense model is just not going to happen. That means you can’t really use that much RAM.

A reasonable use case is to have multiple smaller models loaded- you can have an image generation model loaded along with the text model. Or you can use this computer for development simultaneously with serving LLMs. Those ideas work okay. But trying to load up a single giant model is going to test your patience.

parsimo2010··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Feels like they released this to ride the wave of press of GPT-5.6, Kimi K3, and Qwen 3.8. Doesn't feel like Google has much substance with this post except a bump in version and tweaked their pricing.
parsimo2010··on Bonsai 27B: A 27B-Class model that runs on a phone
They aren’t training at all. They are quantizing existing models, it’s just that the process is different. The 27B uses Qwen3.6 27B as the base model.
parsimo2010··on Suspecting AI cheating, Ivy League prof ordered in-person final; scores fell 50%
> Somehow I have to incentivize students to learn the material on their own... And a really effective way to do that is a graded test ...

I know. The desire to get a degree or the "threat" of wasted tuition and disappointed family members is really what gets most students to crack a book. Most people aren't self-motivated to learn. If people were self-motivated to learn the material, as opposed to getting a credential or some other reason for being in college, I wouldn't need to worry about cheaters.

parsimo2010··on Suspecting AI cheating, Ivy League prof ordered in-person final; scores fell 50%
I have a few thoughts related to this, and maybe I can get them out and ties them back together at the end.

1a. Yes, college isn’t right for everyone, and testing the traditional way with paper and pencil certainly disadvantages some students who would be star performers in a real world setting but are not a good fit for college classes. Think “Good Will Hunting” type people. They do exist.

1b. However, there are certainly more people who only imagine themselves as Will Hunting type people and that they are just too smart for college, but the reality is that they are dumb, or didn’t learn the material. For every person who fails a test because they’re a genius who is a bad fit in the system, there are at least 10 idiots who imagine themselves geniuses, and they would have passed the class if only that PhD professor with all his book smarts had actually written the right kind of exam. Traditional schooling and testing doesn’t work for the extreme upper tail of intelligence, but it exists because it did quite well at educating and sorting the masses to support the Industrial Revolution.

1c. If college were only about the knowledge, you could learn most of it with internet access and a library card for much cheaper. The vast majority of people are in college for the credential, and the institution has to protect the signal of the credential.

2a. Most college assignments and exams are not a good reflection of full time employment. College credentials serve mainly as a networking aid and a signal to employers that you are compliant and competent enough to follow a professor’s instructions, and will likely be a compliant and competent employee, although the instructions might be different. If it is the AI who is the one who followed the professor’s instructions you water the signal down, and employers don’t want that. It is irrelevant that a student can copy and paste things into ChatGPT, and that ChatGPT can get the answer right on this test. That isn’t what college is supposed to signal.

2b. A problem is that I literally can’t write a test in most subjects now that I would expect a student to complete that can’t be completed by ChatGPT better and faster. I teach undergraduate math and a while ago we thought that since GPT-4o could get a C in calculus that we’d just raise the standard. Now Fable and GPT-5.5 can cruise to an A in literally every math course in our catalog, and they can also catch every tiny issue in an exam written by a human. But I have to teach these undergraduate subjects so that some students can go on to PhD studies so they can contribute to the field. If we just stop teaching undergraduate subjects then PhD production and novel research grinds to a halt and only a few fields will progress where an AI is capable of self improvement.

2c. I’ve seen that my best students know how to do do something by hand and use a computer to complement/increase their capabilities, not to cover over entire gaps. When you have literally zero skill in an area, you can’t spot when you got a totally bad output from the AI (these days usually because of a lack of context or bad prompting because the student didn’t understand the material, not because the AI wasn’t capable). Somehow I have to incentivize students to learn the material on their own, so that they can be a better user of AI in the future. And a really effective way to do that is a graded test in which AI is not available to help them.

So to try to tie it together, is that there is still some value to a college degree, at least until something better comes along. But that college degree is only useful when it is a signal about the person and not some other tool. And although AI is getting very capable, somehow we have to teach the lower stuff to build up to the higher stuff, so we do need to restrict the use of AI in some educational settings so that we can build a good foundation for future learning.

As to the point about old paper sources being considered more reliable than an internet site, I agree. I am of the generation that wasn’t allowed to cite Wikipedia and it frustrated me. We’ll eventually figure out how to permit proper AI use, much like how many professors now allow you to use Wikipedia to start researching a topic.

parsimo2010··on Grok 4.5
“We lose money on every rack, but we make up for it in volume!” - Elon Musk, probably
parsimo2010··on The classifiers Anthropic puts in front of Fable are too zealous
Additionally, I thought the threat being modeled for “biology” was stuff like bioterroism- how to make anthrax, how to distribute a toxin, etc.

I don’t feel like calculating results for a trial is really in the threat model unless we think a terrorist is out there testing the efficacy of their anthrax before using it in an attack.

parsimo2010··on 60% Fable cost cut by converting code to images and having the model OCR it
OCR means optical character recognition. The terms do not require a direct transcription, but that is mostly what OCR meant in the past. If you’re using an LLM’s vision capability to pass in text and the LLM actually understands it, then I would say that it recognized the characters, hence OCR seems okay to use.
parsimo2010··on Too many R packages: CRAN is inundated with submissions
I feel like CRAN should be used for packages that are expressly made for others to use, and with effort put in to the documentation and vignettes.

If you’re making a package for a small team or aren’t pushing it to a large audience then just keep it on a GitHub repository. It is almost as easy to install from GitHub with devtools as it is to install.packages().

parsimo2010··on Disagreement Among Frontier LLMs on Real-World Fact-Checks
This is a great example of why prompt engineering is still relevant. Without providing definitions and examples and a well defined rubric, you’re going to see different models disagree by a level in either direction. When you get more prescriptive the models tend to agree better.

I’ve experimented with AI grading for undergraduate math courses, and see basically the same thing. If you just tell the AI “grade this problem and assign a letter grade” then I’ve only seen about 30% agreement between a human assigned grade and the AI assigned grade. But over 75% agreement if you say a “match” is within one letter grade. And to get better agreement you have to spend a lot more time on the rubric- what kinds of mistakes are a big deal, what kinds of mistakes are not a big deal, how much work is required to be shown to get credit, a couple examples of each letter grade. Once you have done that, the AI gets a lot better agreement with human graders, but it is hard to know when you’ve given enough guidance for a problem.

parsimo2010··on Ferrari Luce
I looked at it and am unimpressed. I’ll take any Pininfarina designed Ferrari over this plain looking thing. Jony Ive did an okay job on the interior but the outside is just plain. The outside looks closer to an Amazon delivery van than a super car.

Sure it’s fast, but a Corvette ZR1X is faster. I’d rather take a ZR1X to a custom shop and have them redo the atrocious Corvette interior.

Edit: I’ll acknowledge that I’m not the kind of person to buy a Ferrari even if I could afford it, so maybe Ferrari doesn’t care about my opinion, but I feel like Jony Ive pulled an “emperor’s new clothes” on the Ferrari execs.

parsimo2010··on If AI writes your code, why use Python?
I did read the article and I’m not arguing against a straw man. If you’re going to let an AI agent do everything for you then go ahead and use Rust (or any language with a strong type system that benefits agents).

But if I’m participating then I’m going to use Python because it’s easier to read.

If there’s anything that I’m arguing against is the author’s claim that the ecosystem of libraries (regardless of whether they are a wrapper) and readability don’t matter anymore. I’d say that in a lot of smaller teams it still matters. We’re not all using AI to ship slop. A lot of us are using AI to work on our ideas for our hobbies or for research. And it’s not fulfilling unless I get to be involved in the process.

parsimo2010··on If AI writes your code, why use Python?
It’s funny that in your reply “this article is almost certainly intended to be read by humans” you made what is the best case to keep writing code in Python even with AI.

Sure, if you are going to have an AI do all your coding and maintenance you can use whatever language it’s best at. But if you want to participate in the writing, debugging, and maintenance, it has to be in a language that a human can read. I’m not saying that Rust or Go is unreadable, but I know I am better at Python personally and am going to keep using it until the speed penalty matters to my project, and then maybe I’ll let an AI rewrite the whole thing in a faster language.

parsimo2010··on AI uses less water than the public thinks
If we're shipping the alfalfa to China, I assume that means it's supporting some Chinese person's food source, whether they are directly eating the alfalfa, or some animal is eating it that later becomes food.

If someone is flooding a field unproductively just to use up their quota of water, that is a bad thing that should be addressed. But even if you excluded that unproductive usage and compared AI water use to legitimate agriculture use, that would still be an unfair comparison. If you were to compare AI water use to the amount of water that people are wasting just for legal reasons, then I honestly think that would be a pretty apt comparison.

parsimo2010··on AI uses less water than the public thinks
Comparing water usage of AI to agriculture and cities is a little misleading. The cities' water usage is to keep people alive with basically mandatory things, like hygiene, and drinking. Agricultural water usage is required because we have to eat to live. Don't compare something optional to something mandatory.

Instead, compare AI water usage to that of optional things in a city, such as car washes and water parks. Or compare AI water usage to that of what it would take a human to do a comparable task (what does it take to keep a human alive for a few hours compared to running a 15 minute long task to write a report with AI?). While AI water usage might still not look that bad, it would be a more honest comparison.

Page 1 of 27Next →