HNHacker News
TopNewBestAskShowJobs

MostlyStable

4,891 karma · joined April 28, 2022

submissionscomments
MostlyStable··on GPT 5.6 Sol is the best "vision" model OpenAI ever released
Curious why you didn't try Gemini 3 pro? That is the model I've been using for OCR entry of handwritten datasheets (JPGS of datasheets, structured JSON output). At my scale, the cost of 3 pro is basically not an issue, but if there are improvements in quality, I'd definitely be willing to explore other models
MostlyStable··on Abdominal fat predicts heart disease risk better than BMI
Body fat percentage can't conflict with BMI because BMI is not an estimate of body fat percentage.

And I agree that it's never caused me problems, but that's because I and, as you suggest, doctors who look at the number, have understood that it's not that big of a problem in my case. That's what I mean by not looking at the BMI number naively and ignoring other physiological states.

I don't think anything you said conflicts with anything I said, and it sounds like we mostly agree.

MostlyStable··on A U.S. Strategy to Prevent the Creation of Mirror Life
I'm curious what solutions to Future threats you think qualify as a panopticon, and why the following set of words in this context don't make you worried about the same:

>monitoring, verification, and other early-stage deterrence tools.

MostlyStable··on Abdominal fat predicts heart disease risk better than BMI
At a population level, it's great. For individuals, it can fail in all kinds of ways. For me: it is not very good because my body has very strange proportions. I'm 6'2", but my waist is barely higher than my 5'1" wife, and my torso is nearly the same length as my 6'7" brother.

In other words, for my height, I have a very long torso and very short legs. I think it should be relatively obvious that an inch of leg weighs significantly less than an inch of torso, so at a given level of body fat percentage, I'm going to weigh quite a bit more than someone my same height with more typical proportions, and thus my BMI reads me as more overweight than it otherwise would.

My point is not that BMI is bad or useless or anything else. My point is that it was designed as a population statistic and that it can be fraught when one tries to apply it to any individual with no nuance. A high BMI should cause one to consider and examine your health and weight. But it should not over-ride specific details about your physiology that point in the other direction.

MostlyStable··on Why does Opus 5 feel worse to work with?
I would not overgeneralize from paper. Firstly: forcing JSON output is, in my opinion, a bigger change than asking it to match the above style guidelines, and secondly, as is always the case with these kinds of papers, what was true for the model tested in the paper may either be completely false, or greatly reduced, in later models. That paper is almost 2 years old and models today have been trained in very different ways (or more accurately post trained in very different ways) and are in general far more capable.

Based on that paper, I would maybe try to check if it was true for a modern use case, I would very much not assume it was still true.

MostlyStable··on Nine PBS could lose 70 years of archives after cloud vendor goes defunct
Having that be there only backup solution would have been silly and dangerous. Having it be part of their overall backup solution including with something like the provider would have been smarter and better and prevented this entire problem
MostlyStable··on Show HN: KidScreen, a finite YouTube shelf chosen by parents
YouTube Kids in theory solves all my problems, the issue is that it's garbage software where the most important features are broken. For example I recently went on a long trip and grabbed the kids tablet. Only to discover after leaving the house, that all the previously downloaded videos were gone, and all the new ones that the app had chosen from among my whitelisted channels had "failed to download", despite having been at home on the wifi for weeks, so without internet connectivity, the app had nothing at all.

As others have pointed out, the things I am looking for in a YouTube kids replacement is quite specific and detailed, and the landing page provides not an iota of evidence that I will get that.

MostlyStable··on Timeline of the OpenAI accidental attack against Hugging Face
Yes, the message boards and collaborative hacking occurring during training runs was BY FAR the biggest bombshell revealed, and OpenAI doesn't even seem to realize it. The fact that they continued the training runs, with those rewarded behaviors included, and didn't wind back training to before hand, shows that they fundamentally do not understand alignment and safety (somewhat interestingly, their previous head of safety resigned shortly after OpenAI found about the message boards). I agree that, with that information, it is completely unsurprising that they hacked HuggingFace.....but that is also the Star Wars "You understand how that's worse, right?" meme.

I am flabbergasted at the complete lack of regard for alignment demonstrated here.

MostlyStable··on App Store Rejection of the Week: Dark Hours
There is an easy solution to this: Allow the user to decide whether they want to allow for apps from outside a particular ecosystem. Then Johnny can set up Grandma's phone and (in a deep byzantine menu where she will never accidentally stumble on it, maybe locked behind a pin), Johnny can set it so that Granny's phone can only install software from "Super premium, high class, curated, guaranteed not to hack your phone" walled garden app store. And Johnny, on his own phone, can choose what to download himself, including from the "Buyer Beware, half our shit is malware" dark web app store.

In my opinion, the Apple app store should actually be FAR more curated than it is, and far less curated options should be equally available. The Apple App store (and the Google app store for that matter), should be a relatively small list of apps that have extremely high standards and guarantees. It should essentially be impossible for spammy bullshit to get into these places. User-controlled, parental control like options should determine whether or not other app sources can be installed from (including allowing for specific other sources, so that app source control can be granular and detailed. I should be allowed to add FDroid but nothing else for example).

Done. Problem solved. Everyone is happy. Grandma is safe and Johnny has control over his phone.

But of course, that assumes that keeping Granny safe is actually the reason. Facially obviously, it is not, and is in fact entirely pretextual.

MostlyStable··on Software development with AI is starting to feel like cooking steak
You are protected in two ways:

The first, and by far most important is that unlike ground meat, your bacterial concerns are almost entirely on the surface. Unless something very very very wrong has occurred in the processing and transport chain, bacteria will not have penetrated into the meat. And if it's gotten bad enough that they have, you will almost certainly be able to see and smell it. And the surface is cooked to a temperature far in excess of that needed to kill bacteria.

And even for the interior, rare is not the same as raw. A completely rare steak should be cooked to between 120 and 130 (if you want it "blue" you might go down to 120). That's enough to get a decent reduction in bacterial count on it's own, although admittedly many techniques will not hold it long enough to get the number of log reductions that food safety guidelines recommend, in the case that you somehow did have a large bacterial infection in the interior.

So the answer to your question is that if you are buying your meat from an even halfway reputable source, the only place that bacteria might exist is on the surface, where it has been cut by tools and handled by the butcher/you and in general come into contact with the world. Those bacteria will be 100% killed by the cooking process. The odds of getting sick from a rare steak are low enough that if you are concerned about that, there are innumerable other things in your life that you should be more concerned by.

MostlyStable··on Software development with AI is starting to feel like cooking steak
My point is that steak is very, very, very far from the hardest to get right. I'm not going to contest your argument about how hard it is to get right, I'm going to argue that innumerable other dishes are vastly harder. Because the quality ingredients are harder to find, because they involve far more steps, each of which is harder to judge, because often good recipes or instructions simply don't exist, etc. etc. etc. Every single difficulty that you imagine exists to make a truly great steak exists 100 fold for other dishes.

In my opinion, steak has gotten the reputation that it has because it's so easy to get top tier restaurant quality at home (relative to cooking more generally). It is easy enough that, with a little care, almost any one can master it if they decide to, and therefore lots of people try and end up caring, and discussing, and sharing tips, etc.

There are a whole host of dishes where true mastery, to restaurant level, is borderline impossible for a home chef, and so most people never even try.

I don't disagree that reliably getting the last 5% of quality out of a steak is non-trivial, but the idea that, relative to the entire world of cooking, it's particularly difficult is laughably incorrect.

MostlyStable··on Software development with AI is starting to feel like cooking steak
I take the point, but I think the author picked a poor analogy. Cooking even an excellent steak is actually not that hard. In fact, I'd argue that it's among the easiest things to master/make at a top level quality at home. Does it require some modicum of attention and understanding? Sure. But starting with a high quality cut, owning a meat-thermometer, and knowing about reverse searing is about all it takes to reliably and easily get a near perfect steak every time.

There are far, far better cooking examples out there.

MostlyStable··on Born Against, or why hobby programming communities are against LLM usage
Sure, that's fair. I enjoy the process of figuring out ways software can solve some (usually quite trivial) problem for myself, but then I'm very happy to hand the act of making it to a robot.
MostlyStable··on Born Against, or why hobby programming communities are against LLM usage
Exactly. I have a backyard garden not because I need the food (it would be infinitely cheaper and easier to buy even high end farm-stand vegetables), but because I enjoy gardening. Buying a gardening robot (if such a thing existed) would completely defeat the point.

I use an LLM to make custom software for myself not because I enjoy coding but because I want software that I can't afford to hire a human to make for me. Buying a robot to do it (current agentic coding LLMs) is the only way it gets done at all.

These two things are not the same, and should not be viewed in the same light.

All that being said, I would 100% buy a weeding robot. Weeding sucks.

MostlyStable··on Ten advances in mathematics and theoretical computer science
From a mathematician who was intimately familiar with some of these problems [0]

>I don’t understand it yet. Maybe it’ll take me an afternoon to check all the calculations, but what would still be missing is why this was an approach that would’ve made sense in the first place. Is there some broader context or theory within which this would’ve been the obvious thing to do? What other results can be proven using these techniques? What is it telling us about quantum information or operator theory? I have no idea. I spent about an hour this morning asking ChatGPT these questions, but it’s somewhat frustrating because it speaks with a mishmash of physicist, operator algebraist, quantum information theorist-lingo, plus the usual LLM breezy lilt that annoys everybody.

They certainly seem to have "intuited", in a way that is not immediately obvious to experts in the field, the way to solve at least some of these problems. This was not just simply grinding away at a method that humans already knew would work and just hadn't gotten to yet.

[0] https://nitter.poast.org/henryquantum/status/208362369543662...

MostlyStable··on Deep-sea vehicles spot 'alien' sharks deep beneath the waves in the Pacific
There are loads of unknown (to science) deep sea sharks out there, most of them quite small. A friend of mine in grad school spent several trips on a deep sea trawler out of Madagascar and described multiple new species.
MostlyStable··on Kenji/Serious Eats – 30-Min Pressure Cooker Pho Ga
One of the reasons why I loved The Food Lab was that very often Kenji would write several versions of a recipe at different effort levels, along with detailed reasonings about why and how he shortened a step and what one was giving up by doing so. While he didn't do that for the Pho Ga recipe, he did do that for his Beef Pho recipe, with a full blown 6 hour "traditional" version [0] and a faster (although still not as fast as the pressure cooker pho ga one) 1 hour version of the recipe [1]. My guess is that the write ups of the 3 recipes together probably give one enough information to dial in relatively precisely how much effort you want to give and which steps one can skip or reduce.

[0] https://www.seriouseats.com/traditional-beef-pho-recipe

[1] https://www.seriouseats.com/quick-and-easy-1-hour-pho

MostlyStable··on Kenji/Serious Eats – 30-Min Pressure Cooker Pho Ga
Interesting to see this posted, as it's quite an old recipe (the oldest capture I could find on the Wayback Machine was from 2015 although I'm pretty confident it's older than that; as an aside, can I say how annoying it is when websites only have an "updated on" or "edited on" date and not a date for when it was originally written?). It's also one of my favorites of Kenji's though and a regular in our home.

I think that Kenji's series of pressure cooker recipes are probably my favorites of everything he did when he was doing The Food Lab section. Almost every recipe is fantastic and they are mostly pretty easy. If I had to pick my single favorite, it would probably be the pressure cooker chicken enchilads [0], although that one is on the higher end of the effort spectrum

[0] https://www.seriouseats.com/pressure-cooker-fast-and-easy-ch...

MostlyStable··on Elevators
Given that there is already a simulation, with very clear performance metrics, I wonder if anyone has ever tried making an evolutionary algorithm and if so, how well it would perform. Do we have any idea of how a theoretically perfect algorithm would perform? How much theoretical performance is there to be picked up?
MostlyStable··on Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Yes, I agree that this is still a datapoint for real world utility of AI assisted vulnerability fixing. But it's not as simple as an apples-to-apples comparison with previous years.
MostlyStable··on Tailscale didn't stop the Hugging Face intrusion
I have recently noticed that the words "ad" or "marketing" have become, in and of themselves, with no additional information or context, slurs or dismissals.

I understand why. The modern internet has turned advertising into a morass of constant bombardment and the only sane response is to block as much as possible and ignore as much else as possible.

But it's unfortunate because, in some sense, ever single thing that a company every says that is not legally mandated in some way is a form of advertising.

And in many cases, that "advertising" contains true, useful information that can be helpful.

What is important isn't whether or not something is an "ad", but instead, whether or not it contains true information that is helpful in some way.

Many ads don't reach this bar. They are either misleading, straight up lying, or information that is almost completely useless.

But when I'm searching for a particular product, about the only source of information at all is some form of advertising, and I almost always find at least some amount of it to be helpful in making a product decision.

Ads are more often than not polluting to the informational ecosystem, but that's not because they are ads.

MostlyStable··on Gemini Robotics 2 brings whole body intelligence to robots
The entire arc of human progress is taking things that used to be the domain of only the privileged and making them the domain of the masses. There was a time when if you were not actively contributing to the production of your own food (to the exclusion of almost any other activity), you were extremely privileged. Now, ~ no one (in the West) has to grow/raise/produce their own food and the proportion of most western's productive output that goes to purchasing food is a tiny fraction of overall work.

A new category of task starting to move along this trajectory is to be celebrated. Maybe some day people will clean some small portion of their house the same way that some people now have backyard gardens.

MostlyStable··on Google fixed more Chrome bugs in June than over the past two years, thanks to AI
After doing a bit more research, this fact is somewhat less impressive. Apparently, this was also the first time in nearly 20 years that there was such a large capacity crunch and many researchers weren't able to get into the competition. One of the rejected researchers did apparently have a working exploit, which, upon not getting into the event, they responsibly disclosed, and it was then patched before the event.
MostlyStable··on Google fixed more Chrome bugs in June than over the past two years, thanks to AI
I'm honestly unsure if this is a Poe's law thing or not. I'm going to go ahead assume that you are doing the honorable thing of purposefully not including a /s for the integrity of the joke.
MostlyStable··on Google fixed more Chrome bugs in June than over the past two years, thanks to AI
I just did a search and apparently this fact (the specific one about no payouts for the first time in almost 20 years) has not gotten a discussion on HN. Given the degree of skepticism around the utility of AI bug finding and fixing (this very thread is full of it), I would have thought that concrete evidence that it can help actually make real software more secure against attacks would have gotten a write-up somewhere.
MostlyStable··on Gemini Robotics 2 brings whole body intelligence to robots
This is one of those "tell me you don't have kids without telling me" situations. The process of teaching a kid to be a responsible human is long and arduous. If you expect a 2 and a half year old to successfully be picking up their toys, well, I don't know what to tell you. Go spend some time in a daycare, I guess.

And of course, that's completely ignoring the fact that I also would like to not have to pick up my shit if I could buy a robot that would do it. Yes, I can. And in fact I currently do...mostly. But I'm not happy about it and if there was an even quasi reasonable amount of money I could pay that would mean I would never have to do it again, I 100% would do so.

And if you can honestly say that you wouldn't...well I hope you are enjoying your life in the monastery.

MostlyStable··on Read this before you buy that TV streaming stick
These are the kinds of products I now just straight up refuse to buy.
MostlyStable··on Gemini Robotics 2 brings whole body intelligence to robots
My understanding from parents and friends who have a cleaning service...it pretty much depends on a certain baseline level of cleanliness. Most cleaning services don't (and likely wouldn't be able to provide) basic "pick up around the house". They will come into a house that is mostly picked up and sweep, dust, mop, vacuum, etc. but they won't come into a toddler house and put everything away.

I'm sure that there are services that do this, but I very much doubt they are the $400 per month services you are talking about. Robots may not be able to do that level of task at all right now, but they will get there.

And of course, that's not even thinking about all the cleaning tasks that would be avoided in between the every-other-week visits.

Having my surfaces be somewhat cleaner every two weeks is worth ~nothing to me. It's by far the easiest and least unpleasant part of cleaning a house. Having a robot that will pick up the detritus of everyday life and put it away, pick up and clean dishes, etc. and do so every single day would be worth an enormous amount of money.

I have no idea how long it will take for robots to get to that point, but whichever company cracks that level first is going to make a lot of money.

MostlyStable··on New HIV vaccine shows unprecedented success in preclinical study
It is entirely possible to believe multiple things at once:

1. The drop is fertility is bad for society

2. The drop in teen pregnancy is good, despite it's contributions to 1

3. We should try to fix 1. without undoing 2.

MostlyStable··on Be skeptical of OpenAI's rogue hacker agent story
Again: yes, very many people (just look through this very thread) are saying that. Tons and tons of people are saying some combination of "the whole thing is made up/a lie" or "This is entirely down to OpenAI incompetence and doesn't matter", or for some other reason (often not stated) making it very obvious that they do not think that this story matters very much.

And, again: to the people who aren't saying that: whatever argument is being made, it seems to me like it probably doesn't need to rely on claiming anything about the motivation of OpenAI. It sounds like you are against government regulation of AI. That's a position that a totally reasonable person could have. You should be able to argue that this event does not justify some particular kind of government regulation without reference to OpenAIs motivation for reporting the story.

← PreviousPage 2 of 32Next →