HNHacker News
TopNewBestAskShowJobs

croemer

2,820 karma · joined March 26, 2020

submissionscomments
croemer··on Backblaze drive stats for Q2 2026
This seems like a poorly executed analysis of very interesting underlying data.

With this degree of average age variability they can't compare failure rates.

They should show us Kaplan-Meier (survival) curves not failure in this quarter.

Unless they set some datacenter on fire, not much in Q2 should be different vs Q1 for the same hard drive other than it being older.

Sure, internally, I'd do a better analysis controlled for age that shows if there are common failure rate increases but here they focus on model failure rates which is completely confounded by age differences.

croemer··on America.gov
For those outside US, in Europe I got "Service unavailable".

Switched on US VPN and it works.

Edit: Now it seems to work. Maybe it was a glitch. Or a US government agent is watching this page.

croemer··on We found 24 Android vulnerabilities using our open source AI security agent
The title is misleading, they found 24 vulnerabilities in Android _apps_ not in Android itself. That's a big difference.

In their own words from the body:

> I’ve reported more than 20 vulnerabilities in Android applications

croemer··on Sonnet 5.5
Pretty crazy that the model doesn't know that it needs to stop before it hits 128k output tokens. I guess it has no sense of how many tokens in it is? Wouldn't this be possible to work into the architecture?
croemer··on Sonnet 5.5
This is evidence that Sonnet 5.5 wasn't yet trained on the HN comments from the Opus 5.5 release. Maybe Pelicanmaxing will lead to 127000 thinking tokens being used on Max.
croemer··on Sonnet 5.5
Playing around with it for a few minutes, Sonnet 5.5 feels very fast, much quicker than Opus 5.5. Can't tell yet if it's a lot worse but the speed is definitely welcome.
croemer··on Jev Can't Be Calibrated
Ah, I see, you mean "Jev cannot possibly be calibrated for everyone as there is always missing context".

I understood it as "It is impossible to calibrate Jev".

croemer··on Generate fonts where every LLM token is the same width
I typed some German and it was breaking up words so much more than English. Not really surprising given tokenizers are optimized for most commonly used text.

Here's the token efficiency of a corpus translated into various languages and tokenized with the latest OpenAI one:

  Language              Relative tokens
  --------------------------------------
  English                    1.00x
  Portuguese                 1.23x
  Chinese (Simplified)       1.25x
  German                     1.31x
  Spanish                    1.32x
  French                     1.37x
  Arabic                     1.38x
  Chinese (Traditional)      1.42x
  Korean                     1.47x
  Swahili                    1.49x
  Hindi                      1.57x
  Japanese                   1.66x
  Burmese                    3.16x
  Amharic                    5.78x
  Santali                   13.70x
Source: "Tokenizer Fairness in 2026", a reproduction/extension of Petrov, La Malfa, Torr & Bibi, "Language Model Tokenizers Introduce Unfairness Between Languages" (NeurIPS 2023), using FLORES-200.

https://github.com/partyfly/tokenizer-fairness-2026

croemer··on Revealing the details of how OpenAI agents hacked Hugging Face
TFA seems to be sloppy in writing, they should have kept the "meaning... [Some incorrect assumptions about GET]" out of the paragraph.
croemer··on Revealing the details of how OpenAI agents hacked Hugging Face
You're not missing anything. TFA really state this wrong assumption in their own voice.
croemer··on Revealing the details of how OpenAI agents hacked Hugging Face
> This access seems to have only allowed the agents to make ‘GET’ requests, meaning they could fetch and read websites, but not interact with them, submit forms, or send data to them

The authors of this (very interesting) analysis should really not state the sandbox's wrong assumptions in their own voice.

GET absolutely allows you to interact with sites. And of course GET can also send information. It's all up to the server that receives the GET to decide what it let's callers do with it.

croemer··on Jev Can't Be Calibrated
Article says that it _can_ be calibrated, it just isn't out of the box. So the title seems contradicted by the body.

> If you want calibrated probabilities you’ll still need to recalibrate Jev’s probabilities on your own data. The good news is that is cheap. A few hundred labeled examples from your actual data can be enough to fit a Platt scaling on top of Jev’s scores.

croemer··on Claude Opus 5.5
Hah! You independently picked exactly the same sentences I flagged (I know you posted this 11min before me but the comment only appeared after I had submitted mine).
croemer··on Claude Opus 5.5
That's the standard annoying pattern though: "Rewrites are declared by the publisher, never inferred from overlap." and "NULL means dirty, and DELETE is the fence." - still the same LLMisms. I didn't expect them to disappear, but it's not a radical improvement either.
croemer··on Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
Of course it was exposed - not sure it's explicit or not. Why wouldn't HackerNews comments be part of the training data? And Simon's blog and the many discussions about Pelicans? It'd be hard to miss. Doesn't mean Anthropic has made this an explicit goal in training.
croemer··on Typesafe-computer-use drives a Mac toward a goal for 1/50th of a cent per step
A demo video would be useful, so one can see what it does without having to run oneself.
croemer··on A heap overflow and SSO misconfiguration to compromise OpenAI internal repos
Which is presumably why Debian developers voted to allow responsible use of LLMs.
croemer··on A heap overflow and SSO misconfiguration to compromise OpenAI internal repos
Also had a bunch of typos so wasn't even LLM proof read (at least not final form).
croemer··on Wax motor
I'm not sure I get what you mean.

There are commercial thermostats where the wax element itself is both the temperature sensor and the actuator, eg Thera-100 T1002W0.

croemer··on The American Religion of Self-Storage Facilities
Wendover made a nice video about this 10 months ago: https://youtu.be/uEVv8SOJ6Is?is=Er27clBvXZ4TzAlj

I'm a big Wendover fan, I think many here would like it.

croemer··on A 32-year-old bug walks into a Telnet server
Thanks, makes sense! No need to be sorry!
croemer··on A 32-year-old bug walks into a Telnet server
Redacting what? Did you mean write or edit?
croemer··on A 32-year-old bug walks into a Telnet server
Thanks for maintaining telnet.

  The bug was reported on a public mailing list, which is sadly common nowadays
In defence of the reporter, your Readme only says "Send bug reports to bug-inetutils@gnu.org.", there is no distinction for vulnerabilities. [1]

There is a 3 months old pull request to advertise a private reporting email address but it's unmerged, maybe you could use this renewed interest as a nudge to set it up and merge: https://codeberg.org/inetutils/inetutils/pulls/26

One other thing the reporter could have done to make your life easier is to write a repro script rather than just explain the steps in prose.

[1]: https://codeberg.org/inetutils/inetutils/src/commit/40f19d84...

croemer··on Saving Jet Fuel
If you look at the "optimal" path here looks completely unrealistic/pathological. How is it optimal to zig zag like this? The grid is clearly too coarse and that's possibly not the only issue: https://tech.marksblogg.com/theme/images/flight_planning/msr...
croemer··on Show HN: We built open OpenRouter that turns usage into a better model
Probably shouldn't call it "Open router" in the title as that's a specific brand. Maybe you meant "we built something like OpenRouter".

Rereading I see you wrote "open OpenRouter", which looks a bit like a typo at first glance.

croemer··on Six government passport CAs that Python rejects and OpenSSL accepts
Interesting but would be better if you didn't copy/paste AI generated text here. Also, the link seems to go to a tool and the "Python doesn't parse 6 certificates" is only a sidenote to the repo's main purpose. Either focus on one or the other, don't mix the 2. Your comment I'm replying to here mixes the 2 as well. The default path and open SSL fallback is irrelevant to the invalidity finding.
croemer··on GLM-5.3-Flash Intelligence, Performance and Price Analysis
I see, thanks, then it should state that in the card.
croemer··on GLM-5.3-Flash Intelligence, Performance and Price Analysis
Why does the top card say "Intelligence #1/173" when the bar chart further down shows it only at position 7?

And the model isn't even shown in the speed bar chart just below. Such slop (the artificial intelligence website linked)

croemer··on Over 5,200 Ebola cases recorded in Congo
Newest run from 2026-08-25 shows Rt likely above 1 again unfortunately (percentages are confidence intervals): 30% 1.05–1.19, 60% 0.99–1.31, 90% 0.82–1.6. Daily growth rate: 30% 0.005–0.014, 60% -0.001–0.021, 90% -0.013–0.034 per day.
croemer··on Disruption with Some GitHub Services
Possibly vitess from the latest update:

> primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues

Page 1 of 34Next →