HNHacker News
TopNewBestAskShowJobs

StefanKarpinski

5,587 karma · joined November 14, 2010

[ my public key: https://keybase.io/stefankarpinski; my proof: https://keybase.io/stefankarpinski/sigs/QmBtKeGXG8EFVCErRV-VXRykH-gMg2jTEKGq-2xWMMU ]
submissionscomments
StefanKarpinski··on GPT-5.6 vs. Claude Fable 5 for Physical AI, which performs best?
Yeah, things are moving fast. Those benchmarks will be out soonish. Takes a while to run these things.
StefanKarpinski··on GPT-5.6 vs. Claude Fable 5 for Physical AI, which performs best?
They were all benchmarked, but not included in the post to try to keep the amount of data from being excessive. (Results are also not surprising — Opus good, Sonnet struggles, Haiku fails.)
StefanKarpinski··on Afroman found not liable in defamation case
> Let's not forget Uvalde, where the police department budget was ~40% of the city budget and it resulted in 19 cops standing outside scared while one shooter kept shooting literal children for an hour. Because they were scared.

Not only did they not stop the shooter, but they actively prevented parents—who were willing to risk their lives—from intervening. They didn't just not help, they proactively ran interference for the shooter.

StefanKarpinski··on The Shape of Inequalities
Sure, that would work just as well. Plus, then you get to pick a "good" placement instead of making the user try to find one.
StefanKarpinski··on The Shape of Inequalities
The animated visuals are very cool, but I desperately want to turn them off in order to understand what they depict and reason about it geometrically. A pause button would be greatly appreciated.
StefanKarpinski··on AirPods Max 2
Nope, they are dead. If I happen to catch them before they fully run out of power, they are at 1-2% charge as reported on whatever device they connect to. I can prevent this if I carefully disconnect them from each device they might be connected to. But that is a massive pain and fully defeats the purpose of being able to just put them down.
StefanKarpinski··on AirPods Max 2
Except that they don't. Or at least many, many people report that they don't, including myself and I have tried all the remedies that supposedly help. If you don't provide an off button and you can't build a product this expensive to power itself off reliably for everyone, then you've failed at product design.
StefanKarpinski··on AirPods Max 2
How does headphones not turning off force users to stay in the iOS ecosystem?
StefanKarpinski··on AirPods Max 2
Strange. Are they first gen or later? I did get the absolute first gen of these, so maybe it's a problem they couldn't fix in firmware? Or I just have a defective pair?
StefanKarpinski··on AirPods Max 2
I mean, I regularly leave them on a shelf in my apartment and they apparently do not consider that "down" or "stationary" enough to not just drain the battery completely. Truly a bafflingly bad design from the company that is (was) known for great hardware design.
StefanKarpinski··on AirPods Max 2
I got excited there for a second — free fix for the most annoying problem with my headphones! But no, my AirPods Max have the latest firmware and still have this issue. Any time I leave them for more than a day, the battery is drained.
StefanKarpinski··on AirPods Max 2
Wild. I have been eagerly awaiting this refresh, but this doesn't address either of the main issues with the original AirPods Max:

1. Still just as heavy. The AirPods Max sound quite good, but they are very heavy, to the point of being fairly uncomfortable after listening for any longer amount of time. This release as the exact same weight as the originals (13.6 oz).

2. Still no off button/position. They stay partially on unless you put them in the awkward and useless "case", which means they're constantly out of power when you want to use them. There's even an obvious fix: the ear cups swivel flat, they could just make this the "power off" position. Solved. But they didn't, so presumably these still have the same problem. There's also no mention of magnetic charging via stand, which would be another way to help alleviate this problem.

If these were even a few ounces lighter and powered off properly, I would buy them for sure. Given this announcement, I guess I will look for something else to replace the old AirPods Max.

StefanKarpinski··on Correctness and composability bugs in the Julia ecosystem (2022)
Yeah, probably should be documented.
StefanKarpinski··on Correctness and composability bugs in the Julia ecosystem (2022)
Sad to hear that, CyberDildonics
StefanKarpinski··on Correctness and composability bugs in the Julia ecosystem (2022)
Number isn’t an interface—there are no operations common to all numbers. Subtyping Number is a way to opt into numeric promotion and a few other useful generic behaviors. That’s it. The fact that some abstract types are interfaces with expected behaviors, while others are dispatch points to opt into behaviors is a double edged sword: powerful and flexible, but only explicitly expressed/explained in documentation.
StefanKarpinski··on Beware of Fast-Math
Floating-point arithmetic is non-associative, but it is commutative for the operations that are algebraically commutative: x + y == y + x and x*y == y*x. And x - y = -(y - x) so subtraction is properly anti-commutative.

The only very marginal exception to this is that when both arguments are NaN, the return value will be NaN, but which NaN payload is returned can depend on argument order. But no one ever uses this because it's not specified, so it can't be used reliably for anything useful. The behavior I wish IEEE 754 had specified for this is to define a standard NaN value (or two), and when the return value of an op is NaN, and some of the arguments are non-standard NaNs, then one of those non-standard NaN values must be returned. This doesn't depend on argument order and allows NaN payloads to be reliably propagated, which would let you encode useful debugging information in NaN payloads and know that it will flow through the program.

StefanKarpinski··on Ovld – Efficient and featureful multiple dispatch for Python
The paper "Julia: Dynamism and Performance Reconciled by Design" [1] (work largely by Jan Vitek's group at North Eastern, with collaboration from Julia co-creators, myself included), has a really interesting section on multiple dispatch, comparing how different languages with support for it make use of it in practice. The takeaway is that Julia has a much higher "dispatch ratio" and "degree of dispatch" than other systems—it really does lean into multiple dispatch harder than any other language. As to why this is the case: in Julia, multiple dispatch is not opt-in, it's always-on, and it has no runtime cost, so there's no reason not to use it. Anecdotally, once you get used to using multiple dispatch everywhere, when you go back to a language without it, it feels like programming in a straight jacket.

Double dispatch feels like kind of a hack, tbh, but it is easier to implement and would certainly be an improvement over Python's awkward `__add__` and `__radd__` methods.

[1] https://janvitek.org/pubs/oopsla18b.pdf

StefanKarpinski··on Beware of Fast-Math
Pretty sure that’s not possible. More accurate for some inputs will be less accurate for others. There’s a very tricky tension in float optimization that the most predictable operation structure is a fully skewed op tree, as in naive left-to-right summation, but this is the slowest and least accurate order of operations. Using a more balanced tree is faster and more accurate (great), but unfortunately which tree shape is fastest depends very much on hardware-specific factors like SIMD width (less great). And no tree shape is universally guaranteed to be fully accurate, although a full binary tree tends to have the best accuracy, but has bad base case performance, so the actual shape that tends to get used in high performance kernels is SIMD-width parallel in a loop up to some fixed size like 256 elements, then pairwise recursive reduction above that. The recursive reduction can also be threaded. Anyway, there’s no silver bullet here.
StefanKarpinski··on Higher quality random floats
It’s a little unclear what you mean by that without further explanation. Do you mean that conceptually one selects a real number at random and then rounds that real number to the closest representable float? (And if so, which rounding mode?)
StefanKarpinski··on Modulo of negative numbers (2011)
There are actually as many cases as there are rounding modes, which is seven, and for most of them unlike `mod` and `rem` there's no common name for these remainder operations:

- RoundNearest

- RoundNearestTiesAway

- RoundNearestTiesUp

- RoundToZero — rem, div

- RoundFromZero

- RoundUp

- RoundDown — mod, fld (i.e. floor(x/y) but without incorrect corner cases)

Most languages have no way of doing most of these, but then again, they're mostly pretty useless. They're really only useful when you're pairing division with a specific kind of rounding with a remainder that needs to match. Example of whacky remainder behavior in the "familiar" RoundNearest mode (default rounding mode for floating point):

    julia> [k => rem(k, 4, RoundNearest) for k=-6:6]
    13-element Vector{Pair{Int64, Int64}}:
     -6 =>  2
     -5 => -1
     -4 =>  0
     -3 =>  1
     -2 => -2
     -1 => -1
      0 =>  0
      1 =>  1
      2 =>  2
      3 => -1
      4 =>  0
      5 =>  1
      6 => -2
Wild, huh? Output range for modulus 4 is -2:2 and whether you get -2 or 2 alternates with each cycle around the ring. (Of course, Julia has comprehensive support for all of these because we're a bunch of nerds for this kind of thing.)
StefanKarpinski··on Modulo of negative numbers (2011)
The `rem`, `div` and `divrem` functions also all optionally take a rounding mode argument that lets you have the behavior that matches any rounding mode, where `RoundToZero` matches `div` and `RoundDown` matches `mod`, but there are actually a total of seven rounding modes. Most of them are pretty useless, but if you need some style of division and the remainder to match, this is very helpful.
StefanKarpinski··on Liquid democracy: two experiments on delegation in voting
Was confused by that as well. Sounds like Liquid Democracy doesn't do well compared to the alternatives, which would certainly be an interesting result, but doesn't fit with what the rest of the abstract seems to be suggesting.
StefanKarpinski··on Why Fortran is a scientific powerhouse
Ah, but you see, Julia is both too popular and not popular enough! It's too popular with the wrong people (bad unserious people) and not popular enough with the right people (good serious people).
StefanKarpinski··on Why Fortran is a scientific powerhouse
Ah, what an excellent catch-22! If you don't have a list of notable use cases, people claim that there are no serious users. If you make a list of notable use cases, then the fact that you compiled such a list is "telling". Debunking oft-repeated falsehoods is significantly different from being "hostile to outside opinions".
StefanKarpinski··on Why Fortran is a scientific powerhouse
Ah yes, nobody uses Julia in production... except these companies https://juliahub.com/case-studies/ and many more—it's impossible to keep up. It's 2023, Julia was #21 on the TIOBE index at some point last year. Claiming that nobody uses Julia in production at this point is getting to be rather silly.
StefanKarpinski··on The History and rationale of the Python 3 Unicode model for the operating system
The issue is that when you're implementing something like a programming language or a robust general purpose utility, then simply not being able to open—or list or remove or stat—paths with invalid names is not really acceptable.
StefanKarpinski··on The History and rationale of the Python 3 Unicode model for the operating system
The comment that you're quoting wasn't mine. In the comment you link to says "UTF-8 by convention". If either string is valid, then the result is as expected. If you're concatenating two strings that are both invalid UTF-8, there's not much you can do that's better than just concatenating the bytes together... which is exactly what treating them as byte arrays would end up doing (but it's less convenient). If you're worried about invalid UTF-8 you can check for validity (which again, is exactly what you end up doing if you use byte arrays).
StefanKarpinski··on The History and rationale of the Python 3 Unicode model for the operating system
Unfortunately, neither UNIX nor Windows require path names to be valid Unicode. UNIX interprets them as “UTF-8 by convention” and Windows as “UTF-16 by convention” but both actually allow arbitrary sequences of code units. It would be nice if this didn’t actually occur, but alas, it does, and if you’re writing general purpose utilities that work with files, you don’t want them to simply crash when this happens.
StefanKarpinski··on The History and rationale of the Python 3 Unicode model for the operating system
> That’s not UTF8.

True; I was careful not to call it that, but treating strings as UTF-8 by convention does make sense.

> It’s not, any unicode-aware text processing does it implicitly. This means any such things processing has to either perform its own validation that the input is valid, or it may fly off the rails entirely if fed nonsense.

In theory, but that's just not how most string operations actually work. If you have two UTF-8 strings and you want to concatenate them, you just concatenate the bytes. It would be ridiculously inefficient to decode the code points in each string and then re-encode them back into a destination buffer. If you have two UTF-8 strings and you want to see if one is a substring of the other and at what byte index, you just look for the bytes of one as a "substring" of the bytes of the other. Again, it would be ridiculously inefficient to decode the code points in each and do matching on code points. But what if the strings aren't valid UTF-8?! Both of those operations work just fine even if the strings aren't valid and produce sensible, intuitive results.

If you're implementing a browser or a terminal that has to actually display UTF-8 as characters then sure, you have to actually decode characters. Similarly, if you're parsing text somehow, then you have to decode characters. But many program only do concatenation and search and other operations like that which are actually implemented in terms of byte sequences, not characters.

StefanKarpinski··on The History and rationale of the Python 3 Unicode model for the operating system
It's not at all obvious how it helps, but it does.

First, why is Python unable to represent invalid path names as strings? Because internally it converts strings from UTF-8, UTF-16, or any other encoding, to a fixed-with array of decoded Unicode code points. The width of integer used to represent code points is determined by the largest code point in the string: if the string is ASCII, it can use a byte (uint8) per character; if the string is non-ASCII but all BMP, then it can use a uint16 per character; otherwise it has to use uint32 per character.

Why does Python do all this? So that you can have O(1) character indexing. If you gave up on that, you wouldn't need to convert the string at all, you could just leave it as (potentially invalid) UTF-8 data.

Suppose you get an invalid path on UNIX where paths are UTF-8 by convention? What does Python do with this string? It can't convert it to an array of code points because invalid UTF-8 doesn't correspond to a code point (well, it can if it's just illegal, not malformed, but in general, we have to consider completely malformed strings that don't even follow the basic UTF-8 format). So Python is stuck: it can only replace the invalid data with something like the Unicode replacement character. But then you can't do anything useful with that because it's not the correct name of the path you're trying to work with.

How does using UTF-8 to represent strings help? Because you can represent invalid strings: just leave them as-is and don't try to decode them unless you have to. Sure, you can't decode them as code points, but that's actually a pretty unusual thing to do. If someone asks for decoding, _then_ you can give an error. What about Windows where paths are UTF-16 by convention? You can convert them to WTF-8 and everything works out. (Described in way more detail here: https://news.ycombinator.com/item?id=33984308).

Page 1 of 32Next →