HNHacker News
TopNewBestAskShowJobs

orlp

8,833 karma · joined March 27, 2013

Developer at https://pola.rs/.

Publish a blog at https://orlp.net/blog/.

Other socials:

    http://github.com/orlp/  
    https://stackoverflow.com/users/565635/orlp  
    https://linkedin.com/in/orson-peters/
submissionscomments
orlp··on DraftKings is using AI to behaviorally target chronic gamblers
Unless proven otherwise I assume any delete button in 2026 is just

    UPDATE tracking_data
    SET shown_in_interface = false
    WHERE account_name = 'm463'
orlp··on Fearless SIMD v1.0
You can do it with just 3 instructions for IEEE 754-2019 minimumNumber (ignores NaN):

        vminpd          ymm2, ymm1, ymm0
        vcmpunordpd     ymm0, ymm0, ymm0
        vblendvpd       ymm0, ymm2, ymm1, ymm0
If you want proper IEEE 754-2019 minimum (propagate NaN, -0.0 < +0.0, NaN bitpattern picked in the usual way) you can do it in 6:

        vminpd          ymm1, ymm0, ymm1
        vbroadcastsd    ymm2, qword ptr [rip + .LCPI0_0]
        vandpd          ymm2, ymm0, ymm2
        vorpd           ymm1, ymm2, ymm1
        vcmpunordpd     ymm2, ymm0, ymm0
        vblendvpd       ymm0, ymm1, ymm0, ymm2
I personally find this a load of nonsense I don't care about.

If you want propagating NaNs but don't care about signed zero or NaN payload/sign, you can use

        vminpd  ymm2, ymm0, ymm1
        vminpd  ymm1, ymm1, ymm0
        vorpd   ymm0, ymm1, ymm2
What I do in Polars is a bit different, there for propagating NaNs I do

    if (self < other) | self.is_nan() { self } else { other}
this isn't fully optimal on x86-64 but it's fairly simple and autovectorizes decently on various platforms, here's AVX2:

        vcmpltpd        ymm2, ymm0, ymm1
        vcmpunordpd     ymm3, ymm0, ymm0
        vorpd           ymm2, ymm3, ymm2
        vblendvpd       ymm0, ymm1, ymm0, ymm2
orlp··on Early rogue AI agent activity and attempts to hack found on urlquery.net
How do you define malicious?
orlp··on Claude Code now reads AGENTS.md if there is no Claude.md
If you want this it's trivial to add an AGENTS.md that simply says "if you're Claude read CLAUDE.md, if you're Astra read ASTRA.md". A common entry point is good regardless.
orlp··on Claude Code now reads AGENTS.md if there is no Claude.md
I wonder when they'll finally fix the VS Code plugin to not constantly dump your current file into the context.
orlp··on google.com/goto: Google's anti-scraping update
1. It adds extra latency due to an extra hop.

2. It is non-bypassable tracking of every single click.

3. It does not allow you to inspect the URL to see where it goes to before visiting.

4. It breaks the feedback signal of extensions which redirect sites. E.g. Fandom wikis are shit so I have them redirected to equivalent much better wikis like. But now any such redirection has to be done post-click tracking meaning Google still believes I want to see the Fandom site.

5. It breaks extensions which hide certain shit search results based on URL.

orlp··on We Replaced MMAP with Io_uring in Our Rust Query Engine. It Got Slower
We had a similar experience in Polars trying to use io_uring.

Rust isn't inherently bad at io_uring but at least Tokio currently is. I'm not the one who implemented and benchmarked it so this is second-hand information but if I recall correctly Tokio shares one buffer pool for all threads so as you scale to 100+ threads the whole thing grinds to a halt.

Migrating our I/O to a different async runtime than Tokio was rejected. So we'll wait until it's fixed in Tokio and now use regular blocking reads instead.

orlp··on All grown-ups were once children, but only few of them remember it
That's a nice story and all but it isn't a reality for most people. For most people on this planet it's work or starve on the street.
orlp··on Show HN: Compute polynomials twice as fast
It is applicable to fast universal hashes like Poly1305 and Polymur (the latter of which I'm the author). However it's not clear to me whether this work improves over the state of the art for that purpose, see some questions here: https://www.reddit.com/r/programming/comments/1wbgcke/comput....

This purpose is however much easier/flexible than actual polynomial equivalence since the requirement here is only that the polynomial is injective, not identical.

WyHash and xxh3 do not have polynomial structures.

orlp··on We have a year to fix security everywhere
> If I look at actual incidence involving memory safety issues compared to supply chain issues in general, it is the later which is much a higher risk to me.

Again, can you even just name a single supply chain attack that was *shipped* in Rust software? Against the thousands and thousands of known memory vulnerability bugs throughout time?

> And yes, there were successful supply chain attacks on Rust developers, even just recently: https://blog.rust-lang.org/2026/08/20/supply-chain-attack-on... despite this being a "solved" problem.

No, your linked blog post predates the brand-new min-age requirement. The minimum age would have prevented it, since it was detected by AI within an hour. If anything it supports my point.

Plus, as I already mentioned, that is a build.rs supply chain attack that targets developers, not shipped software.

orlp··on We have a year to fix security everywhere
The above is a general rule protecting you against supply chain attacks by default. If there is an important CVE published with a patch you can manually review that patch and bypass the minimum-age requirement for that dependency specifically.
orlp··on We have a year to fix security everywhere
Supply chain risks are essentially a solved problem.

    1. Set a minimum age on dependencies: https://github.com/rust-lang/cargo/issues/15973
    2. Scan all dependency code with AI
Even if you don't do #2 yourself as long as anyone does in the age window you've set, you're protected. In the age of AI the "you can't read all dependency code" argument doesn't work anymore.

On top of the above modern age argument, let's compare the amount of vulnerabilities found in shipped Rust software due to supply chain attacks (0 to my knowledge) against memory safety vulnerabilities (the majority of all vulnerabilities).

There have been successful supply chain attacks against Rust developers due to build.rs but those were quickly dealt with, and should be a thing of the past once min-age hits stable (next release).

orlp··on Discovery of a new OpenAI agent message board
Also Wiki4D, a D programming language dev wiki: https://prowiki.org/wiki4d/wiki.cgi?action=browse&id=RecentC...

Found by searching for wiki + texas poverty.

orlp··on New type of dice guarantees no tie when deciding who goes first
If all you have a coin there is a simple algorithm that's equivalent to sorting by random real numbers in [0, 1].

    1. All players flips a coin.
    2. Players that got heads go before players that got tails, forming (up to) two groups.
    3. If a group has more than one player go back to #1 to determine the order within that group.
It's not a finite process though - it could go on forever if really unlucky. But this is unavoidable, since the number of permutations on n players with n > 2 has factors not divisible by 2 there is no finite series of n coin tosses that could without any bias create a permutation, as the number of outcomes is 2^n.
orlp··on Pre-Release of Polars 2.0
Well... once my recent work on out-of-core lands the batch could be on disk when we run out of memory budget ;)

But no, that's not what I meant. I meant that the batch is meant to be of a size that fits in your CPU cache. This can be a huge throughput improvement as each bit of data stays in cache as it moves from data source to sink.

Compare this to column-at-a-time execution: by the time you start the next operation on this column the start of the column will be out of cache again, meaning you operate at RAM speed (or worse, disk speed) rather than cache speed.

I gave a (fairly surface-level) talk on the streaming engine a bit over a year ago: https://pola.rs/posts/talk-polars-meetup-1-streaming-engine/.

orlp··on Pre-Release of Polars 2.0
Streaming here has a different meaning than perhaps what you're used to. It's not referring to online processing where you maintain aggregates/state while an endless stream of data comes in.

The name was chosen early on to contrast with the old execution model, which was essentially all-data-in-memory, column-at-a-time. That engine still exists, we use it as a fallback mechanism for things that aren't supported yet in the new engine (or if you explicitly ask for `engine="in-memory"`).

The new execution model first constructs a computational graph of nodes which communicate in streams of in-cache batches (morsels) of data, meaning the full dataset will never be held in memory if not necessary. This was called the streaming engine for that reason in an early prototype and the name stuck. In hindsight I do admit the naming choice is somewhat confusing.

orlp··on Pre-Release of Polars 2.0

    from polars import col as C

    df.select(C.x, y = C.w / C.z)
orlp··on GPU World
If there was a continuous 500 watt GPU per person we'd increase energy consumption by more than double.

The continuous energy footprint per person globally averages to 356 watt currently.

orlp··on Rust project goals: Immobile types and guaranteed destructors
This is the opposite, it is further opting out of flexibility.
orlp··on RFC 9851: TLS 1.2 is in Feature Freeze
Isn't kind of the point of TLS that your communication doesn't get 'inspected'?
orlp··on How Our Rust-to-Zig Rewrite Is Going
I think if you interpret it charitably it means that any bug in the emitted machine code is already a likely memory-unsafe miscompilation if it is ran.

The compiler itself might be perfectly "memory safe" but the generated binary fundamentally is always at risk (besides WebAssembly I suppose).

I'm fully aware of the separation of compiler and binary, and being able to compile untrusted code safely is nice, but a perfectly safe compiler that generates vulnerable binaries isn't that much better.

orlp··on Almost Always Unsigned
That's a GCC skill issue. You can do it in five branchless instructions for unsigned by splitting the unsigned up in two 32-bit halves, converting those to floats simply by inserting their values as mantissa into constants 2^52 and 2^(52 + 32). This conversion is exact.

Then to finish the conversion you subtract 2^52 and 2^(52 + 32) respectively from the halves and add them together.

    vmovq       xmm0, rdi
    vpunpckldq  xmm0, xmm0, xmmword ptr [rip + .CONST1]
    vsubpd      xmm0, xmm0, xmmword ptr [rip + .CONST2]
    vshufpd     xmm1, xmm0, xmm0, 1
    vaddsd      xmm0, xmm1, xmm0
Here CONST1 = [0x43300000, 0x45300000, 0, 0] and CONST2 = [0, 0x43300000, 0, 0x45300000].
orlp··on SpaceX wants to launch 100k more Starlink satellites for 100x the bandwidth
Don't give these ghouls ideas.
orlp··on Orasort: 5x faster column-sorting with an expired patent from Oracle
First, this article is mostly (AI?) regurgitation. This is much better: https://smalldatum.blogspot.com/2026/01/common-prefix-skippi....

Second, I have independently invented this (quicksort on string prefixes) at my time at CWI, although I didn't end up publishing it, because...

Third, this was already published in the original 1961 Quicksort paper by Hoare: https://www.cs.ox.ac.uk/files/6226/H2006%20-%20Historic%20Qu.... Near the end, the section on "Multi-word keys" describes a quicksort that partitions on just the first word, and only accesses the next word for the equality partition. And funnily enough this paper credits P. Shackleton for this, thus this idea was thought of even before the Quicksort paper came out.

So as is usual for software patents, this patent never should have been awarded.

orlp··on Professor denounces mass AI fraud on an exam at Brown
This is a dumb take. It's like not teaching kids 1 + 1 because a calculator can do it for them.
orlp··on Ford hired AI and sacked humans. It backfired badly
If your data is sufficiently noisy or your relationship sufficiently simple a linear regression will outperform a SOTA LLM.
orlp··on Apple announces significant price increases for MacBooks, iPads, more
It seems like there aren't extra duties (anymore), but then again it's all very confusing and hard to navigate so who knows.
orlp··on Apple announces significant price increases for MacBooks, iPads, more
You save a lot less after paying import duties.
orlp··on MSG Made Dossier on Activists Who Opposed Facial Recognition
Actually, they're making an effort to force your business to not do something.
orlp··on Shall we play a game? My AI nuclear simulation
> Is there really anything about them that's bad? Or any worse than other things?

A full-on nuclear war will literally make a large portion of our planet uninhabitable for anyone for centuries, and leave the rest severely crippled and contaminated.

Sorry I know we're supposed to be kind and whatnot in these comments but I can't help but explicitly state that your comment is one of the dumbest things I've read on this site in a while. I hope you otherwise have a good day.

Page 1 of 23Next →