HNHacker News
TopNewBestAskShowJobs

orlp

8,870 karma · joined March 27, 2013

Developer at https://pola.rs/.

Publish a blog at https://orlp.net/blog/.

Other socials:

    http://github.com/orlp/  
    https://stackoverflow.com/users/565635/orlp  
    https://linkedin.com/in/orson-peters/
submissionscomments
orlp··on ICPC 2025 World Finals Results
It doesn't seem that hard to solve to me either. It's solvable with basic linear programming.

    1. Add a variable for each node, and a variable for each output edge from stations.
    2. For each reservoir add equality constraints to the sum of incoming edges with the coefficients given in the problem.
    3. For each station add equality constraints between the weighted sum of its inputs (which is 1 for the root station) and its outputs (which are the variables we added).
    4. Add an out_edge >= 0 constraint for each output edge on stations to forbid illegal negative flows.
    5. Add a variable m which is constrained to be less than all the output station variables.
    6. Maximize m.
orlp··on Polars Cloud and Distributed Polars now available
> The creator of duckdb argues that people using pandas are missing out of the 50 years of progress in database research, in the first 5 minutes of his talk here.

That's pandas. Polars builds on much of the same 50 years of progress in database research by offering a lazy DataFrame API which does query optimization, morsel-based columnar execution, predicate pushdown into file I/O, etc, etc.

Disclaimer: I work for Polars on said query execution.

orlp··on I Was Wrong About Data Center Water Consumption
> The width increase (10-100x) completely overwhelms the depth increase (maybe 3-10x), so surface area increases substantially.

No it doesn't.

The only thing that matters (in this oversimplified calculation which only takes into account surface area) is average depth of the freshwater while it is on land. If the reservoir is on average deeper than the rivers the freshwater otherwise would be flowing in, there will be less evaporation per liter of freshwater available for use.

Now a dam also increases the total amount of freshwater that's kept on the land in a steady state situation compared to if the water flowed free into the sea. It would be absurd to count this as "extra evaporation" when this extra freshwater otherwise would've simply be lost when it would flow into the sea instead of being kept in the reservoir.

orlp··on I Was Wrong About Data Center Water Consumption
> On the one hand, a huge dam reservoir does increase the level of water evaporation relative to an undammed river by increasing the amount of water surface area.

That depends entirely on the depth of the river and the depth of the reservoir. If the average depth of the reservoir is deeper than the average depth of a river there is less surface area.

orlp··on Tesla said it didn't have key data in a fatal crash, then a hacker found it
The marketing doesn't even matter. It either needs to be full self driving, or nothing at all. The "semi self-driving but you're still responsible when shit hits the fan" just doesn't work.

Humans are simply incapable of paying attention to a task for long periods if it doesn't involve some kind of interactive feedback. You can't ask someone to watch paint dry while simultaneously expect them to have < 0.5sec reaction time to a sudden impulse three hours into the drying process.

orlp··on Uncertain<T>
Not sure why this is being upvoted as the article is not describing interval arithmetic. It supports all kinds of uncertainty distributions.
orlp··on God created the real numbers
Actually in math it's very common for the more general system to be simpler. Compare for example the prime numbers with the integers, or general groups with finite simple groups and the monster group.
orlp··on macOS dotfiles should not go in –/Library/Application Support
I assume they've made up their mind and are now just tired of discussing it. I don't know why they refuse to even consider an option for it.
orlp··on macOS dotfiles should not go in –/Library/Application Support
I and others have brought this up with the dirs Rust crate maintainer but they refuse to see it this way: https://codeberg.org/dirs/dirs-rs/issues/64. It's very frustrating.

I now use a combination of xdg + known-folders manually:

    [target.'cfg(windows)'.dependencies]
    known-folders = "1.2.0"

    [target.'cfg(not(windows))'.dependencies]
    xdg = "2.5.2"
to get the config directory:

    use anyhow::{Context, Result};

    #[cfg(windows)]
    fn get_config_base_dir() -> Result<PathBuf> {
        use known_folders::{KnownFolder, get_known_folder_path};
        get_known_folder_path(KnownFolder::RoamingAppData).context("unable to get config dir")
    }

    #[cfg(not(windows))]
    fn get_config_base_dir() -> Result<PathBuf> {
        let base_dirs = xdg::BaseDirectories::new().context("unable to get config dir")?;
        Ok(base_dirs.get_config_home())
    }
orlp··on A German ISP changed their DNS to block my website
Are you using a third-party DNS like 1.1.1.1 or 8.8.8.8?
orlp··on 4chan will refuse to pay daily online safety fines, lawyer tells BBC
> will solve 90% of the problem

Remind me again, what the problem they're trying to solve is?

orlp··on Going faster than memcpy
That is something I can agree with, but I can't in good faith just let "it's just a hint, they don't have anything to do with correctness" stand unchallenged.
orlp··on Going faster than memcpy
> Non-temporal instructions don't have anything to do with correctness. They are for cache management; a non-temporal write is a hint to the cache system that you don't expect to read this data (well, address) back soon

I disagree with this statement (taken at face value, I don't necessarily agree with the wording in the OP either). Non-temporal instructions are unordered with respect to normal memory operations, so without a _mm_sfence() after doing your non-temporal writes you're going to get nasty hardware UB.

orlp··on Rotring 600 Ballpoint Pen
My favorite pens are the Frixxion pens. You can erase them, and it actually works well.
orlp··on Itch.io: Update on NSFW Content
This is useless. You can't stop Collective Shout (their campaign almost surely falls under First Amendment rights), and even if you could, 30 minutes later a new group pops up. Plus your message would fall completely on deaf ears for anyone who agrees with Collective Shout.

Bring attention to the fact that payment processors are acting as active censorship of legal content, rather than neutral infrastructure. Emphasize that if they can censor legal content, anything could be next, including but not limited to political donations of a specific party.

orlp··on Top DNS domains seen on the Quad9 recursive resolver array each day
> Wow, that's smart. I was wondering whether there is a way for the bots to generate "unpredictable" domains such that security researchers could not predict them efficiently (even with source code), but the botnet controller can.

There is a fairly simple method which achieves the same advantage for a botnet controller.

1. Use a hash of the current day to derive, for that day, an infinite stream of domain names. This could be something as simple as `to_human_readable_domain(sha256(daily_hash + i))`.

2. A botnet slave attempts to access servers in a diagonal order over (days, domains), starting at the first domain for today and working backwards in days and forwards in domains. An image best describes what I mean by this: https://i.imgur.com/lcEbHwz.png

3. So long as one of those domains is controlled by the botnet operator (which can be verified using a signed response from the server), they can control the botnet.

This means that the botnet operator only needs to purchase one domain every couple of days to keep controlling their botnet, while someone trying to stop them will have to buy thousands and thousands every day.

And when you successfully purchase a domain you can publish the new domain to any connected slaves, so this scheme is only necessary for recruitment into the network, not continued control.

orlp··on FP8 is ~100 tflops faster when the kernel name has "cutlass" in it
Intel's C++ compiler is known to add branches in its generated code checking if the CPU is "GenuineIntel" and if not use a worse routine: https://en.wikipedia.org/wiki/Intel_C%2B%2B_Compiler#Support....
orlp··on FP8 is ~100 tflops faster when the kernel name has "cutlass" in it
GenuineIntel moment.
orlp··on Nvidia Becomes First Company to Reach $4T Market Cap
Literally pasting "Wikipedia list of companies by market cap" into Google gives me https://en.wikipedia.org/wiki/List_of_public_corporations_by....
orlp··on A compact bitset implementation used in Ocarina of Time save files
My comment was worded a bit too harshly, sorry about that, I certainly wouldn't have worded that way in a personal message.

As I mentioned the hexadecimal printing coincidence is a neat fact, I was just excited when clicking the link to find a novel bitset idea. In my disappointment to find the standard bitset (albeit with 16-bit limbs) I reacted a bit too harshly. And as per https://xkcd.com/1053/, just because something isn't new to me doesn't mean it's not new to anyone.

orlp··on A compact bitset implementation used in Ocarina of Time save files
> The most novel aspect of OoT bitsets is that the first 4 bits in the 16-bit coordinate IDs index which bit is set. For example, 0xA5 shows that the 5th bit is set in the 10th word of the array of 16-bit integers. This only works in the 16-bit representation! 32 bit words would need 5 bits to index the bit, which wouldn't map cleanly to a nibble for debugging.

There is nothing novel about this really. It's neat that it works with hexadecimal printing to directly read off the sub-limb index but honestly who cares about that.

Outside of that observation there's no advantage to 16-bit limbs and this is just a bog-standard bitset where the first k bits indicate the position within the 2^k bit limb, and the remaining bits give you the limb index.

orlp··on Creating fair dice from random objects
How to create a fair coin from an arbitrarily biased coin:

1. Toss the coin and remember the answer.

2. Toss the coin again, if it is different from your previous toss then your result from #1 is fair. Otherwise, go back to step 1.

If p is the probability of getting heads, there are four possible outcomes with their associated probabilities:

    TT -> (1 - p)^2   (rejected)
    HT -> p * (1 - p)
    TH -> (1 - p) * p
    TT -> p^2         (rejected)
Needless to say, p * (1 - p) and (1 - p) * p have an equal probability, so if we don't reject our two tosses, we have a fair outcome.
orlp··on How Cloudflare blocked a monumental 7.3 Tbps DDoS attack
Economic fraud detection is like trying to find a needle in a haystack.

Blocking DDoS is like trying to separate the shit from the bread in a shit sandwich.

It's a completely different problem.

orlp··on Why Pandas feels clunky when coming from R (2024)
> DuckDB is faster than polars

Whether or not DuckDB is faster than Polars depends on the query and data size. I've spent a large portion of the last 2 years building a new execution engine for and optimizing Polars, and it shows: https://pola.rs/posts/benchmarks/.

orlp··on Math Symbol Frequencies
You won't like bra-ket notation then :)
orlp··on Beware of Fast-Math
I don't believe so, no. Currently these operations only set the LLVM flags to allow reassociation, contraction, division replaced by reciprocal multiplication, and the assumption of no signed zeroes.

This can be expanded in the future as LLVM offers more flags that fall within the scope of algebraically motivated optimizations.

orlp··on Beware of Fast-Math
No, there is no guarantee which (if any) optimizations are applied, only that they may be applied. For example a fused multiply-add instruction may be emitted for a*b + c on platforms which support it, which is not cross-platform.
orlp··on Beware of Fast-Math
I helped design an API for "algebraic operations" in Rust: <https://github.com/rust-lang/rust/issues/136469>, which are coming along nicely.

These operations are

1. Localized, not a function-wide or program-wide flag.

2. Completely safe, -ffast-math includes assumptions such that there are no NaNs, and violating that is undefined behavior.

So what do these algebraic operations do? Well, one by itself doesn't do much of anything compared to a regular operation. But a sequence of them is allowed to be transformed using optimizations which are algebraically justified, as-if all operations are done using real arithmetic.

orlp··on Trump administration halts Harvard's ability to enroll international students
It's not about the percentage of Harvard international students who fall into this category, it's about the percentage of students in this category who go to Harvard.
orlp··on Writing into Uninitialized Buffers in Rust
It is as easy as it looks to add `freeze`. That is, value-based `freeze`, reference-based `freeze` while seemingly reasonable is broken because of MADV_FREE.

Some people simply aren't comfortable with it.

Currently sound Rust code does not depend on the value of uninitialized memory whatsoever. Adding `freeze` means that it can. A vulnerability similar to heartbleed to expose secrets from free'd memory is impossible in sound Rust code without `freeze`, but theoretically possible with `freeze`.

Whether you consider this a realistic issue or not likely determines your stance on `freeze`. I personally don't think it's a big deal and have several algorithms which are fundamentally being slowed down by the lack of `freeze`, so I'd love it if we added it.

← PreviousPage 4 of 24Next →