The whole mantra of avoiding "premature optimizations" was applicable in an era when "optimizations" meant rewriting C code in assembly.
The whole mantra of avoiding "premature optimizations" was applicable in an era when "optimizations" meant rewriting C code in assembly.
You need to be thinking about performance from the very beginning, if you're ever going to be fast.
Because, like the article said, "overall architecture trumps everything". You (probably) can't go back and fix that without doing a rewrite.
(Though it can be OK to have particular small parts where say "we'll do this in a slow way and it's clear how we'll swap it out into a faster way later if it matters".)
But if your approach is just "don't even worry about performance, that's premature optimization", you'll be in for a world of pain when you want to make it fast.
but you'll also be in a world of pain if you took so long to architect performance into your app/service that you miss your market window, or fail in some other important metric that causes your entire business to topple over.
It's better to have a product/market fit really early, but poor performance (which can be fixed), than to miss your opportunity to ship in time to get that marketshare and thus fail entirely! Try fixing that!
Sadly, most software today is insanely slow(†), compared to what it could be, which I think reflects the unfortunate business reality that's it better to get anything out to market quickly than actually build fast software.
But you have to be careful or you'll end up in the situation of e.g., Microsoft, which recently had a video bragging that the latest version of Teams did something like take startup time down from 10 seconds to 5 seconds (video seems to be gone now, alas). They're apparently working heroically to try to improve the performance, but it's really really hard since slowness was baked into the architecture.
(† Just typing this comment in Safari is lagging pretty terribly. This on a Macbook M1 pro. Though of course just entering text in a textbox should be snappy even on a potato.)
It's the same thing with security, or for that matter portability. Changes in fundamental goals are hard to 'bolt on' after the system has been built.
This also means you have to understand the customer’s problems before trying to optimize or before starting to architect your system.
Modern servers seem to have about 100ns latency to main memory. The speed of light (actually electrical signals) delay is maybe 1-2ns.
I found multiple papers that also simply use 100ns for RAM latency (e.g. this one[1]) but in reality it is a lot more complicated than one number. E.g. the time to first word for DDR3 is just above 6ns[2]. I'm not a CPU engineer but it looks like CPUs don't necessarily wait for the entire cache line and can utilize the first word straight away.
Also, the signal needs to travel both ways and the traces are not straight lines. I'm not an electronics engineer either so the best number for the maximum DRAM trace lengths I could find are 12-30cm, which would be 1-2ns indeed. This means that there is some room for improvement (in theory).
But even with ideal RAM it is physically impossible to beat register access at 0.17ns (i9-13900KS @ 6GHz) or even L1D cache at 0.7ns[3].
What is causing the extra 4-5ns latency of the RAM? Would it be physically possible to remove it?
[1] https://www.usenix.org/system/files/conference/hotcloud18/ho...
[2] https://en.wikipedia.org/wiki/CAS_latency
[3] https://chipsandcheese.com/2022/11/08/amds-zen-4-part-2-memo...
The CAS latency for RAM is considerably higher and has been decoupled from this - as clock speeds go up CL multipliers go up and the CAS latency ends up being roughly the same - ~10 ns for the best RAM.
I'm not sure why CAS latency can't be driven lower. Row selection (tRCD) could be limited by the physics of charging up sense amplifiers, but I don't see why the logic for column address shouldn't be able to run faster. It's probably just something that hasn't been an area of optimization for the simple reason that there are many other much larger sources of latency.
And on that note, a CPU won't just jump to accessing some random word from RAM and copying it into a register. It will first look through L1, then L2, and then L3 before seeking the data from RAM. It will then copy an entire cache line into all three caches and load the target word into a register. This entire process results in ~100 ns latency.
I've seen code that does "fast" searches of a tree in a dumb way come out O(n^10) or worse (at some point you just stop counting), and the solution was not to search most of the tree at all. Find the relevant node and follow links from that.
Meanwhile in my day job performance really doesn't matter. We need a cloud system for the distributed high bandwidth side, but the smallest instances we can buy with the necessary bandwidth have so much CPU and RAM that even quite bad memory leaks take days to bring an instance down. Admittedly this is C++ with a sensible design (if I do say so myself) so ... good design and architecture means you don't have to optimise.
Yep, completely agree. I worked at a company with a poorly architected high-throughput system that was written in Perl. It got to a point where no more optimisations could make it scale, so it was rewritten. Of course the rewrite in a "faster language" was touted as the reason for its success but the truth was the new architecture didn't pound the database anywhere near as much.
That sentence was always taken out of context
I guess that makes sense in the context of large-scale data research, where you should be thinking big if you really want to have outsized impact.
Even in ML, which I work in, optimizing a data pipeline or model code to be 3x faster means you can run 3 experiments a day instead of 1.
There's a huge chilling effect from slow code in research iteration. But par for the course as academic code tends to be absolute trash.
Of course, there is plenty of low quality academic work that only compares to some existing system that has incomprehensibly bad performance, then shows a 10% improvement. That sort of research work should definitely be avoided.
My point is that you should be writing fast optimized code by default in ML research because you tend to throw large volumes of data at that code multiple times a day, and slow code reduces the number of experiments you can run.
This isn't what I see however, lots of research code is horribly slow.
It makes the least sense in that context imho, as any optimization leads to larger differences the more data you process. Any performance difference is multiplied by the amount of times that code runs, that's why people look closer at hot loops than init code that runs once at startup.
One need only look at the programs he typically writes to see that he is far more obsessed with performance than most of us even know how to be.
(I am going off memory here, I should say. Mayhap I have the source wrong. Pretty sure that is the discussion, though.)
Also, within the last few years I have written assembly for an application originally in Java, so that's still on the table. =)
Merging two loops into one is an optimazion that should be postponed until it's necessary and the system have been profiled.
Joining two large lists by a hash join instead of a nested loop is a design decision that should be made at the outset.
Well, these days the equivalent to that seems to be rewriting {{ $HIGHER_LEVEL_LANGUAGE }} code in Rust.
I would still be conscious of "premature optimizations" and wouldn't rewrite all the things in Rust (or something else) before profiling my existing code.