Do you know how much your computer can do in a second? (2015)
computers-are-fast.github.io
computers-are-fast.github.io
The way to estimate this is that light takes 3ns to go 1m. If the room is 10m wide, that’s 30ns. At typical frequencies that’s 100 clocks. With, say, 8 cores that’s 800 steps. At each step a core can retire about 8 multiplies using AVX vector instructions. The estimate of 6,400 multiplications is actually very conservative because you also have the GPU! Even low end models can put out hundreds of thousands of multiplications in 30ns. High end models can do about a million.
When asked, most people estimate “1 to 10”. Even programmers and IT professionals do not have a mental model of how fast computers really are.
The other day I was watching The Twilight Zone, and one episode involves a crew "655 million miles" from Earth. I know what miles are, I know the rough layout and scale of the solar system, galaxy, and universe, but I had no idea how far away that actually is because it mixes these two units from very different contexts and intuitively it just doesn't make sense to me; is it still in the solar system? No idea. Converting it to AU showed it's about 7 AU, which put it in the appropriate context: a bit beyond Jupiter.
It's always fun to see the repeated scale analogies as you build your way to enormous sizes and distances, starting with something familiar, zooming to something 100 or 1000 times bigger, doing it again and again. It engages our physical understanding while counting zeroes 2 or 3 at a time, but I don't think it actually gives us any physical understanding. It really is just counting zeroes.
I just don't think there's any chance you or anyone else has any idea how far Jupiter is except as a calculation.
I can wake up in my hometown, drive west for 12 hours, and still be in the same state. I know how big my state is in those terms. How many miles is that? How many football fields is that? How many times does light cross the room during that travel distance?
I agree, it's very hard – almost impossible – to really "understand" things at scales that are so far outside of our human experience and intuition. However, "near Jupiter" does put things in to relative proportion, whereas "X million miles" doesn't.
I'm not saying I like GPs puzzle, but I could kinda answer that question (my estimation was a bit higher, assuming GP's answer is exactly correct) only because he combined these contexts. Ask me, how many multiplications a computer can do in 30 ns and I'll say, well, I dunno, 10? Well, ok, I'm kinda used to log-times in μs, so maybe I'll say 100. I have no idea how much is 30 ns, but it sure seems a short time.
Also, I don't remember what is the speed of light. Probably, I can figure it out, if I try, but I can't just say from the top of my head.
What I can say instantly from the top of my head, though, is that individual computations are performed with a speed of light, literally, and that a single transistor in my PC is about a billion times smaller than my room, so assuming a single multiplication takes about a hundred of logic gate operations (again, I cannot really tell without thinking, but seems legit) I'll say: well, about 10 million, I guess? And it doesn't matter that I don't know how much time it'll take for light to travel across my room.
I don’t think understanding the speed of light is something that can be considered intuitive by most people.
There is no generally observable way to understand the speed of light, it’s just instant to most people.
So without physics training, people just have no idea.
I always like the comparison of more computing power than it took to put a man on the moon!
Which used to be along the lines of a phone or car had more computing power than NASA in 1969.
But these days it’s probably a toothbrush ;-)
So it’s not as useful anymore.
I also have no idea of what level of computation converts to something practical. A million calculations a second means nothing on its own; what can that do? Draw a picture on the screen? Print a PDF?
So it has to convert to practical understanding.
Still, it’s a nerdy and fun comparison but one that isn’t understandable or relatable to the layperson.
Anything involving networks is very much affected by it. The internet runs on fiber, where the speed of light is a fundamental constant! Even electricity in copper obeys those same limitations.
For example, when I ping microsoft.com, I get something around 100ms (Aside: what does that time mean? Round-trip or one way? I read a few man pages for ping, but couldn’t find out)
There also is the weirdness that, when pinging apple.com or google.com I get about 6 ms. They must be playing tricks with DNS.
Then again, I do know as a dev that unless it's happening every frame (ie for games), as long as you're not introducing non-scalable complexity, the computation time of a bit of reasoning is often negligible.
Units definitely matter but anything outside of our day to day experience is difficult to communicate widely. I think most people don't know more than Jupiter is actually really far away and really really big compared to the Earth. 7 AU means nothing to my brain just like 655 million miles. It's a computed unit based on mean distance from the center of the Earth to the center of the Sun; I also have no idea how far we are from the Sun because my brain just experiences a hot circle in the sky. The only reason these things seem to make sense is because we have seen really vague maps of the Solar System. Other falsely but widely-accepted measures include using state sizes to describe other regions. People readily accept it despite never having been to said state or potentially any other state ... or even across their own state.
I think these are accepted because of a false confidence in understanding (think Dunning-Kruger) based on the fact that we have some information.
Personal anecdote: A friend flew from the East coast to Arizona and asked if we could just drive to the other side of the state to see the Grand Canyon in the evening. I had to explain that it's 6 hours one-way despite being in the same state. Had I instead said "Arizona is the size of 22 Connecticuts," the idea of how long it would take to get there would still not make any sense.
PS: I said "vague map of the Solar System" because generally nothing is always the correct scale. Either the planetary distances are the correct scale from each other or the planets are the correct scale from each other for pragmatic reasons. Sometimes, none of those things are scaled properly. Some animations attempt to show things correctly but that's still difficult to comprehend because of the non-linear velocities required to keep the attention of the viewer.
7 AU has at least some meaning. Further than the sun. Probably closer than Jupiter if I had to guess? Definitely closer than Pluto. Would take on the order of 100 minutes for light to travel that distance. We could send a space vehicle there in a single-digit number of years.
My brain can't comprehend either number in the context of my morning commute, but at least it can comprehend 7 AU in the context that that number was given. I can't tell you how long it would take to walk that distance, but that doesn't matter since nobody will ever walk that distance. What I am able to do is make inferences about the implications of that distance (assuming no fiction/sci-fi stuff): we could have a conversation with a few back and forth messages over the course of the day, and they aren't getting home any time soon but they probably won't die of old age.
That's the point being made about combining contexts.
Also, if I were asking that sort of question in an interview I would fine with a candidate getting a conversion wrong as long as they explain their thought process, and I’d help them correct the error early. I’m this sort of question is t meant to be a trick, it’s about whether you can connect different variables which may not appear connected at first glance.
It's about absolutely nothing at all. You can project whatever you want on pointless questions like that. Whatever you think you're measuring in a candidate by asking them, well, you're not, lol.
The impression I had was that the question was asked in a social context: over coffee, beers, or something like that.
> What units do you think you know the scale of the solar system in?
AU, as I mentioned: https://en.wikipedia.org/wiki/Astronomical_unit
AU
What you are saying is technically true but with HUGE caveats. At this point, the issue isn't one of clock cycles but rather memory bandwidth and latency. Even assuming the best case, a hit in L1 cache, you are looking at +1ns to move the data from memory into the registers. But now talk about L2, L3, or worse system memory? Oof. And then you mention GPUs, but that problem is compounded 10x due to the bandwidth constraints of PCIe.
I mean, sure, if your CPU is doing nothing other than multiplying the same value by 2 over and over again, it can do that wickedly fast. However, once you start talking about large swaths, gigs even, of data to operate against and those numbers start to take a precipitous drop.
Heck, one of the benefits of AVX isn't the fact that you can do 64 lanes of multiplication at once, but rather the fact that you can tell the CPU to grab 64 lanes worth of memory to work against at once.
This is why when you start looking at what is talked about when people talk about next gen math machines, it's not the FLOPs but instead the memory fabric that gets all the attention.
That’s 64 bytes per nanosecond, or 16 floats per ns, and hence 480 in the time needed for light to go 10 meters.
That’s the “base” performance. Caches and registers would improve on this significantly in realistic scenarios.
Again, this is why I like this thought experiment! It’s not a trick question. It’s testing intuitions.
People just think it’s a trick because their intuitions are so ridiculously out of whack with reality that they’re looking for the sleight of hand instead of reexamining their own notions.
Again, terms and conditions.
The slow part for memory access isn't how much memory can be piped over the wire but rather how fast a request for the next chunk of memory can be issued. So while it's possible to get those 480 multiplications that relies on the cache being warm and the memory access being predictable to the CPU. If either of those two constraints are violated, then the number of multiplications is severely decreased.
In other words, the more frequently you have to tell the memory controller to get more memory, the less work you can do.
To put it in perspective, it costs around 100 cycles to load something from main memory. Imagine you are talking about something like an n body simulation that may be accessing memory all over the heap. In the worst case, you are looking at 4 multiplications (if multiplying by a scaler) and perhaps even just 2 multiplications if you need 2 memory pulls per multiplication.
Caches don't just help, they are necessary to come anywhere near achieving that 480 number.
Just a little history lesson.
Once upon a time, the memory controller didn't exist on the CPU, it was on a separate chip (the north bridge). Whenever the CPU wanted to load memory, it'd send a request to the north bridge and the north bridge would send that request to main memory.
One of the huge improvements to CPU performance for both Intel and AMD was integrating the memory controller onto the CPU. Why? Because they shortened the distance a request for memory, and the memory itself, had to travel significantly. Electrical signals travel at about 2/3 the speed of light and if your memory is 0.1 meter from the CPU, you can see how many requests for a new chunk of memory can really slow down everything.
This is fundamentally why nVidia's big GPGPU processors they sell to datacenters come chock full of memory. Because they want to limit the amount of time and times the GPGPU has to reach back into main memory to process stuff. It's far quicker to load up the 20gb of ram on the card in one go then to need to constantly pull that data over the PCIe line.
To use an analogy: I'm not asking how far a car can go on the highway if accelerating from zero, but how far it can move if moving at highway speeds.
Nonetheless, the latency to main memory is also a great example of how bad programmers' intuitions are! A processor can perform thousands of operations in the time it takes for it to wait for just one memory access.
RAM is the new disk, and disk is the new tape.
PS: There was an article on here just a couple of days ago about how "handles are better than references". I didn't see many comments pointing out that this doubles the number of memory requests required when chasing pointers, halving the overall performance!
> To use an analogy: I'm not asking how far a car can go on the highway if accelerating from zero, but how far it can move if moving at highway speeds.
The terms and conditions keep applying. The everyday real world conditions implied by "highway speed" include regular cache misses and branches.
It's not about "start" or "acceleration". This massively impacts the steady state speed.
If you want a car analogy, tight vector packing compares to normal code about as well as looking at the speed and width of the flame front across a cylinder and using that to measure fuel consumption.
It sounds more like testing whether you know the speed of DDR5 without looking it up, and doing a bunch of basic unit conversions.
And your analysis doesn't even cover latency (which GP mentioned), which apparently according to my first google result is ~90ns. So unless your data is already in a cache it might not even be fetched within the requisite time.
Your question sounds like a trick question because there are so many nuances and gotchas, and if you're asked this question without a concrete context and the problem one's trying to solve, it's basically a check of whether the interviewee's mental model of computation matches the interviewer's. And as this thread shows, even among (presumably) competent people, the factors to consider can differ quite a bit.
What I’ve seen, intuition wise, on where slowness is in computation it’s usually along the terms of “how many FLOPs can this do?” or “How much bandwidth does this have?”. But that’s a tiny fraction of the story. Very little computation is boring math operations.
It’s true, CPUs are absurdly fast at multiplying 2 numbers together. That’s not what trips them up. What trips them up is getting those two numbers and determining when those 2 numbers should be multiplied. Those two problems are so hard that CPUs are often guessing, running the calculation, and then backtracking if they guessed wrong.
And that action has lead to side channel attacks (spectre, meltdown) which can leak information about data to malicious software.
This is why such a question doesn’t have either an intuitive or non-intuitive answer. It’s very strictly “it depends”. And if you boiled it all the way down to “no, what’s the absolute max it can do” you are now talking about software problems that don’t exist/aren’t valuable.
Consider, what execute the add instruction faster, a 3.4 ghz CPU from 2004 or a 3.4ghz CPU from 2023? The answer is they are the same. Both do the add in 1 cycle. Did you learn anything about the performance of a 2004 CPU vs a 2023 CPU? Heck no. That’s because how fast the CPU can execute instructions has almost nothing to do with how fast it can process data. (Yes, I’m aware that even the above statement is somewhat incorrect as x86 CPUs do weird things with adds where they’ll run multiple additions per cycle with some clever reordering… Not even that example can be simple. Assume they didn’t).
The trivia question is a boring question. The actual interesting discussion is “so WHY aren’t we doing multiplications that fast?”. That is what can actually lead to deeper understandings about what slows things down in computing. There’s a reason the 10ghz CPU never surfaced. It’s not because we couldn’t make it, it’s because we couldn’t keep it busy.
Or of how slow light is!
I’ve been to America a few times. The internet is so fast there: not because of bandwidth differences, but because of latency, and especially sites carelessly loading chains of resources, which amplify the effects.
As soon as you deal in intercontinental stuff, you realise just how slow light is.
(As for studies in the effects of page load abandonment, those are never in the slightest bit relatable—at the time those studies suggest half the people are giving up, most sites still haven’t rendered anything at all, here.)
Then you notice a small button at the right bottom corner, that button allows you to auto-scroll at the speed of light, wow, now that’s slow!
[0]https://joshworth.com/dev/pixelspace/pixelspace_solarsystem....
https://www.visualcapitalist.com/visualizing-the-speed-of-li...
[1] https://en.wikipedia.org/wiki/Space_travel_under_constant_ac...
Yes, this is why it won't happen, with our current understanding of physics, and the practical limitation that implies, with the resources we have. ;)
https://files.mtstatic.com/site_4539/12414/0/webview?Expires...
This means it's required infinite energy to reach c speed for any mass.
And c speed is quite slow for space travel.
Just for the fun of it, let's imagine humanity has assembled, in space, an aircraft-carrier sized vessel, fully equipped to function and nurture the little humans living inside of it.
Weight: 100 ktons.
Power is infinite ok, because badly rewarded nerds discovered new physics. Good for them they are now immortals of human history.
With Lorentz kinetic energy equation, one can estimate the kinect energy this vessel would have while traveling at say, 1% of the speed of light.
Energy: 4.5*10^20 J.
This is about two thirds of the total energy Earth receives from the sun in one hour. Or close enough to the energy the world consumed in 2017.
Now we have humans inside a vessel hurdling through space at 0.01 c, in addition to the pre-existing humans in a planet swirling through space. But there is a problem! It would take 424 years to reach the nearest star system. So we need to go faster and maybe break things. Hopefully not the hull, though.
F*** it let's go 0.5 c and reach Andromeda in about 9 years - long enough to write a book.
Energy: 1.4*10^24 J.
That's 3x the energy released by the Chicxulub meteor impact. Or 30+ times the 2003 world's total fossil fuel reserves.
Which raises the question what is the fuel being used?
Doesn't matter ok because new physics, we are transforming mass literally in energy no constraints 100% efficiency lol.
By e=mc^2 that fuel would weight - at least - 14.9 ktons.
There is margin for error, since the vessel would be shedding mass, and getting lighter. That would allow engineering to run the global process at 80% efficiency, which is a very realistic metric and maybe miss a turn or two on the way to the neighboring star.
Returns not included.
sad piano plinks
I suppose the question is where do you want to go and how much luggage do you want to take with you?
Speed of light is not exactly a relatable point of reference.
The performance of everything is limited in one way or another by ‘c’.
Not just WAN links, but the data centre Ethernet as well. The distance to the disks matters. The physical size of the motherboard. The placement of caches, etc…
When people say things like “putting the compute near the data” they’re implicitly talking about overcoming the limits imposed by the speed of light.
When you hear about an N+1 performance issue in some ORM, that’s bad because of the speed of light.
When you test your LAN with “ping”, you’re lying to yourself because it won’t show measurements below 1 ms, which is an eternity.
I just told you how much compute can occur in just 10 nanoseconds.
Go ping something. Look at the “1 ms” in the output. Go back to my post and work out what can occur in 1,000,000 nanoseconds. Go look at the 1 ms again.
Repeat until you have an epiphany about your zone redundant Kubernetes-hosted cloud native microservices architecture.
Your ping is 156,000 nanoseconds. You just saw that you can in principle do about 6,400 computations per 30ns, so... that's about 33 million arithmetic calculations per round-trip, within the data centre.
I hope this makes you see every unnecessary network hop in a different light.
PS: It typically takes 3 round-trips to establish a TCP connection, and 5-7 for a TLS connection. A database connection over TLS needs a few more. And then you have load balancers, firewalls, proxies, envoy, ingress, dapr, and, and, and...
Datacenter Ethernet has too much overhead to make a difference in almost all cases. Disks have an even higher overhead:distance ratio. The size of a motherboard only matters for signal integrity, not that half a nanosecond extra. Cache location inside a chip can matter, but even then size is a significantly bigger factor than location.
Oh group velocity.
And also the fact that we eat up just enough of that performance for it to be slow sometimes but usually no more, because we can’t
Make about 25% progress toward the start menu popping up after the click on the button.
The problem is that it's filling an array in freshly allocated memory. While you might expect `malloc(NUMBER)` to give your process a crap ton of space in your RAM, that's far from the truth. First of all, glibc will just translate that into an `mmap`, and even if it didn't, the Linux kernel still wouldn't allocate the whole buffer due to "optimistic memory allocation." Instead, you'll receive ownership of this chunk of virtual memory, but you won't actually allocate anything at first. Only when you dirty each page will it actually allocate the backing physical memory. And even then, the Kernel has various heuristics to preemptively page in memory that it thinks you're going to use soon.
I'm sure the author was aware that the example was more nuanced than just "muh CPU cache", but reducing spatial locality to just "muh CPU cache" for the article's sake does a disservice to the reader.
(edit: Not the 0th index, but the same index as the base of the array.)
I don't need a website to tell me my computer is fast. What I need is a website that spoonfeeds "how to make webapp feel like Q3A" to the average developer and somehow keeps them on objective.
You really should be able to constantly shame yourself into focus on this. "Is rendering this report table more complicated than a scene from Overwatch?" You will likely answer "no", hang your head in shame for a moment, and then admit to yourself you need to throw away your stack of 20+ 3rd party js libs and break out the MDN bible.
Hell, even just asking if it is more complicated than DOOM should illicit an even more appropriate level of shame. DOOM ran on a 386 with 4MB of RAM and 12MB of disk.
Do you know how much your computer can do in a second? - https://news.ycombinator.com/item?id=10445927 - Oct 2015 (174 comments)
(edit: was supposed to be a pun)
The tags ( something like env=prod, app=auth ) will narrow things down to say 1% of the data in the time period. The software then just greps though a few Gigabytes to find the exact lines you are after.
the startups using grep on aws are undercutting those doing slower things on aws.
i wonder why aws architects never talk about grep.
I will never go back to not using it.
Was also very surprised to see how slow Bcrypt is.
edit: Yes, BCrypt, is designed to be slow. I just had no idea how slow. I assumed it was 100-1000 times faster than it is/can be.
Though unlike all the other examples here, this one is actually intentional. If/when computers become considerably faster, the cost factor of typical bcrypt implementations will be raised so that it stays slow, to keep it difficult to throw brute-force attacks at it.
Basically the important argument to bcrypt is literally the amount of time you want it to take to run on your hardware (ie the number of rounds).
So I would guess in C++ you could make 240 million calculations in a second.
The only thing I learned is apparently default Python JSON lib sucks on speed
Benchmarking json_read_struct: Collecting 100 samples in estimated 5.2012 s (81k iterations)
json_read_struct time: [64.126 µs 64.376 µs 64.660 µs]
change: [+0.2646% +0.7525% +1.2538%] (p = 0.00 < 0.05)
Change within noise threshold.
Found 7 outliers among 100 measurements (7.00%)
6 (6.00%) high mild
1 (1.00%) high severe
My benchmark code[0] takes advantage of knowing the shape of the data, and also knowing that I can avoid allocations for the strings so we're not measuring the allocator performance. With 64 microseconds per iteration, that comes out to about 15,600 parses per second.[0] https://gist.github.com/Measter/acbae474ba8e1451946630da2a2c...
- you don’t seem to specify the size of the input. This is the most important omission
- you are constructing an optimised representation (in this case, strict with fields in the right places) instead of a generic ‘dumb’ representation that is more like a tree of python dicts
- rust is not a ‘moderately fast language’ imo (though this is not a very important point. It’s more about how optimised the parser is, and I suspect that serde_json is written in an optimised way, but I didn’t look very hard).
I found[1], which gives serde_json to a dom 300-400MB/s on a somewhat old laptop cpu. A simpler implementation runs at 100-200, a very optimised implementation gets 400-800. But I don’t think this does that much to confirm what I said in the comment you replied to. The numbers for simd json are a bit lower than I expected (maybe due to the ‘dom’ part). I think my 50MB/a number was probably a bit off but maybe the python implementation converts json to some C object and then converts that C object to python objects. That might half your throughput (my guess is that this is what the ‘strict parse’ case for rustc_serialise is roughly doing).
-> ᛯ ls -s -h big.json
101M big.json
[11:53:30] ^ [/tmp]
-> ᛯ cat /tmp/1.pl
#!/usr/bin/perl
use v5.24;
use JSON::XS;
use File::Slurp;
my $j = decode_json(read_file('/tmp/big.json'));
say scalar @{$j->{'a'}}
[11:53:39] ^ [/tmp]
-> ᛯ time /tmp/1.pl
34952534
/tmp/1.pl 0,81s user 0,13s system 99% cpu 0,939 total
so just around 100MB/sI get over 3.3 billion additions per second. 550 million is way too low.
Not much if you are running Electron apps.
1. They will not be rewarded for going the extra mile
2. They will not be punished for making the site slower
3. They will be punished for shipping the feature slower
Point out "X will be slow" to a lot of devs and they go "I know, but I don't have bandwidth to prioritize the extra work"
Part of why it's so important to make the right thing easy.
Most just want to get paid and dont give a shit. Which is normal.
IME >90% of non-junior engineers understand most of the slowdowns they're introducing and choose to add them anyway.
> Most just want to get paid and dont give a shit.
My point is that you can't get anyone to give a shit about doing work for which they're not only not rewarded, but actively punished. Expecting people to actively hurt their careers for the sake of shipping good product is not a reasonable expectation.
You need some mechanism to fix those incentives.
If you needed more than just adding a single second, you were stuck with it.
But size-wise it was ginormous.
One machine counted clock cycles between two key presses using the same finger. From what I recall the smallest answers were in the … thousands?
https://youtu.be/DlEa8kd7n3Q?t=12m52s
In one second I can type about 7 ASCII characters, read maybe 20, or aim and click like twice. Each typed character is worth like 600 million CPU instructions of a single core of a 4GHz CPU. It's actually this fat, precious thing, the product of a giant computation from a "computer" (that was optimized for something else) that we aren't harnessing at anywhere near its real value. Instead, we type semicolons.
Edit: Apparently I forgot to account for the milliseconds part. My bad.
loop.py does 68 million operations per second, and the text below says "we know about the most we can expect from Python (100 million things/s)", but write_to_memory.py writes 2 billion bytes per second writing one byte at a time.
How can writing to memory be faster than doing nothing?
That, and Python is probably the most accessible language for the most number of people.
Most people who are tech adjacent, learn declarative langs like SQL, HTML first. Python doesn't look like that at all.
It's like when people say "Go is so easy to read!" I learned to program in SML then OCaml, __insert other FP langs__ then Ruby and ES6, Go looked pretty alien to me the first time I saw it.
In my point, I rebuttal that non programmers would find python "easy to read", kinda not true.
I then use my exposure to Go as antidotal evidence, of the above observation.
Could have been worse. Only a few were off by more than one order of magnitude (mostly I underestimate just how fast md5 is).