If you hold up a sign with, say, a multiplication, a CPU will produce the result before light reaches a person a few metres away.
If you hold up a sign with, say, a multiplication, a CPU will produce the result before light reaches a person a few metres away.
The latency on multiplication (register input to register output) is 5-clock ticks, and many computers are 4GHz or 5GHz these days.
5-clock cycles at 5GHz is 1ns, which is 30-centimeters of light travel.
If we include L1 cache read and L1 cache write, IIRC its 4 clock cycles for read + 4 more for the write. So 13 clock ticks, which is almost 70 centimeters.
------------
DDR4 read and L1 cache write will add 50 nanoseconds (~250 cycles) of delay, and we're up to 13 meters.
And now you know why cache exists, otherwise computers will be waiting on DDR4 RAM all day, rather than doing work.
3
Also the pipeline length is certainly not 5 stages but more like 20-30.
> The dependency chain length is what is normally intended as instruction latency.
Yes, the way I read the original post and others was that you actually your response back in 3 cycles, which isn't correct. It doesn't get comitted for a while (but following instructions can use the result even if it hasn't been committed yet). You're not getting a result in less than 20 cycles basically.
It makes sense now! :D
> You're not getting a result in less than 20 cycles basically
But the end of the pipeline is an arbitrary point. It will take a few more cycles to get to L1 (when it makes it out of the write buffer), a few tens more to traverse the L2 and L3 and hundreds to get to RAM (if it gets there at all). If it has to get to an human it will take thousands of cycles to get through the various busses to a screen or similar.
The only reasonably useful latency measure is what it takes for the value to be ready to be consumed by the next instruction, which is indeed 3-5 cycles depending to the specific microarchitecture.
I assume you are talking about from fetch to hitting the store buffer? That would be the aabsolute min time before the data could be seen elsehwere I would think. It can still potentially be rolled back, and that would be higher than reciprocal be way too fast to sustain, but for a single instr burst, I'm not sure. So much happens at the same time, An L1 read hit will cost you 4 minimum, hut all but 1 of that is hidden. can't avoid the multi cost of 3 or add 1. the decoding and uop cache hit, reservation, etc will cost a few. I have no idea.
If you know of anything describing it in such detail, I would be comopletely curiouis.
I used to play 3D games on Pentium-based machines and I thought of them as a "huge upgrade" from 486, which in turn were a huge upgrade from 286, etc...
Now, people with Ice Lake CPUs in their laptops and servers complain that things are slow.
However there is definitely still less intrinsic optimisation from a dev perspective I think - people will iterate over the same array multiple times in different places rather than do it once.
I guess our industry has decided moving faster is better than running faster for a lot of stuff.
Are there any directives to the Operating System to say - “here keep this data in the fastest accessible L[1,2,3] please”?
Some architectures (e.g. GPUs) provide local "scratchpad" memories instead of (or in addition to) caches. These are separate uninitialized adressable memory region with similar access times to a L2/L1 cache.
I'm probably the worst person to explain this.
Long long ago, I took a parallel programming class in grad school.
It turns out the conventional way to do matrix multiplication results in plenty of cache misses.
However, if you carefully tweak the order of the loops and do certain minor modifications — I forget the details — you could substantially increase the cache hits and make matrix multiplication go noticeably faster on benchmarks.
Some random details that may be relevant:
* When the processor loads a single number M[x][y], it sort of loads in the adjacent numbers as well. You need to take advantage of this.
* Something about row-major/column-major array is an important detail.
What I'm trying to say is, it is possible to indirectly optimize cache hits by careful manual hand tweaking. I don't know if there's a general automagic way to do this though.
This probably wasn't very useful, but I'm just putting it out there. Maybe more knowledgeable folks can explain this better.
As you said, being aware of this lets you optimize away cache misses by controlling the memory access pattern.
I used that approach once on a batch job that read two multi megabytes files to produce a multigigabyte output file. It gave a massive speed up on at 32-bit intel machine.
For x86_64 there are cache hints, no pinning/reserving parts of the caches (as far ad I know).
I wonder if Apple M1 or M2 cpu with unified CPU/GPU memory has anything like pinning or explicit cache control?
If the data is not contiguous it could make the CPU's life much harder.
There's also the matter of program size (the amount of instructions in the actual program) and whether the program does anything which forces it to go lower cache levels or RAM.
There are intrinsics for software prefetching such as __mm_prefetch, but those are difficult to use such that they actually increase you're performance.
Not for general purpose programs, because L1 caches change so quickly each year there is no point.
For embedded real-time processors, yes. For GPUs, yes. (OpenCL __local, CUDA __shared__).
This is because Microsoft's DirectX platform guarantees 32kB or something of __shared__ / tiled memory, so all GPU providers who want a DirectX11 certification are guaranteed to have that cache-like memory that programmers can rely upon. When DirectX12 or DirectX13 comes about, the new minimum specifications are published and all graphics programmers can then take advantage of it.
-------
No sane Linux/Windows programmer however would want these kinds of guarantees for normal CPU programs, outside of very strict realtime settings (at which point, you can rely upon the hardware being constant). Linux/Windows are designed as general purpose OSes.
DirectX 9 / 10 / 11 / 12 however, is willing to tie itself to the "GPUs of the time", and includes such specifications.
They allow you to buffer everyone's playing, at a user specified interval, then replays the last measure of music to everyone.
It's definitely not the same as live playing, but it's still pretty fun, and actually forces you to get creative on different ways.
The weird side-effects are why I guess the parent said that "it's still pretty fun, and actually forces you to get creative on different ways."
It works pretty well with well structured music like the blues. You probably couldn't play well if the piece was changing tempo or key all the time.
Big downside is you're stuck playing to a metronome, which would be enough for me to skip it, but it depends on the kind of music you're playing.
I could imagine that if the music is rhythmically slow and vague and improvised, big latencies are OK, and actually might yield some pretty interesting creative results.
Another model I've thought about is to structure players in a rooted DAG, and players can hear only people upstream of them.
E.g., you could build an orchestra by having a conductor and section leaders in a room together (or at within very low latency of each other). Other players could hear the leaders and play along, and then an audience could hear everyone. You could also do something more complicated like build things out in linear or power-of-2 layers, where each layer can hear everything upstream of it, and therefore many players would get a partial sense of the orchestral effect.
This could work nicely for improvised music, too, with causality preserved.
OTOH can the heart perform equally good if human get miniaturized?
Bay Area, 2020s: "Soon we will have rooms the size of a single computer"
The trips themselves of course consist on transmitting the mind to be downloaded on a new body on the other side.
And you're a mind running on a computer! You were already going to experience the heat death! This is just a scheme to get the most subjective time out of it as possible. (Running slower is more efficient.)
It's interesting to consider this paired against how technologically primitive we ostensibly must be, given that digital computers didn't even exist 90 years ago.
The only weird one is quantum entanglement, but even then information transfer doesn't travel faster than light and that's about all I know on that subject.
For anyone wondering why that is, I found this explanation super interesting:
https://www.forbes.com/sites/startswithabang/2020/01/02/no-w...
In short, the only information you gain is about the outcome of the other side's measurement. They cannot introduce information into the particle, and thus can't transmit information. The only thing you learn is what you already knew: the other particle had a 50% chance of being in one state or another, with the added fact that it's now correlated to your particle with a 75% (depending on experiment) chance. This is information that didn't exist until that moment, so it couldn't have been sent out and reach you before you measure your particle, which would inform you with 75% certainty how your particle would act and break causality.
The universe expands faster than light. So this expansion is not a physical effect?
So the processor in my hand can compute a multiplication fast than light can cross the room?
Hence the ratio between latency and compute has been changing exponentially. Even a linear or quadratic change would be dramatic, but exponential is something people just can't wrap their heads around. They're unable to really internalise it, in much the same way that in the early days of COVID people couldn't quite fathom how it is possible to go from 3-5 cases per day to tens of thousands.
HDD random I/O latencies are about 10 to 100x slower than a network hope. These days local SSD latencies are about 100x better than a typical network hop and this is just going to keep going. It'll soon be 1,000x better, then 10,000x, etc...
Any architecture using "remote storage" or "remote database calls" will be absolutely hamstrung by this. It'll be the equivalent of throwing away 99.99% or even 99.999% of the available performance.
People will eventually wise up to this and start switching over to distributed databases that run in the same VM/container as the application tier. So instead of "N" web servers talking to "M" database servers, it'll be N+M nodes with both components deployed into them.
Whatever argument can be made against this new architecture will become exponentially invalidated over time. Putting everything together is "too many GB of software to deploy"? Bzzt... we'll have 1 TB ram soon in typical servers. The CPU load of both together is too high? Bzzt.. the next EPYC CPUs will likely have 128 cores! Cache thrashing a problem? Bzzt... 1 GB and larger L3/L4 CPU caches are just around the corner.
Sometimes typos contain wisdom
Your estimate is off by 3 orders of magnitude.
No thanks.