Latency Numbers Every Programmer Should Know
eecs.berkeley.edu
eecs.berkeley.edu
Note: This is a huge gif, zoom in to the top.
[1] HN discussion: http://news.ycombinator.com/item?id=702713
If you had your swap partition set to a raid array of SSD's. Like two 64GB inexpensive SSD drives... I wonder how it would perform compared to RAM.
Getting 500-600MB/sec is "cheap". Getting 1200MB/sec gets a bit pricy but isn't hard (the setups we've tested "just" involved us popping 4x OCZ Vertex drives of various models into a server with SATA III, without much regard for anything else).
on an intel sandy bridge processor, each L1 access takes 4 cycles, each L2 access takes 12 cycles, each L3 access takes about 30 cycles, and each main memory access takes about 65 ns. assuming you are using a 3 GHz processor, this would make you think that you can do 750 MT/s from L1, 250 MT/s from L2, 100 MT/s from L3, and 15 MT/s from RAM.
now imagine 3 different data access patterns on a 1 GB array:
1) sequentially reading through an array
2) reading through an array with a random access pattern
3) reading through an array in a pattern where the next index is determined by the value at the current index.
if you benchmark these 3 access patterns, you will see that:
1) sequential access can do 3750 MT/s
2) random access can do 75 MT/s
3) data-dependent access can do 15 MT/s
you might guess that sequential access is fast because it is a very predictable access pattern, but it is still 5x faster than the speed indicated by the latency of L1. maybe you'd think it's prefetching into registers or something? but notice the difference between random access and data-dependent access. this is probably not what you expected at all! why is random access 5x faster than data-dependent access? because on sandy bridge, each hyper threading core can do 5 memory accesses in parallel. this also explains why sequential access seems to be 5x faster than the speed of L1.
what does this mean in practice? that to do anything practical with these latency numbers, you also need to know the parallelism of your processor and the parallelizability of your access patterns. the pure latency number only matters if you are limited to one access at a time.
Slides: http://research.google.com/people/jeff/Stanford-DL-Nov-2010....
Video: http://www.youtube.com/watch?v=modXC5IWTJI
Also, a previous thread on latency numbers:
CPU Cycle ~ 1 time unit Anything you do at all. The cost of doing business.
CPU Cache Hit ~ 10 time units Something that was located close to something else that was just accessed by either time or location.
Memory Access ~ 100 time units Something that most likely has been accessed recently, but not immediately previously in the code.
Disk Access ~ 1,000,000 time units It's been paged out to disk because it's accessed too infrequently or is too big to fit in memory.
Network Access ~ 100,000,000 time units It's not even here. Damn. Go grab it from that long series of tubes. Roughly about the same amount of time it takes to blink your eye.
Local gigabit network ping time is 200 μs and an SSD typically reads a 4K block in 25 μs. I imagine that converts to 200,000 time units and 25,000 time units, respectively.
Both transmission delay and propagation delay are measured in units of time -- I don't understand how it's inconsistent. If you'd like to change the payload size, I've made it very easy to do so -- simply change the html at the bottom of the page and it will be reflected immediately up top.
(1) Mutex performance isn't constant by OS. My own measurements tell me FreeBSD and Linux are around 20x slower than Linux locking and unlocking a mutex in the uncontended case.
(2) Contention makes mutexes a lot slower. If a mutex is already locked, you have to wait until it unlocks before you can proceed. It doesn't matter how fast mutexes are if there's a thread locking the one you're waiting on for seconds at a time. And even if other threads aren't holding the mutex long, putting the waiting thread to sleep involves a syscall, which is relatively expensive.
This lame comment comes up every time I see something like this posted. You're right. Knowing exact latency numbers is probably not going to change how you program, or how I program, or how most program. But why does it hurt to know? Why refuse to learn a few numbers that give some perspective on your system's limitations?
The other end of this is, for example, a programmer who is unaware that disk seeks take something like 100000x as long as accessing main memory.
(Anyway, the posted link is interesting since it shows these numbers over time.)
Depends on the usage pattern. Database servers generally have more memory than end user machines, but not proportionally more: A database serving a million users won't have a million times more memory. (This matters primarily when each user stores unique or low popularity data, otherwise the shared database has the advantage.) Moreover, data stored "on disk" on user machines will, with modern long-uptime large-memory desktops (and before long mobile devices), have a high probability of having the data cached and retrieved from memory rather than disk.
Thus, if the data is accessed a relatively short period of time after it is written (i.e. within a few hours), storing it locally on disk may be faster. (And storing it nominally on disk even though it's all cached may be preferable to storing it in strictly memory either because some minority of users have insufficient memory to store all necessary data in memory, or because the data should survive a loss of power or other program termination.)
Edit: It's also worth pointing out that latency isn't the sole performance criterion. Local disk generally has more bandwidth than network. If you're bandwidth constrained (either at either endpoint, the database or the user device), or either is getting charged by the byte for network access, that can be an important consideration.
The number on a round trip is actually kind of incorrect, because the number of routing stations it must traverse will dramatically impact the round trip time, so it isn't just a fixed value. It is a good ballpark though. I think the real impressive thing is how ethernet hardware has been designed well enough to be able to process billions of packets per second where they have to read the IP destination header, compute the next node in the transaction, and send it off. You can often end up doing that a half dozen times in relatively small trips but the individual time commitments are so much smaller than disk IO :P
Visualization for a nanosecond: a piece of wire 11.8 inches long, which is the maximum distance light/electricity can travel in 1 ns. A microsecond is a coil of wire 984 feet long. So if you're trying to understand why it takes so dang long to send a message via satellite or whatever, just understand that there's a whole bunch of nano/micro/milli-seconds between here and there. This also explains why computers must be small to be fast.
(Background: Rear Admiral Grace Hopper started programming on the Harvard Mark I in 1944, and wrote the first ever compiler in 1952 on the UNIVAC I. She earned the nickname "Grandma Cobol". Info via https://en.wikipedia.org/wiki/Grace_Hopper , which also summarizes this video in the "anecdotes" section.)
https://en.wikipedia.org/wiki/Betty_Holberton
>> in 1997, she received the IEEE Computer Pioneer Award from the IEE Computer Society for developing the sort-merge generator which, according to IEEE, "inspired the first ideas about compilation."
Whereas Hopper coined the term "compiler" and wrote the first one that's recognized as such:
https://en.wikipedia.org/wiki/History_of_compiler_constructi...
https://en.wikipedia.org/wiki/History_of_compiler_writing#Fi...
[1]: http://www.wolframalpha.com/input/?i=as+the+crow+flies+from+...
http://www.wolframalpha.com/input/?i=2+*+sin+%280.5+*+0.22+*...
Latency to screen display: approximately 70ms (between 3 and 4 display frames on a typical LCD).
http://en.wikipedia.org/wiki/Display_lag
Obviously, you only have 1 display latency in your pipeline but it's still typically the biggest single latency.
It would be nice if the source code of the page were not shown in a frame below though. That screen space is served better to show the rest of the actual diagram. You can already see the source code of a page with your browser, why waste screenspace and annoying extra scrollbars on that?
If you have any suggestions, a pull request would be much appreciated! https://github.com/colin-scott/interactive_latencies
We need somewhat accurate measurements on modern hardware as a base. I would not be surprised if some values have gotten worse over the years.
See the html at the bottom of the page for explanations of how I extrapolated.
If you have any improvements to make, a pull request would be much appreciated! https://github.com/colin-scott/interactive_latencies
What is the license of the code? I'd like play with visualization some day. Some of the bars is not very comprehensible due to height.
Pull requests would be much appreciated!
This almost seems like promotion.
His comment made me feel like maybe he thought I should feel bad for not knowing the information already.
Edit: seems like they can be produced with a high-intensity laser http://iopscience.iop.org/0954-3899/28/6/324/
In the experiment, they can generate pulses of 3 nanoseconds, each with up to 524 nanosecond gaps. This gives a 2MHz signal. However, not all of those pulses were detected.
If I read this right, "We report here on results based on part of the data taken during 2008 and 2009 runs, which amounted respectively to 1.78 and 3.52 × 1019 [protons on target], respectively. Among the 10122 and 21428 on-time events taken in 2008 and 2009, respectively 1698 and 3693 events were identified as candidate neutrino interactions in the target bricks."
Elsewhere, I found that the experiment ran from 27 May - 23 Nov 2009. I don't know if the beam was in continuous use for all 6 months. Let's assume it was full-time for only 1 month of that, which is unlikely.
There's a signal of some 3000 events per month, giving a transmission rate of 100 bits per day, or about 1 bit every 15 minutes. Rather a lot lower than 2MHz.
We would need some way which is about at least 1,000,000x better at detecting neutrinos before some company is willing to fund the Sydney/Wall St. neutrino beamline.
(If I'm wrong please correct me!)