You mentioned below that you have templated FAQ docs, I'm assuming these templates are the approximate structure of the memo? Would you be willing to share? This would be very helpful for my startup when we are planning our strategy.
46 karma · joined April 6, 2011
You mentioned below that you have templated FAQ docs, I'm assuming these templates are the approximate structure of the memo? Would you be willing to share? This would be very helpful for my startup when we are planning our strategy.
1) Actually your earlier test confirms our point that the CPU does not saturate the 10Gbps network with small value sizes. For example, in your 10B value example you got 15M req/sec with 16 cores. This rate is 1.2Gbps (15M x 10B * 8), well below the network limit of 10Gbps. The FPGA would still be ~8X faster at line rate.
2) Regardless of the get/set ratio the FPGA will hit line rate. However, thanks for pointing out that GET requests are much faster in the latest version of memcached, we didn't know that. Looks like the FPGA is only 3X faster for this 10:1 Get to Set workload.
3) Users would prefer if there was no pipelining at all. We only used pipelining to get around packet/sec limitations on AWS. If the FPGA was connected directly to the network we could hit line rate without any pipelining.
We aren't trying to mislead, we are just showing what's possible with the FPGA on AWS: line-rate processing of incoming requests at close to 10Gbps. The cool part as I mentioned is that the FPGA is still under-utilized so we could add encryption without affecting requests/sec at all because hardware cores execute in parallel. Another idea is to compress the data on the fly.
Agreed that RAM is the expensive part, which is why we picked a CPU instance that had similar RAM to the FPGA instance. Yes we have heard of users caching data to SSD to save cost.
First a few significant differences:
1) Your value size is 10B which completely changes the results. Let’s keep the value size at 100B, which is more realistic.
2) The ratio of gets to sets significantly affects the requests per sec. We were assuming 1:1 ratio when we did our measurements. Increasing the percentage of gets really speeds up req/sec. We didn’t observe this effect on elasticache. Is this a recent improvement in the github version of memcached?
3) Your benchmark is using multiple keys in the same get command. What memtier does is pipeline multiple get commands each with one key. This seems more realistic.
4) We pipelined 16 get commands per packet while your configuration had 50 keys per get command.
I was able to reproduce the same setup as we had with ~1.2M req/sec with your mc-crusher benchmark using the following config. This has 1:1 get to set ratio with pipeline 16 and value size 100B.
send=ascii_set,recv=blind_read,conns=50,key_prefix=foobar,key_prealloc=0,pipelines=16,value_size=100 send=ascii_set,recv=blind_read,conns=50,key_prefix=foobar,key_prealloc=0,pipelines=16,value_size=100,thread=1 send=ascii_set,recv=blind_read,conns=50,key_prefix=foobar,key_prealloc=0,pipelines=16,value_size=100,thread=1 send=ascii_set,recv=blind_read,conns=50,key_prefix=foobar,key_prealloc=0,pipelines=16,value_size=100,thread=1 send=ascii_get,recv=blind_read,conns=50,pipelines=16,key_prefix=foobar,key_prealloc=1 send=ascii_get,recv=blind_read,conns=50,pipelines=16,key_prefix=foobar,key_prealloc=1,thread=1 send=ascii_get,recv=blind_read,conns=50,pipelines=16,key_prefix=foobar,key_prealloc=1,thread=1 send=ascii_get,recv=blind_read,conns=50,pipelines=16,key_prefix=foobar,key_prealloc=1,thread=1 send=ascii_get,recv=blind_read,conns=50,pipelines=16,key_prefix=foobar,key_prealloc=1,thread=1
I used the github memcached on an r4.4xlarge. I ran memcache-top on the server instance to measure the requests per second, showing about 750k gets/sec and 600k sets/sec.
With a ratio of 10:1 gets to sets I’m seeing about 3.5M req/sec which seems better than elasticache.
One main constraint here is that we are using AWS virtual machine instances on the cloud. My guess is your previous experience is with physical servers. The FPGA performance is also significantly better when you can use the physical board with a direct ethernet connection, pipelining isn't required in this case the FPGA can handle minimum sized ethernet packets at line rate.
Another question, in your experience is compression/encryption used much with memcached? Because this is another area where the FPGA can compute much faster.
In the latest version of memcached have you added support for batching/pipelining multiple requests per packet? Because this was crucial for achieving high requests/sec in this example.
Were the 55M requests/sec coming from another machine? Even with small 100B values you would need a minimum of a 44 Gbps network link. How many cores were required? In our benchmark we wanted a fair comparison between instances of similar price and RAM size.
100-500 byte values are the majority of requests at companies like Facebook and Lyft for their key value clusters. For large value sizes the network interface becomes the bottleneck so FPGAs won’t be able to help.
We have been told by cloud providers that the FPGA cannot be directly connected due to network security concerns. Since there is no easy way to control how the arbitrary hardware programmed by users on the FPGA will interact with their network. Microsoft has been using FPGAs directly connected to their network (called a bump-in-the-wire architecture) for the FPGAs used in their datacenter (see Project Catapult for details). But these FPGAs are not programmable by Azure users yet.
The interesting part is the FPGA could still do much more computation (for example, compression or encryption) while maintaining the same throughput due to hardware pipelining. We described this concept further in the blog post I linked to.
Using a single AWS F1 (FPGA) instance, our Memcached accelerator achieves over 11 million ops/sec at less than 300 microsecond latency. Compared to ElastiCache, the AWS-managed CPU Memcached server, our Memcached accelerator offers 9X better throughput, 9X lower latency, and 10X better throughput/$.
We need to batch multiple requests per Ethernet packet to get around packet per sec rate limiting on AWS. See more details here: https://www.legupcomputing.com/blog/index.php/2018/05/01/dee...
If anyone is interested we would love to hear from you, we will be showing off an online demo later this week.
FPGAs are great for processing data at 10Gbps line rate with low latency. They are also good for compute tasks like compression and encryption.
<script type="text/javascript">
function matrix() {
// initialize variables
for (s = window.screen,
w = q.width = s.width, // on my monitor: 1920
h = q.height = s.height, // on my monitor: 1200
m = Math.random, // random number from 0-1
p = [],
i = 0;
// i ranges from 0 to 255, one element for each character horizontally
// this is enough characters to fill the entire screen horizontally
// canvas won't let you draw off the screen - so I could set this to 1000
i < 256;
// initialize p (the y coordinate of each character) to start at 1
p[i++] = 1);
setInterval(
// every time we call this function we draw the entire screen a very faint black (with a high transparency of 0.05)
// this means every 33 milliseconds the screen is getting slightly darker
// this also acts to darken and fade the green characters - when they are first printed they are dark green, then they slowly fade to black
function() {
// draw black (0,0,0) with alpha (transparency) value 0.05
q.getContext('2d').fillStyle='rgba(0,0,0,0.05)';
// fill the entire screen
q.getContext('2d').fillRect(0,0,w,h);
// #0f0 is a short form for color green (#00FF00)
q.getContext('2d').fillStyle='#0F0';
p.map(
// this function will be called 256 times - once for each element of array p,
function(v,i){
// map over the array p
// v is the value in the array p, which represents the y-coordinate of the text going down
// i is the index of the array p, which represents the x coordinate
// start from unicode char code 30,000 (0x7530) then add a random number from 0-33
// from wikipedia: http://en.wikipedia.org/wiki/List_of_CJK_Unified_Ideographs,_part_2_of_4
// U+753x 田 由 甲 申 甴 电 甶 男 甸 甹 町 画 甼 甽 甾 甿
// U+754x 畀 畁 畂 畃 畄 畅 畆 畇 畈 畉 畊 畋 界 畍 畎 畏
// U+755x 畐
randomNum = m()*33;
// note how the asian characters are slightly different shades
// of green, this depends on their line thickness etc, and doesn't
// really happen for english characters
randomAsianChar = String.fromCharCode(30000 + randomNum);
q.getContext('2d').fillText(
randomAsianChar,
i*10, // x coordinate - each character is 10 x 10
v // y coordinate
);
// draw at least 758 characters down before reseting to the start
minimumHeight=758
num = minimumHeight+m()*10000;
p[i] = (v>num) ? 0 : v+10 // increment the y coordinate by one character (10 pixels), reset when y-coord gets too big
})
},
33) // call every 33 milliseconds
}
</script>
<body style=margin:0 onload="matrix()"><canvas id=q>