HNHacker News
TopNewBestAskShowJobs

andrewcanis

46 karma · joined April 6, 2011

Email: andrewcanis@gmail
submissionscomments
andrewcanis··on Amazon's empire rests on its low-key approach to AI
I'm very interested in this 6-page memo, I also found during my PhD that writing can really help clarify thinking. You may think you have a very clear idea in your mind but then when you write this to paper, many gaps can suddenly appear. I've looked all over online for an example of a memo but I couldn't find anything beyond a basic narrative structure.

You mentioned below that you have templated FAQ docs, I'm assuming these templates are the approximate structure of the memo? Would you be willing to share? This would be very helpful for my startup when we are planning our strategy.

andrewcanis··on An FPGA-based In-line Accelerator for Memcached (2013) [pdf]
Nothing was dialed back. My previous post just confirmed that our elasticache was not misconfigured. I used the latest github memcached with 14 worker threads as you suggested and your benchmark script gives the same results that we reported for the r4.4xlarge.

1) Actually your earlier test confirms our point that the CPU does not saturate the 10Gbps network with small value sizes. For example, in your 10B value example you got 15M req/sec with 16 cores. This rate is 1.2Gbps (15M x 10B * 8), well below the network limit of 10Gbps. The FPGA would still be ~8X faster at line rate.

2) Regardless of the get/set ratio the FPGA will hit line rate. However, thanks for pointing out that GET requests are much faster in the latest version of memcached, we didn't know that. Looks like the FPGA is only 3X faster for this 10:1 Get to Set workload.

3) Users would prefer if there was no pipelining at all. We only used pipelining to get around packet/sec limitations on AWS. If the FPGA was connected directly to the network we could hit line rate without any pipelining.

We aren't trying to mislead, we are just showing what's possible with the FPGA on AWS: line-rate processing of incoming requests at close to 10Gbps. The cool part as I mentioned is that the FPGA is still under-utilized so we could add encryption without affecting requests/sec at all because hardware cores execute in parallel. Another idea is to compress the data on the fly.

Agreed that RAM is the expensive part, which is why we picked a CPU instance that had similar RAM to the FPGA instance. Yes we have heard of users caching data to SSD to save cost.

andrewcanis··on An FPGA-based In-line Accelerator for Memcached (2013) [pdf]
Thanks for the details on how to use your benchmark script and for taking the time to investigate this. I hadn’t heard of your benchmark before and mc-crusher seems to work a bit differently than memtier_benchmark.

First a few significant differences:

1) Your value size is 10B which completely changes the results. Let’s keep the value size at 100B, which is more realistic.

2) The ratio of gets to sets significantly affects the requests per sec. We were assuming 1:1 ratio when we did our measurements. Increasing the percentage of gets really speeds up req/sec. We didn’t observe this effect on elasticache. Is this a recent improvement in the github version of memcached?

3) Your benchmark is using multiple keys in the same get command. What memtier does is pipeline multiple get commands each with one key. This seems more realistic.

4) We pipelined 16 get commands per packet while your configuration had 50 keys per get command.

I was able to reproduce the same setup as we had with ~1.2M req/sec with your mc-crusher benchmark using the following config. This has 1:1 get to set ratio with pipeline 16 and value size 100B.

send=ascii_set,recv=blind_read,conns=50,key_prefix=foobar,key_prealloc=0,pipelines=16,value_size=100 send=ascii_set,recv=blind_read,conns=50,key_prefix=foobar,key_prealloc=0,pipelines=16,value_size=100,thread=1 send=ascii_set,recv=blind_read,conns=50,key_prefix=foobar,key_prealloc=0,pipelines=16,value_size=100,thread=1 send=ascii_set,recv=blind_read,conns=50,key_prefix=foobar,key_prealloc=0,pipelines=16,value_size=100,thread=1 send=ascii_get,recv=blind_read,conns=50,pipelines=16,key_prefix=foobar,key_prealloc=1 send=ascii_get,recv=blind_read,conns=50,pipelines=16,key_prefix=foobar,key_prealloc=1,thread=1 send=ascii_get,recv=blind_read,conns=50,pipelines=16,key_prefix=foobar,key_prealloc=1,thread=1 send=ascii_get,recv=blind_read,conns=50,pipelines=16,key_prefix=foobar,key_prealloc=1,thread=1 send=ascii_get,recv=blind_read,conns=50,pipelines=16,key_prefix=foobar,key_prealloc=1,thread=1

I used the github memcached on an r4.4xlarge. I ran memcache-top on the server instance to measure the requests per second, showing about 750k gets/sec and 600k sets/sec.

With a ratio of 10:1 gets to sets I’m seeing about 3.5M req/sec which seems better than elasticache.

andrewcanis··on An FPGA-based In-line Accelerator for Memcached (2013) [pdf]
I'm still seeing similar results (~1M req/sec) after compiling your latest version of memcached from github and running with 16 worker threads. I just spun up two r4.4xlarge instances (one for client and one for the memcached server). I'm using memtier_benchmark with pipelining of 16 requests, 100B values, 10:1 get/set ratio. I compiled mc-crusher but you'll have to let me know the command to run because the readme wasn't clear.

One main constraint here is that we are using AWS virtual machine instances on the cloud. My guess is your previous experience is with physical servers. The FPGA performance is also significantly better when you can use the physical board with a direct ethernet connection, pipelining isn't required in this case the FPGA can handle minimum sized ethernet packets at line rate.

Another question, in your experience is compression/encryption used much with memcached? Because this is another area where the FPGA can compute much faster.

andrewcanis··on An FPGA-based In-line Accelerator for Memcached (2013) [pdf]
Our assumption was that elasticache would be highly optimized by Amazon. Remember that these are virtual machines which means limitations such as packet per second throttling. What specific configuration options do you think are missing?

In the latest version of memcached have you added support for batching/pipelining multiple requests per packet? Because this was crucial for achieving high requests/sec in this example.

Were the 55M requests/sec coming from another machine? Even with small 100B values you would need a minimum of a 44 Gbps network link. How many cores were required? In our benchmark we wanted a fair comparison between instances of similar price and RAM size.

andrewcanis··on An FPGA-based In-line Accelerator for Memcached (2013) [pdf]
Do you have a source showing elasticache running faster than this? For example, Redis labs was only able to achieve 10M req/sec by using 6 m4.16xlarge instances which are double the price of the CPU instance we used: https://dzone.com/articles/10m-opssec-1msec-latency-with-onl...

100-500 byte values are the majority of requests at companies like Facebook and Lyft for their key value clusters. For large value sizes the network interface becomes the bottleneck so FPGAs won’t be able to help.

andrewcanis··on An FPGA-based In-line Accelerator for Memcached (2013) [pdf]
Sounds interesting, are there any papers or public details that you could point me to with more technical information about this Alibaba RDMA-based memcached project? Alibaba also has FPGA instances available and we have been investigating their cloud offering.
andrewcanis··on An FPGA-based In-line Accelerator for Memcached (2013) [pdf]
Although there are Ethernet transceivers on the AWS FPGAs (these are Xilinx UltraScale+ VU9P FPGAs) they are unused and not connected directly to the AWS network. Instead the FPGAs are connected over PCIe to a host server, which has a standard NIC. This required us to use DPDK to bypass the kernel and pass raw network traffic directly to/from the FPGA over PCIe.

We have been told by cloud providers that the FPGA cannot be directly connected due to network security concerns. Since there is no easy way to control how the arbitrary hardware programmed by users on the FPGA will interact with their network. Microsoft has been using FPGAs directly connected to their network (called a bump-in-the-wire architecture) for the FPGAs used in their datacenter (see Project Catapult for details). But these FPGAs are not programmable by Azure users yet.

andrewcanis··on An FPGA-based In-line Accelerator for Memcached (2013) [pdf]
I wouldn't characterize Elasticache as running slow, a single instance in this case is handling 1.3M request/sec. But we can be 9X faster by batching multiple requests per packet and then offloading the TCP network stack and memcached operations to the FPGA. The FPGA allows us to handle the requests at network line-rate, even with small 100-byte requests. On Elasticache, past a certain point these small requests start to overload the CPU.

The interesting part is the FPGA could still do much more computation (for example, compression or encryption) while maintaining the same throughput due to hardware pipelining. We described this concept further in the blog post I linked to.

andrewcanis··on An FPGA-based In-line Accelerator for Memcached (2013) [pdf]
Interesting, what synthesis settings have you found have the most impact? I have also seen FPGA designers trying different seeds when closing timing. In this case, AWS provides an FPGA shell for external interfaces that has a maximum clock frequency of 250MHz. We have been able to meet this timing constraint without many issues. But we will keep you in mind for the Intel FPGA boards we are working with now.
andrewcanis··on An FPGA-based In-line Accelerator for Memcached (2013) [pdf]
Our startup is working on accelerators using FPGAs on AWS including memcached.

Using a single AWS F1 (FPGA) instance, our Memcached accelerator achieves over 11 million ops/sec at less than 300 microsecond latency. Compared to ElastiCache, the AWS-managed CPU Memcached server, our Memcached accelerator offers 9X better throughput, 9X lower latency, and 10X better throughput/$.

We need to batch multiple requests per Ethernet packet to get around packet per sec rate limiting on AWS. See more details here: https://www.legupcomputing.com/blog/index.php/2018/05/01/dee...

If anyone is interested we would love to hear from you, we will be showing off an online demo later this week.

FPGAs are great for processing data at 10Gbps line rate with low latency. They are also good for compute tasks like compression and encryption.

andrewcanis··on The Matrix in JavaScript in less than 600 bytes
If anyone is curious how the code actually works, I've commented up a version below:

    <script type="text/javascript">

    function matrix() {
        // initialize variables
        for (s = window.screen,
             w = q.width = s.width, // on my monitor: 1920
             h = q.height = s.height,  // on my monitor: 1200
             m = Math.random,   // random number from 0-1
             p = [],
             i = 0; 

             // i ranges from 0 to 255, one element for each character horizontally
             // this is enough characters to fill the entire screen horizontally
             // canvas won't let you draw off the screen - so I could set this to 1000
             i < 256;

             // initialize p (the y coordinate of each character) to start at 1
             p[i++] = 1);

        setInterval(
            // every time we call this function we draw the entire screen a very faint black (with a high transparency of 0.05)
            // this means every 33 milliseconds the screen is getting slightly darker
            // this also acts to darken and fade the green characters - when they are first printed they are dark green, then they slowly fade to black
            function() {
                // draw black (0,0,0) with alpha (transparency) value 0.05
                q.getContext('2d').fillStyle='rgba(0,0,0,0.05)';
                // fill the entire screen
                q.getContext('2d').fillRect(0,0,w,h);
                // #0f0 is a short form for color green (#00FF00)
                q.getContext('2d').fillStyle='#0F0';

                p.map(
                    // this function will be called 256 times - once for each element of array p, 
                    function(v,i){
                    // map over the array p
                    //      v is the value in the array p, which represents the y-coordinate of the text going down
                    //      i is the index of the array p, which represents the x coordinate
                    // start from unicode char code 30,000 (0x7530) then add a random number from 0-33
                    // from wikipedia: http://en.wikipedia.org/wiki/List_of_CJK_Unified_Ideographs,_part_2_of_4
                    //      U+753x 	田 	由 	甲 	申 	甴 	电 	甶 	男 	甸 	甹 	町 	画 	甼 	甽 	甾 	甿
                    //      U+754x 	畀 	畁 	畂 	畃 	畄 	畅 	畆 	畇 	畈 	畉 	畊 	畋 	界 	畍 	畎 	畏
                    //      U+755x 	畐
                    randomNum = m()*33;
                    // note how the asian characters are slightly different shades
                    // of green, this depends on their line thickness etc, and doesn't
                    // really happen for english characters
                    randomAsianChar = String.fromCharCode(30000 + randomNum);

                    q.getContext('2d').fillText(
                        randomAsianChar, 
                        i*10,   // x coordinate - each character is 10 x 10
                        v       // y coordinate 
                    );
                    // draw at least 758 characters down before reseting to the start
                    minimumHeight=758
                    num = minimumHeight+m()*10000;
                    p[i] = (v>num) ? 0 : v+10   // increment the y coordinate by one character (10 pixels), reset when y-coord gets too big
                    })
            },
            33) // call every 33 milliseconds
    }
    </script>

    <body style=margin:0 onload="matrix()"><canvas id=q>
andrewcanis··on Support requests: How software gets better
Any recommendations for keeping track of email support requests? I'm getting beyond the point where I can just reply straight from gmail. Ideally something I can just self host to avoid a monthly fee, but I'd be open to paying for a really good solution. Has anyone had any luck outsourcing support questions? Especially for support emails in the middle of the night, it would be nice to have someone answer basic questions and then escalate the difficult ones to me.
andrewcanis··on Show HN: Themes for Bootstrap
Great idea! I would love a valentines day bootstrap theme.
andrewcanis··on Notch's New Game - Minicraft (And An Android Port)
It's actually very easy to hack the code. I've created a "God-mode" version that you can play on github: http://acanis.github.com/Minicraft-God-Mode/
andrewcanis··on Notch's New Game - Minicraft (And An Android Port)
Try this version on github, you can compile and run the source with 'ant run': https://github.com/skeeto/Minicraft
andrewcanis··on Trader: I dream of another recession (and Goldman Sachs rules the world)
Whenever I hear doomsayers like this I try to remember Warren Buffett's advice, "Be fearful when others are greedy and greedy when others are fearful."
andrewcanis··on Google lobbies Nevada to allow driverless cars
This technology has the potential to save so many lives. Does anyone know how much these cars cost? I would expect the radar system and servers to be pretty expensive. Presumably the higher cost could be offset by lower insurance rates.
andrewcanis··on Building a Web Application that makes $500 a Month – Part II
This is really good to know. I always seem to come up with ideas that have already been done. But it's possible my new approach would be better than the competition, like TweetingMachine.