Powering CRISPR with AWS Lambda
benchling.engineering
benchling.engineering
It costs them less than a $100 a month.
It was written by an intern.
In America, this tech will be patented and regulated to within an inch of its life. I'm hoping China will be more sensible.
Actually, this technique is about 3 years old.
But in this case, it does seem incredibly advantageous to be able to scale up to any number of parallel searches, and to be able to search arbitrary new genomes.
For example, you cannot send dynamic response headers using the AWS API Gateway (the complementary service to expose HTTP endpoints). In my case I wanted to change the mime-type depending on JSON vs JSONP response.
It's also not possible to connect Lambda directly to ElastiCache and mostly you are expected to work with S3 or DynamoDB (Amazon's proprietary JSON store and what was mostly responsible for the data outage recently in US East). ElastiCache would allow easy persistence which is why it's surprising it can't be connected to given that it's an AWS service (you can connect to it by creating an EC2 proxy but that would defeat the purpose of a serverless architecture).
Some other oddities were sniffing the response body to set HTTP headers as opposed to just allowing your Lambda function to set the HTTP header directly or parsing the JSON response as opposed to doing a regex match.
API Gateway tries really hard to HIDE things from you. For instance, you can't see what the requested URL was without using a fair bit of VTL to put it back together from some other variables. Any only lately can get you a full list of query parameters, without having to specify them at the time of API creation. In fact, it seems like most of the work on API Gateway, since it's release, has been to let end-users have more access to data they hid in the first place.
I'm a huge fan of CRISPR. I've been following it closely since I heard Radiolab's podcast about it.
I'm also the founder of the JAWS framework, which is an open-source application framework built entirely on AWS Lambda and AWS API Gateway: https://github.com/jaws-framework/JAWS
I would LOVE to grab a coffee with you or anyone on your team some time, and chat about lambda or CRISPR, or anything really :) I live in Oakland and my email address is austen[at]servant.co
Also, will you be at Re:invent? I'm doing a breakout session on JAWS and I'll be there all week.
Good luck to you!
Austen
To improve performance, AWS Lambda may choose to retain an instance of your function and reuse it to serve a subsequent request, rather than creating a new copy. Your code should not assume that this will always happen.
Here's a template if it helps anyone:
I'm the creator of the project. We currently support 11+ programming languages.
We're adding support for more languages (JS & Python currently) but also have an entire build chain that let's you specify any Linux OS and language packages you want, and it's also trivial to embed your binaries and shell out to them if needed.
Using the new Lambda infrastructure, we pay for the number of Lambda invocations, the total duration of the requests, and the number of S3 requests. This comes out to $60/monthfor hundreds of thousands of CRISPR searches!"
Well, how much of that money you spent on EBS storage for your copies of genome data?
EC2 instances could read from S3 directly as lambda does, maybe that could alleviate the cost a lot.
Using AMI S3 backed instances could save a lot too.
But great work, nonetheless!
My friend is refactoring an app at his company right now, using only Lambda via JAWS and we ran some numbers on the cost savings. He's retiring 2 EC2 c3.large instances which were costing $2.97/day. On Lambda the app will cost $0.05/day.
We don't hear about it nearly enough yet, but the cost savings of building apps on Lambda are huge. Then you add in the time saved on devops... and you realize how seriously disruptive this tech is.
It's cheaper, because more efficient use of resources, because finer-grained. Each component of an app only gets what it needs; and the vendor can sell that unused capacity to someone else.
Geometrically speaking, finer grains pack tighter, wasting less space.
They also utilize multi-core effectively.
Anyway, benchling wants to avoid genome indexes from the sounds of it, in case users upload their own genomes. Having said that, if someone is doing multiple searches, it would quickly become more efficient to just index the genome. I would have thought most people seriously concerned about off target CRISPR hits would be using high quality reference genomes though.
http://www.technologyreview.com/news/537916/rebooting-the-hu...
Can't tell if that ticks your boxes or not. If you're interested we're free to try.
https://github.com/motdotla/node-lambda
I write/test locally and then deploy to AWS with a single command. The lock-in is helped by the fact that (a) the touchpoints and interface/interaction of AWS Lambda are pretty simple and (b) I could spin up a production version of node-lambda too.
The code deployment side is helped by using S3 as an interim place to upload and deploy packages from. The CLI makes that nice and easy once set-up.
On my projects that have been open sourced, we mostly open sourced to make debugging easier and make extension authoring easier. I've gotten comments from people that that makes them feel easier about vendor lock in, but honestly, I haven't seen many people try and stand up their own service. Would you say that matches your own expectations?
Most code is open-sourced at http://www.github.com/StackHut and we're working on making it easily deployable on your own hardware.
Take it from s/xxx/yyy/ into being /bin/sed. And then run the search in wetware.
Not even remotely close.