Show HN: Zappa – Serverless Python/Django on AWS Lambda
gun.io
gun.io
Valve mentioned them on the official CS:S page, and things went haywire. The team restructured into a lambda friendly architecture, they scaled without breaking a sweat, and ended up paying pennies in costs.
Link: https://www.reddit.com/r/webdev/comments/3oiilb/our_company_...
But honestly, the big problem is the API Gateway. The product is a mess, the docs are a mess, and the whole system is kind of insane. Lambda is awesome, but API Gateway still seems half baked. 90% of my development time was fighting with APIGW. There are some kind of crazy hacks in there (base58 encoding cookies, regex based on b64 encoded status codes) that shouldn't have been necessary.
Still, now that the system works and it's easy to use, I think this is ready for real usage. And I'm sure Amazon will get their stuff together for future releases of APIGW. They probably weren't completely anticipating that people would use it in this way.
To head of a few questions at the pass:
Here are the hacks necessary to make this work: https://github.com/Miserlou/Zappa#hacks
Here's how to avoid the cold-start problem: https://github.com/Miserlou/django-zappa#keeping-the-server-...
To ensure that your servers are kept in a cached state, you can manually configure a scheduled task for your Zappa function that'll keep the server cached by calling it every 5 minutes.
The cost of running the warmer (300s * ~3M * 512MB) comes out to about $18 and that's not counting the number of actual requests from end users. It's interesting but costly as a substitute.
The monthly compute price is $0.00001667 per GB-s and the free tier provides 400,000 GB-s.
Total compute (seconds) = 3M (1s) = 3,000,000 seconds # ( roughly equals to 8640 * 300 ~ 2.6M seconds in the case of this app)
Total compute (GB-s) = 3,000,000 * 512MB/1024 = 1,500,000 GB-s
Total compute – Free tier compute = Monthly billable compute GB- s
1,500,000 GB-s – 400,000 free tier GB-s = 1,100,000 GB-s
Monthly compute charges = 1,100,000 * $0.00001667 = $18.34
The idea is that something like actually distributing the code to a front-line server seems to add to boot time, so if the lambda is not warm, it will be a little slower. If you keep it warm by having at least occasional traffic headed to the server, you're able to avoid this penalty.
If you look above, you'll see someone else did the math for you, and you're actually only talking about 864s of execution time, not 3M.
It also sort of looks like you just pulled up the pricing example which describes 3M requests that take 1s and worked back to it, because when you do the math you describe it works out to ~ 2.5M
Would love an AWS engi to shed some light on this though.
You will wind up with multiple containers in use at once if you have enough traffic so a call to the general pool that is running your Lambda function will most likely only keep one of the containers from recycling.
Disclaimer: I am a Lambda fanboy.
how are static assets served? through django?
To me, this sounds very much like reinventing the plain old CGI, just with different names ("Web server" -> "API Gateway", "CGI script" -> "server").
Am I missing something here?
Obviously there are still ways that this can be optimized since Django wasn't really designed to be used like this, but it still seems performant enough for me, and the other advantages gained are major.
Response times for a warm server are almost always <200ms, averaging just over 100ms (we did some tests on reddit yesterday.) In my own tests just now, I was getting <80ms response times consistently. And I'm certain there are ways that we can shave this down further.
* CGI is the closest UNIX thing you can find, fork a process, write on STDIN, read from STDOUT. Lambda actually is a framework for running javascript/java/python code, from which you can call actual binaries.
* Lambda's attractive point is that containers are reused, which means that instead of paying the full price of a new process your context is already "hot", even for the actual binary you run. That means less latency and less CPU usage, allowing you to scale far more easily
I'm no Lambda user so I'm possibly wrong, but the idea and execution behind sure looks nice.
Disclaimer: I am a heavy Lambda user.
However, being distributed and virtualized is a pretty big deal. In practice this means Lambda apps perform differently than traditional CGI, and also require different app architecture to support them.
Feel free to contact me on the address on my profile...
Anyone who wants to share their thoughts about this?
And it now supports Python. Would be interesting to see someone using AWS Serverless framework with Django Plugin support.
I used to use Django-nonrel for GAE so I wouldn't be logged in, and guess what: By the time I wanted to move away, I had so much GAE-specific behaviour that I pretty much had to rewrite the app.
I'm not guaranteeing that everything will work out of the box on the first try, but I bet you'll be very close. You will have to make a few design decisions, but if you're making the right ones it should just work. At the very least it should be far, far easier that GAE.
(The one thing you'll have to watch out for are C-extensions, which, for now, require that you do your deployment from an x86_64 machine.)
The two aren't mutually exclusive. My CDN points to Django for retrieving the static media.
> if you're making the right ones it should just work
"If you're making the ones that Zappa requires", you mean.
> The one thing you'll have to watch out for are C-extensions
Eh, that's to be expected, though.
Have you checked if Postgres works via the py-postgresql driver? It's implemented in only Python 3, with all C optimisations optional.
https://github.com/jkehler/awslambda-psycopg2
https://www.reddit.com/r/aws/comments/3on09a/using_psycopg2_...
2. what the hell is this http://imgur.com/a7FevNF? where's at least the close button? thank god ESC works.