Serverless LAMP stack
aws.amazon.com
aws.amazon.com
All I want is user-isolated (suexec) shared web hosting on scalable web servers and SQL servers that someone else runs for me, where the extent of what I need to do is put PHP/Python/shell/whatever files in a directory on some network filesystem, and optionally wire up FastCGI if I want. It works well, and it's been doable for years because I used to run infrastructure like this. In the modern era you can even imagine extending suexec to create containers and not just switch user IDs if you want more isolation - as long as you have the container image locally, the individual steps of making a container (make some namespaces, make some cgroups, set a seccomp policy) shouldn't be noticeably more overhead than exec / process startup itself, and, again, FastCGI is an option.
That's "serverless" just as much as anything else is (i.e., the person running the website doesn't need to think about launching and operating servers, the platform takes care of it) and it's way simpler than using all these fancy tools on top.
Scalable no-downtime MySQL is still a total bitch though.
I feel like everything that was cool in 2010 has reached the point where its functionality either converged with "old tech" or just died because there were too many problems with it.
Examples:
- JSON wanted to be simpler than XML. Now it has schemas, JsonPath, comments (but only supported by some parsers), and is still pretty verbose and has way too few data types (datetime anyone?)
- MongoDB wanted to be a document DB with shards and eventual consistency, now it also supports ACID-like operations
- JavaScript was cool because it was untyped, now TypeScript is the new cool kid on the block. It's only a matter of time until untyped programming is the new hype again
- SemVer got somehow rediscovered
- People made fun of Maven, nowadays everybody and their mother are pulling directly from npmjs.org
I don't think JavaScript has ever been cool. Before TypeScript, it was CoffeeScript that was cool.
Unfortunately however this mid-game between "traditional" and "scalable" makes it difficult to get people to move to a few things that would help which includes always using S3 so we don't need a central NFS server (which you lose a lot of performance to, even if it was just for some folders like the main code directories). And the reality is the majority of our customer base is on wordpress... and there are a lot of terrible wordpress developers which lead to all sorts of fun stuff.
Also PHP in general reloading the entire code base every page load is stable though somewhat intensive compared to a Rails-style application for people that want really high hit-rate sites but don't want to make everything static - for example someone with a shopping cart. (We have had customers exactly like that, wanting to hit a badly designed wordpress site with thousands of people at once in a rush sale).
So it is a thing.. but there are challenges. And honestly, surprisingly, getting people to move the needle a little from their basic cpanel setups is surprisingly almost harder than just changing platforms entirely because you have to educate people on how you want to be different to everyone else rather than selling a fun new solution.
It also wouldn't help much with the primary PHP challenge I've run across, that of a stateful Wordpress box (and to be fair, the article doesn't claim to help with that. It assumes a 12-factor app). Most PHP I've had to deal with is not stateless, thus making horizontal scaling a challenge. Luckily you can scale pretty far vertically with PHP as long as you haven't made Big O mistakes.
Usually, the state of PHP app is in the DB. Other "state data" are: 1) sessions. By default saved in files, can be moved to DB, making PHP stateless 2) storage (i.e. user's attachments). To decouple from PHP server, - need to switch to object storage service.
So, it's not hard.
But most popular PHP application - WordPress if failing to be stateless, because of number 2: it uses local storage for plugins, media, user's data.
If you are the developer of the app, that's true. If you're the ops guy to whom the app was lobbed over the wall to, it's not quite that simple without making code changes to the app. In the case of Wordpress also, it's been my experience that many of the "developers" are just marketing people and don't actually write code. These people are not equipped to remove state from the app either.
The RDS Proxy isn’t GA yet and has huge disclaimers on using it in production.
The only alternative is to use the RDS Data API which requires sending raw SQL queries over HTTP.
If you’re using an ORM getting the actual bound SQL query from the database request can be a bear.
And to make matters worse the RDS Data API doesn’t return JSON results with the tables column names. It puts the column names in a separate key that requires developers to map to generic column names using the index of the value.
AWS should be embarrassed about Aurora Serverless. No question in my mind.
Most of the required addons aren't in the "cloud model" where you pay for usage, but instead you pay to have them on regardless of usage.
You don't need to use DynamoDB Accelerator though, and it provides value you won't automatically get for free by using another database. If you don't use it you will be managing your own redis/memcache instance with all cache invalidation logic.
> Global Tables (cross region replication),
Again, you don't get this for free in any other DB. Setting up your own multi master cross regional database is not free.
Dynamo DB works fine without any of the above two features.
>It is best called "walletscaling.
And which DB out of curiosity scales without any load on your wallet? What mythical DB can one run which needs neither horizontal nor vertical scaling, thus not impacting the wallet.
Is multi primary the new term? Has an agreed upon term been decided yet? Django seems to be switching to primary/replica.
BigQuery includes all these things "for free" in the base price. It is pretty straightforward, with a price for storage and a price for querying (which you can choose to be usage or fixed). DynamoDB started off the same way, and added nickels and times by the roll.
DynamoDB gives you redundancy out of the box (your tables are replicated across the three availability zones in a region), the scale is available to you if you have sudden traffic in on demand mode or you can set a limit if you wish to manage costs; your queries may receive errors about being throttled at some point if you approach those limits.
For OLTP workloads, DynamoDB (and a lot of other NoSQL-style, cluster based databases) cannot be beat for performance, capacity, scalability and costs. Which is exactly what you want on the front line of a workload that can receive large amounts of traffic.
For OLAP workloads with unknown query patterns across a variable set of data that can change over time and large table scans, a relational database is king because the actual volume of traffic is low but the size of queries are a lot larger usually.
As to how that autoscaling performs, its instant and always available so you have nothing to worry about there. Rather spend your time focussed on optimisation of queries than managing the scalability of your datastore which is as it should be.
Can you give more details on what you mean by this?
>> RDS Proxy
TIL about RDS Proxy. Looks interesting for newer projects, hoping that it becomes GA in the coming months.
Separately, all applications don't scale linearly with the volume of open connections. There's typically a sweet spot of open connections that provides maximum throughput. Exceeding this sweet spot will actually reduce aggregate throughput. To the extent that serverless applications are stateless and are unable to pool connections as well as a stateful application, you should expect to see the volume of connections to grow proportional to load, potentially knocking over the backing DB (or in Aurora Stateless's case, returning max connections errors).
I'm not sure why this product hasn't become their premier RDS offering, it looks like it has the foundation to offload a good deal of operational complexity.
Looks like it costs $1 per 1 million HTTP calls. That does not sound very costly if you are talking about just side projects which may not generate close to a million requests?
You can also hook up an Application Load Balancer to Lambdas if that works for you: https://aws.amazon.com/elasticloadbalancing/pricing/
It still increases the cost of running Lambda a whole lot, but that seems pretty reasonable for the projects I'm running.
It’s a really nice setup, and obviously has tooling on top of AWS to make the management and deployment process handled nicely.
Zappa (python) has a deployment configuration that allows this. It's basically a Lambda that keeps itself alive all the time and for each request, fetches the SQLite DB from S3, does its transaction, and then puts the modified database back on S3.
The upside is it's basically free for low traffic read-only apps, the downside is the obvious problem of write conflicts if you have more than one write-capable user at any given time.
If you were to use the django test framework to generate a new SQLite DB on each request, you'd have what you're talking about.
One fewer network hop compared to DynamoDB, and for something that might get an update once a week or even once a month I get low latency without having to oversubscribe to another service.
Comparing a read of a database driven app (going to the URL and getting the admin login in Django via Lambda/SQLite/S3) to an async javascript submission (submitting a form on a cloudfront static site that POSTs to DynamoDB), the javascript/DynamoDB round trip is faster by about a full second.
I suspect it's because of the simple bulk of the Django deployment. Putting the whole bundle on a Lambda with all of its dependencies was about 45-46 megs of crap, whereas a simple Node DynamoDB insert is a couple dozen lines.
So ultimately, while spiffy to play with, I didn't bother to use it much due to performance. Although it has been awhile, I heard that Amazon made some Lambda changes recently to address initial wake-up request performance.
Even AWS salespeople will not give you a price because it heavily depends on your usecase.
I think for a lot of use-cases, a $5 droplet in DigitalOcean will be cheaper than paying for all these cloud APIs separately. (Although AWS does provide free-tier for many APIs your first year).
Now that I can afford more than $5 for side projects, I am looking for things like ease of deployment, reduced burden of maintenance etc. IMO serverless gives me this (once you have a proper setup of course). There are no machines to login to, no need to worry about patching hosts etc.
There was a time when my wordpress blog (RIP) used to get hacked every 6 months. A purely serverless model reduces the surface area for attacks.
I guess what I am trying to say is that for me personally, serverless is not about the immediate infra cost, but the TCO long term. And if/when your side projects take off, you can always go back to bare-metal servers as required.