Microservices without the Servers
aws.amazon.com
aws.amazon.com
Also, Lambda-like options are available with most PaaS providers now.
It is not much different and might be easier than migrating web app from a custom PaaS. The only issue is, and has always been, the migration of data. And I don't see that getting solved until some startup writes a bunch of layers on top of a bunch of providers. It's very tricky, for good reasons.
Which other ones are there? I used to use PiCloud until they were bought out by dropbox and it atrophied. Shame, it was exactly what I wanted in a service.
We launched a month before Amazon Lambda, and have better features like full support for streaming HTTP.
Still in the process of establishing our service tiers, but our basic hosting plan starts at $5.00 per month.
IBM Bluemix - IBM Workload Scheduler, RabbitMQ (this is weak, but similar functionality can be achieved)
Microsoft Azure - Apache Storm (generally available now)
I was trying out a Gateway API + Lambda + DynamoDB setup in the hope that it would be a highly scalable data capture solution.
Sadly the marketing doesn't match the reality. The performance both in terms of reqs/sec and response time were pretty poor.
At 20 reqs/sec - no errors and majority of response times around 300ms
At 45 reqs/sec - 40% of responses took more than 1200ms, min request time was ~350ms
At 50 reqs/sec - v slow response times, lots of SSL handshake timeout errors. I think requests were throttled by Lambda but I would expect a 429 response as per the docs rather than SSL errors.
My hope was that Lambda would spin up more functions as demand increased, but if you read the FAQs carefully it looks as though there are default limits. You can ask these to be changed but that doesn't make scaling very realtime.
Given that DynamoDB can reliably write in the 4-5ms range, a kinesis queue may not be necessary. Unless the point of the Kinesis layer is to keep the cost of DynamoDB provisioning low?
If you are using "Node.js", you may be seeing slow times if you are not calling "context.done" in the correct place, or if you have code paths that don't call it.
Not calling context.done could either cause Node.js to exit, either because the node event loop is empty, or because the code times out, and Lambda kills it.
When node exits, the container shuts down, which means Lambda can't re-use the container for the next invoke and needs to create a new one. The "cold-start" path is much slower than the "warm-start" path. When Lambda is able to re-use containers, invokes will be much faster than when it can't.
Also, how are you initializing your connection to DDB? Is it happening inside your Lambda function? If you either move it to the module initializer (for Node), or to a static constructor (for Java), you may also see speed improvements.
If you initiate a connection to DDB inside the Lambda function, that code will run on every invoke. However, if you create it outside the Lambda function (again in the Module initializer or in a static constructor) then that code will only run once, when the container is spun up. Subsequent invokes that use the same container will be able to re-use the HTTP connection to DDB, which will improve invoke times.
Also, you may want to consider increasing the amount of RAM allocated to your Lambda function. The "memory size" option is badly named, and it controls more than just the maximum ram you are allowed to use. It also controls the proportion of CPU that your container is allowed to use. Increasing the memory size will result in a corresponding (linear) increase in CPU power.
One final thing to keep in mind when scaling a Lambda function is that Lambda mainly throttles on the number of concurrent requests, not on the transactions per second (TPS). By default Lambda will allow up to 100 concurrent requests.
If you want a maximum limit greater than that you do have to call us, but we can set the limit fairly high for you.
The scaling is still dynamic, even if we have to raise your upper limit. Lambda will spin up and spin down servers for you, as you invoke, depending on actual traffic.
The default limit of 100 is mainly meant as safety limit. For example, unbounded recursion is a mistake we see frequently. Having default throttles in place is good protection, both for us and for you. We generally want to make sure that you really want to use a large number of servers before we go ahead and allocate them to you.
For example, a lot of folks use Lambda with S3 for generating thumbnail images, sometimes in several different sizes. A common mistake some folks make when implementing this for the first time is to write back to the same s3 bucket they are triggering off of, without filtering the generated files from the trigger. The end result is an exponential explosion of Lambda requests. Having a safety limit in place helps with that.
In any case, if you are having trouble getting Lambda to scale, I'm happy to try and help.
I'm using Node.js, this is a gist of the Lambda function: https://gist.github.com/paulspringett/ec6d3df65e977342d6ea
I'm initialising the DDB connection outside the function as you suggest. However, I'm calling context.succeed() not context.done() -- would this be problematic?
I'll try increasing the "memory size" and requesting an increased concurrent request limit too, thanks.
I'll take a look tomorrow and see if I can reproduce what you are seeing. I'm not super familiar with API gateway, so there could be some config issues over there.
If you want to discuss this more offline, feel free to contact me at "scottwis AT amazon".
You can't have it both ways.
I suggest making that limits request, then retesting and reposting; otherwise, you've sold this benchmark short.
5 years ago people used to debate whether we should use a virtualized server vs. a physical one. You still can see similar discussions but rarely - we all have more or less agreed that using AWS/Rackspace/etc. is good for a business in majority of use cases.
I think 5 years from now we'll still be debating servers vs. services, but the prevailing wisdom will be that "services" have won.
There's still a ton of creating zip file artifacts of your lambda payloads (instead of pushing to a magic git repository that amazon controls, say), so there's a bit of "build monkey"ing to do instead of "devops"ery. But I think a lot of shops will be happy to make that trade, as "build" is closer to their core experience than "devops".
You gain vendor lock-in. You are now tied to the Amazon platform. If they shut down or suspend your account, for any reason, you are out of business. You are also paying premium for the platform, with the cost of devops built in.
I'll take an open ecosystem that gives me options to migrate my business anytime over a proprietary solution.
I think the open ecosystem approach will work well for apps/systems that are more established and have a predictable and sustainable user base and related revenue stream to maintain the system.
But for new ideas, early stage startups, the open eco-system approach will be a lot more work than using the high-level services providers like amazon are developing ...basically a no-stack approach can help one quickly find some product / market fit. Once some level of baseline level of utilization is understood then switching to a more difficult, but less-locked in approach would be smart.
all that said, there's probably a middle ground here too (but arguably not) -- where one deploy's using docker containers, private registry & image repo, s3 for static immutable storage and some open database like PostgreSQL on RDS. All that could be layered on top of AWS and other providers -- but it will still be a PITA to move the whole thing.
BTW -- take a look at the elasticbox product -- they really have an interesting approach to cloud that sort of obviates the lock-in issue by allowing one to build apps across clouds from the start using their "boxes" metaphor.
http://recode.net/2015/03/18/amazon-will-shut-down-amazon-we...
As shown in another thread here, this service does not infinitely auto-scale (the recommendation was to use Kinesis), so you still have to know which services to choose, which is pretty much a full-time job with the number of different services offered these days.
Edit: well, unless you run your php script in a PaaS like Google App Engine.
The smaller the scheduled unit of code is, the more densely they can pack the workloads and make more efficient use of their system ...squeezing out the pennies at their scale makes a lot of sense.
I've seen some vendor lock in comments here -- but certainly the big 3 or 4 service providers will have these features figured out soon - and seems to me some open source to schedule functions across compute will appear in no time that could be used by the smaller providers. replicating all the other features is much harder -- this is a great moat amazon has built.
Serverless implies that a server is not required after an installation step and that you only need your local device. Or perhaps too that you only need a collection of like clients to do something P2P.
Just because you as the developer don't have to think about the server doesn't make it serverless.
They say after all:
"and then unit and load tested it, all without using any servers."
...with a title of "Microservices without the Servers" when what they mean is
"without provisioning any servers yourself" which is obviously much less attractive.
My curiosity was peaked by the title the title, inserting "w/o provisioning" would have not gotten me anywhere near as interested in the topic.
The whole notion of a "virtual server" is basically like saying, "radio with pictures" or "horseless carriage". It only makes sense in a specific historical context. We're still figuring out what comes next, but I think it's safe to say that 20 years from now the main building block will be something different than the simulation of a 1970s university department minicomputer that we're all using now.
So you either spend very little time on it and build servers adhoc ("snowflake" style - SSH in, install some stuff, etc), or you spend precious time doing "the right thing" - which right now is a huge universe of options (Chef/Puppet/Ansible, Docker/other containers/no containers, etc).
If you're part of a larger team, not having a properly structured infrastructure is a nightmare - specially when it comes to scaling or dealing with failures of all kinds.
TLDR; - yes, I'd say it's somewhat hard...
I don't think it's hard, it's just time consuming. And we all know that "time is money", specially for small teams or solo devs (as you pointed out).
During the years I've used ssh, puppet, fabric, ansible, capistrano, cloud formation, etc for managing servers and infrastructure. And I think that the main benefit of any PaaS, AWS Lambda or AWS API Gateway is (obviously) that they're time saving and abstract the internals. In fact I use them in several small projects.
With virtual machines, we started to share physical resources which were otherwise wasted in dormancy. But even virtual machines were a waste of resources, as you need someone to set them up, keep them online and working. What if you could share the money you spent with those people and their knowledge?
Amazon just did that to you.
All some people want is to run code somewhere on the internet.
Amazon gather them together to share a server and an admin team.
In my point of view, Amazon Lambda is a huge improvement in efficiency for people around the world, as I (or someone else) can pay for a very tiny part of the work of a very specialized sysadmin and focus on my code.
If your scale is large, you have enough work to make full use of however many servers you rent.
Reducing infrastructure waste at small scale seems pointless to me. The actual cash money saving is just lost in the noise.
The only way i can see that i could be wrong is if there's a suppressed demand for huge numbers of tiny-scale services. Is there?
Investing in things like API Gateway and Lambda gives me the ability to keep my programming desires satisfied, while reducing the friction of tedious, time wasting operations, like automatically encoding a video of a class, extracting thumbnails, and setting permissions via API calls back to our home-brewed DAM system to make the video available for our students. Sure, I could do it myself each time. Six times per day of teaching. And turn down the process priority so it won't interfere with the rest of my job.
The costs of using AWS at that scale, while certainly nothing like the scale some of you are dealing with, are a considerable savings over the hourly rate I can charge for the time spent in my subject domain, and would have otherwise lost. We manage VMs for classes already, and most of that is now automated by our Hubot. In the coming weeks, the services/sites/apps we do have running will be migrated over to the docker-based infrastructure I'm building (in my copious free time, ha!), and the services therein will be interconnected with API Gateway and Lambda functions. It's a beautiful thing. I'm really quite pleased with it, and proud of what we've been able to do.
I'm know that I'm not alone, and maybe this post will encourage others in similar situations to share their experiences too. Maybe my peers don't frequent HN - that's an unknown to me, I guess. But there are far more entrepreneurs like me who enjoy the job we do, and will do everything we can to offload the time sucking administrivia to some other system, especially if we get to flex our programming muscles along the way.
Agreed. I think what it avoids though is setting up ansible + puppet + firewall + updates + monitoring + blah blah stuff you do to secure and dev-ops a prod app (for a baby app).
You can half-arse this stuff in a day for a new company. I assume the defaults of lambda will be a little more secure... but amazon security is pretty complicated to set up right, so not totally sure what the max win will be.
If this was close to true, AWS would not have existed. The thing is, when you're big, you're planning for spikes, and you're paying for your maxima per month. (Spare capacity aside, as that has to be a percentage of your desired monthly capacity.)
You could correctly argue that usage increases with scale.. so while you with 1 app can maybe only keep a server 20% busy, Amazon can afford to keep their servers 70% busy.. which is part of the story.
But the other part is you are paying for the convenience. You can get raw CPU way cheaper than what lambda gives it out at. Even if you only kept your CPU usage at 10-20% - for example Digital Ocean would still be cheaper than the Amazon Lambda app version. You are literally paying a surcharge to not have to set up your own servers, or deal with maintenance. I find it very unlikely you will save big money with lambda over your own servers.
However when talking about microservices, any non-trivial app quickly gets extremely complicated to setup and actually more inefficient in the amount of time it takes to execute each request. And probably more expensive since there still have to be servers somewhere, always ready to run your app and with all the management and orchestration layer on top.
Also VM technology is sufficiently advanced that they can be multitenant on the same base hardware very efficiently and consume almost no resources while in a sleep state. With live-migration and auto-scaling advancements it's not a big deal.
And that without having to set up virtual servers on an IaaS like EC2 isn't new, lies of people, big and small, were doing PaaS offerings long before Lambda.
Like when you say you have no carbon footprint because you don't own a car, even though you call a taxi every time you want to go somewhere?
Are microservices different from SOA? Or is it just a more modern, streamlined buzzword?
You say "microservices," but all I see is "omg, you realize inter-node latency isn't a trivial component to ignore when building interactive services, right?"
You'd be surprised how many people taking the "micro service" bait have never heard of an n-tier architecture or even SOA. I pity the next generation of developers that will have to maintain all that micro-servicing mess developed today.
Instead of relying on a simple architecture with messaging and workers, micro-service evangelists are turning maintenance and deployment of applications into a nightmare for the sake of being hip.
I won't even talk about amazon lambda which is the "acme" of vendor lock in .
Amazon's "Lambda" page (esp scroll down to the "benefits" https://aws.amazon.com/lambda/ ) shows it's more like offloading some tasks (like worker threads in the cloud).
I had a play with the second (linked) app, SquirrelBin http://squirrelbin.com/ which can edit and run javascript snippets. The latency is awful, 2-3 seconds for me (I'm in Australia, but that should only add 200ms roundtrip or so). They seem to spin up (reuse?) an entire instance for each request - it's incredible that it's as fast as it is.
But the problem is the architecture of this specific app: the delay would be fine if you could edit-run-loop code locally, without the cloud. But they wanted to demonstrate quick development (for them) by just making a CRUD app, using AWS Lambda existing http endpoints for PUT, POST, GET, DEL. So after editing you have to save, load and run - and each one interacts with the cloud. BTW the article about SquirrelBin https://aws.amazon.com/blogs/compute/the-squirrelbin-archite...
There are some clever platforms running on bare Xen (no direct OS) that can spin up an entire instance and destroy it on every request pretty quickly. http://erlangonxen.org is a great example. 100ms to boot your entire "system" for production usage.
cf push myapp
It figures out the language/runtime I'm using (Java, Ruby, Go, NodeJS, PHP), builds the code with a buildpack, then hands it off to a cloud controller which places it in a container. My code gets wired to traffic routing, log collection and injected services. I can deploy a 600Mb Java blockbuster using 8Gb of RAM per instance or I can push a 400kb Go app that needs 8Mb of RAM per instance.I don't need to read special documentation, I don't need special Java annotations.
I just push. And it just works.
I'm talking about Cloud Foundry. It runs on AWS. And vSphere. And OpenStack. It's opensource and doesn't tie you to a single vendor or cloud forever.
I worked on it for a while, in the buildpacks team, so I'm a one-eyed fan.
Seriously: why are we still talking about devops? It's a solved problem. Use Heroku. Install Cloud Foundry. Install OpenShift. And get back to focusing on user value, not tinkering.
Disclaimer: I work for Pivotal Labs, part of Pivotal, which donates the largest amount of engineering effort on Cloud Foundry (followed by IBM).
Granted I only spent a minute, but if this is a typical experience, I'm unsure how anyone would come to the conclusion that there's any software worth using there.
Unless you know where to find the docs[0], they're not obvious. There's a single master repo[1], but it's oriented at deployment and works by aggregating dozens of sub-projects[2] into a BOSH release and BOSH deployment.
... which requires you to know what the hell BOSH[3] is ...
So recently we started trying to make it easier. The best place to start tinkering is Lattice[4], which is a cutdown extract of Cloud Foundry. or Pivotal Web Services[5]. Or IBM BlueMix, I guess[6].
[0] http://docs.cloudfoundry.org/
[1] https://github.com/cloudfoundry/cf-release
[2] https://github.com/cloudfoundry and https://github.com/cloudfoundry-incubator
Cloud Foundry is a bear to install because you will probably wind up needing to wrap your head around BOSH, the IaaS orchestration tool. Once you get past that hump it's relatively obvious. Getting past the hump is tough.
Bear in mind that it's a complete PaaS. The kind of thing you bet your company on (and our customers do). BOSH is a heavyweight system that predates a lot of later tools like Terraform or Cloud Formation. On the other hand, we use BOSH to update Pivotal Web Services to the latest cf-release every 2 weeks or so and basically, nobody ever notices. It just works.
The easiest way to start is either Lattice or a public Cloud Foundry installation. The former has the advantage of being easy to install on a laptop, and it's intended for developers to tinker with. The latter has the advantage that someone else ran `bosh deploy` and is provisioning the VMs that Cloud Foundry runs on. Pivotal Web Service (based on AWS) and IBM BlueMix (based on SoftLayer, I think) are the two main ones.
So... it's not a solved problem after all? :)
You only have to install CF once, not every time you deploy. After that it's easy to upgrade. We do so on Pivotal Web Services every time cf-release is incremented, which is approximately fortnightly.
PaaS just makes life easier by reducing the amount of communication needed. It's more like a fancy wall that makes it easier for developers to throw code over it.
It makes Ops happy by isolating the damage Devs can do. It makes Devs happy by removing Ops roadblocks and allowing Devs to see, immediately and directly, what their apps are doing.
Classic ops culture emerged because computing was an expensive shared resource. Early tooling favoured utilisation over isolation because the latter is expensive. So it became possible for devs to accidentally or deliberately step on each other's toes. Ops was like the mature adult in a shared house: making sure everyone kept their junk out of the living room, restocking the bathroom, insisting that dishes get washed.
With a PaaS that changes radically. Devs can't burst out of the box they assign themselves, subject to Ops-set resource pools. Ops hands Devs keys to personal apartments that are created on demand. What they do inside is up to them, Ops needn't worry or care.
My main worry is not on the technical side but on how things are charged for. If I build something that starts to get used I am covered in terms of scalability, but not in a way that protects me from 'cost scalability' so to speak. I know I can set up billing alerts and hit a big 'shutdown' button in response to high load, but what I don't think I can do is throttle these services based on the money I want to budget/spend. With my own services I have a hard cost limit, with a hard scalability limit, or rather I just accept that my response times will go down or fail once I've allocated all I can afford.
If there something for AWS in terms of 'cost throttling'? It may be a gap in their services, especially for people want to build things that might get traction?
For Lambda style services it's odd because if I was hosting them in my own container (or an AWS one even) then they would still accept requests but their responses would start to slow (but still work). Trading through-put for cost/scalability?
The trouble with AWS services for that 'fear of success' type of budgeting (different from losing a key or malicious calls) is they are either on or off, with no in-between cost/latency/resource allocation ratio.
I'm surprised this hasn't come up more actually (or I just missed it), considering the overlap between devs that would like less devops and devs with limited budgets..
Is Amazon trying to extract a profit? Unless they are using specialized hardware, can they really execute this service any cheaper than you could? Or do they pad the billing to fund Amazon?
Why would I go with a likely much more expensive option (I'd be pretty surprised if Amazon's billings for this service scale linearly with their costs) instead of the option that provides more traditional cost scaling?
http://docs.aws.amazon.com/awsaccountbilling/latest/aboutv2/...
"I actually had a billing alert set, and I did get an alert, but it looked like this: "You are receiving this email because your estimated charges are greater than the limit you set for the alarm "awsbilling-AWS-Service-Charges-total" in AWS Account XXXXXXXX. The alarm limit you set was $ 10.00 USD. Your total estimated charges accrued for this billing period are currently $ 1050.95 USD as of Saturday 18 July, 2015 17:34:36 UTC." So, it came a bit too late to take action."
* Mostly OSS to avoid lock-in
* Git integration
* Full stack specification (OS, dependencies, etc.)
* Python/ES6 support (Ruby and PHP coming)
* Client libs so you can call your functions 'natively' in other languages.
It would be awesome to hear what people would like us to build for them. Here is a blog-post on how to build a PDF -> image converter: http://blog.stackhut.com/it-was-meant-to-be/
I know if it's free, then it's going to be under some kind of fair usage policy, and you're going to rate limit me or have some kind of restrictions eventually. There's no way it can be sustainably free if I start to push it really hard. I'd prefer to just know the limits upfront, or have some kind of usage based pricing.
We're going to add some better pricing. How would you like this to work?
- per month, flat rate, ups w/ usage - per request - per compute / storage
We really like the idea of only paying for the compute you actually use a la lambda; one of my gripes with Heroku was having to pay $x when the server was only in use for short bursts. Why should I pay for downtime?
That said, we've actually had many people say they would prefer per month, as it is more predictable and they are worried it could spiral out of control.
I would be super interested to hear your thoughts.
I personally think that is the correct kind of pricing for something like this; but monthly plans including X time/requests would likely be a good idea.
If you're on the public cloud, I'd argue that this is an even bigger problem as you're then relying on VPC (or the equivalent) to always work correctly.
Why not ignore the networking and just build in robust security? Pubkey authentication where possible, random long passwords where not? Retry limits for clients, network intrusion detection, etc. To me, relying on the network to keep you secure seems a bit like a crutch.
However realistically nearly all persistence services such as MySQL, Postgres, MongoDB, Memcache, ElasticSearch, etc either have been insufficiently hardened as a public service or flat out are not intended to be used on a public port and depend on the network for security.
There is not currently an option to connect an RDS database instance to a Lambda function without opening said database instance up to the public. It's a problem.
You are correct that SSH tunneling could be used to provide security but such usage is not yet a standard approach.
With traditional VPS you just point ansible/salt/puppet to new servers and you're good to go.
It took about 2 months with support (Business support) and finally they chose to close the account.
We created a new account with a new card and migrated our AWS infrastructure. Unfortunately we still have to use AWS..for now.
Plus, while it might be cool for some microservices, it seems like it would be a lot more unwieldy trying to do a full scale application.
API Gateway could be replaced by something like Open Resty.
S3 can be replaced by any other file storage solution.
You can also call native binaries so you therefore you can use languages like Go or C or anything else that statically compiles.
Check this out instead:
A write up and audio from that session is also available.
http://thenewstack.io/the-comparison-and-context-of-unikerne...
Here are the details on pricing:
https://aws.amazon.com/lambda/pricing/
You get up to 1M requests / month and 400,000 GB-seconds of compute time per month.
A default lambda function uses 128 MB of ram (0.125 GB), which gives you about 3.2 M seconds of compute time (time actually spent executing requests) every month for free.
Thus if you have functions that take 500 ms on average, and use the default amount of RAM, you can process about 6.4 M requests in a month for a total bill of $1.20.
Above the free tier limits you pay $0.20 per million requests, and $0.1667 per million GB-seconds.
The pricing is fairly attractive.
The site is served up via S3, and the back-end logic is a Lambda module that wraps a SOAP API.
It's a mapping template for the AWS API Gateway you can use to convert both HTML form POSTed data and HTTP GET query string data to JSON.
what is even more interesting is that you felt it worked nicely for expressing event handlers. Can you talk a bit more about that - very interesting to see why not something like python or ruby. I know that nodejs is a callback-oriented framework... was it the fact that you can test locally on nodejs consistently versus what would be the expected output on Lambda ?
For example: Say you've got an AngularJS app sitting in S3 or something, and your backend is a Node.js app running in Lambda. Google finds a link to "random-new-social-network.com/profile/drinchev" somewhere and tries to index it -- their request is routed to "random-new-social-network.com," where Angular recognizes "/profile/drinchev" as a route to a profile for some user named drinchev, pulls in the "profile" template, and spits out your profile, where Google can read it and call it a day.
If you're talking about search engines getting along with Javascript-reliant sites, that's a different story, but I don't think I see the problem.