Anti-Patterns to Avoid in Lambda Based Apps
blog.basimhennawi.com
blog.basimhennawi.com
Some counterpoints:
1. Micro-lambdas
How many routes does your application have 3, 30? Trying to maintain 30 different codebases and deployments would be a nightmare.
Package size.
Package size is unlikely to be much different by adding a little extra code, especially if the code for each route uses the same set of libraries(which will make up the bulk of the package size).
Least privilege.
If it makes sense to limit a routes privileges, you can simply deploy the same lambda code separately, with a different role.
Upgrading.
On the contrary, upgrading a lambda that uses a single codebase, is much simpler than having to deploy multiple systems and deal with the issues of interoperability and compatibility between versions.
Reusing code.
Also much much simpler with a monolithic codebase, the code is right there, ready to use.
Testing.
Testing a monolithic app is very straight-forward, you don't need to setup a complex integration test environment, all the functionality can be tested locally.
4. Lambdas calling lambas
I tend to agree with this one, at least for synchronous calls, but you can also call another lambda asynchronously, actually this is something you tend to find yourself doing if you go down the path of breaking down a monolithic lambda.
5. Synchronous waiting.
Yes, from a cost perspective, you should minimize the amount of time your lambda is spending waiting on external IO. However, from an application complexity perspective, you should make things as simple as possible, building overly complex event based lambda flows, can make it really difficult to debug and maintain your application.
You have a single codebase, with modules for each route, and modules for shared infra/domain code.
Maven builds a JAR for each Route with only the resources that route needs, you deploy a single lambda for each JAR - but - you deploy them all together as part of the same build pipeline.
Then you have, non-monolithic lambdas, that can easily be performance tuned and privileged individually, but a single codebase that represents a whole 'microservice' application in one place, with shared code, that can be upgraded and kept consistant.
Sometimes it bites you in the ass though and you get cryptic errors when trying to update the app in cloudfront, but overall I think its worth it.
I was surprised to learn that Google Cloud Functions deploy a 1.5gb docker image. After that I was like fuck it, I don’t need to think about saving a few megabytes payload here and there.
I loathe blog post "anti-patterns" for this reason. Our job as engineers is largely about managing tradeoffs. And trade-offs don't come pre-packaged in a blog of dos and donts, they're specific to the constraints at hand.
Just about all of these "anti-patterns" are actually freaking awesome when what they offer (and do not offer!) lines up with what you need.
1. monoliths are awesome when that's all you need. Simple to write, maintain, and deploy.
2. Orchestration lambdas are A-OK when it's all you need and your workload fits comfortably in Lambda's constraints. I swear, the 'cloud' is rotting people's brains. There's a complexity threshold where Step Functions start to pay off. It's fairly high imo. The complexities in testing and debugging it introduces are huge compared to just dropping a break point in your control script.
3. This one is sane
4. Calling this an anti-pattern is again wrong imo. Goodness of fit depends on context. Lambdas calling other lambdas can squeeze a helluva lot of burst CPU performance for not a lot of money. There are a lot of workloads for which this makes sense.
5. This one is stepping over dollars to pick up pennies. I agree with your take entirely. Whatever minuscule amount is saved on idle CPU time with be dwarfed by the additional dev time spent writing, deploying, instrumenting, and debugging. Classic Solutions Architect wankery.
I believe the reason why Lambda has seen so little adoption is because cloud architects suggest to break existing services into a Lambda function for each endpoint. This causes an explosion of complexity and code that runs in the cloud in one way but it can't be reproduced locally on developer machine easily.
My advice is the opposite: DO USE MONOLOTHIC LAMBDAS.
Do you think monolitic Lambdas still make sense in that case?
Yes in my opinion the right division is
GOOD 1 lambda = 1 service
BAD 1 lambda = 1 endpoint
I'm eager to chat with you about CDK since it's kind of new and I don't know anyone else using it. Feel free to write me at giorgio DOT zamparelli AT gmail DOT com to chat about CDK, Lambda, serverless, Infrastructure as Code.
It's based on CDK and has a great local development environment for Lambda. It allows you to set breakpoints and test it locally: https://serverless-stack.com/examples/how-to-debug-lambda-fu...
I believe "4. Lambda functions calling Lambda functions" speaks to what I believe are the author's shortcomings for #2, with the caveat that "Cost" is definitely a consideration, as waiting for a waterfall of lambda calls will incur more cost. Call-it-and-forget-it is the more cost-efficient method for lambda-to-lambda. I trend more towards pub-sub for this scenario if the call-stack is more than 2 lambdas (e.g. lambda calling one lambda, end of stack). The author doesn't mention SNS curiously, which would suggest some inexperience in this theater. SQS can be used if payload process explicitly calls for the capabilities of SQS, but bare SNS is much more suited to light pub sub duty.
Overall, this seems like a rather reductionist article written for the SEO keywords rather than content. Anyone happening across this should deep dive "the why" behind each of the claims before taking it at face value. And I wish the article stated that at the very top.
Good practice for using external APIs is to NOT use any default http client settings and always provide your own timeouts to what you consider reasonable for connections and responses, as well as using a context system with deadlines so you can time out any requests that are taking more than a reasonable time to complete. Making your described surprise long expensive requests into nothing more than short errors (which hopefully you'll pick up on after a while, as long as you've got your alerting system setup right).
I agree, but that doesn’t solve the problem here. The remote API was asynchronous, I was just waiting for it to ack my instruction. Because of issues out of my control (network congestion, maybe?) the time to get a 200 OK from the server shot up.
> always provide your own timeouts to what you consider reasonable for connections and responses
Agreed here as well, and I was providing my own timeout. The problem is that (cost-wise) it’s fine for 1% of requests to hit a 5-second timeout, but gets expensive when 100% of requests do. And lowering the timeout means that during normal times, requests that have latency in the tail of the distribution but ultimately go through would fail, which is undesirable.
Presumably anything that can generate a "high bill" also has some reasonable level of on-call, alerting, SLAs etc ... such that if your Lambda invocations were all taking 30x longer people would notice.
The reason I went for serverless over a server in the first place was so that I wouldn't have to wake up in the middle of the night because of server issues, so alerting is not a satisfying solution. Especially where the problem is more an artifact of how the service is billed than how much it consumes actual server resources.
After moving from traditional servers to Lambda, we had lower hosting cost. After switching, it is easier to deploy new back-end features. It is easier to provision a sandbox for development.
Of course, there are many ways of achieving these goals. Kubernetes has great ways of doing the same with less vendor lock-in, but requiring more knowledge and care with the system components.
What's interesting about learning to push the right buttons on one company's black box, especially when I know it's likely powered by or equivalent to familiar F/OSS?
In practice, you spend more time writing lambda specific code and changing an obvious workflow to avoid DB access becoming a bottleneck or debugging why published events did not trigger the correct function, etc. that you wonder whether throwing cpu/memory/disk resources at a single instance tuned for the workload with dedicated local SSD storage might have been a better option especially as tasks around logging, persistent storage, debugging, profiling, error handling and getting stack traces are so much easier.
But clearly you do, except now the resources, permissions, versioning, and costs you have to think about are all specific to various parts of AWS, which are locked in, probably more expensive, and probably less familiar than their counterparts on operating systems or in containers.
Seems like a lot of work for a kind of scalability that leaves you with little insight into how it works, idk
I don't know a better dev environment than a (possibly scaled down) personal replica of production environment in the cloud. With proper tooling (e.g. Serverless Stack or SAM) you can achieve very fast code updates so the old argument of slow feedback due to having to deploy changes to the cloud on each iteration is getting less and less true as well.
With more traditional models already keeping your OS, possible container images, web server and any other middleware secure and up-to-date is pretty expensive if you want to do it properly.
Going all-in on serverless might not make much sense for a large software product company but when building bespoke business software it allows small teams to do wonderful things very cost-effectively.
Like, suddenly you can't just debug your code; debugging is magic and requires special tools. And you can't just use a recursive solution: recursion might literally cost you per call.
Not that the shift is bad, but just that our tools and thinking haven't quite caught up to it.
I get the same mismatch in a tiny way, when using Google colab to teach Python. It's very easy to get started, it solves lots of "local install" and versioning issues; but the keyboard shortcuts aren't quite there yet, and debugging might be scary.
Your lambda application should be runnable locally. Doing otherwise is an implementation choice. Some errors related to underlying OS considerations are harder to catch unless you spin up a Lambda container...but, that's just it. Spin it up locally and do your debugging. Not magical at all.
Recursion in an Lambda application is MORE powerful than recursion locally. With lambda, you have as much horizontal scale as you can pay for. Spin off 10,000 threads if you want. You're upset you have to pay for it? You always pay for it with hardware and electrical bills and time if not directly to AWS. My application uses a recursive lambda strategy as its core design.
What's magical is that all of your logs exist cleanly in Cloudwatch and retention policies are easy.
With EC2 or ECS backed by EC2 there is much less impedance mismatch between local dev environments and prod environments which results in less surprises.
Java, C#, etc. are also interpreted languages; they just-so-happen to have a separate compile-to-bytecode step (unlike e.g. Python, which compiles to bytecode when a file is first imported); and their interpreters just-so-happen to have very slow startup times.
"Compiled languages" can have very fast startup times, if we avoid languages like Java (which sacrifice a lot of startup time in favour of steady-state throughput; which is obviously inappropriate for a Lambda). If we do compile down to a native (x86_64 Linux) binary, it requires a custom runtime (e.g. if using Haskell, Rust, etc.).
Could you expand on that? I understand that both the JVM and Node are JIT compilers, but I don't understand why the JVM usually takes longer to start and what are the trade-offs involved here compared to Node.
There are many JVMs, each operating slightly differently, but the HotSpot JIT seems to be the most widely used (e.g. https://cl4es.github.io/2019/11/20/OpenJDK-Startup-Update.ht... )
Node runs on V8, which is described in more detail at https://v8.dev/blog/ignition-interpreter
Another big diffence is that Java does lots of special handling for classes, e.g. hooking up 'reflection' machinery, allowing dynamic 'class loaders', etc. which happens during an application's startup.
In contrast, Javascript 'classes' are basically just functions (constructors); and Javascript functions are values just like anything else; hence there's less of an up-front cost to JS code. Python works in pretty much the same way from a language level (but its CPython interpreter works very differently to V8).
V8 is highly optimized for fast startups because that's necessary in loading web pages.
JVM startup time is generally not an issue anywhere but Lambdas, also, JVM apps run for longer and so they want to take advantage of performance data to optimize over time.
If Oracle was desperate to make JVM start fast as a matter of life-or-death for an important revenue stream, then JVM would start quickly.
What the trade-offs would be I don't know, but keep in mind there might not have to be any, at least materially. Products are made to do certain things and not others, if something was never a requirement, nobody ever cared to think about it hugely material terms, well it's unlikely to happen.
FYI one simple tradeoff is that JVM might not compile code until it's run a bit and optimized whereas V8 might immediately compile some things meaning the later you get 'pretty fast right away' but miss out on the runtime characteristics that make it optimally fast, so in the former you pay a little bit of 'learning time' up front and then get better optimization.
[1] - https://www.graalvm.org/docs/getting-started/#native-images
But the point here is that JVM (and V8 actually) can do more than simple compiling by understanding the nature in which the program runs and therefore be quite fast often making up for the fact they are VMs.
I was surprised to find out that Common Lisp (SBCL) is so incredibly fast to start up that it can run faster than Rust even starting from source. And Lisp is the archetypical dynamic langauge! Made me question my entire career working with JVM-based languages.
I feel your comment fails to take into account the main reason teams stick with Java in AWS Lambdas: code and dependency reuse, and consequently turnaround time.
Arguing that nodejs and Golang are the ideal runtime in problems involving peeling tasks out of a monolith and into AWS Lambdas is a complete waste of time. Your goal is to offload processing tasks out of your service asap without wasting time developing and testing code written from scratch. Thus, you just create a project for your lambda, pick up your Java code that works and is well tested, offload the battle-hardened code and it's delendencies to the lambda package, add a handler to the lambda, add a lambda client to the monolith, sprinkle integ tests, and you're done: you have a production service. You're running a long-lived background task which is invoked only from time to time, and you don't really care if it takes 1 or 2 minutes to run.
Additionally, do you really want to force your team to manage two distinct tech stacks just because you want to offload a process out of a service? What will that cost you?
1. Python version: https://github.com/robhowley/lambda-warmer-py
I'm sure this is not an unpopular opinion anymore (not sure if it ever was) but:
1. if you build something new, starting with a monolith lambda is OK. Your customers care about a secure, user friendly, performant product, if a monolith lambda gets it, great!
2. your code can be in a monorepo but you can still have have multiple lambdas (running the same code but each with it's own least privilege IAM policy, it's own SLA / reserved concurrency / memory settings)
3. As soon as you have something that is indeed used by more than one service / team, that it can have it's own persistence / authentication / authorization, e.g. something if there was an API/SaaS out there you would "delegate" it to it, this is when you extract a microservice.
If you start with designing your microservice architecture for your side project / MVP or even seed stage startup, I applaud you, as you are succeeding in places many failed.
You grow into complexity, not start with it on day one.
I'm sure many agree with this sentiment but I keep seeing people getting an allergic reaction seeing any hint of mono* in any project.
> use AWS Step Functions to orchestrate these workflows using a versionable, JSON-defined state machine.
I think a better approach would be to use a proper workflow engine [0]. Using a standard modelling language like BPMN and the ability to monitor, audit, and optimize make this a better option than an Amazon-specific approach.
[0] https://www.infoq.com/articles/events-workflow-automation/
https://docs.aws.amazon.com/lambda/latest/operatorguide/anti... https://aws.amazon.com/blogs/compute/operating-lambda-anti-p...
For people with an established codebase, not wanting to go all in on lambda/serverless, or wanting to use existing frameworks, the trade-off might be worth it.
The author would problaly argue to use Step Functions in this case, but using "Event" can also work fine.
[1] https://boto3.amazonaws.com/v1/documentation/api/latest/refe...
Or trying to learn about the CLI (Command Line Interface or Common Language Infrastructure)
A group is trying to propose a standardized language for querying graph databases, ala SQL. They decided to call it Graph Query Language. This has nothing to do with the GQL spec, which itself has nothing to do with graph databases.
The biggest obstacles in this field really feel self erected sometimes.
The bigger problem comes from naive practitioners using a word and insisting to novices it has one exclusive meaning when it plainly doesn't.
"Lambda" is not a protected mathematical term. It's a Greek letter (and 500 other things besides)! The fact that 'lambda calculus' is a recognizable concept is itself a consequence of the language evolution you're complaining about.
Seems to do the trick.