The reason they do this is to keep the alb running efficiently by not having to architect for random large responses.
The reason they do this is to keep the alb running efficiently by not having to architect for random large responses.
API Gateway: 10mb
Lambda: 6mb
ALB: 1mb
Yet if your target is an ec2 instance, there's no limit on the load balancer. So IMO the limit for a lambda target should be the limit imposed by Lambda itself. 6mb.
ALB to containers or servers is a different beast - here the entire request and response need not be buffered at all (there might still be a very small buffer, mostly negligible), so streaming responses, websockets etc become possible.
We use lambda to resize images, so we do push against these limits a bit, but it's a fair tradeoff for the advantages - no worries about CPU throttling from too many requests, no waiting for servers to start for spiky loads etc.
The two styles are an impedance mismatch.
Each Lambda invocation gets a dedicated VM for the duration of the request. It is a great match for synchronous code.
Lambda does reuse VMs, so I hope you aren’t relying on containers being discarded for any integrity or security outcomes.
All the responses in this thread illustrate to me that AWS needs to put more effort into socialising how the product works. Since I was physically in the room for Lambda’s AWS internal launch this is twice disappoint because the technical messaging then was very clear and compelling.
Also request/response is not inherently synchronous or asynchronous. It's just a protocol design pattern.
Buffering, overcommit, etc are also kust normal facts of life in both sync and async messaging.
Request/response is fundamentally synchronous. If you want to nitpick about other layers not blocking, that’s missing the wood for the trees.