In the end, we came to the conclusion that Amazon is smart and won’t let you hack together the equivalent of a cheaper EC2.
On the other hand, their incentive to solve the problem is relatively weak vs an on-premise alternative.
Life on a lambda is too short to pay 6-8 second startup penalty over and over millions of time.
[1]: https://github.com/libmir/mir-algorithm [2]: https://mratsim.github.io/Arraymancer/
Often it is a sign that the problem is not "big" enough (eg: not crunching truly large data sets) OR data science team gets disproportionate amount of goodwill (thus money) to spend on its foibles. :)
I came to this thread specifically to find out about numpy and pandas on lambda.
A similar method is described here: https://serverlesscode.com/post/deploy-scikitlearn-on-lamba/
Layers should make this entire process easier.
package: exclude: - venv/
would reduce the size considerably (to 50 MB in my case)