This is an awesome writeup, thanks for sharing!
A couple thoughts occurred to me as I read the post:
- Lambda functions deployed using Docker images can be up to 10GB.[1] Would that change your math here? I'm curious what the tradeoff would be vs parallelizing more function executions searching smaller datasets on both cost and performance.
- Great notes on the anti-competitive nature of the current market. If there was an open standard on crawling, maybe we'd see more innovation here.
- Cool use of a bloom filter!
1. https://aws.amazon.com/blogs/compute/working-with-lambda-lay...