We have a similar application (parallelized C++ code operating on large files, for bioinformatics even) and ended up down the reverse path: started on Lambda, moved to our own RPC system. Lambda got super expensive (in part because there was no way to reuse a worker while it was doing async work like an S3 download), couldn't parallelize nearly enough without hitting AWS hard-limit quotas, had significantly lower CPU perf and didn't provide a way to cache files (like a genome in this case). Spot/preemptible instances keep our costs down while letting us keep a few hundred servers up at a time.