> The 100 Gbyte file doesn't fit in the 48 Gbytes of page cache, so we have many page cache misses that will cause disk I/O and relatively poor performance.
This is the kind of thing that is becoming more and more common as literally no one wants to think about how to process anything that does not fit in memory.
> The quickest fix is to move to a larger-memory instance that does fit 100 Gbyte files. The developers can also rework the code with the memory constraint in mind to improve performance (e.g., processing parts of the file, instead of making multiple passes over the entire file).
It is not trivial for a team which has never thought about why things that can be done in constant memory footprint ought to be done in constant memory footprint to make this change. Ideally, your team will adopt the view that while slurping files might be OK in toy examples, if you start with a constant memory footprint goal, you eliminate a whole huge range of issues at the outset.
Your biggest problem then becomes trying to get everyone to see the value of this approach because they will never have experienced crashing machines, corrupted processing pipelines, sleepless nights, missed deadlines because every thing every step of the way wants to read everything into memory.
The whole encrypt/upload thing could be done in a single pass in fixed size chunks here, reading the file linearly (I am not sure why the S3 bucket is not encrypted or if it is what the advantage of the "double" encryption is). Incidentally, this would not even require that much extra programming.
The cache being full is exactly what you want. During a linear read, most of the reads will be satisfied from the cache. And, the OS will do a much better job of deciding how much of each file should remain in the cache etc.
Say you move this thing to a server with 128 GB memory. What happens when the service actually has to handle four uploads at the same time?