> Amazon S3 attempts to stop the streaming of data, but it does not happen instantaneously.
...which doesn't really explain it. It shouldn't send more than a TCP window after the connection is closed, and TCP windows are at most 1 GiB [1], usually much less, so this completely fails to explain the article's observed 3 TB sent vs 130 TB billed.
The article goes on to say:
> Okay, this is half the explanation. AWS customers are not billed for the data actually transferred to the Internet but instead for some amount of data that is cached internally.
In other words, how much they bill really isn't bounded by how much is sent at all. This is unacceptable.
I interpreted that to be that their code was doing this over and over again, so in total they retrieved 3TB over a set of requests. Still horrifying, but mildly more explainable.
One could disprove this explanation with a packet capture; just add the remaining window.
S3 is a block storage, so retrieving an object for such a high availability and high perf service means it tries to pull some X block of data and cache it before sending through the socket.
That X block of data is out of internal S3 storage, just not sent through the bigger Internet egress subsystem.
So technically aws may argue this is egress for s3, just not for aws
if you think egress is expensive, well storing data in RAM for cache purposes is 1000000x more expensive
a lot of stuff could be happening. Main problem is AWS (i think) is charging for egress out of S3 system, but customers are looking at their ingress at client side and there is mismatch
Charge for the buffer you fill: sure.
Charge the theoretical maximum: silly.
But egress fees only apply to S3 transfers outside AWS?
So which is it? Data transferred to the Internet? Or data processed internally?
Right now some class action lawyer is drooling.