A quick estimate based on prices quoted by Google Cloud Engine for Egress transfers ($0.08/GB) gives $183 million in yearly savings.
Surely Google pays much less than that for their own transfers, but even if we assume they are paying 1% of that price it is still respectable. Pretty sure they are paying more than that though.
EDIT: Found better estimates - Amazon Cloudfront is charging as low as $0.02/GB (on demand pricing), which translates to $46 million. Given very good knowledge Google has about their users and their vast infrastructure I'm guessing their costs are fraction of this. Anyone with domain knowledge that can comment on this?
It wouldn't surprise me if Google was scheduling the majority of updates to take place in off peak times, when the bandwidth would simply be otherwise unused. If so, the cost to Google would be a rounding error of your estimate based on GCE traffic prices.
Edit: bsdiff is documented to use memory 17x the file size.
It's especially a problem with AWS's tN.micro instances: when something thrashes them for long enough, they ratchet down to such a low CPU allocation that it will likely take hours before SSHD gets context-switched to enough times to start your process. By then, all sorts of timeouts will likely have killed your session.
Considering the number of AWS credentials out there, I can see malicious users logging in, forkbombing the instance, getting it to throttle back, and then walking away whistling until the person with dashboard login wakes up and can reboot the machine.
The system can easily slow to a crawl as memory accesses and allocations are constantly triggering reclaim and causing the system to thrash. The kernel has no way to know that it should prioritize the terminal you are trying to open to kill things rather than try its hardest to keep the offending application running.
The best way to prevent this, if you know an application will use lots of memory, is to run it in a memory cgroup. With a memory cgroup you can limit the applications total allowed memory, reserving some for the rest of the system.
Unfortunately, I don't know of a super nice CLI, but here's an example using cgcreate/cgexec (https://packages.debian.org/stretch/cgroup-tools):
$ # Create a memory cgroup named 'mygroup', allowing $USER to change the parameters and run tasks in the cgroup.
$ sudo cgcreate -a $USER -t $USER -g memory:mygroup
$ # Set the memory limit to 10MB
$ echo 10000000 > /sys/fs/cgroup/memory/mygroup/memory.limit_in_bytes
$ # Run a program in the cgroup
$ cgexec -g memory:mygroup lsThis is one of the problems multitask OS's were invented to solve. The answer is always let other tasks run too.
Speaking of: is there a good way to ask an OS to "reserve memory and CPU-time" for emergency interactive maintenance (e.g. spawning a console on a thrashing box), in a similar way to how ext2/3/4 "reserves space" for the root user for emergency maintenance?
The only real way I can think of to do this (that doesn't involve already knowing what processes could thrash and cgroup'ing them) is to run a hypervisor and treat the dom0 as the emergency console. That still doesn't let you get inside the thrashing instance to fix it, though.
* https://engineering.linkedin.com/blog/2016/08/don_t-let-linu...
It's HN. Do you even have to ask?
Thanks, truly, for creating such a tremendously useful tool.
Really awesome utility.
The key observation is that when performing updates, large runs of machine code will be identical, with only relative offsets and absolute addresses changing. Code changes tend to shift huge pieces of binary code in parallel (by the same offset), so the bytewise subtraction expresses these shifts as many repeats of the same bytewise difference, which is then more easily compressed by zlib DEFLATE (or maybe another compressor... I forget the details).
The flag will modify the storage format, -kinda packing files/chunk together instead of just compressing into a huge block- so that tools can detect diff and same sections over the full archive.
It's not well know but it is an absolute must known. e.g. for backup archive that goes over rsync, and change little over time.
I couldn't find the equivalent for 7z in a quick glance. It supports at least 5 compression algorithms so there is probably not a general option. Maybe one has to play with the advanced settings per algorithm, if they support it.
However, the resulting differences don't seem nearly as amenable to the bsdiff algorithm as the changes to the underlying uncompressed data. In almost all cases, you'll likely get better results with bsdiff on the uncompressed data (and bsdiff internally compresses its result).
Good deduplication and diff algorithms have adaptive block size detection.