How to estimate disk space
lethain.com
lethain.com
I was expecting a deep dive into how different OSes handle storage and indexing, which file systems/drive types make it easier or harder, the tradeoffs between truly random sampling versus a sampling scheme that takes into account typical drive fragmentation patterns and speed of access, and was very excited.
I'm hoping this comment will nerd snipe someone who likes to write.
Edit: to comment on the article, 2**10=1024 is handy to know so then 2**20 is about a million and 2**30 is about a billion. That then helps you estimate common log2 values which allows you estimate sorting, searching, and tree type structures that have some logarithmic aspect to their time or space complexity.
It makes no sense to count it in 1000-based SI units. You'll perpetually mis-align the data in whichever underlying storage technology you're using, and it's bad for performance and resource consumption.
Ask for disk sector alignment; we don’t have disk sectors any more, we use SSDs with blocks. And the sizes don’t need to have a neat number of kilobytes. Just the same as a litre of water doesn’t have a neat number of H2O molecules in it.