sort -u | wc -l but faster, with less memory, and slightly wrong
github.com
github.com
If your log lines look like
010.100.111.132 - - [30/Jun/2021:15:30:11 +0000] "GET /rest/... HTTP/1.1" 200 0 "-" "User Agent" "-" 212.031.212.003
010.100.111.132 - - [30/Jun/2021:15:30:10 +0000] "GET /rest/... HTTP/1.1" 200 0 "-" "User Agent" "-" 212.031.212.004
010.100.111.132 - - [30/Jun/2021:15:30:09 +0000] "GET /rest/... HTTP/1.1" 200 0 "-" "User Agent" "-" 212.031.212.003
010.100.111.131 - - [30/Jun/2021:15:30:08 +0000] "GET /rest/... HTTP/1.1" 200 0 "-" "User Agent" "-" 212.031.212.006
Then you may be tempted to just use some `awk '{print $1 " " $15}' | sort -u | cut -d" " -f1 | uniq -c ` to back out unique IPs per server over the last hour for a rough estimate.By using sketches, we can scale this solution to cases where the number of log lines would make `sort` fail due to OOM or take forever. Here, you'd run `dsrs --key`.