This is really cool, thanks for the write up. I did the same thing in my homelab but much simpler, my shell script only fetched from a few sources and converted it into unbound format so it was much shorter. And I didn't care about statistics, but the idea was the same.
Also, why use ftp instead of curl?[1]