Thanks for sharing that!
Just skimmed through it and seems pretty interesting. I'll read it more in depth later.
Just skimmed through it and seems pretty interesting. I'll read it more in depth later.
I noticed that you used 5000 buckets to store the frequency of 7000 non-unique words in the section on 'Counting Bloom Filters'. How is that better than using 7000 buckets and a uniformly distributed hash function, which would maintain frequencies perfectly? We would be using fewer buckets by an order of magnitude in a real-world implementation to save memory.