When you are going to fetch data from a remote location, you don't want to make a trip in vain.
So you end up storing them as a fixed size metadata chunk that lets you guess whether to go fetch it or not.
The neat trick is that the bloom filters have no false-negatives - the data might not exist (false positives), but it will never say "no" if it does exist.
But unfortunately, it doesn't really support deletion neatly - so it's really useful for scenarios where it's a first point of lookup before overloading a central source-of-truth.
The best use case I've seen for it is in Chrome, where the "Safe browsing" list is actually a huge bloom filter, which is used to decide whether to ask Google if this domain is safe.
So the list of banned URLs might be in the millions, but the bloom filter is a few megabytes and when it has a false positive, it goes & checks upstream whether it is indeed still banned/problematic.
[1] - http://www.slideshare.net/Hadoop_Summit/orc-2015-faster-bett... [2] - https://issues.apache.org/jira/browse/HIVE-11306
I believe browsers also store their Safe Browsing (anti-malware/phishing) blacklists in bloom filters.
When you check the Bloom filter it tells you:
1) it might be there
or
2) it definitely isn't there.
In the case of 2, you don't need to look it up. In case 1, you'll need to do the actual lookup.
It is commonly used to filter high volume / frequency requests for something. For example, if you have a list of banned IP addresses, user accounts, etc, you can quickly go through the bloom filter without hitting the database.
I can't actually remember what data was being looked up in the tables, though.
The service kept an array of 7 filters, rotated daily - the oldest would be cleared and reused for new items, giving us 6-7 days of history. Each individual header was low-value, and a few false positives every week wasn't a big deal - Usenet servers lost a lot more during their normal course of operation.
A certain product of ours keeps track of certain urls visited. We're talking millions of (unique) urls. We use bloomfilters to quickly check if a url was visited or not. If the bloomfilter search is positive a more expensive search inside a log file begins that gives a conclusive result (since bloomfilters have a (very) small false positive-rate, but we want to be perfectly sure).
One of the neat things about Bloom filters is that you can choose your own false positive rate, by tuning the number of hashes and the storage size.