The Beautiful Math of Bloom Filters
nyadgar.com
nyadgar.com
The basic idea is that instead of table[hash1(x)] & table[hash2(x)] & ..., you calculate table[hash1(x)] ^ table[hash2(x)] ^ ... basically substituting XOR instead of AND. To construct the table, you need to solve a big system of linear equations. The various filter types change the parameters (mostly load factor of the filter) and indexing function (turning the hash into something where the bits you look up have correlated positions) in order to make structured equations that are easier to solve.
Or, worse, when people let the names of things keep them from learning them. Imaginary numbers being high on that list.
The real/imaginary dichotomy is the confusing part, having a rotational component doesn't make them any less real. O that reminds me, I'll just drop this link for my favorite video lecture on the matter full of visualizations, 13 parts, "Imaginary numbers are real" https://youtube.com/playlist?list=PLiaHhY2iBX9g6KIvZ_703G3KJ...
If I had it my way the distinction would be straight and twisted.
It's more organic. They just grow out of colleagues, speaking amongst themselves, referring to a fellow colleague's particular idea or elaboration. They all share a common base understanding of their field, many know of the colleague directly, and all know how to look something up if they know the colleagues name and gist of the idea. And so that's how they refer to it.
Once in a while, these so-named insights prove really important or lasting -- after the fact -- and the name continues to stick because it's the one everybody was using. Meanwhile, most of the time, the insights just kind of fade back into the baseline body of knowledge and either don't break out at all or evolve through some collaborative work that earns a more formal name.
Jokes apart, words are symbols that even if they have some semantics through etymology, in general they are quite arbitrary. I’d rather go with outlandish names that help mnemonics, if I were to choose. Names from people can serve that purpose; I still remember what a Kohonen map is, back from Uni, because of the childish resemblance with “cohone” (Andalusian for cojones), and a silly joke from a close friend.
It's like you get a choice between math hell or cartoon hell.
We do have another name for it. "Bit." You could probably roll out a new programming language today that uses something like `let shouldUpdate: bit = true;` or without blowing too much of your novelty budget. Or `u1`, if you wanted to allow arbitrary integer sizes.
> If you don't give your creations good names, they might name them after you.
From his tone it was clear that this was something to be avoided. I don't know whether too late to retcon existing names but let's try to do better going forward.
I don't care about the justification for the term "dynamic programming" nor for the term "memoization". They are just plain wrong.
Honestly compared to these two, "Bloom filter" sounds reasonable.
Heck, if "dynamic programming" had been called "Bellman programming" and "memoization" had been called "Mitchie memo" (by the name of their respective inventors), it would be less confusing.
So maybe, after all, that "Bloom filter" isn't that bad. Had we let the author pick a name, maybe he'd have picked "ephemeral spectrum" or something like that.
Yeah. I think naming after the name of the inventor is far from the worse actually. And it kinda gives credit where it's due too.
We can leave it as named after someone, but that is a lot easier to understand if you keep the possessive. Plank's constant is a good one, in that vein.