Just to mention a modification on the "simple and intuitive cardinality estimator" that's far more accurate and actually used for Google's SZL:
Given N objects, hash each object and store the lowest M (s.t. M << N) hashes, then calculate what percentage of hash space is covered and compare to how much of the hash space would be expected to be covered.
This removes the issue of having to assume the N objects are distributed evenly as hash(object) should exhibit that property.
[1] My naive Python implementation for tests on the NLP WSJ corpus -- https://github.com/Smerity/Snippets/blob/master/analytics/ap...
[2]: Szl's implementation -- http://code.google.com/p/szl/source/browse/trunk/src/emitter...