Do you plan to document/opensource you work ?
Do you plan to document/opensource you work ?
Also we pack each individual hash before storing in Elasticsearch and gained at least 50% storage space this way.
Our Fingerprint data is quite different from theirs(unreliable ID3 tags, N versions of same track) which is why we needed some tweaks. So far the matching is still far from perfect...
Whether we will open source the whole thing at some point we don't know yet.
I know at some point we did adapt the Solr end (for example, we removed the N most occurring codes) for speed optimizations.
Many users of Echoprint in the wild have adapted the python matching logic for their use case as well as changed the hash update rate on the codegen. A great modification to watch was Sunlight Labs' "Ad Hawk", which ID'd commercials: https://github.com/sunlightlabs/adhawk