71 karma · joined November 14, 2014
- Exact time for analysis can be passed
- Added additional third party fingerprinting technology(doreso) to find snippets echonest can't identify
- false positives reduced drastically as we can be stricter about thresholds due to above point
- works much better on mixtapes with altered BPMs due to above 2 points
- Finds and embeds additional sources for identified songs
Our matching algorithm is based on the open source echoprint-codegen fingerprinting method, which we have built our own stack around:
- Replaced Solr/Tokyo Tyrant with Elasticsearch
- Reimplemented matching-logic
- Crawlers search multiple sources for audio files to be indexed (mp3s arent stored long term, only fingerprinted then deleted)
- Indexing about 1 new track per second
- Found method to verify unrealiable ID3 tags (in progress, current database also includes unferified)
- mogilefs as primary data store for fingerprints
- perl everything
We also provide a free music identification API.
Any feedback would be much appreciated!
I try to code as much as possible as early as possible. I throw away lots of stuff and recode it. Besides that I have an eye for stuff that is "similar" and can be abstracted. If someone wants an estimate, I guess as good as possible.
Big code is idealy split into one-person-chunks each with a documented API, but sometimes many people have to work on the same "files". Then big code is split between multiple people that sit nearby and communicate personally while discussing implementations based on technical arguments.
How to make a product of software is a different story. But I guess it works when you design your product in estimateable pieces and adapt fast to changing requirements.
Also I am pretty sure I forgot one or two things...