Show HN: Imagededup – Finding duplicate images made easy
github.com
github.com
It includes several hashing algorithms (PHash, DHash etc) and convolutional neural networks. Secondly, an evaluation framework to judge the quality of deduplication. Finally easy plotting functionality of duplicates and a simple API.
We're really excited about this library because finding image duplication is a very important task in computer vision and machine learning. For example, severe duplicates can create extreme biases in your evaluation of your ML model (check out the CIFAR-10 problem). Please try out our library, star it on Github and spread the word! We'd love to get feedback.
The idea is that the activations within an image recognition network are similar for images that are similar and so you can measure the distance between two images in a space that has some semantic meaning.
Can you elaborate a little bit on the performance you've observed? Can it work iteratively, basically asking one image at a time "have we already seen this?"
Would it work for text documents of varying quality aswell or is this unfeasible?
What is the scaling like? E.g. what if it was 10 million?
How much memory would you need for ~2000 images, how slow does it get, etc.
Thx
"Doesn't sound too bad."
I doesn't? As a word of caution: If a 10k dataset takes a few minutes I would be careful how long 10MM pictures take. I predict it does not take 10MM/10k times a few minutes.
We also have an image heavy production use case that would be able to yield some nice metrics from this tool.
cdef extern int __builtin_popcountll(unsigned long long) nogil
dist = __builtin_popcountll(key ^ phash)
It would only take a couple of minutes to fill out the rest.
Anyone can just simply fork this kernel and add their dataset to it and perform the task!
install_requires=[
'numpy==1.16.3',
'Pillow==6.0.0',
'PyWavelets==1.0.3',
'scipy==1.2.1',
'tensorflow==2.0.0',
'tqdm==4.35.0',
'scikit-learn==0.21.2',
'matplotlib==3.1.1',
],
A while ago, I asked about sth like this (or more about the underlying methods) here on SO:https://stackoverflow.com/questions/4196453/simple-and-fast-...
There are some interesting discussions. (Nowadays, such a question would have been closed...)
As this is being consumed by larger applications that may have dependencies that conflict with these, they should be much more liberal.
I don’t have the background in imaging these people likely have but mine works by breaking an image into an X by X map of average colors and comparing, written specifically because I needed to find similar images of different aspect ratios and at the time I couldn’t find anything.
I made a script to calculate the hash of every file and if it found a double it would move it to another duplicate folder. This worked reasonably well but I couldn't stop thinking there should be more than 1 solution already made for this.
Reddit, Imgur, and any other site that uploads significant amounts of images from significant amount of users.. do they attempt to do this? To de-dupe images and instead create virtual links?
At face value it'd seem like a crazy amount of physical disk space savings, but maybe the processing overhead is too expensive?
There's no way to make the UX work out for images that are only similar. Would be pretty wild to upload a picture of myself just to see a picture of my twin used instead.
But I do wonder if it's possible to deduplicate different resolutions of an image that only differ in upscaling/downscaling algorithm and compression level used (thereby solving the jpeg erosion problem: https://xkcd.com/1683/)
Not perfect, but it worked pretty well for images that were exactly the same. Of course it isn't as advanced as Imagededup.
Legacy style: https://66.media.tumblr.com/tumblr_m61cvzNYF81qg0jdoo1_640.g...
New style: https://66.media.tumblr.com/76451d8fee12cd3c5971e20bb8e236e3...
I wonder how much of it could be adapted to finding duplicate documents, e.g. homeworks, CVs, etc. Presumably, the hashing would have to be adapted slightly. But how much?
Now if only Apple hadn’t repeatedly broken PyObjC over the years...
G