Show HN: I made a site to catalogue 10,000 CC0-licensed stock photos
finda.photo
finda.photo
Very nice site! Since your site is so much based around search, I thought I would pass on a few suggestions based on what I saw. If you happen to be using a search based engine for your content such as ElasticSearch, SOLR or maybe Azure Search :-), there are a few simple things you could add to make the experience a little smoother. Suggestions in the search box are nice to allow people to quickly see results as they type. You could even add thumbnails of the images in the type ahead such as you see using the Twitter Typeahead library (http://twitter.github.io/typeahead.js/). I also noticed that your search does not handle spelling mistakes or phonetic search (matching words that sound similar). Finally, through the use of Stemming, search engines can often help you find additional relevant content. For example, if the person is looking for mice, but your content has the word mouse in it, this will bring back a match. Since you don't have a lot of content, this can really help people find relevant content.
Hope that helps.
Unfortunately, it's no longer maintained [0], plenty of unfixed issues. You could try a recent fork [1] :
I'm particularly uncomfortable with Flickr's "no known copyright restrictions". What if people infer PD from that and upload it somewhere else under CC0? Then it gets sucked into this finda.photo? Yuck.
As for finda.photo, why are you truncating the source down to just a domain name?! Many of the sources include proper uploader details so why aren't you copying those over and displaying them?
I know you're not required to, but attribution isn't a bad thing if you can give it. I for one would be much happier using a photo if I knew exactly where it came from.
There are potential ways that you could fix this from a technology perspective, e.g. have a process to create a new JPEG with the credit below the original photo. But anything like this is going to be a bit clunky and potentially ugly graphically.
Once the photo metadata is gone, it is too easy for others to claim it is an 'orphan work' and avoid liability under copyright law. At the opposite end of the spectrum, people like me who release most images as CC0 are annoyed that that license tag was stripped from the metadata, preventing others from freely reusing them. I use and rely on Exif tags a lot but they are fragile and you cannot rely on them staying embedded with your images once they hit the web.
What IPTC fields have over the typical ways of handling attribution is that they are not left behind when the image file is copied--so they should be more resistant to accidental removal of attribution metadata.
On most websites, the attribution is a line of text that is displayed next to the image. Anyone copying the image, who wishes to preserve attribution, must also separately copy the attribution text. Then they need a way to store that text, and keep it associated with the image. Not easy, actually!
> Once the photo metadata is gone, it is too easy for others to claim it is an 'orphan work' and avoid liability under copyright law.
You cannot avoid liability this way. Under the law, it is the responsibility of the person using an image to know that they have the right to use it. Just claiming "I thought it was orphaned" does not work if you are being sued by the actual image rightsholder.
> I use and rely on Exif tags a lot but they are fragile and you cannot rely on them staying embedded with your images once they hit the web.
Yes, this is my point! They're fragile because web services don't preserve them--but theoretically they could.
The cynical side of me thinks that a lot of web services don't want to know all the rights data for the media they carry. Ignoring rights gets them more traffic and engagement, and under the relevant law (the DMCA), they are allowed to. All they have to do is remove infringing images when the rights holder requests it.
Then there are all the issues with the NC and ND license variants and what they even mean exactly. But that's another rant.
EDIT: I'd just add that clearing rights and giving credits have been an issue for ever. On more than one occasion, I've gotten a semi-panicky email (and I think once actually a phonecall in pre-email days) securing permission to use one of my photos that was clearly on the verge of going into production. Presumably, someone came along and asked "You do have rights to this, correct?"
They often collect images en masse from a bunch of sources without further inspection. If someone uploads a copyrighted image to these sources and marks them CC0, they will end up in these CC0 aggregators. And, if your use of this image is discovered, you will be held liable for the damage caused by your action (well, at least here in Sweden).
I would do some research before using these images in a professional context. Look up the photographer and confirm that the image is a work of her/him. If this site included proper uploader details, it would make this work easier.
For example, http://finda.photo/search/?q=--aspectratio+%3C+1 would give you portrait images.
Examples: https://twitter.com/kamy22/status/479040852028051456 https://twitter.com/kamy22/status/472517258418606080
Have a good day!
Here you can find awesome publications (http://rodrigob.github.io/are_we_there_yet/build/classificat...). It's something like a bible of neural networks :P
Seems very good but I'm still researching. Hope it might help you.
[1] http://alana.io/
[4] https://www.pexels.com/photo-license/
They currently have over 5000 photos (~600 new images are added every month)
I've tagged the images with singular terms to make them easier to search, so it will change terms like "bridges" to "bridge", or "men" to "man", unless I override each term. If anyone can suggest a better way, I'd be very grateful.
I'll add "Australia" as an override, for now. Thanks for pointing it out.
https://en.wikipedia.org/wiki/Stemming
Most major languages should have a library available to handle this for you.
Alternatively, you can use the Levenshtein distance to find words close to the search term (like its singular form).
Most of the images are automatically tagged, and can sometimes be incorrectly labelled. I'm aiming to work through and manually check them all. In the meantime, I might add the ability to flag incorrect keywords.
They don't appear to make it into the top 10 pets of 2014: http://www.pfma.org.uk/pet-population-2014
Perhaps a ferret adoption drive is necessary! http://ferretshelters.org/
The selection is mostly based on the source sites I chose, like Unsplash, which all have only good-quality photos. The aim was to show all of the images from those sites.
The actual download and analysis of the images is done on my local machine, and each image has its own JSON file. These are then used to populate/modify the database, so I can track any changes to each image's data (if I add/remove keywords, for example) using Git.
Disclaimer: I work for the company behind GraphicStock. Oh, and we're hiring!