Google Broke Image Search for Creative Commons
cogdogblog.com
cogdogblog.com
Image search is another story. I've been less and less satisfied with Google Image search results, and visual search has been totally neutered. It only returns low res results that rarely match the original as well as it used to. I used to be able to plug in a 400x400 image and find a dozen copies of it at a usable size. No more. Too many copyright complaints, I assume. I've started using Bing for image and visual search now. It's not as good as old Google Images, but marginally better in some cases than the current iteration.
One used to be able to just right click on an image and copy it. Now, I only get a still. I try clicking into the website that hosts the image and it's click, click, click just to get anywhere close to the size I want and often times its not even an animated gif anyway because they do some sort of media query and serve me up an uncopyable movie instead.
The web sucks now. People work around it with bots and the like on Reddit, but I feel like the economics have been figured out and it's not fun anymore.
Instead, we use /giphy in our Slack and hope the algorithm finds something that kinda-sorta was what we were thinking.
(yes, im kidding)
That said, I do see the sentiment. It would be nice to have the old convience of being able to look up old forum posts (especially with summaries in the OP via edit). Stackoverflow often fits that role now although I dread the answers even worse than old forum posts. I guess what I want is my cake and the ability to eat it too, I dont know why we can't just have the ability to search forums and have the general kinder attitude that modern media tend to have, they shouldn't be mutually exclusive.
in all honesty, isnt this a bit of "beauty is in the eye of the beholder", "sticks and stones" etc?
yes, you might be called various words, told to RTFM. but seriously here, is it really so bad? do you really want to drag down actual verbal abuse to something so absolutely trivial?
I don't want the internet to be kind everywhere, some artificial culture. If I don't like the tone of the place, I can go elsewhere. Strong moderation comes with too high a price in my opinion.
It's Yahoo Groups all over again, except that open groups could have their messages indexed, unlike on Discord.
The thing is, this doesn't matter anymore. You have very little control as Google tries to be smart. It's very hard if not impossible to find something older, obscure, things from other regions, languages, etc...
"Google-fu" is not progressing. It's struggling to hang on by the decreasing number of threads before it's utterly ineffective.
Where we used to say “rome fall why”, you’d now write “why did the roman empire fall”. Because the AI likes that phrasing and produces better results.
Soon you’ll write a 300 word description of what exactly you’re looking for, like you would when asking a trusted expert, and Google will figure something out. The days of keyword searching are long gone.
I think this is actually an example of the benefit of their AI. Despite the big difference in the "style" of phrasing (simple english vs more formally naming the subject noun), both seem to map to a very similar representation in their embedding space. I've run into frustrations with this myself, but for basic questions like this it seems like the search works Pretty Good.
[1] https://www.google.com/search?q=why+rome+fall&oq=why+rome+fa... [2] https://www.google.com/search?q=why+did+the+roman+empire+fal...
Humans instinctively know this, so we are able to construct queries like "why rome fall".
But that "why rome fall" query, which we think of as a purely mechanical keyword search, already requires quite a bit of sophisticated processing in the search engine. The system has to recognize that "fall" is synonymous for "collapse" or "wane in power" and not synonymous for "autumn". It also has to recognize that "rome" means "the (Western) Roman Empire" and not the modern city of Rome in Italy or "the Holy Roman Empire" or the city of Rome, NY, USA. It furthermore needs to interpret "why" in such a way that it emphasizes results with "reasons" or "explanations", rather than something like a "timeline" or "summary".
Personally I find it really weird that Google is interested in pushing users more to interact with its digital librarian / AI assistant, instead of continuing to improve keyword search.
I have a few guesses as to why they are going this way:
1. It makes the user interface simpler from an engineering perspective (fewer user-facing buttons and options to implement and test).
2. There is strategic benefit to making search more of a black box. Maybe they are specifically trying to "educate" users to expect and be comfortable with such black boxes. Maybe the plan is to get people so accustomed to "AI assistant" search that they see keyword search as outdated, and thereby secure a competitive advantage for the next several years over other search engines, by having the biggest and best AI models.
3. They are trying to increase the amount of rich "natural language" user search inputs in their data. Making keyword search worse will encourage people to use queries that more closely resemble natural language. I assume that this has strategic benefit related to Guess 2 above.
Personally I'd suggest "collapse of rome" which does deliver you a rich embedded result specific to the fall of rome.
I agree that Google's search parsing peaked a while back though, it seems to be getting weaker and weaker and now partially relies on the fact that search term autocompletion on mobile devices will supplement it by helping present an array of options near what you might want.
If I want to find a particular user-run forum on some obscure bit of some hobby, "<hobby name> <forum topic>" brings it up. But if I type out "Forum for <hobbyists> discussing <topic>" I get... a random selection popular of fora where someone has mentioned <topic>, often in passing or with minimal information.
> you’d now write “why did the roman empire fall”. Because the AI likes that phrasing and produces better results.
"the AI likes that phrasing" is exactly the problem here. How is anyone supposed to know what the AI "likes", other than painstaking trial-and-error in the unbounded and arbitrarily high-dimensional search space of human language?
Even the people who built the model probably don't know. Language models (and deep NNs in general) are extraordinarily complicated things, and there are problems with pretty much every technique that purports to provide visibility into their inner workings. There are just too many parameters and too many "information paths" in such a thing for regular people to wrap their heads around it. The ability to incorporate a high amount of complexity is a big part of why those models are so effective to begin with, but it also makes them really hard to reason about.
"AI" is currently in a weird spot where it's starting to kinda-sorta behave like an intelligent human in some limited settings, but in general is nowhere near as smart as a human. Most models still have a very shallow conceptual understanding of anything, even if they're becoming uncanny in their ability to match sophisticated patterns. It might not even be possible to teach some concepts to language models as they currently exist today, if only because there is only limited conceptual understanding available to be learned from corpora of text and images, even huge ones. Humans are still tremendously more effective than our best language models at understanding meaning and intent. Can an AI ever learn about love, regret, fear, or bliss, by reading millions of news articles and books and looking at millions of images?
Thus AI right now is in a kind of "worst of both worlds" situation, where it is complicated enough to be hard to reason about precisely, but still mostly unsophisticated and therefore highly sensitive to how inputs are crafted. Therefore it's hard to formulate inputs that provide useful outputs. It's still alpha-level technology at best, and there might be one or several conceptual innovations remaining between what we have today and something resembling general intelligence.
Consider also that "AI assistance" is complementary to keyword search, not a replacement for it. Google search AI is becoming something like a "digital librarian", a creature that can understand your queries and guide you to a starting place in the relevant literature. But much like in a real library, the digital librarian is going to be most useful as a starting point. At some point, if you already know what you're looking for, you still are going to want to search on "structured" criteria, as well as, yes, keywords embedded in text.
And finally, do you really want to type a 300-word description in order to get good search results? I was already getting good results with 3 keywords. I have already done the sophisticated pattern-matching and concept-graphing in my own brain, and now I know exactly what terms I want to look for. Why should I be forced to coach an AI on how to redo all that work for itself, instead of just letting me do a damn keyword search? Not to mention wasting my time and giving me carpal tunnel typing it all out.
The results quality has gone down and it's noticeable only after you switch to another independent search engine, like Brave Search.
For example, search for the term: "javascript undefined vs null" on Google and Brave Search. Brave Search gives way more information in the sidebar and Google doesn't at all.
The discussions feature on Brave Search is great, you don't even need to append queries like 'stackoverflow' or 'reddit' for searching discussions.
On top of that, let's say you're trying to search for an npm library like 'react-select', if you search that term on Brave Search, it gives you a button to copy `npm install react-select` right below the npmjs.com link.
It's crazy how good Brave Search is compared to Google sometimes, haven't used Google Search in a long time because of it.
Big things are not the same as small things, and should not be treated the same.
So I guess you're advocating that this dynamic equilibrium and "circle of life" is just panglossian optimal?
So far, I've noticed Brave Search only shows sidebar results for Stackoverflow and a few other popular forums that do not advertise on their pages directly, nothing else. So they're not really taking any revenue away. As for npmjs thing, I'm not sure if their revenue is hurt in any way because to view the package documentation you still need to open the link. Brave Search just provides you a copy command button extra for convenience.
As for discussions, they do not give full context so you always need to click the link so that's great for website owners as they get more exposure and traffic too.
At this point, Brave Search features are more on the UX side of things than anti-competitive so I personally will hold off the tinfoil hat for now.
More on topic, they still use Bing's near-useless image search, so even if it's getting worse Google Images is seemingly the only decent option in that category.
OK, I'm definitely switching to Bing Image Search for images.
Per OP... they do seem to have a public domain/CC search limit feature too. I don't know how well it works. (I'm not sure how either it or Google identify CC/public domain content. Is there an opengraph tag?)
https://www.bing.com/images/search?q=dogs&qft=+filterui:lice...
I couldn't find any reference on tineye.com but it seems like it has to be.
I just get page after page after page of "content" that appears to be either GPT written or written by somebody who has no idea about the topic.
They all seem to follow a pattern, they have a table of contents, and they take sentences from real sites regurgitate them and put them together into semi-random paragraphs.
If you know nothing about the topic it appears on the surface to be legitimate. And I bet to any quality engineers it all seems totally legitimate, because they're not experts in these fields.
Google search quality is in severe decline in my opinion. I have many experiences where I am searching for stuff that I know exists and that I know Google of yesteryear would have found, and Google comes back with garbage results and spam. Personally, I am hopeful that this means a Google-killer will be coming along soon.
Google search is severely broken.
Now I installed OpenSearchServer and see how far a local index gets me. The simple query "a" gets me at most 394 results at google, after letting OpenSearchServer crawl/index for just 2h I already get over 900 results for "a". Well.. I guess it really is not hard to beat that meager 394 results.
Initially read this as, well, "FU Google", and thought "yep, FU to them too". I guess I will acknowledge I have a bias. Then curious about this term, I googled-on-bing "FU Google". Top result was Google-fu, not the expletive.
I have no Google-fu.
'More than 1 result is a bug, citizen.'
* Reverse Image Search (sometimes no matches although image is for sure out there). I wonder if Reverse Image search sometimes broken because of copyright?
* NSFW images (I'm really old enough and I don't need to be protected by Google or anyone else, I think it's a kind of censorship)
* And now also Creative Commons as pointed out by Op
The best alternative (tested them all) is in my opinion
Chrome extension for reverse image search supporting Yandex
https://chrome.google.com/webstore/detail/fast-image-researc...
Android app supporting Yandex for Reverse Image Search
https://play.google.com/store/apps/details?id=com.thinkfree....
P.S.: I know Yandex is based in Russia, I only use it for specific image searches and I'm happy we have a good alternative based outside of USA... I wished there were more in the World.
This one is for Firefox with Google, TinyEye, Bing and Yandex.
In theory, site reliability engineers and software engineers work together at Google to maintain existing functionality and check for regressions. In practice, regression checks can only cover what they've been told to cover, and if breaking a feature doesn't cause a metric to crash, Site Reliability doesn't have the information to know something is wrong.
The result is that generally speaking, Google prioritizes existing features by how popular they are (i.e. "Maintenance by Popularity") modulo how much of a stink someone influential can make if they screw it up (i.e. "Maintenance by Twitter."). If almost nobody uses the CC filter (and I bet, in the grand scheme of things, they don't), Google may not have enough signal to know they broke the filter unless someone had the foresight to add a metric to check search results on that query separate from the rest of the search result data (and at Google's scale, you can't just add a metric for free; in addition to eng-hours to build and tune it, the data has to be stored somewhere and teams have finite budgets for space that can only be grown by negotiation with the relevant teams managing the monitoring services or the project as a whole).
Both google image search and "normal" google search are becoming more and more a pain to use, where you have to use quotemarks on pretty much everything plus a few excludes to find anything at all.
It makes some sense; take a step back, and the story is "If Google software engineers don't care enough to make it good, and Google users don't care enough to keep using it in spite of its flaws, why is Google throwing money at it at all?"
Google is a weird company because so many priorities are set by software engineers, not managers; some projects die because they literally run out of passionate engineers to work on them and management isn't incentivized to force engineers to work on projects they hate, instead asking the question "If nobody wants to work on this, is it worth it to keep doing it?"
Atleast I think so... what else can explain so many chat platforms that came out of google and died soon after?
I mean someone type something on google the result is not what is expected how do you troubleshoot that or know that it's a regression?
[0] https://kagi.com/search?q=sur%3Afmc
Kagi is good, but get a proper PR person and don't let imbeciles ruin the experience.
It's terrible. The entire stock photo industry is so bad at creativity you can basically instantly tell if a photo is a stock photo, making anyone using them look like a complete fool.
Anyway, they either have some serious legal issues with image search or they are becoming precautionary, but it's becoming almost impossible to find decent images.
Only the bad, outdated stock photos. There is still a whole market for the "obviously corporate corporate website" corporate website, but that's falling out of fashion.
What you're thinking of are the bland, white backgrounded photos and photos shot in stale, generic office settings so they could be worked into any design. That's not really how it's done anymore. Modern stuff doesn't look posed and staged, and often isn't.
Sure, someone could take a candid image and then after the fact attempt to gain releases. However, that's not workflow with a high margin of success. At that point, the "model" has all of the power. Also, crowd shots in public streets blah blah.
Even if the lawsuit is won, the product is still doomed.
I find this to be a highly interesting take...
The whole point of stock photography is that you can buy images to use various projects, quite often one-off or short-term ones. Because for the project, taking your own photos or paying a photographer to take them would be too expensive or time-consuming relative to the estimated value and return of the project.
I feel like once someone has made the decision to buy or use a stock photo, they have already decided where they want to stand on the scale of authenticity and originality. Deliberately seeking out "stock photos that don't look like stock photos" just sounds too much like trying to be something you're not.
What if one day every Google search has an "auto-generated" panel of images as well?
I think copyright trolling is more prevalent with images, and I think it's generally easier to determine the canonical origin of software. But yes, it's absolutely a risk and a reason why many companies have a legal review process before any new libraries can be used.
Of course, for many casual purposes it's widely ignored and for photos of people used for advertising and marketing, you need a model release anyway.
https://helpx.adobe.com/stock/contributor/help/known-image-r...
Note: I get that adobe might be wrong here. Whether they are right or wrong on particulars is beside the point. I just linked there because it was clearer than most other search results I found.
Here's another
(So, watch out for CC licenses older than 4.0 for that.)
Google Images, on the other hand, seems to look for the nearest monetizable domain cued by the uploaded image, and gives you adjacent results about that. It certainly does no facial recognition, etc. (well, none that it will feed back to you in results, anyway).
Then they disabled the Google web cache which was hugely useful since it allowed to open dead website/webpages or allowed to browse something when your government restricts access to it. I guess the copyright lobby and China forced the removal of this feature.
Google search is still unmatched in terms of being able to find text but other features have seen a huge cut. :-(
And don't remind me about iGoogle. I loved it. https://web.archive.org/web/20160314122329/http://linuxfonts...
Anyways, I'd argue that's the most comprehensive database of CC-licensed works.
Also, I just checked that flickr.com still allows you to filter their searches by license.
(see https://en.wikipedia.org/w/api.php?action=help&modules=query for reference)
edit: formatting
I notice a slight difference between
https://commons.wikimedia.org/w/api.php?action=query&generat...
and searching for 'cat' on canweimage.com
Is there some query processing needed?
You cannot find anything.
Personal anecdote - I was at a specialist doctor to discuss the results of my tests. I was shocked when he typed a health issue to Google Image search, clicked on (what seemed to me) a random table matching his search terms and compared it with my results. He then said that everything is fine, the parameters are normal.
That got me scared, what if someone intentionally put a doctored table and used SEO to promote it to the top to mislead doctors and causing bad outcomes to patients?
Followup story, my return rate for CC licensed images of "dog" went from 3 to 13. More than that, license info is not displayed (only linked), is frequently wrong, and photo credits often given to the site, not the creator of the image.
https://cogdogblog.com/2022/10/google-cc-image-search-better...
Try Openverse for much more accurate and plentiful results https://wordpress.org/openverse/
I'm going to assume that there are hundreds or thousands of products, tools, hobby projects, ect, that direct to Google searches; none of which have any mechanism to know and break gracefully when the API changes. Furthermore, Google is under no obligation to coordinate with anyone who just arbitrarily send queries their way. (I've had a few hobby projects use Google Queries.)
Seems like the most we could really ask is to put some kind of version stamp into the query parameter; and Google could optionally support old parameters or simply return an error. Otherwise, we have to accept that sending browsers to other pages via query parameters is inherently fragile and has a high probability of breaking at any time.
If I'm a content provider that wants to maximize the chances that an image search will flag my images as CC or public domain, what should I do? Are there open graph or other meta tags? Or what?
[1]: https://www.bing.com/images/search?q=dogs&qft=+filterui:lice...
Here's Matt Cutts' tweet about it in 2014: https://twitter.com/mattcutts/status/422944316458168320
But I think it dates back to 2009? https://googleblog.blogspot.com/2009/07/find-creative-common...
If I click on all the other dog pictures bar ("german shepherd, puppy, baby, rottweiler, police", etc. right below the option to select size, color type, time, licenses) additional CC images show up of the selected sub-type. The selection is additive (and) filters.
It is more that Google deciding what I really want.
https://flickr.com/search/?text=dog&license=2%2C3%2C4%2C5%2C...
Support alternate search engines, we used to have a bunch of viable options. Let's get back to that.
Says the person who puts their stuff into Flickr, which requires an account.
I do somewhat unusual things, like search for parts of a joke or a string of semi-randomly generated words for part of an AI type thing I'm working on creating to see if these are either unique or original and how they may appear in the context of the internet. Often times, if anything, there's something like a banned twitter (bot?) account only available through a cached backup that said it once in some bizarre context.
I've noticed it significantly can change search results depending on if you're logged in, or the country you're searching from (via a VPN). Different countries have different levels of success for different types of searches, but I don't have any sort of solid guide to map this out in any shareable way.
Boot up a virtual machine on a VPN if you're curious to take a look yourself. You may need to manipulate it so you don't bring in any suspicious cookies or other identifying information to show your actual country. Some VPN IPs are well known, and your results may be manipulated anyway.
Some results literally will never show up if you're searching from the USA. If you switch to Hungary for example, suddenly things could start appearing. Even if the matching result is a Chinese site that should have relatively equal relevance to both countries.
Sometimes I use Bing. It's not exactly better, but it's also not really worse. It's just different. In my non-scientific opinion after seeing so many of these differences, it's because it feels like they just forgot to enable (or haven't yet gotten to enabling) the kind of filtering Google has.
If something is no longer showing up in Google search, using Bing feels like going back 1 relative year in time before Google nerfed your active search. Sometimes I suspect Google is breaking down searches to parsable keywords and then sometimes adding those to a blacklist.
DuckDuckGo seems to sometimes filter things too, and some of the other commonly recommended alternative search engines. I don't know if they're actively doing this, or if it's a byproduct of forking off some other engine. I don't have much information here because I've largely given up bothering with these.
There's some other large engines not widely discussed in the US I am currently looking at as well. I don't have enough experience to form a solid opinion yet, and I have suspicion of their privacy so I don't want to be loosely associated with recommending it until I know more.
Miscellaneous other thoughts around this topic:
- Google has been heavily pushing some results more than others obviously. Pinterest and Quora are always at the top of searches now. I think this is pretty common knowledge.
- Chrome has a right click -> Search with Google Lens button now. Are they working on AI object detection of images more so than a visual match now? Could this factor into image matches?
- TinEye - When looking into this question myself, a lot of people recommend TinEye. I've literally never had TinEye actually match a picture by the way. Am I using it wrong?
Tineye is also useless. They have a very outdated library or not crawling spaces they should be crawling.
Content thief steals contents anonymously, posts it as Creative Commons content anonymously using free hosting, archives the content on Way Back, and then uses content claiming it is Creative Commons if anyone asks?
Basically it is content laundering - and something that impossible for Creative Commons to address as is.
(If reasoning is not flawed, also interested in possible solution to the issue.)
It's also flawed in the same, albeit weak, manner: the source it was stolen from predates the stolen copies, and so can show its true provenance. Of course, if it wasn't published or otherwise registered, no one would know.
CC is just a usage license that authors can apply to their works in order to share them freely with the world. It's strange to me that you might think it is the job of CC to police how people (mis)use it?