Google Deactivates Web Search API
ajax.googleapis.com
ajax.googleapis.com
What's interesting about this is that the API has been officially deprecated since 2010 (and "offline" since September 2014), but today is the first day that it's actually become unavailable.
Edit: Additionally, there are probably going to be big repercussions to the web/average person's browsing experience as a result of this. A massive number of sites (and programs, bots, etc) were using this API because of its simplicity (and absence of registration/authentication).
http://googlesystem.blogspot.ca/2008/04/google-search-rest-a...
You could actually go cool things with the original API. I used it to power a predecessor to https://friskr.com, but we had to switch web search engines after that change.
Yep. My harmless little IRC bot can no longer Google search anymore :(
https://developers.google.com/custom-search/json-api/v1/over...
https://support.google.com/customsearch/answer/2631040?hl=en
A free search engine would enable API calls and also boost privacy and freedom from the likes of Google. We have built a lot of experience about search engines since 2000, we have access to scientific papers, cheap cloud servers and a huge interest in freeing search, so I think the open source community do it.
I hope "OpenSearch" becomes a thing, like OpenAI.
Writing an "objective" ranking function (for any values of "objective") in an open, and distributed manner is structurally not favoured by humanity's current incentive structure. As in:
* a dev team have to agree on signals, and weights: "SERP quality" has dedicated teams of people assigned for specific verticals @ Google; replicating this in a distributed manner will be played politically
* Assuming any significant usage, the second you submit ranking code to public github repo, the algo will be played by thousand SEO scammers to their advantage
* Executing custom ranking function on other people's computer not only introduces security risks, but will have scammers setting up honeypots for collecting other people's ranking signals, and playing accordingly.
only open source the framework for the server and client
then companies / communities / etc, can make their own algos, and buy their own servers. The reward for proving your servers / crawlers is more people use your algo (higher chance of hitting your nodes)
and then allow the client to have configurable automatic node filtering, along with manual node filtering, so if a person feels that a specific node set is just full of bs, they can filter them out (and also prefer certain node sets in turn, to which they can donate if they are consistently happy with the results)
its like, choose your own filter bubble.
Just a thought : Ranking code could itself learn & adapt to each individual user (the learned "weights" could be sync'd online across your devices). Weighted signals from users can be fed back to the mother ranking algorithm (un-customized one). Basically millions of distributed deep minds[1], instead of a single one.
I can imagine there are a lot of holes in my theory, but we can't simply accept that open sourcing the algorithm implies that it can't be done.
[1] : https://deepmind.com/
Cue blackhat SEOs creating millions of subverted "users" on AWS spot instances/Lambda
Why not just take humanity out of that picture, then?
With the current AI/deep learning hype everywhere, why not start developing an AI-driven search system?
I think for it to produce the most relevant results, it will need access to your browser (or be a browser) or better yet, work at the OS level, so it can have a better idea of the current context you're working in, and learn from your habits and preferences. Say I'm coding and have an IDE and a bunch of dev-related websites already open, so the AI gives more weight to development-related results. If I've been playing a certain game a lot then it should assume that I'll be looking for stuff related to that game. And so on.
So, the index would be globally accessible to all computers, but the ranking will be unique to each individual user.
Something like this could very well be the actual beginning of a true A.I. "butler," more so than Siri and whatnot.
Sadly we couldn't use DDG because their API won't let you filter for video results only (Source: Asked on IRC, got a reply from DDG staff + haven't found any documentation mention this either).
We ended up using searx [0] which takes little more time for search results (about 1-3 seconds) but we gain more video-sources (such as vimeo or dailymotion).
I'm aware that youtube-dl got a search functionality build in as well, but the requirement that the client needs to get results without being prepared by our embedded computer wouldn't fit.
I've migrated to Bing, myself -- no volume limits, generous free tier, and much cheaper than Google.
Bing has a decent search api? Interesting, somehow I missed that. Want to provide a URL to a page for the service you're talking about? MS is really bad at documenting their services and prices. I eventually found this page, at this ridiculous URL: https://datamarket.azure.com/dataset/5BA839F1-12CE-4CCE-BF57... . I think that's what you're talking about? Note the link on that page to "Bing API FAQ" -- is broken. If this wasn't MS, the poor docs would make me think it was surely a terrible product.
You can, refer to this SO answer: http://stackoverflow.com/a/11206266/3354209
> WARNING: we did development using the free version, but to upgrade to the paid version (to do more than 100 searches), google forces you to turn off the "search the entire web but emphasize included sites"
I remain not very confident that you can really search the whole web with google custom search, or if you can that it's not some kind of a loophole that google might close without warning.
But if someone has actually done this succesfully, with a paid account for more than 100/queries a day, I would definitely be interested in hearing about it!
The billing system is confusing. You just have to sign up for billing, and you can raise the limit.
The upgrade in the google system, is only if you want site search with no ads.
You can get key right from there and an API console.
Should I chalk up to conspiracy that this page is nearly impossible to find on, um, Google, searching terms that you would reasonably use to find it?
But if you have an in with MS, it would probably help an awful lot if they removed the old pages with broken links, unhelpful content, and possibly out of date inaccurate content too.
The page you link to has "coming soon" for pricing for Search (which was kind of confusing to find) -- does that mean that the pricing info on the page I found is not accurate, or soon won't be, with no indication on the page I found that that will be the case?
MS is really really bad at documenting this service. Only because it's MS am I willing to entertain that the quality of docs may not reflect the quality of service, generally I'd give up on something that markets and explains itself so poorly and figure that if they can't get that right, they probably didn't get the service right either.
Some folks mentioned this page: https://datamarket.azure.com/dataset/bing/search. I want to clarify that the page includes the old Bing Search APIs that are still in use, but will be deprecated in the future.
You can send feedback or questions here: https://cognitive.uservoice.com/
[Edit] Here is also a blog post with useful resources: https://blogs.msdn.microsoft.com/bingdevcenter/2016/04/13/re...
I don't mean to rain on your efforts, but I would personally never consider using an API with such limits, even for trivial hobby projects.
I just migrated to the Bing API when Yahoo BOSS closed 31 March :-(
Will there be a new image search api similar to the one you offers now? Can you share any timeframe for the deprecation?
But when between one fairly poorly documented API that you've told me will soon be deprecated, and another fairly poorly documented API that doesn't have pricing beyond 1K calls a month....
...my response is still "Okay, that could be good in the future, I guess I'll sit back some more and wait for pricing and better docs to show up before paying more attention to it, I hope it does soon!"
I need to know pricing before deciding to use something. And I need good (and google-able! Sorry, even if I'm using Bing API, I'm using Google to look for info on it) docs, along with of course a well-designed API that works well, but I need the first two things before I even get to evaluating the next.
That being said, I'm still looking for alternatives as well. This broke my IRC bot and a few scripts rather unexpectedly.
"With Google Custom Search, add a search box to your homepage to help people find what they need on your website." (emphasis mine)
Their video also seems to indicate the same thing: https://youtu.be/Qd9z48Bo8ZA
They say you can search "one website or even a specific topic across multiple websites". This is specifying multiple domains, not a full-index search.
The results are their zeroclick info snippets and it is very useful for pulling an answer from wikipedia since they have already parsed it.
Those are the only web scale indexes that have an API I am aware of these days.
Prev discussion (few weeks back) https://news.ycombinator.com/item?id=11281700
Credit : @sylvinus
(PageRank never actually had an official API, but it was exposed for the Google toolbar. They had announced that this would go away a while back but just dropped the axe in the last day or two.)
The Google REST API was a great way to provide audiences with a known tasks in order to connect them to a historic setup or interface. RIP.
[1] http://www.masswerk.at/googleBBS/ [2] http://www.masswerk.at/google60/
For example: https://www.google.com/#q=related:http:%2F%2Fagiliq.com%2Fbl...
It returns pages related to any given page, and works at the specific page level (not just at the domain level, like SimilarWeb's API).
I've migrated to Bing for all other web search needs (way cheaper and without volume limits like Google Custom Search), but Bing does not offer a "related:..." equivalent as far as I can tell.
I had success with Authority Labs. You might want to check them out. They'll raise the query limit if you ask them.
For feature requests, it would be awesome if you posted them to: https://cognitive.uservoice.com/ (I'm also forwarding your post to Bing).
If you want to get in touch with the team directly about a question/issue, the "Contact Us" widget at the bottom of the Cognitive Services page is the way to go: https://www.microsoft.com/cognitive-services
http://www.bing.com/partners list all the services/APIs/partnership opportunities Bing offers.
Very good product. I'm curious how it works though.
These seem to be for products in a slightly different domain than plain old search, though.
Simply put, the API never returned the results in the exact same order as the actual search results did.
The best / most reliable rank trackers did (and still do) simply use proxies to get around the google captchas. I've scraped millions of pages from google over the years, and with enough proxies, the correct proxy delays/timeouts, and other little tricks you pick up along the way; you can actually scrape google pretty easily.
This is especially true with the dropping cost of proxies. I'm obviously an exception considering I run several scraping SaaS and have generally specialized in it for years, but I can he hitting x.com with 30,000 different IP's within the hour.
Try putting a site: in the search and it becomes suspicious fast.
It was highly effective and very scalable because you can stimulate the captcha pages easily and out-of-band generate a very large cookie pool by captcha solving. Then when you go to actually scrape google you can balance your pool of IPs and cookies, chilling them before they trigger another captcha, to handle spiking demand. Contant, sustained demand was very easy to plan for.
Anyway, we stopped using the API after the third day and realized it was not only not accurate and you couldn't turn many knobs (want to scrape results as they look for a different geographical region?)
Were you scraping with real browsers or something like Mechanize/Curl? Rate limiting at all? Proxies or real servers?
We didn't rate limit, we would just increase the size of the cookie pool if a captcha was hit, which was rare because we would scrape n-pages till a threshold was met to prevent that session from being captcha'ed so we wouldn't have to captcha solve it. We had two pools, the primary pool and the "chilling" pool, cookies near their captcha life would cool off for a few hours before returning to the active pool which behaves just like any other resource pool, every page scraped would "borrow" a cookie out of the pool, customize the encrypted location key, and make the request with a common user agent string.
Scaling it was difficult but once we had it figured out, Erlang was invaluable to us and our dependence on IPs dropped once we figured out the cookie methodology.
Solving captchas is cheaper than renting IPs.
Parsing it wasn't hard but it wasn't fun...
Most of their revenue is ads.
For example, Microsoft Silverlight only returned 200 or 404 status codes: https://msdn.microsoft.com/en-us/library/system.net.httpstat...
No sign of any replacement
It seems pety, but I really feel the duck look is holding it back.