Crawler Hints supports Microsoft’s IndexNow in helping users find new content
blog.cloudflare.com
blog.cloudflare.com
The idea is that you can push a notification to a search engine when content changes instead of waiting for the crawler to notice.
https://<searchengine>/indexnow?url=url-changed&key=your-key
You can also submit more than one URL with a POST.
You can notify Bing at https://www.bing.com/indexnow?url=url-changed&key=your-key
If you notify the IndexNow API endpoint it notifies Bing plus other search engines on your behalf:
https://api.indexnow.org/indexnow?url=url-changed&key=your-k...
This announcement is about how CloudFlare can now do this automatically for sites it hosts.
Some other hosts and CDNs support IndexNow, eg Akamai. See: https://blogs.bing.com/webmaster/october-2021/IndexNow-Insta...
The default being at the root seems… stupid.
Another interesting feature I saw in the standard is that you can host keys in subdirectories too.
"the location of a key file determines the set of URLs that can be included with this key. A key file located at http://example.com/catalog/key12457EDd.txt can include any URLs starting with http://example.com/catalog/ but cannot include URLs starting with http://example.com/help/."
https://www.rfc-editor.org/rfc/rfc8615#section-3
Registrations MAY also contain additional information, such as the
syntax of additional path components, query strings, and/or fragment
identifiers to be appended to the well-known URI, or protocol-
specific details (e.g., HTTP [RFC7231] method handling).
So it could be: /.well-known/index-now/<key>IndexNow would need to change the semantics of how they handle directories, as a key authorises only subdirectories.
I also notice there is an option for changing the filename of the IndexNow key file, but there is less flexibility about the directory it's hosted:
https://<searchengine>/indexnow?url=http://www.example.com/product.html&key=af4c4e043c7d42afad6bdeeda948527d&keyLocation=http://www.example.com/myIndexNowKey63638.txt
This seems like a potential vulnerability as if an attacker knows a text file path that contains a known (hex?) string it looks they could use it as a key?Here on HN we've been seeing posts of alternate search engines. How will those small bespoke engines make use of IndexNow unless the website participates?
The way I see IndexNow, I'll still get crawled relentlessly by the bots I don't want crawling my site (because robots.txt never seems to apply to them unless there's a special listing explicitly for them)
So, unless you're a participating search engine, a website will still be getting crawled by low hanging fruit, not alleviating the problem.
A good compromise would be something like an RSS feed, which a site can publish, and crawlers can hit for updated changes. It would also allow easier management for those domains that have many moving parts: individual search engines can be pinged, but the search engine just grabs the changes.xml file... Or something.
There already is such an "RSS" feed, its called a sitemap available at /sitemap.xml or you can alternatively list your url in the robots.txt file
The lack of trust means a search engine needs to know if what it's being presented in metadata is actually what's being served to the browser!
That's why we can't have nice things! :-)
Petalsearch (Huawei) indexing rate is half of Google, Apple and Bing half of Petalsearch, Yandex half of Bing. The other search engines respect the robot.txt.
I wonder how much energy and Internet capacity is wasted because of Google and PetalSearch indexing.
"Only you and the search engines should know the key.. so obviously, we want you to host it in plain text, in the root directory."
it is however much easier to serve static content than evaluating headers. the benefit of significantly increased compatibility in how you can serve the content probably outweighs the risk of logging the secret in many cases, as static content serving is compatible with virtually anything, adding additional logic to be evaluated at runtime through other means than URL contents is not as widely supported.
It would be cool to be able to push the update signal to a bunch of search engines when I publish a new page (even if all of my websites get virtually no traffic and don't even come up with the appropriate, highly unique keyword combos in Google, Bing, or Marginalia - they don't even have any ads or anything terrible, perhaps terribly boring or SEO unoptimized, haha).
I wonder if there could be a market for something which collects website change information and offers an API to query for new pages / updated pages over the past X time interval across Y set of [interesting] websites. I could see this being useful, sort of like RSS but 100% general purpose.
I'd call it something like Invertdex.
I created an issue for it as a reminder. Probably not gonna implement support for this type of active crawling hints in the short term, because a lot of my model is based on full-site crawls at 8 week intervals. While something like this would help with identifying new content, it doesn't do much to identify when links go dead, so you still need to crawl passively.
Although, on some level I'm a bit uneasy, I have a hunch this may be a bit of a vulnerability. In general giving websites the tools to control the crawling process beyond like robots.txt and so on seems a bit sketch. Maybe it's possible to build checks and balances to prevent that though.
Would take some work to support real-time updates. Not impossible, but I've got a lot of work to do with regard to search result accuracy that I feel is more important.
I've been looking at using RSS feeds as a signal for when to re-index sites before, not doing that now but the general idea works I think
Eg, would Ahrefs or Semrush qualify to join the party?
I've had staggering success with just ending emails to people, including businesses.
That said, I can't find any contact information on the IndexNow website.
This is an exciting problem to solve, and we look forward to working with others that want to help the Internet be more efficient and performant while reducing needless energy consumption. We plan on having more news to share on this front soon. If you operate a bot that relies on content freshness and are interested in working with us on this project, please email crawlerhints@cloudflare.com.
Having to have existing market share as a search engine is going to really limit the number of participants.
i usually have a whole-sites-inventory-sitemap.xml which gets updated once per day/week/x and a limited update.rss (rss is a valid sitemap.xml format) that gets pinged either real time with every update or if-changed-5-minute interval.
wrote a book about it and other distribution concepts
to my knowledge google indexing api is still whitelisting only (if not jobs search) and yoast does sitemap.xml update & ping.
Yoast updates the last modified field in the sitemap and pings Google that the sitemap changed
Slice off a few percent of the savings to make the crawled/pushed/IdenxedNow pile free for the public to supplement other archival efforts, train ML models on, use for small-change operator recovery, seed torrents of public-domain content, etc. and I'll call it fair. All the (relevant) actors involved are so heavily peered that they could store it on 8TB M2 sticks and it would vanish into the noise of 1/6th of gross Internet throughput.
My initial instinct was to be all cynical like "great, a new model for SEO competitiveness that locks the heavy-hung dark fiber incumbents deeper into the fabric".
But I'm trying hard to be more positive and there is a really positive outcome here.