Isn't that one more data sets for ML other research purposes instead of a highly up to date search index (For example with news from a few minutes ago).
A question to ask could be: how often do users care about information from a few minutes ago, compared to information that has been available for a longer duration of time?
- a few thousand news-sites (like nyt.com, bbc.co.uk),
- a few thousand very popular blogs (based on what influencers people search for),
- a handful of social media sites (e.g. Twitter),
- a few hundred databases in areas like weather, airlines, sports (like ATP for people who look for Wimbledon results today)?