After stripping any past statistical data from each entry, it shouldn't be that much of data per URL...
After stripping any past statistical data from each entry, it shouldn't be that much of data per URL...
Giving the domain to a 3rd party is not going to happen.
The point is whether Google considered any other options than keep operating it or burning everything to the ground. Google could also keep the domain and let users reach a intermediary landing page of the Internet Archive first
It would be a serious breach of trust for them to publish the database. It likely includes links to non-public YT video URLs, for example.
> there are about 230 billion* links that need visiting
> * Thanks to arkiver on the Archive Team IRC for correcting this number.
Also when running the Warrior project you could see it iterating through the range. I don't have any logs handy since the project is finished but they looked a bit like
https://goo.gl/gEdpoS: 404 Not Found
https://goo.gl/gEdpoT: 404 Not Found
https://goo.gl/gEdpoU: 302 Found -> https://...
https://goo.gl/gEdpoV: 404 Not FoundThey could possibly provide a GCP service where you make an authenticated request to look up the value of a given goo.gl key. That would mitigate fishing concerns, eliminate the pressure of running a productionized legacy service, and allow the to do use quotas etc to tamp down on abuse. But that also would be covered by the regulatory laws and I don't know what they say about such a thing.