My comments here are informed by experience with three systems in particular:
- A 1990s era online publication's discussion forum which used sequential post IDs for its content.
- Google+, which utilised a form of UUID for both user IDs and submissions to the site.
- The somewhat notorious recent Parler web-scraping incident.
During the 1990s I was one of several participants in a forum that was being decommissioned, and which would be taken off-line entirely. Several years of what had seemed to be critically sigificant discussion history at the time (and in fairness, there are still bits I'd like to be able to call up now) would be lost.
The content management system (CMS) assigned a sequential post ID to each post on the site. Scraping the content was (mostly) as simple as running a bash loop of the form:
for i in [1..$MAX_POST_ID]; do wget $BASE_URL/post_$i; done
... which occupied a modest-for-the time laptop on a dial-up connection overnight.The recent Parler content archival used a similar characteristic:
donk_enby managed to exploit weaknesses in the website’s design to pull the URL’s of every single public post on Parler in sequential order, from the very first to the very last, allowing her to then capture and archive the contents.
Google+, by contrast, assigned 20- or 21-digit numeric sequences to users and content on the site (along with "vanity" text-based user identifiers for a small but generally active set of users, which turned out to confound ultimate archival efforts). This was larger than the actual populated target space by a factor of billions to quadrillions. Exhaustive search of either user or content space was simply infeasible.
But there was an interesting side effect.
As a search company, Google is (or at least was) remarkably conscientious about providing comprehensive sitemaps for many of its properties. This included Google+, and (when I accessed them), some 25 gigabytes of sitemaps for every single one of the then 2.2 billion user ID assigned at Google+. These were broken out into files of about 50,000 records each (the maximum permitted under the sitemap protocol). Yes, roughly 45,000 sitemap files.
Because reasons, possibly the algorithm used to generate UUIDs, the contents of any one sitemap appeared and tested via several methods to be a random sampling of user IDs. So when I got the bright idea after one too many fruitless arguments with someone that based on my sense, Google+ activity was nowhere near the levels the company was claiming, I realised that I could simply grab any arbitrary sitemap and sequentially scrape user profile pages, and look for the date of the most recent public posting activity, if any.
The key to valid sample-based statistical inference is having a random sample. And Google had just handed this to me. After only about 100 profiles, it became clear to me that at best about 9% of accounts had ever posted to the site. Another Very Simple Bash Script chugged through all ~50k profiles in the file I'd chosen, again on a Laptop of Very Modest Proportions, though over a rather nice broadband connection of the time, and again, a night's data pull and some crude awk scripts for reporting revealed the hard truth about Google+ activity: https://ello.co/dredmorbius/post/naya9wqdemiovuvwvoyquq
But I totally punted on the sampling, trusting (and yes, doing some rough checks) that any one sitemap file would actually be a validly random sample.
Eric Enge of (then) Stone Temple Consulting replicated the methodology on a much larger sample of 500k profiles, and doing some more robust resampling of the data as I understand it, validating my own rough estimates though providing additional detail. (A larger sample does not increase the accuracy of a statistical analysis by much, though it can increase resolution of that analysis, especially for rarely-occurring phenomena, such as, in this case, people actually using Google+, roughly 0.3% of the total registered profiles.) https://blogs.perficient.com/2015/04/14/real-numbers-for-the...
(I was impressed by the work, unaware of it prior to publishing, and somewhat relieved I hadn't made any stupid huge errors in my own effort. A key point in publishing my own findings was to note that pretty much anyone could quickly come up with a rough assessment of true G+ activity, a fact that made Google's own increasingly ludicrous statements, openly mocked in the trade and business press, easily disprovable and ultimately hurting the company's overall credibility.)
When G+ was in the process of shutting down, I accessed those sitemaps again, this time to find the list of categories Google had used to classify data (interestingly: these existed for English, as one might expect, and Spanish. Only.) and for the Communities (groups) feature of the site. In that case, an immediate interesting find was that rather than the 5 million or so groups that was generally claimed for G+, there were nearly 8 million when I first pulled the data, with that count growing to over 8 million by the time new Community creation was finally disabled in January of 2019. Again, by drawing samples of the full population, and being aware that I was looking for rare phenomena (here, large or active communities), and still undergoing large amounts of both creation and deletion activity. In the few hours between pulling a new communities listing and being able to scan a portion of that for size and activity data, a substantial fraction no longer existed. This suggested some interesting dynamics going on, likely around spam or disinformation. Another interesting finding of this research was that Google had managed to police G+ commmunities pretty effectively against extremist groups, with little apparent collateral damage. But that's another story.
But I got some pretty pictures of community membership and activity: https://joindiaspora.com/posts/23acfd90f05d01368b220218b79d8... https://joindiaspora.com/posts/ab6a5470f57001368d400218b79d8...
The methods aren't universally applicable. I haven't spent much time digging into, say, Facebook or Twitter numbers, though the simple approaches effective on G+ don't seem applicable to them. There might be other avenues to assessing activity or other characteristics independent of published statistics.
At the same time, the UUIDs also meant that the naive approach of archiving site content by simply iterating (or randomly poking) through the namespace would result in only one successful request for every few billions, or quadrillions, or worse, of requests. Without some means of identifying populated values, that approach was useless.
The upshot: UUIDs may protect user privacy. But depending on how deployed or exposed, could also provide a powerful tool for imputing true characteristics of site activity or behaviour.