Right now you're probably wondering, yeah, but there's this one problem that wouldn't have been solved if x or y....
But really it's just hoarding behavior. They're trying to collect it all. Statistical significance is reached very quickly and after that point they're doing harm to society.
Things like emoji usage, page navigations, feature uses, etc. Ideally anonymous; no IP, user agent, etc, just a small byte or two packed properly can go a long way.
The problem is that it's actually quite hard to reliably anonymize data especially once you start to begin combining data sets from multiple places. That's the problem differential privacy is trying to solve in a mathematically rigorous way.
See for example how researchers partially de-anonymized Netflix Prize data by cross-referencing it with IMDB reviews.
DDG:
{ used_advanced_search }
{ used_country_toggle }
{ tabbed, *tab_maps }
{ filtered, *filter_date }
{ os "iOS", *ver "13.5", browser "Safari" }
iOS: Mail
{ disabled_remote_images }
{ flagged_mail }
Keyboard
{ emoji_keyboard_via_globe }
{ *emoji_use "100-1000", *emojis [ ":)" ":P" ":(" ] }
Each of these could be stored separately without metadata then aggregated no problem. Things marked * could be left out, and some things could be randomized up or down buckets and such."Everything you say can and will be used against you"
All they study is ways you are bad, and all they research is ways to keep you down.
If anyone reaches the wrong conclusion "maybe they are innocent.." Then it only takes two seconds, then they are fired.
Privacy from social scientists is one of the most important forms of privacy.