A Face Is Exposed for AOL Searcher No. 4417749 (2006)
nytimes.com
nytimes.com
Search has exposed so much data about ourselves to the services we use with very little regulation on what they are permitted to do with it inside their own walls.
My fear with AI is that we are moving toward sending even more data to party services. Tools such a co-pilot (which I enjoy using) are a gold mine for behavioural analysis. The profiling that will be possible with these tools is extraordinary and we don't yet fully understand the implication.
It's because of this that I'm a massive proponent of "Local AI". We need to be pushing for the industry to adopt a local inference architecture asap. It needs to become the standard pattern as early as possible to reduce the risk of the AI revolution being a repeat of the invasive internet search and advertising industry.
Google knows alot about your behavior - they can and have been able to correlate online behavior with health and meatspace actions to identify budding extremists or people at risk of addiction, etc. AI will bring that capability to business processes.
With the number of little companies that are springing up, it will become much easier for outside parties to figure out how instituions work. This capability exists, but it's gated by Google and Microsoft and they have drawn lines to protect the overall business. Some jackass will install a creepy AI tool to scrape outlook and salesguys will be able to get a profile of who makes what decision in a company, for example.
And a large infrastructure to ensure you can scale. No easy feat with the current GPU stack.
Smaller probabilistic models line linear regression, I’d call machine learning.
Yet cloud providers happily sell resources and APIs to unethical companies. They rightfully don’t insert themselves into most legal business matters, with exceptions.
not in the EU it won't.
if you can de-anonymize people from the data it's not anonymous, and collecting this data at all would be illegal in the EU without user consent, unless it's being used solely for the purpose of delivering the service.
You need a personal "AI" that just does random searches unconnected with your life, constantly, in the background, and then injects this data into all the portals that are watching you.
Ultimately, their data will become dominated by noise, and ultimately useless to the point of severely destroying the value of the entire enterprise and data collection mechanisms in the first place.
No matter how many tools you make "local only" you're only a forgotten "send telemetry back to the mothership" checkbox away from being right where you started.
It’s hard to imagine any population-scale solution that doesn’t involve regulation.
The biggest problem with regulation (in my eyes) is that it thwarts competition between countries. E.g. if the US imposes restrictions on technology, innovation is incentivized to happen elsewhere. The EU has been bold on the privacy regulatory front with GDPR and the like, and has probably lost out on immeasurable monetary gains as a result. There’s a huge cost to regulation, but it works.
I don't believe there's any worthwhile way to regulate this technology nor is there any particular reason to. The imbalance comes about because of spare electricity, which is truly the crux here, that we have so much to waste on dippy language models while also having so little to spare that some people have none... or possibly worse... electricity with random availability.
This you could actually regulate. I think the game is clear.
This was a plot point in a Neal Stephenson novel, "Fall [...]". It's been a while so I'm fuzzy on the details, but one of main characters floods the Internet with constant AI-generated content, but I can't remember if it was to drown out bad PR or break the anonymous identification mechanism ubiquitous in the book's world.
In my eyes this just strengthens the case.
The "60s lady with the dog that kept peeing her sofa" got her hour of fame, and the whole thing became a case study in de-anonymization.
A few pointers:
https://en.wikipedia.org/wiki/AOL_search_log_release
https://www.researchgate.net/publication/233390862_Privacy_P...
https://github.com/wasiahmad/aol_query_log_analysis
https://www.technologyreview.com/2006/08/15/100592/who-benef...
https://www.sciencedirect.com/science/article/abs/pii/S00200...
https://isquared.wordpress.com/2014/04/24/mining-search-logs...
https://web.archive.org/web/20130822124159/http://www.someth...
https://web.archive.org/web/20130822124204/http://www.someth...
------------------------------------
https://web.archive.org/web/20131024112051im_/http://i.somet...
https://www.mercurynews.com/2015/03/23/san-jose-man-convicte...
And a documentary series about User 711391: https://www.imdb.com/title/tt1455044/
1. Click on the link
2. Find out it's behind a paywall
3. Go back in the browser
4. Click on the "comments" link.
5. Look for the post that has the archive.is version of it.
6. Click on that.
Surely that could somehow be collapsed into just a single click?
javascript:(function()%7B%0A%20window.location%20%3D%20%22https%3A%2F%2Farchive.li%2F%22%20%2B%20window.location%3B%0A%7D)()Sites are usually archived already.
* Compile a list of domains like nytimes.com that have soft paywalls.
* When a link like https://example.com/ is submitted and its domain is on the paywall list, insert [archive](https://archive.is/timegate/https://example.com/) after it in the title area. Just prefix the timegate part and it's a working link.
[1]: https://en.wikipedia.org/wiki/Private_information_retrieval
I'd say yes, throw disinformation at at, as once I started noticing someone was really doing that level of tracking I did throw disinfo at it... a lot. The problem was with their interpretations from there of searches intentionally made to mess them up - as I neither had consent yet realized I was being stalked.
That went on and on and on - until people died - my father thought I was swatting him (I had a welfare check made - not a swat - as they were disclosing his financial advisor and info - he's white and safe presumably) . Yet yes, people did die. It is evident no one understands it is never just them in research or in a channel for conversations. The result of the whole thing was my childhood best friend also being left homeless and run over (by someone opportunistic in the whole process). My brother de-housed by a gang member a party kept sending about (one actually poisoned my dog). My father alienated. My sister used as a puppet for extortion instead of being rehabbed.
It's very very very ugly the string of deaths and alienation that came from what they thought was funny research into AI.
This old data needs to be located and disposed of or put into proper custodianship. It's grown teeth that cost lives. Throwing misinfo at it is just crapping up the Internet more to push to web 3.0, where the same problem will thrive. Requires legislation. Not quite relaxed enough to articulate how and what right now due to the extensive harassment that came about from it all. Maybe some day. It's been pretty terrorizing.
For now
Storage was expensive, and data wasn't seen as a goldmine as now, so most long-term logs went to /dev/null.
That the normality now is to ask users to create an account, have data-scientists (whose goal is precisely to find needles in haystacks), etc.
From their settings page:
> Save My Search History
> Currently this option can not be turned on. Kagi does not save any
> searches by default. In the future we may add features that will
> utilize your search history and then we will allow you to enable this.
It sure seems like it will always be opt-in, even if they add query saving in the future.20 years ago, when this leak happened, the situation wasn't like that.
Gmail was barely born, so Google accounts didn't make sense.
The article was like "wow we managed to deanonymize a search query", but that's actually the norm now.
Essentially, this scandalous AOL-leak, became a legitimized every-day routine (sadly).
Today when you type anything inside a ChatGPT-like AI app (this is the case also for many search engines), you get tons of contractors, workers and partners who have access to such dataset:
researchers, engineers, support, advertising platforms, technical intermediaries, legal, etc.
Though the future isn't gloomy; in the short-term, with the advent of LLMs, we may actually see a really good solution private-wise: fully local answers.
Which means that for the first time, queries and questions may not leave the device or sent to whoever you need to trust.