The AI Trust Crisis
simonwillison.net
simonwillison.net
There is a consent crisis as well, though I agree it's a smaller issue related to trust. There needs to be an actionable, legal definition of consent as it applies to website privacy—my naive assumption is that there was, but clearly that's not true, or it's not good enough or actionable enough—and it needs to preclude implying users must positively grant consent to harvest, process, or transfer data to third parties, when in fact the dirty deeds have already been done in secret.
It should contains hundred of thousands of entries.
Like fines for speeding.
Imagine if you had a page for parking tickets since 2018 (!) and it contained only a few hundred penalties.
This is ridiculous, especially since:
- it's not hard to find most cases. Just take any cookie banner that makes harder to opt out than opt in, and that's it, you have a violation. It can even be automated.
- it can bring a lot of money to a system that is dying for it.
Also the jurisdiction is in question. Who should enforce? Country where user is? Country where domain is registered? Country where data center with servers is? Country where nominal owner of the web page is? Country where final beneficiary is? Country where creator of the banner is?
In theory anybody can report to their country's data protection officer, although at least in Finland they don't care about individual citizen complaints.
The NGO noyb is sending out batches of these reports. They have some minor wins, but as you can probably tell from all these nags around, by and large there's not much effect.
All these questions are sorted, on paper.
https://techcrunch.com/2022/08/08/noyb-gdpr-cookie-consent-c...
If you trick someone into signing a contract, then that contract is fraudulent.
If you tell someone that you will ask their permission before doing something; then silently claim you already obtained that permission in a prior contract, you are committing fraud.
I don't know when our judicial system lost all of its teeth, but you sure as hell can't blame that on the citizens they are failing to defend.
See. No fraud.
Dave is a prostitute. He is hired by Sue via contract. In this contract, Dave agrees to have sex with Sue. Also in the contract is the agreement that Sue may introduce any sex toys that will not physically harm Dave.
When they get together, Dave notices a camera pointed at them. Sue tells Dave that, "I only record sex with people who consent to my use of a camera."
Later, Dave learns that Sue did in fact record the encounter. Sue's defense? The camera was a sex toy! It didn't physically harm you! This was all in the contract!
Is Sue's defense valid? Of course not! Sue has committed fraud.
---
We should hold corporations every bit as accountable as we would fictional prostitutes.
Or, framed another way: consent, in both a moral sense and a legal one, requires awareness from both parties of what you are consenting to. Each party is, to some extent, responsible for using the resources at their disposal to ensure they understand the contract (e.g. a blind person must find someone to read them the contract), but intentionally undermining the process of mutual understanding is at best very good evidence against enforcement of the contract and at worst fraud.
Easy. It's when we started treating the virtual reality of the internet as though it is actual reality.
Digital contracts shouldn't be binding, any more than mowing down pedestrians in Grand Theft Auto constitutes murder. These video-game contracts are mutable and ephemeral; how the hell do you prove what version of the contract/TOS you agreed to, when companies can change it arbitrarily-- and without any revision history? Tech makes it easy to gaslight the fuck out of anybody.
There'd be an actual cost to them were they do play this game with paper contracts.
These agreements often have millions of copies, so it's easy to find the applicable version of the T&Cs.
If all you have are digital T&Cs at the end of a web link, and those T&Cs can be updated at any time, and you are only ever sent the link, and not the updated T&Cs - you have nothing.
You could argue that everyone should download every T&C update. In theory that's true. But in practice a problem that could be solved relatively simply - keep a single letter - now requires many more steps, backup strategies, and so on.
You could also argue that banks should keep copies of their T&C versions and supply them on request. Which they do - except that in practice a bank that's trying to forge a document won't have a problem forging T&Cs.
This not hypothetical. I've known people win court cases because they kept a single piece of paper.
The actions of the government have been and likely always will be done for the interests of the wealthy and powerful. Policing itself was a concept first introduced by the wealthy to protect them and their interests. I mean just look at George Washington - he was the richest US president of all time after Trump at a present valuation of around $700 million.
i just wrote a WTF email to their support, but most likely i will be discontinuing my account. can't imagine what they can possibly say that will make this OK
I had some words with their support department, here is what they said:
>Hello,
>Thank you for contacting pCloud's Technical Support.
>This is a pCloud banner for Black Friday and it's not a notification. Unfortunately, you won't be able to remove it manually and you should wait until the end of the Black Friday promo - 30.11.2023.
>Should you require any further assistance, do not hesitate to contact us.
>Regards,
>George Lewis
>pCloud's Technical Support
Like it matters that its not a notification. You still need my consent regardless of what kind of ad it is.
Hi there,
Thanks for taking the time to write in to Dropbox Support. My name is Ross, and I will provide assistance with your case.
From my understanding, you are inquiring about being opted into Dropbox AI by default.
Thank you for alerting us to this problem.
Dropbox engineers are aware of the problem and are working on a solution.
Also, please note that no files were shared in this case.
Sorry for any inconvenience this is causing.
We'll update you shortly on this issue.
Don't hesitate to reach out to me again for further questions!
Best regards,
RossThis might seem like splitting hairs, but I think it's extremely counterproductive to act like the battle for users' data privacy and sovereignty is lost just because most users can't tell the difference. I see this a lot from the other side: extremely cynical, at least somewhat tech-savvy people reacting to each new corporate abuse with an "old news" mentality as if we've already reached the endgame—if you haven't been using tails linux for at least a decade you may as well just zip up your home directory and email it to every shady tech corp and data broker you can find contact info for. These people ought to know better, and they ought to set a better example for those who don't.
This inculcated helplessness definitely hurts trust, but it also gives people the impression that a better world isn't possible—that there aren't better and worse choices for who to trust with their data and privacy. This Dropbox snafu seems to be that mindset coming back around: users won't give a shit if we imply we're sending their private files to a third party without asking, right? Lunacy.
Incidentally, I've had most of my data out of Dropbox for a while (in favor of a self-hosted solution), but yesterday was the kick in the ass I needed to cancel it for good. Thanks, Dropbox!
Meanwhile cloud data access is completely on a “trust me” basis, and plenty of companies have been proven to have abused that trust.
> One interesting difference here is that in the Facebook example people have personal evidence that makes them believe they understand what’s going on.
> With AI we have almost the complete opposite: AI models are weird black boxes, built in secret and with no way of understanding what the training data was or how it influences the model.
Completely agree that our biggest threat right now is complacency. If people form incorrect mental models of what's going on and then shrug their shoulders and accept that's just how it is, we won't make much progress in improving the actual problems.
People who are really hardcore about data sovereignty would probably take issue with my choice of Synology, but you could roll your own if you prefer. Synology seems to have a decent record and reputation, and my experience with it has been pretty good overall.
For backup, I'm currently using Synology's proprietary thing to do snapshots to Backblaze, and mirroring all my critical stuff to various other places. I'm planning to set up a Restic backup for more redundancy (and a bit of obscurity), and I would also like to figure out a cold backup scheme for the whole NAS though I haven't thought very hard about that yet.
I have no choice but to believe that a better world is possible. The current state of things isn't tolerable, and if tomorrow can't be better then what's the point of anything?
Also, I'm sure that there are better and worse choices for who to trust with data and privacy -- but it's literally impossible for me to know who that would be (or if anybody is "trustworthy" in a broad sense), so I have to work on the assumption that nobody can be trusted.
I want to be less cynical, but the way things have gone over the last decade or two makes cynicism look entirely justified.
Is my attitude incorrect? If so, how can I correct it?
I think it's very possible to get a decent sense of who cares more and who cares less. For instance, look at Apple's recent pushback against scanning private user data for CSAM, and compare that with the recently publicized case of a guy having his Google account permanently closed (and being referred to the police) because he took pictures of his baby to send to a pediatrician.
"Normies" think these companies are the same, that they're all selling your data and that trusting one is just as bad as trusting the other. If we can chip away at that misapprehension, maybe said normies will start giving more of their money to companies that do a better job keeping private data private, and maybe that will drive long-term trends in the market.
Or maybe not! Things are bad, I agree. But I still don't believe everything is equally bad everywhere. I don't believe it's all over.
I don't think that's a good example, because Apple was very much in favor of client-side scanning until they encountered a popular uproar about it.
My impression of Apple is that they're very sensitive to privacy intrusions by anyone who isn't named "Apple".
I always find it curious when I see "company changes plans due to consumer feedback" framed as a bad thing rather than a good thing, especially when they are one of a field of companies that are largely already doing the thing people were upset about. Personally, I think it's great when a company has a pragmatic motive to act in the consumer's interest rather than (or possibly in addition to) an idealistic motive, because a pragmatic motive is much more likely to stand the test of time.
So, I disagree with you that it's not a good example. In fact, I think it's a great example of one of these big platforms being materially better wrt. consumer privacy than their competition. Frankly, I think your "impression of Apple [...] that they're very sensitive to privacy intrusions by anyone who isn't named 'Apple'" is exactly what I was talking about in my first comment in this thread: a cynical, unsubstantiated implication that all these platforms are equally as bad as each other, and that we've already lost the battle in the consumer space.
I wasn't framing it that way at all. Being responsive is a good thing. But in the context of this issue, that they did so means they're responsive to outcry. It doesn't seem to me to be an indication that they care more about privacy issues generally.
> implication that all these platforms are equally as bad as each other
As I stated in an earlier comment, I don't think this at all and certainly wasn't intending to imply it. But just because one company is better than another on these issues doesn't mean that company is good on these issues.
And now we have ChatGPT/OpenAI and their competitors.. if the other players eat data like a secret midnight snack, current-gen AI is like zombies (the fast and twitchy variants) starving for blood and brains. Both because data serves a more direct role in the product, but also because of the typical hype-train-race psychology of tech VCs has been woken from slumber by the first potential paradigm shift in decades.
All circumstances point towards zombie apocalypse/ gold rush / ask for forgiveness later / etc etc. I strongly believe that’s why they’re (all of them) doubling down on the safety/responsibility rhetoric now, before the inevitable reputational PR crises (plural). Get ammo to muddy the waters early.
Meanwhile we techies are lulling around like we didn’t deeply just experience the last 10 years and thinking it’ll be different this time because.. AI is rooted in academia? Flashy new companies? The safety rhetoric? Edgy Twitter takes from “down-to-earth” founders?
I don’t pretend to know exactly what goes on behind the scenes but I’ve been around long enough to know how people work. And they haven’t changed for the better.
You see, the new giant company promises not to be evil...
These companies have already stolen everyone's data and techies are bitching about IP law and saying they don't need to ask permission to use anything on the public internet. That might be the legal reality but you still look like tech douchebag for doing it.
I'm a working professional and I have clients that are governed by confidentiality agreements and regulations regarding where information goes, and I would just prefer using a service where my data rests on a server instead of having more and more points for a data breach to be introduced.
I don't really understand why my data isn't fully encrypted at all times and only I can view it in the first place, but the idea that they are actively sending it across the internet to be ingested by other companies and processed without my consent or interest is so terrible.
I often use AI features when I opt into them, but having a company just sending my personal files all over the internet without my consent is insanity.
Honestly, OneDrive has a migration tool and I got a trial for dropbox business and moved all my files automatically last night. It's just the last straw in their company constantly doing things I don't ask them to do like introducing crap and popups into my desktop interface and never offering the feature I constantly ask for... end to end encryption.
If you want a couple click migration from Dropbox Business to an Office 365 Onedrive account, it's right here: https://learn.microsoft.com/en-us/sharepointmigration/mm-dro...
If anyone has suggestions on a reliable end-to-end encrypted solution with a lightweight cross platform sync software that doesn't force you to download all the files to your device (my Mac's HD is too small) and is generally fully featured. That won't take very much time to migrate with a trustworthy migration vendor. I'm willing to pay a premium price for it and I'm all ears.
They could also just lie. Having a company claim end-to-end encryption still means that I have to trust that the company isn't being sleazy.
The only encryption you can really have some measure of trust in is the encryption you apply yourself.
I'm glad to see that fantasies about omniscient AI taking over the world are giving way to a better grasp of the more mundane realities; AI just accelerating the already obscene power imbalances present in our world. Keep your private stuff private.
It's why I use dropbox to begin with.
In today's world just championing sanity is already an "ideology" :)
It is that easy to buy a few Tb disk and just run a program, if not for the: 1. Routers that doesn't allow easy port forwarding (or even ban certain ports) 2. Dynamic IPs 3. People selling domains pushing their own VPS services. 4. The amount of steps you have to take to allow
A lot of small organizations I know didn't need a system administrators to configure and run programs like this. A lot of them are beginner friendly! They needed them to configure the network.
Same applies for self-hosting sites. If there was a program that just hosted on your PC address any html-page you put into it, a lot of people would self-host something. But you can't unless you can wrangle your router and figure out how to buy static IP – two tasks that are way harder than basic html.
The stack for easy and secure self-hosting is here, but the network changed too much. Hopefully, ipv6 will help to solve this problem.
For now it's funny how Tailscale in their "How Tailscale works" post asks you to pretend that everyone uses ipv6 to simplify explanation, and then acknowledges that almost no one uses ipv6 and a lot of clients are behind NATs and firewalls.
The tricks they deploy to pass through NATs are fascinating to me though. A great service and a great blog, thank you for sharing!
Sure, it's not end to end encryption but it prevents the company from using the encrypted data as a training corpus.Are shared folders and files created by co-workers and family not tech savvy enough to know about encryption?
I don't think this is rogue actions, I think it falls into the category of perfectly reasonable expectations that are not actually met by a wide variety of cloud services.
https://techcrunch.com/2022/11/29/dropbox-acquires-boxcrypto...
Nit, but what you’re describing is e2ee (except for the metadata). If you encrypt your files before uploading and decrypt it only locally then only the logical sender and recipient have access. That it goes through Dropbox is not important (and also the beauty of e2ee).
This is a bit unusual, otherwise it’s typically people (and shady service providers) who say that something is e2ee when it isn’t.
[0] https://apps.apple.com/us/app/cryptomator-2/id1560822163
Also searching files was impossible unless you downloaded everything, decrypted it, and searched locally.
I quickly realized it was adding huge delays in my day-to-day work and a lot of stress during time-sensitive tasks.
Have these e2ee overlays improved in usability since then?
I rarely use indexed file contents search (filename search is usually enough and that works well, and tools like grep work transparently), however the Boxcryptor drive can be added to the Windows Search Index (or whatever search software you use), and I assume it’s similar on Mac. You don’t have to decrypt manually. Indexing causes more system load due to the necessary decryption, of course.
On desktop systems I always store all relevant data locally, exactly for the redundancy. I’m not sure what you mean by “big encrypted files”, because each original file is encrypted individually and thus has basically the same size as the original.
Cryptomater on top of any of the cloud storage providers is a great setup for home / personal use. I have been doing this for the past 3 years with minimal issues. Google Drive + Cryptomater on Windows + Cryptomater on ios, working pretty seamlessly.
No affiliation, but it does exactly what you're asking for, and I've been very happy with it.
I know they have a cheap solution, but it's not exactly something that checks my box for stability and high character for hosting my very important files.
"If you’ve used any of Dropbox’s AI tools, some of your documents and files may have been temporarily shared with OpenAI."
Good luck thinking a cloud provider has YOUR best interest at heart. This is Hacker News, I feel like trust should be earned, never implied. Dropbox’s practices aren’t unprecedented, but customer documents do pass through OpenAI’s servers and are stored there for up to 30 days, and the “third-party AI” toggle is turned on by default in account settings.- They gave their chat interface the ability to run regular full-text searches against Dropbox - when you ask a question that can be answered by file content, it searches for relevant files and then copies just a few paragraphs of text into the prompt to the AI.
- They might be doing this using embeddings-based semantic search. This would require them to create an embeddings vector index of all of your content and then run vector searches against that.
- If they're doing embeddings search, they might have calculated their own embeddings on their own servers... or they might have sent all of your content to OpenAI's embeddings API to calculate those vectors.
Without further transparency we can't tell which of these they've done.
My strong hunch is that they're using the first option, for cost reasons. Running embeddings is an expensive operation, but storing embeddings is even more expensive - to get fast results from an embeddings vectors store you need dedicated RAM. Running that at Dropbox scale would be, I think, prohibitively expensive if you could get not-quite-as-good results from a traditional search index, which they have already built.
If they ARE sending every file through OpenAI's embedding endpoint that's a really big deal. It would be good if they would clarify!
Bing Chat (based on ChatGPT), Copilot, etc. As a GitHub user, I never got a checkbox to opt-out of GitHub Copilot's training on my code. At least Dropbox provides a checkbox.
This is only part of it. I don’t want my data being sent anywhere unless I authorize it, regardless of what it’s being used for.
In this case, we not only have to worry about OpenAI training on our files (I have no reason to doubt that they are truthful when they say they won’t), but we also have to trust that they can securely handle our files.
Especially since they've had a few documented security problems in the past.
It's a real issue.
If you don’t want third (or second) parties reading your data, make sure it is e2e encrypted clientside.
This means no Dropbox, but Syncthing instead. That means no Slack or Discord, but Signal.
Use clientside encryption. Technical measures, not legal ones.
you may win the court case... 3 years from now. your secrets have been on the dark web for that whole time, and have filtered to the regular web by then. even if you know who did it, and can get them extradited they may not have cash to pay for the damages, and seeing them in jail, though nice, won't make your data private again.
Seems quite sensible for people not to trust that they fully understand the small print, since (i) they probably don't and (ii) the one thing most AI companies have made clear is that they think they should be able to use whatever material they like however they like, regardless of whether the creator of that material gives them their blessing or not.
Even if you have a deep technical understanding of how this stuff works and how these tools are built, you can still have very legitimate questions and doubts about how your data is being used beyond pure model training.
Facebook literally takes your data from their apps and internet, tracks your behavior on the internet, and feeds this data into models of you. These models are so accurate they can sometimes basically predict what you're thinking. Hence the layman jumping to the conclusion that they must be spying though the mic.
A LLM company like OpenAI, and their partners, employs almost literally this exact model. Grabs data from whatever sources to improve their models, to increase the likelihood you'll keep clicking where they want you to click, to monetize you.
Right, and in a larger sense, the laymen aren't exactly wrong. They're technically wrong about the mechanism, but they're exactly right about the extreme intrusion into their private lives. That the intrusion comes in the form of accurate models rather than the microphone is just a technical detail. The end effect is the same.
and all of this just to show me shitty ads for online games that I will never, ever, EVER play, college-themed dating services I'm not going to use, yoga shit, and money remittance services. I live near a big university, so I'm guessing it's simply by IP.
occasionally get ads for Lexus or Jags tho. that's nice.
Even though I don’t believe Facebook is spying on anyone surreptitiously through their phone’s microphone, I find that specific argument entirely unpersuasive.
Facebook’s reputation is dogshit, at least with regular non-technical people I know. I’m in the US. People know they helped foment the January 6 insurrection in 2021 and saw how they deflected any and all responsibility while not fixing a thing afterwards. The reputational damage they’d absorb from actually doing this thing that so many people already think they’re probably doing pales by comparison.
Yet I believe OpenAI isn't using data from Dropbox to train their models without users' consent.
BUT I don't think that's the problem here. The problem is data in transit; data sent to third parties who can actually read it, and who may have rogue employees that Dropbox has no control over; data that can appear in logs or subject to different policies, etc.
If I send private data to Dropbox they can't send it to anyone for any reason, including "improving" their product, without my explicit and informed consent. I'm not sure how this is even debatable.
If Dropbox wants to house models and offer RAG search themselves, to consenting users, that's one thing.
If Dropbox sends all data of all users to third parties without telling anyone before the fact, that's another thing. A terrible thing.
Well, I'm a paying Dropbox customer, and I would not pay for this feature. I would like it if they encrypted my data in a way that made offering this sort of feature impossible. I do want my data recoverable, but the fact that they can offer this AI "feature" at all, it seems like they've made zero effort to prevent malicious internal employees or third parties from accessing my data.
If they've ever signed a BAA (business associate agreement) with any enterprise customers using Dropbox for documents bound by HIPAA, this would get them into a lot of trouble very quickly. The financial penalties are _high and per exposure per employee involved_. They also hit employees doing directly (if you disclose/share HIPAA info then you personally are liable).
So I'm certain that even if they did share documents with undisclosed 3rd parties without notice, it wouldn't be "all". Enterprise data is likely safe. Those contracts get heavy scrutiny before signing.
why? they trained on my code without my consent, why is user data any different?
training is either fair use or it isn't
and high growth Silicon Valley companies aren't known for the adherence to the spirit of the law
1. How do you know for sure it was trained on your code?
1.1 if you saw it reproduce you code verbatim how do you know it didn't just hallucinate it ?
2. Was you code publicly available?
3. What licence was it under?“ GitHub Copilot is trained on all languages that appear in public repositories.”
Note it says “public” and not “open source”
https://docs.github.com/en/copilot/overview-of-github-copilo...
Even if training on copyrighted, private code _was_ fair use, training on PII without consent is _still_ problematic. It's a huge violation of privacy and adds all kinds of legal/regulatory/ethical risks that potential consumers of said models won't be on board with.
The language and spirit of the law certainly hasn't caught up, but I feel like in some cases (unconsenting use of PII, medical information, etc), the laws and ethics of how such data is used is pretty well established.
You can certainly argue that generative AI systems are different than previous AI systems and should be treated differently. But the current situation is basically that you are allowed to train an AI on any data you have, regardless of copyright or consent. I wouldn't be surprised if that ends up being considered legally and ethically okay, because it's the status quo, and because it's hard to define "what counts as AI".
For example, it's possible to extract training data from an LLM—which could include PII/medical data/etc. Those risks don't exist with spellcheck as far as I'm aware.
To your point about what is "AI", I'd state that AI is a misnomer. What we're really talking about are generative large language models (LLMs). What an _can_ be considered an LLM is definitely up for debate, but if you were to describe one in general terms I think we could reasonably say that most (or all) things we consider LLMs are:
1. A probabilistic model of a natural language
2. Have the ability to interpret and generate natural language text
3. Typically encode a large volume of training data
I'd love to hear other thoughts on how one would define an LLM in practical, simple language. I imagine doing so would be a pre-requisite to any effective legislation.But my point is, this (training) isn't the main problem.
It's different. Code and other content that you shared online - no matter what license you shared it under - is still fundamentally different from user data that you never shared anywhere at all.
That difference is really important, to me at least.
I would be absolutely furious if I found that someone had trained an AI model on my private data in Dropbox. I personally have no problem at all with someone training an AI model on content I have posted to my blog.
It'd debatable because your data isn't "private" when you foolishly hand it over to some third party without encrypting it first and their policies say they can basically use it for anything as long as they can claim that it was "in furtherance of its legitimate interests in operating our Services and business". In fact their policy says they can update their policies any time they feel like it.
Privacy policies aren't even legally binding. If you're in the US and didn't sign a contract with dropbox you have near-zero rights, and any attempt to assert whatever rights you think you have will require going to court which is basically a pay to win system and you'll be up against a company with billions in assets so good luck with that.
Yes, it'd be a very shitty thing for dropbox to outright violate the trust you put in them, and it might be a terrible business decision that means no one ever trusts dropbox with their data again, but if they one day decided to go fully evil and start handing your data over to anyone willing to pay for it I doubt there'd be much of anything you could do about it.
Don't put any data you care about in the cloud without locally backing up and encrypting it and you'll never have to worry about what the cloud provider does with it or who they give it to.
No its called unethical behavior and has become the norm for big tech.
Unethical behavior definitely exists in AI companies. Personally, I'm agnostic about whether it's higher or lower than the base rate of unethical behavior in the general population. Anyway, if we're going to talk about bad behavior, we should use specific examples with cited evidence, not fear-mongering.
I don't think that AI companies are less trustworthy than our industry in general, but I absolutely think that our industry is far less trustworthy than the general population.
AI companies engaging in the wholesale harvesting of web data to train their models, though, was a particularly egregious violation of trust.
I've no dog in this hunt. I'm not partial to either Altman or OpenAI. And I have considerable reservations over where this Brave New World may be taking us. Or whether there is any credible option to stop riding this merry-go-round, no matter how unattractive the destination(s).
The DropBox behaviour described is only one in a long, long, long line of trust-violating practices by tech firms.
This doesn't represent reality. Take this quote:
> A society where big companies tell blatant lies about how they are handling our data—and get away with it without consequences—is a very unhealthy society.
So, what are the consequences of lying? If there are any, I'm curious what they are. The author is arguing that "big companies" are going to do the morally correct thing even though it hurts their profit. That's just a silly argument.
In a post-trust society, companies have to prove they're doing what they say. There's no more making a statement and having the masses act as if that's what's going to happen.
Honestly, is there less than a 100% chance that down the road there will be a disclosure that there was a "mistake" where your data was used differently from what they claim? Fool me once, shame on you, fool me 5,872,328 times, shame on me.
That wasn't the message I was trying to communicate at all.
My point with this piece is that people don't trust AI companies, and companies need to figure out what to do about this - that's why I called it a "crisis".
I certainly didn't argue that big companies would "do the morally correct thing even though it hurts their profit".
> "In a post-trust society, companies have to prove they're doing what they say. There's no more making a statement and having the masses act as if that's what's going to happen."
I think you and I might be in strong agreement there. My goal in writing this company is to get AI companies to take this problem seriously and, like you said, "prove they're doing what they say".
> Use artificial intelligence (AI) from third-party partners.... Your data is never used to train their internal models
Qualifiers like "their" and "internal" make me deeply skeptical. Are they saying that my data is used to train some "non-internal" models, whatever those are? Also, is my data being used to train any models in Dropbox other than one exclusively used by my account for the features I'm expecting? etc.
The EU is well established now in setting penalties as a percentage of revenue. These really do scare big tech because their revenues are huge.
But I think your whole point post-trust is highly aligned with the article. It's exactly that: people don't trust what they are told by anybody now (even government), and that's a crisis.
Is it really useful to think in black and white teleological terms? “Post-trust” just sounds a bit sophomoric/marxist. Like “late stage capitalism” or something similarly presumptuous.
FB has been/is spying on people and making us collectively more stupid, but using different, more esoteric/harder to explain methods (e.g. cross-device behavioural targeting).
> Those companies need to earn our trust. How can we help them do that?
This is a wicked problem and such a broad question (+ a bit of a weird take imho), that I don't see anything actionable. So, here's a half-baked list off the top of my head (so, poorly expressed, incomplete, incoherent):
- slow down the ill-conceived progress at all cost and understand that moving fast in the wrong direction ≠ progress
- force stricter transparency rules through legislative action
- understand and accept that no-trust should be the default
- invest more in open and offline models
- educate people
- accept that some of those companies in their current shape don't deserve our trust
Edit: This is a better take: https://news.ycombinator.com/reply?id=38643792&goto=item%3Fi...
How can I trust a company about my data, when their whole technology's performance depends on scraping the whole internet and feeding to a model? Also consider the fact that the data I'm providing them is way cleaner than what they scrape.
Moreover, we (at least the older ones amongst us) have seen what Microsoft has done in the past, and said company is almost one with the same Microsoft, which is ready to do anything for platform dominance.
OpenAI, esp. when combined with Microsoft and the whole A.I. hype paints a very untrustworthy picture in my eyes.
I'm not against the technology, but how it's trained, developed, hyped and how the researchers in the ecosystem behave about the data they handle makes it very off-putting in every sense, esp. in privacy and ethics.
It's like looking into a proverbial sausage factory, but this one is way worse.
While there are perfectly rational, non-hot-mike explanations for the phenomenon of talking to a friend about something and then having it show up in an ad, the experience itself is damn spooky, no matter how it’s occurring.
And hypertargeted ads are just the earliest ripples of the wave of AI-driven applications for predicting human behavior. Those predictions—about where we’ll decide to go on vacation, when we’ll quit our jobs, what we’d like to buy, who we’ll marry and divorce—will be more right than wrong to an incredibly spooky degree.
I have no idea how society will react to that.
I think this is always the case. Unless the companies are forced to actually simplify and limit their TOS -or even divide it by features-. No one is realistically expected to read hundred of pages daily in legalize. I don't think many vendors do it our of necessity but as a workaround consent issue. This needs to be addressed by legislation that ban those practices.
In practice, I think the PR playbook is too entrenched, but we can dream.
The larger you are and the more people that are communicating, the bigger the chance that someone says something terrible.
Particularly if you are a large public company and there may be legal requirements around specific parts of disclosure.
Besides, companies don't want their secret new feature announced before their competitors know about it, in a way that is poorly communicated and confuses customers, and then find that the person who announced it has a profile picture which is them at a far-right protest / far-left protest / made some sexist jokes / has a bio advertising their onlyfans account etc or a million other things that could be seen as against the corporate image.
I certainly don't. I don't trust any of the AI companies at all. Their behavior and statements have given me quite a number of reasons not to, and no reasons why I should.
If the AI pioneering companies are able to push back against the entire publishing industry and have nothing more than a conversation about it, then why would I ever expect any company to respect me or my personal data?
In fact, I think there is a large and growing portion of the population who have watched big business scoff at laws and regulation (see Wells Fargo, Uber multiple crypto exchanges, etc) and are realizing that profit is the only motivator.
That's not even talking about the chaotic political landscape.
I think we will soon see a crisis in institutional trust across the board.
Facebook absolutely has a lot of information that the average person considers personal and private, and big tech companies work hard to mantain the illusion that their data collection isn't as pervasive as it is.
With OpenAI, most non-technical users understand that the company has ingested a LOT of creative works done by individuals, most of which weren't expected to be ingested by a giant AI company and potentially regurgitated to millions of users.
The better corporate messaging this article argues for will help ameliorate concerns to some degree, but won't address the root of the problem. Most people might have the details wrong, but they are directionally right about the massive data being collected and used by these companies and they have every right to be concerned about that.
That's also my frustration with it: I want people to understand what's actually happening with targeted advertising so they can get angry about that. Having them get angry about a fake microphone conspiracy theory is a distraction that makes it harder to focus on solving the underlying issues.
In dubio pro reo? I don't think so.
>With AI we have almost the complete opposite: AI models are weird black boxes, built in secret and with no way of understanding what the training data was or how it influences the model.
In principle, I think we could have significant visibility into how OpenAI trains its models by doing something like the following:
* Generate a big list of "facts" like: Every fevilsor has a chatterbon. Some have a miralitt too, but that is rare. (Or even: Suzy Q Campbell, a philosopher who specializes in thought experiments, lives in Springfield, Virginia. Her phone number is 555-8302.)
* Generate a big list of places where OpenAI could, in principle, be harvesting data: HN threads, ChatGPT conversations, StackOverflow, Dropbox files, reddit posts, etc.
* Randomly seed facts in various places.
* Check which facts get picked up by the next version of GPT.
If you really don't trust OpenAI, be subtle about where and what you seed (e.g. avoid obvious nonsense words) to make it hard for their engineers to filter out your canaries even if they tried.
It's important to seed some facts in places they're known to scrape, as a control. If those facts aren't getting picked up, then you have methodological issues. You might have to repeat a fact a number of times in the dataset before it actually makes its way into the model.
If you think about modern society and the levels of trust needed to do certain things : I spent nearly a lifetime of earned income on a house and I completely trusted the agent and bank who facilitated that. Without that trust, the whole transaction wouldn't be possible. Why did I have that trust? These are the hidden value of all those laws and regulations that we often hate so much. As much as anything they create the stable platform for high value transactions to take place. The price of not having them is not being able to access the value you can get if you can successfully execute high trust transactions. There's no way I would gamble my whole life's income if I didn't trust the process. I would have to buy a smaller house.
So AI is playing out along these lines. How much of the value we can access from AI is eventually going to be limited by how much trust we can grant it. Which ironically may derive from how much regulation / legal framework is successfully put around it, which may in the short term slow it down. The self-hosted open source model route is a substitute / fallback for not having trust in hosted models. We can't really pretend that it's possible to self-host a model that will be as good as a hosted one, the scale is always going to be a material difference. But we may be to hit the point where for targeted applications there are diminishing returns, and that may be good enough.
Made me feel like I was going a bit crazy TBQH. Surely I didn't misremember?
Annoyed because it was a convenient file storage solution and now they have proven themselves untrustworthy so I have to set up my own thing. My fault for trusting to begin with, I suppose.
Hundreds of millions of people use ChatGPT, billions Google and Facebook. The value wins. If the service is good enough people just tend to "hope" more than trust that their data is save enough.
For the dropbox backlash people just didn't see the value.
This isn’t the disincentive you think it is. VW manufacturing vehicles that purposely deceived emissions testing was textbook corporate malfeasance, perhaps one of the worst examples of it. They took a loss on the ledger and came out of it relatively unscathed
Generate a unique string, store it along with my data, notify me if it shows up anywhere.
I wonder if it's possible to generate a canary of that sort that can be detected in AI models.
Local models also solve a lot of the trust issues
Threat models yada yada.
I'm in love with how fast models like insanely fast whisper, ggml llama.cpp, Mixtral, YOLO v5 e.t.c have come along.
I very much predict devices in 2024+ that have powerful energy efficient GPUs and do multimodal inference locally without any connection to the internet.
I remember buying the first comma.ai pre-panding which was essentially a modded android LG phone driving the car. No connection to internet, a freaking mobile phone driving the car without any human input on the highway. That totally blew my mind.
I'm awaiting Apple to release a version of Siri that runs totally offline.
Smart phones are the perfect target.
Where there's smoke, there's fire. Public trust in large tech companies is at an all-time low. Nobody thinks tech companies want to help anybody but the few rich people running them anymore.
I wouldn't trust a word OpenAI says, because even their name is a lie. The code isn't open, it's not auditable and you have dishonest people running the company.
Bullshit. Consent means consent.
There is no confusion here. Nothing is vague. This is explicit fraud.
I feel like every day now, I'm reading an article where the problem is obviously just fraud.
Fraud has always been illegal. When did we stop prosecuting it? Why aren't we talking about that every day?
It's not free to read his articles on Bloomberg, but his newsletter, Money Stuff, is free and one of my favorite daily reads.
AI is very cool technology. The incredible overstepping of any and all ethics with regard to getting training data, just all of the data from any source as fast as possible right now, in this mad dash to create it at all costs is not and should mandate a complete restart on behalf of the people behind it. For this one, and for so many other ethical lapses on the part of OpenAI, the models as they exist are tainted beyond any ethical use as far as I'm concerned.
If your product uses this stuff, you are not getting one red cent from me for it. Period, paragraph.
This is what I mean by the "trust crisis".
Dropbox very clearly deny that Dropbox content is being used to train AI models, by them or by OpenAI.
You don't believe them, because you don't trust them.
No shit.
> Posting something online you are literally inviting people to look at it,
Yes, PEOPLE. Posting things for people to see is why the Internet exists, is why USENET was created, is why web forums were created, is why social media was created, is why 9/10ths of the Internet as we know it today was brought into creation.
It was not put there so people who do not know any of those people and do not give a rats ass about what they made could take millions of images, writings, and sounds and shove them into their product without their consent for purposes it was never meant for so they could automate art. That is categorically not what any of that is for, and you, and everyone else making this tired point damn well know that.
If you have no issue at all with your creative output being used to train data models, more power to you! That's how consent works! You consent and that's completely, 100% fine. That consent should not have been presumed as it was, and even if you assume complete and total innocence on the part of the AI creators, once it became extremely fucking obvious that tons, and tons, and tons of creatives absolutely would not have consented if asked, then their data models should've been re-trained with that misused data removed. That is the ETHICAL thing to do: when someone says "hey I really don't like what you're doing with my material, please stop" you STOP, not because that's legally binding, not because you'll be sued, not because you're infringing copyrights, but because you have a fundamental respect for your fellow human being who says "I don't like this, please don't do it" and then you, you know, don't do it.
Unless of course what you actually are is an ethics-free for-profit entity that needs to get to market as soon as possible with a product to sell that you probably can't be sued over, in which case you tell those people who's work your product could not exist without to eat shit, and proceed anyway. Which is basically exactly what happened and continues to happen.
And before you even go there to the "well how could they ask for the entire dataset's contents" I DON'T CARE. I'm not the one doing this, this is not my problem to solve, just because the ethical way to do a thing is hard, time consuming, and/or expensive or otherwise difficult, you don't get to just waltz past the ethics involved, even if you're a research project! I personally wouldn't want to get permission from a few million artists to use their work in this way, I don't think most of them would be comfortable with it, and even if they were, I don't really want to do that, it sounds like a ton of work. SO I DIDN'T.
If a corporation argues that it can ignore your AGPL because it didn't have to blow the bloody doors off to get hold of your code and its training process is "just like your browser cache" or "a person learning" and the derivative stuff is completely novel, why would you trust them not to deploy the same "but it's not exactly copying" arguments when given access to other stuff that has third-party "no copying" agreements wrapped around it, like your Dropbox?
And I do not buy the "just like a person learning" argument. At least, not fully.
I could see that, if you have a fully-functioning AI system, then handing it a new article to ingest could be "just like a person learning".
But many people graduate high school having read just a few dozen books (or less), having been around maybe a few dozen people (or less), and watched a few dozen movies (or less). A person does not need to ingest a nontrivial percentage of the entire wealth of human knowledge just to be able to be intelligent enough to read an article in the first place.
There may be people who do not care about this distinction. That's fine. But I am quite convinced that the distinction exists. And thus I do not believe that training an AI system is just like teaching a person -- and making copyright decisions on the basis that the two things are identical does not make sense.
100%. I love the way you put this and just wanted to expand on it a little bit, to remind everyone that there have been numerous, flagrant examples of various creators of various media who have their names/handles put into these models, who have work that is ludicrously similar to theirs in style produced from the model, and despite the fact that it's technically original, it is not original in any way meaningful to the topic, or defensible by anyone debating in good faith. That you need to put things like "unreal engine" "featured on artstation" and the like proves this. You're telling the machine to aim for works you know are of a higher quality in the dataset to get a better result.
Now if you're just fine with that and content to fuck over artists like that for no other reason than you can, I can't stop you. But please spare me the righteous indignation of objecting to the characterization of such behavior. It's fucking obvious, do not insult the intelligence of your opponents by insisting otherwise.
Sure maybe you care more about whether OpenAI has stuff derived from the contents of your Dropbox on their servers which is technically neither "training a model" nor the actual "copy" they were required to delete after 30 days than you ever did about copyrighted stuff. But why would OpenAI?
The problem here is that these corporations are given carte blanche to make any derivative works they want. They get the exclusive freedom to ignore copyright law.
The rest of us don't.
The worst part is that they get to turn around and say their models are protected by copyright!
This is copyright laundering. There are only 2 reasonable avenues for us to react:
1. Make "AI" companies respect existing copyright law when compiling training datasets.
2. Get rid of digital copyright for everyone.
I vote option #2.
In a conversation online about this recently, someone said (paraphrasing) ‘Government should regulate this!’ - but there are already regulations about all of this! It’s fraud! Fraud is literally paying for YouTube (and to varying extents a lot of the rest of the web) and nothing is being done about it.
I fear that it’s just considered so normal now that it will take a very long time to stop. If existing regulations about fraud had been enforced when the fields were nascent, establishing norms that it was just as unacceptable ‘with tech’ as it was before, I think it would be a lot less widespread now.
We all expect the courts to give the benefit of the doubt. The problem is that we echo that expectation in our discussions. We are not the courts! We should be loud and direct with our criticism.
There is no doubt here. I refuse to give corporations that benefit.
Well, unfortunately, we have a government trust crisis also.
They'd enslave you and charge you a subscription for oxygen if they could – and insist it was a feature.
The key question is: Who is the actual customer?
That drives everything else. The FAANG's customer isn't the "users", but rather its advertisers, et al. Anyone who will pay for the data their "users" generate.
This is the Faustian bargain we all made when we decided that free stuff from the internet was a good idea. Someone else pays, and they get our data in exchange for that payment. The same data is likely sold as many times as they can do so, of course.
Also, I have a feeling this could be another case of the EU having to fine an US company, which means half of HN will cry again how the EU is always so mean to US companies. Or maybe this time HN acknowledges that the reason could be that the companies trying to use private data without permission are often US companies.
I agree it's a faustian bargain but I don't think we had much agency here. Our society now depends on these services to a high degree, which is why these companies are (and probably should) face increased scrutiny and regulation.
https://insight.kellogg.northwestern.edu/article/shareholder...
> There are a lot of misconceptions about maximizing shareholder value, even among economists. But talk to a legal scholar or a corporate lawyer: a CEO or board is not legally obliged to maximize shareholder value. They need to maximize the value of the corporation and act in its best interest. Only when there is a change in legal control, such as a merger or imminent hostile takeover, do they have to maximize shareholder value.
Concluding "if we respect people's feelings and do one weird trick, people will suddenly trust AI" is naive beyond number.
This theory has been floating around for years. From a technical perspective it should be easy to disprove:
* Mobile phone operating systems don’t allow apps to invisibly access the microphone.
* Privacy researchers can audit communications between devices and Facebook to confirm if this is happening.
* Running high quality voice recognition like this at scale is extremely expensive—I had a conversation with a friend who works on server-based machine learning at Apple a few years ago who found the entire idea laughable.
The first point seems believable, but the second and third do not. Obviously a nefarious Facebook would be toggling these features off when they were being inspected, and even if they were not, they could be using sneaky exfiltration techniques. With the rise of app attestation, they would be able to do so in a way that would be 100% undetectable by reverse engineers. The relevant code would only be delivered to the app when it saw that it was running on an L1 trusted phone (with hardware security intact, unrooted). Additionally, they could embed some whisper-tiny on the device and force it to do its own speech recognition and hotword detection, only sending a list of "ad topics" back to their servers.I don't think Facebook is doing this, and the social reasons hold more weight with me, but I don't think it's technically impossible or easy to prove that they did not _ever_ do this.
Yes, they could do that today. People have been assuming they've been doing this since way before 2017, long before Whisper.
Restore the trust? Pay us.
I used to believe this, but I do not understand how anyone can say this after the Snowden revelations. The vast majority of engineers are not martyrs. Given a choice to blow the whistle and then seek politcal asylum in Russia, most will not blow that whistle. Instead, they will live their lives and work to support their families.
"But this is not about Goverment Surveillance!"
Maybe this is true - I don't want to argue it here but I do think large tech companies are willing accomplices in many surveillance states - but an engineer at Open AI or Dropbox will probably not be facing felony charges for blowing the whistle. They will be facing the end of their career in big tech, and that's enough to dissuade most of them.
I'd guess maybe 200 engineers would have to be in loop before there is an 80% chance that one blows the whistle in any five-year period, and if they are careful they can keep the number of engineers in the know MUCH, MUCH smaller than that.
The idea that conspiracy is impossible ignores the conspiracies that happened and were covered up for significant amounts of time, long before we have the aid of modern technology. Why does the argument receive much attention despite that?
Because government national security secrets and dumb corporate secrets are different.
Is there a big black binary blob of source-not-available functionality in there that they are told to include in their builds without question, and for which a competent iOS or Android engineer wouldn't be able to tell if it has access to the microphone?
I'm not worried, because I use Cryptomator. Great app, acts as a file encryption layer on top of any cloud storage.
(I'd very much like them to disclose the full scope of their training set, but it's not possible for them to do that on a prompt-by-prompt basis in the way you describe.)
word
It's only annoying if a blog post is using the structure to pull a "here's a list of reasons why X sucks; this is why my monetized project Y is better, please use it!" which is not the case here.
I added the bit about local models mainly because I knew that if I didn't 90% of the conversation about the article would be "yeah but he didn't talk about the obvious solution, which is local models".
I agree that “opt in by default” is a dark pattern - but how else would any cloud service provide a RAG enabled service? Perhaps it should be made explicit during the onboarding process like geolocation (opinions as a former location PM at faang)
Having a "enable AI features" checkbox seems to me like it would be a smart approach here.
Not only that, it's a nonsequitor. It's ridiculous on its face to say someone else making a decision for you is you "opting in".
No, that’s a thing that should never have happened. Ever.
The arguments made in favour of bigco trustability are not great though, esp the parallel between AI data ingestion and Instagram's rumored use of mics to capture voice data for ad targeting:
> Facebook say they aren’t doing this. The risk to their reputation if they are caught in a lie is astronomical.
FB has already paid the biggest fine in FTC history for privacy breaches. It has no goodwill, it's running on user inertia and indifference.
The Facebook Spying idea is the common place to agree and complain about phone spying in general. It is the poster boy for communicating an implicit belief about a larger topic.
We have begun to use language to communicate implicit and shared beliefs without any evidence, facts, agreement or qualification.
We come to our own beliefs privately and then choose a poster boy to carry the conversation. There's not much surprising about the continuation of conspiracy theories, or topics that look like conspiracy theories. We shredded the common cultural references apart...
Earlier this year Microsoft was fined $20M for collecting PII from children, which it "shared" with advertisers[1]. And this is _Microsoft_, not even an adtech giant. The abuses from those have their own Wikipedia articles[2,3]. FB was famously fined $5B in 2019[4]. It would be beyond naive to believe that these were just "oopsies", and that somehow these companies have changed their ways.
_This_ is why it's not implausible that the software that runs on our most personal computing devices, built by these same companies, is hoovering up our data in the sneakiest ways possible. Equating this to conspiracy theories is dishonest at best.
When Mark Zuckerberg puts tape over his webcam and microphone, it would be foolish to think you know better not to.
[1]: https://www.ftc.gov/news-events/news/press-releases/2023/06/...
[2]: https://en.wikipedia.org/wiki/Privacy_concerns_regarding_Goo...
[3]: https://en.wikipedia.org/wiki/Privacy_concerns_with_Facebook
[4]: https://www.ftc.gov/news-events/news/press-releases/2019/07/...
Ah, bullshit. They've lied about other things, repeatedly and constantly. It's a risk, sure - but a calculated one, and one they've taken before.
Facebook doesn’t explain how their ad targeting works at the individual level, but they thought about it and tried—barely. It was perceived as an obstacle to black-box ML, even when there were clear opportunities to show customers (and partners): here’s how the algebra works, here are key dimensions that we have identified explicitly and that the advertisers wanted us to target (age, gender), here are more dimensions and the feedback from showing that ads to other tells us (people who like board games respond more to the ad, and we think you do).
That training, in clear, accessible language that anyone would understand exists: it’s part of onboarding.
The result of not having it is a large chunk of ads are shown to people who literally don’t speak that language (something Antonio GM mentions in his book and that was identified, corrected, broken, and re-identified at least three separate times since). But even benefits would be enormous: it would free what looks like 60% of my ad inventory by not showing gambling ads.
But they went the other way, treating users asking questions that their own employees ask on their first days with suspicion—assuming they knew better than the people they are talking about.
OpenAI is less adversarial, but they must explain how training and testing work with user-relevant examples: “Here’s a conversation you had last week. Here’s the feedback you gave. Here’s how we detected it was relevant to refinement. Here are examples of questions whose answers were changed by your contribution.”
Having a full-text search on their training set might be difficult, but something along those lines could be implemented and reviewed by a neutral third party. Those conversations would lead to insights about where to find more relevant data.
As Simon W. writes, this matters, and you want to respect privacy because it’s an inherent good, but you also want to preserve trust because it’s commercially essential (even though people are quick to forget if they get free stuff). I think you should do it because security by obscurity, or at least the ML equivalent of that, isn’t working.
https://xkcd.com/1838/ is funny, but the truth is that we can offer some tools for exploration and feature breakdown. Every time I did, we learned so much from those.
The author refers to the theory that FB is always listening as “laughable”, but this statement is even more laughable. If FB actually got caught doing this, at most there might be a few articles about it, and maybe a class action lawsuit resulting in a few pennies distributed to individual users, and then things would go right back to business as usual.
The only convincing argument that FB isn’t always listening is that it would be too expensive to be worth it. Arguments about “reputation” risk are beyond absurd. They don’t care.
The individuals who never had to go through a second of adversity in their lives cannot be expected to understand. In fact, their sociopathic tendencies only skyrocketed through the years of isolation from the real world, with COVID compounding the effects.
Casually talking about how robots should replace humans as a species because of "efficiency"...
By the way, did you know that Sam Altman was VERY active in politics until the latest drama?
He was backing a candidate to primary Joe Biden, but now he had to go take care of his own stuff. The outcome of this would be pretty obvious to those who follow American politics:
https://www.theatlantic.com/press-releases/archive/2023/12/a...
My point is, a person that cannot draw a straight line from F---ing Around to Finding Out cannot be trusted with something infinitely more complex.
The only solution I see is to make it illegal to use personal data. Especially make data brokers illegal. When you give a company access to your data, you often agree that they will pass the data to third parties who then pass it on and suddenly everybody has your data because you have permission to use your data.
But the whole „facebook listening to you” os absurd - on iOS you would see a system notification that the microphone is active, I assume on android as well. Also, it would take crazy amounts of bandwidth/battery/processing power to pull off.
It would also be trivially discoverable - a ton of people are listening to what apps are sending out, and even if encrypted, it would be noticeable by the outgoing traffic volume being high.
It is virtually impossible to pull off even today, as anyone who ever developed anything on mobile can tell you
The ELI5 version is that it can watch everything and then keep the bits that happen just before the person hits the button.
That’s how they can catch things that explode or do something and you aren’t sure when it will happen. Like popcorn popping.
Siri is basically doing the same thing.
Siri also misfires. And that data can be sent in.
What's funny about it is that counterclaims will often be that this information doesn't necessarily need to be coming from reading your chat messages. It could be that they just know you happened to be chatting to that person -- despite "metadata" supposedly not being personally identifiable information -- and then one or more parties started to look up on google the topic they were discussing. As if that's supposed to be less invasive.
We might chat about all sorts of random things, but many times there's something that happened before, let's call it event X, that leads us to talk about a certain topic, say topic Y (like when you mention a cool shirt a friend just bought).
Then, if they see an ad about that later, folks jump to the conclusion that "Facebook was listening" to their chat. But what's more likely is that this topic was already trending among people like you, and that's why it popped up in your ads.
So, it's not that talking about it made it appear in your ads. It's more about a common event that sparked your conversation about it and also made it show up in your ads.
Simon does a great job highlighting why this doesn't make sense from the technical and business standpoints: https://simonwillison.net/2023/Dec/14/ai-trust-crisis/#faceb...
It's played out over and over again, especially with Facebook. They don't deserve the benefit of the doubt, even when "logically" the ROI for the company doesn't make sense. Facebook specifically can't be trusted at the executive level either; that culture trickles down.
You can’t use net profit as a yardstick of rationality in the corporate world, but especially Corporate America. People will seek revenue that results in negative returns on investment. There’s nothing rational about it, which is the problem.
They can increase revenue by making ethically terrible decisions. Full stop.
"We're not actually transcribing every word you hear and say, we're just using capabilities that the Stasi could only dream of at rates of speed they couldn't even conceive of, and it's just that these coincidences happen so often that it looks like we're transcribing your speech. Come on, be less paranoid."
If people think "they listen to me through the microphone" it distracts them from understanding what's actually happening - the whole sequence of things you listed there.
Which means they can't take measures to protect themselves, or campaign for better practices - because they're working off the wrong mental model of how this stuff works.
I told my mom i wanted a vacuum cleaner for christmas - guess what, facebook advertises to all your friends what you're shopping for. "lookalike audience" - to your point.
It's not a common event. They were totally random conversations. That's why I noticed that something is going on. I am not sure if they really listening on the microphone or what else they are doing but it's super creepy and should not happen.
Technology good enough is indistinguishable from magic.
What we need is basically GDPR except that consent can only be collected in the form of a signed and notarized contract.
> Trust is really important. Companies lying about what they do with your privacy is a very serious allegation.
> A society where big companies tell blatant lies about how they are handling our data—and get away with it without consequences—is a very unhealthy society.
> A key role of government is to prevent this from happening. If OpenAI are training on data that they said they wouldn’t train on, or if Facebook are spying on us through our phone’s microphones, they should be hauled in front of regulators and/or sued into the ground.
I find it very difficult to believe that big technology could be held accountable to anything by legal means by regulators. They have too large legal teams, and too well written legal agreements, and eulas.
I find it impossible to believe that I, as a person, could challenge them in any way though legal means.
I would argue big companies care more about losing customers. Сompetitive pressure is more effective than laws and fines. And laws are not without downsides, including becoming barriers to entry into competition.
Corporations run America.
* My Documents can be trained on
* My Instagram images are legally not mine and can be trained on
* My voice can be cloned without consent
* My complete motion capture data can be replicated to another image
* My analytics and usage data is owned by either US or Chinese corporations who either side will train AI on to feed me more ads
* A green colored hardware giant is leading the AI warfare
* The internet is turning highly anti consumer