Meta Is Probably Training AI on Images Taken by Meta Ray-Bans
macrumors.com
macrumors.com
Note that the tiny bit Macrumors added is converting TechCrunch’s accurate “Meta declined to say” into a claim of probability.
TechCrunch has updated the story [1] with concrete answers: Meta trains when the user asks for recognition of an image, but not passively in the background.
1. https://techcrunch.com/2024/10/02/meta-confirms-it-may-train...
I have a slightly different interpretation of this: I think Meta want to keep their options open.
I think that’s true of many of these “will they train on your data?” stories.
People tend to over-estimate the value of their data for training. AI labs are constantly looking for new sources of high quality data - but quality really matters to them. Random junk people feed into the models is right at the bottom of that quality list.
But what happens if Meta say “we will never train on this data”… and then next week a researcher comes up with some new training technique that makes that data 10x more valuable than when they made that decision not to use it?
Safer for them to not make concrete promises that they can’t back out of later if it turns out the data was more valuable than they initially expected.
OpenAI has 200M users, and solve over 1B tasks per month interactively. This amounts to 1-2 trillion mixed human/AI tokens. The fact is that every user has their own unique life experience, and a reservoir of tacit knowledge they didn't communicate or write down anywhere else. The LLM can elicit that tacit knowledge that would be otherwise lost, it can crawl our minds for ideas and problem solving choices.
LLMs are in a situation of indirect agency. When they propose a solution, the human usually takes it out and implements it in the real world, and comes back for more help communicating the outcomes. Across many sessions it becomes possible to check what AI ideas worked out and what ideas were bad. This is a huge resource, it collects experience from every user. The LLM becomes an experience flywheel, people are attracted to the best models, and they will get the lion share of this experience.
And yes, you can do it with privacy in mind. You can train just a preference model instead of supervised training on chat logs. Just a model that would pick the right answer from a lineup. This way PII and user specifics don't leak.
1) I strongly double if even one in a thousand people actually bothers to report the outcome of their result. I and every programmer I know certainly don't (if we give up or it works we both just stop asking). And I'd think its one in a hundred thousand who gives detailed feedback that is useful for training. And one in ten million an expert that sits down and talks about its deep knowledge with it (and the problem of how you deduce this from chats that it isn't some crazy person).
2) AI's are extremely confident in their answer, which fools many people, especially those in a void of knowledge. Even if people did tell OpenAI whether every solution worked or not, I would heavily discount the accuracy of such data.
3) AI output autocannibalism does not lead to better outcomes, AI companies avoid using AI data for training like the plague. Mixed tokens I doubt would be much better.
The situation in reality is something like one in some huge number - maybe hundreds of thousands - of mixed tokens can be useful. Of those, those that are repeats of high quality sources like textbooks, dictionaries, man pages - have no value. Of the remaining, there is huge problem of how do you extract these needles in the haystack with high confidence. Given the incredibly lopsided confusion matrix (with the massive amount of actual negatives to actual positives) and this incredibly unstructured data set, I doubt its even remotely possible to find a way where you don't end up with a totally unacceptable ratio of actual to false positives. Letting this kind of garbage data in is how you get Gemini's gasoline spaghetti.
I think not. Say you ask the model to help solve a coding problem. It gives you an idea, you try, it fails, come back and iterate. They can save a note for later finetuning - what worked and what didn't work, using you the user as a validation system for the LLM.
But you might also have your own experience and help the model where it struggles, and finally achieve the task. That is how the model can borrow both your experience and your manual validation work to improve itself.
Some tasks are spread over multiple sessions, or multiple days. They can cluster and look at your progress over time. The latter steps provide rich feedback on the quality of the former steps. Hindsight is 20/20.
Even in chats where the user doesn't perform validation there is rich feedback, people share some of their tacit experience. It's a form of delayed feedback, humans act as caches of unique experience.
The way I conceptualize this is as a search process - problem space search. LLMs can search better with assistance, and humans also search better with assistance. LLMs collect experience from millions of people, they funnel experience into their logs.
That’s not accurate. All of the big LLM training labs are leaning increasingly into deliberately AI-created training data these days. I’m confident that’s part of the story behind the big improvements for tasks like coding in models such as Claude 3.5 Sonnet.
The idea of “model collapse” from recursively training on AI-created data only occurs in lab conditions that very deliberately set up those conditions, from what I’ve seen. It doesn’t seem to be a major concern in real-world model training.
I do not think this is at all accurate. Sure quality is increasingly important, but that is just the pendulum shifting only slightly back from the fact that quantity is the primary thing you need.
https://twitter.com/karpathy/status/1797313173449764933
> Turns out that LLMs learn a lot better and faster from educational content as well. This is partly because the average Common Crawl article (internet pages) is not of very high value and distracts the training, packing in too much irrelevant information. The average webpage on the internet is so random and terrible it's not even clear how prior LLMs learn anything at all.
Further, you can generally at least bootstrap on shitty data a model that lets you pan through your data for the higher-quality gem.
> it doesn't seem like there would be a reason for Meta to be ambiguous about answering, especially with all of the public commentary on the methods and data that companies use for training.
There is another possible reason and it's not that hard to think of: they want to keep their options open. It's also possible that they don't want to play the game of "anything not explicitly
My guess is that they are currently (or will in the future be) training on it, but I don't think we should take these statements as evidence of that.
In general I think big tech is atrocious with private information, and if the average person knew the depth of the data they have, they would not stand for it. Certainly Meta doesn't have a great track record on the data front so there's no reason to think they'd be different. Unfortunately the average person thinks people like me are paranoid and/or crazy when I try to tell them about it. At best they just feel powerless and shrug and keep using the product anyway.
I would assume that since you can enable AI and “improve the AI” that that’s data that’s fair game for training.
But when you take photos using the capture button or saying “Hey Meta, take a photo” the photos don’t even use cloud storage by default. You have to specifically turn on Meta’s temporary cloud storage feature to sync data while charging with your iPhone to work around iOS rules. If you don’t, those photos are just local only.
There are instances where questions can be answered by just using the product and reading the legal documents. I think this is clearly just lackluster research from TechCrunch carried to MacRumors.
The real question is, who's paying $500 for the privilege of being a willing mule for Zuck's surveillance empire building dreams?
Now it is just: 'we care about your privacy, allow us to sell it'.
Thats core of their market value, just think for a second what is valuable to them and what they couldn't care less about.
As seen with their social media products targeting less savvy consumers, will this product cross into the mainstream, after early tech savvy people (innovators)?
Mark is betting they wouldn't, like with other tech using their data (ex: Android).
In what universe would Meta not use the data it collects to improve its AI models?
Most consumers don't care about privacy implications in the abstract, so they won't even think of asking Meta to stop.
Most tech people working with AI want Meta to continue to improve its open-source, open-weight Llama models, so they will be reluctant to ask Meta to stop.
If they do, then of course they will.
Do they? You be the judge. https://www.meta.com/legal/smart-glasses/
I judge anyone who made their money there because they simply have made the world a much shittier place just by existing as some low quality 21st century nicotine dealership.
These glasses will 100% be used to ID, track, and ad bucket tag people without their consent. I'll be slapping them off anyone's face who looks in my direction and you should too.
Versus mobile phones, Meta are making a better go at, and seem to be the leader on, the bet that fundamentally ML/AI-driven glasses are going to be the next default modality for UX on the Internet.
Of COURSE they will use whatever data they can get.
https://www.theguardian.com/technology/2018/apr/17/facebook-...
I'd say its a worthy battle to wage, even if ie via EU but they need to step up the fines t be really punitive and demotivating for clearly long term amoral / illegal practices, ie tens of % of global revenue (income can and is trivially gamed for those of such size).
It's not been possible for a corporation to collect images and videos from thousands of sources and use facial recognition and AI to track the movements of many people in public and be able to associate that with their online activities, credit reports, and other records. Information that will then be subject to subpoena or worse, depending on the desires of governments that may eventually be in power in the future.
I mean it’s Facebook, we get it. But they make WhatsApp and they make affordable and actually working Quest and now glasses… I’m not going to get triggered by opinion posts just because the Vision Pro was a flop.
This would sound exactly the same if we say Apple is training their AI on everyone’s Photos and content from notifications (because Apple Intelligence).
The current status regarding electronic devices is:
- If you have a pager or walkie-talkie, assume that it might blow up.
- If you have a smart phone, assume that it records your conversations.
- If another person has these RayBans, run don't walk.
More and more people know this. Even non-technical people are beginning to wake up.