X/Twitter has updated its terms of service to let it use posts for AI training
stackdiary.com
stackdiary.com
AFAIU neither of those are relevant to GPT-like architectures but it's not inconceivable to think there might be a model architecture in the future that takes advantage of those. Purely from a information theoretic POV, there's non-zero bits of information in the timestamp and relative ordering of tweets.
In which the distinction between "data" and "information" is crucial. Especially now that the "floodgates" have been re-opened regarding misinformation, bots, impersonators and the likes.
Data is crucial when in need of training body. But information is crucial when the training must be tuned, limited or just verified.
Current LLMs are trying to predict typical human prose from samples pulled from the internet. So it isn’t as if they are sacrificing quality for quantity. A bunch of text from the internet is a very good representation of typical human prose. Whether it is well written or the descriptions contained in the prose accurately represent, like, actual physical reality is another issue.
Maybe they want to predict something with, like, less dimensionality but more utility than a paragraph of fiction.
I guess X is harder to scrap without permission.
We introduce new datasets derived from the fol- lowing sources: PubMed Central, ArXiv, GitHub, the FreeLaw Project, Stack Exchange, the US Patent and Trademark Office, PubMed, Ubuntu IRC, HackerNews, YouTube, PhilPapers, and NIH ExPorter. We also introduce OpenWebText2 and BookCorpus2, which are extensions of the original OpenWebText (Gokaslan and Cohen, 2019) and BookCorpus (Zhu et al., 2015; Kobayashi, 2018) datasets, respectively.
From https://arxiv.org/abs/2101.00027 (The Pile: An 800GB Dataset of Diverse Text for Language Modeling)
And there's my incentive to stop posting on HN.
It's been a blast, guys. I'm going back to lurker mode.
Everybody smile for the camera, or we could just moon them, or both!
I don't think Musk is the type of person to make the same mistake, so we'll either end up with a Twitter LLM that accurately represents the sum total of the Twitter firehose, and/or many derivative LLMs each having a set of, possibly orthogonal, biases. Honestly, I think the later is preferable and would represent the diversity of opinions in reality more accurately.
Given the data source, I think it will be important to be able to switch between LLM personalities in the future to get the "crowd truth".
We need an xkcd showing a conversation between twitter, reddit, and hacker news based LLMs. Political rage meets memes meets pedantry.
1) Facebook Posts/Comments, 2) Instagram Posts/Comments, 3) Youtube Comments, 4) Gmail content, 5) LinkedIn Comments, 6) TikTok contents / comments
X and Reddit are definitely valuable, but they're definitely not unique. I think Meta and Google have inherent advantages because their data is not accessible to LLM competitors and they have the actual capabilities to build great LLMs.
Unless X decides to tap AI talent in China, they're going to have a REALLY hard time spinning up a competitive LLM team compared to OpenAI, Google, and Meta, which I think are the top three LLM companies in that order.
I mean, could it be that it's just that the platforms you're familiar with are similar quality? There are major quality differences. Consider, for example, HN vs Instagram. Do you really see no difference in the quality of discourse, or do you just not use Instagram?
By bulk/raw data volume, I'd say that the vast majority of internet communication is the same quality, yeah-- I'll stick by that assertion. That's not at odds with acknowledging there exist locations where intelligent communication happens. My position is just that the signal to noise ratio is pretty bad in the majority of places.
Is probably LLM poison.
> 4) Gmail content
Is huge but also has enormous privacy issues. Most people by default assume their emails are reasonably private, whereas most people wouldn't assume their comments on these platforms are private.
https://www.youtube.com/@HyperspacePirate
https://www.youtube.com/@scottmanley
That's high quality content, timestamped and about current events.
There is very little content on Twitter that compared in quality to one will written news article.
There’s not a lot of data in Twitter today resembling long-form content: essays, news articles, books, scientific papers, etc. That’s probably why Twitter/X expanded the tweet size limit, to be able to collect such data.
Its strength is freshness and volume, but I guess these can be achieved without Twitter if you have a strong web crawling infrastructure? Also, the current generation of LLM is not really capable of exploiting minute-level freshness... at least for now.
You can bet Google/Gmail/YouTube, Amazon, Microsoft, TikTok, and every other Internet platform that works with user-generated content will soon do the same ... if they haven't done so already.
If you pay, corporate versions of Google Workspace won't train on your data. That's very much by design, since companies don't want anything internal ever being exposed.
But with the free version, that's part of what you're "paying" for it to remain free.
https://9to5google.com/2023/07/03/google-privacy-policy-ai-t...
Although judging from that, it sounds like Google may have changed policy so that even private data in the free versions is no longer used for AI training.
They stopped targeting ads in free Gmail based on your e-mail contents years ago because of the bad press. So maybe they've stopped training AI on free Gmail/Docs data out of similar precaution, now that LLM's are everywhere in the news.
Individual lawsuits are better for everyone.
"Contract formation is increasingly scrutinised. Following Concepcion and its progeny, some courts have focused on issues of contract formation to determine whether the consumer in fact agreed to arbitration and the class action waiver. This inquiry is largely confined to online transactions, where a consumer is deemed to have consented to arbitration by using the business's website to purchase goods or services. These contracts fall within the rubric of "clickwrap," "browsewrap," or "webwrap" agreements and their enforceability is beyond the scope of this article. However, it is important to note that the courts will refuse to enforce class action waivers and arbitration agreements in such agreements when the arbitration provisions were insufficiently conspicuous to ensure the consumer objectively agreed to their terms."
The article also mentions that non-negotiable consumer contracts are viewed with more suspicion by some courts.
1. https://content.next.westlaw.com/practical-law/document/I9f1....
How do you know that the AI won't be used to sell people's attention to advertisers?
The US doesn't have loser pays and has some of the most expensive litigation in the world, which has created all kinds of problems. Someone can file a lawsuit against you knowing that they're unlikely to win, but in so doing they could cost you hundreds of thousands of dollars for lawyers, so why don't you just go ahead and settle for tens of thousands of dollars? It will cost you less to settle than to win in court.
This flaw was made to scale by class action lawsuits, which more than any other should be loser pays, because there is little question that thousands of people who have each been harmed to the tune of $100 could each front $10 for a meritorious lawsuit. But instead you get opportunistic lawyers signing up anyone they can find for questionable claims, so they can reach a settlement where the plaintiffs each get $7 -- or a $7 gift certificate -- and the lawyers get millions.
This was rightly regarded as a problem but the lawyers had enough political power to prevent a good solution, so what we got instead was to make it easier to force binding arbitration and opt out of class action suits.
Lawyers ruin everything. They even ruin lawyers.
Facebook will be known as y.us.
"Replacing some shoddy heuristics with a massive AI model" seems like product improvement to me.
Therefore, they were already allowed to train AI models with your data.
This often works to his advantages; the 'always believe' are loud.
1. didn't know what he was talking about
2. was over-optimistic about timelines to the point that there would be very little difference if it were a lie
3. objectively lied on purpose in order to further self interests
4. exaggerated or over-promised more than could be attributed to salesmanship
5. flip-flopped and refused to acknowledge the change in stance
6. stated things entirely to be vindictive, and let the truth of the statements be irrelevant
The problem isn't that people 'never believe', the problem is 'whatever is said should be suspect to the point that if it is actually true then that is entirely a side-effect of the statement and not the point of it'.
You seem to be unaware or Elon time or how Elon has made predictions in the past, generally not that "this will happen by date X" but that "this cannot possibly happen _before_ date X". Which are very different statements.
There is a contingent of people who want to retcon any statement ever to paint Elon as a fraud, when reality is much too subtle for trivial blanket labels like you want to apply.
At this point, there are shades of him being E Lon Hubbard to the remaining believers; for most of them it seems totally unshakeable.
(Humanity was extremely fortunate that the existence of L Ron Hubbard and the existence of the Internet did not significantly overlap.)
Which, as I recall, he tried but the courts said 'no'.
How many times have Trump been caught lying or doing a 180 from one day to the next? And there are still tons of people believing him.
CEO can use it on Elon to keep him placated
I'm honestly surprised that anyone on HN wouldn't understand this.
Maybe I'm overthinking it and this is just a rhetorical gotcha.
After all, there's definitely no market for a product that has been finely-trained on social media posts to the point where it can perfectly ape social media posts. There's no money for Twitter to make there, so the technology it produces from this analysis shall certainly never find its way into spam bots that create content indistinguishable from genuine humans on Twitter.
It doesn't matter though. It's not just a matter of lock in. There are other reasons.
For one thing, data isn't really data in private chunks. It's only valuable collected.
What if the consumer action was poisoning the dataset ?
"All capabilities are built from data analysis and the data they are analyzing is you. Whether by direct intent or not, all advancements are encroaching on consuming every knowable fact and inference about you. Your soul must be sacrificed to the machine to grant the powers it manifests."
Not that I'm particularly worried but this is a good reminder that I need to blank my twitter history.
[0]: https://redact.dev/
(In fairness, the AI trained on LinkedIn data might be worse.)
Any ideas how we, the users of the internet, should act to prevent those steps? or building a much better connected and not controlled by mega-corporates INTERNET social platform?
Don't use services with ToS you don't agree with.
Understand that nothing is more sacred than the Holy Dollar to an entity whose only purpose is to extract money from its customers.
Always assume you'll be betrayed by corporations and store your possessions (data in this case) accordingly.
Diversify your interests and hobbies such that a betrayal by a corporation doesn't significantly impact your life.
At this point, trusting corporations without putting any mitigations in place seems like begging desperately for trouble.
These things need to be shorter and written at a 8th grade writing level, and I don’t mean that in any pejorative way. It needs to be clear to the average person (and I’m not even sure if I’m being generous there by suggesting 8th grade writing/reading level).
This is like Cold-war between companies, that we as users are the cannon fodder...
The prompt should read “So, you cool with that brah?”, and the word “Accept” needs to be replaced with “Yeah sure why not”, and “No” with “Nah, no thank you”.
Current TOS’s don’t provide enough clarity.
That way any semantic issues down the line can be boiled down to “brah, users said they were cool with it” (to the judge).
There's plenty of places where you can pay a nominal free to have your data 100% controlled by you alone. Plus, you can still use X, Insta or whatever to drive traffic to your site. Its not that hard folks!
Do you happen to have a list handy? If I had to do this today the only clear way would be a personal cloud setup.
What I want to know is, what's the endgame? We've proven that AI learning eventually plateaus, so you keep feeding it data and...? Seems like an underpants gnome problem or companies trying to find another stock bump for all of this data lying around (that should be considered toxic instead)
X and other social networks will do with anyway and you're just waiting for all the other social networks, like Facebook, Instagram, Threads, LinkedIn, Reddit, TikTok etc to follow suit and train their AIs on your data.
The solution? If you have a problem with it and you haven't done so already, just delete your accounts on ALL social networks.
It casts the dismal level of governance and regulation of online life in the unforgiving light of an X-ray. Let those skeletons be clear to all.
If you need to write X/Twitter to make people understand what are you writing about, this means the rebranding was not good
In twitter's case, they built a lot of culture around "tweeting" and the bird theming. The fact is that none of those are going to change soon.
Source?
Obviously, there can be extrajudicial consequences though.
Aaron Swartz was not "imprisoned for years".
Weev "exposed a flaw in AT&T security in June 2010, which allowed the e-mail addresses of iPad users to be revealed.[39] The flaw was part of a publicly-accessible URL, which allowed the group to collect the e-mails without having to break into AT&T's system.[40] Contrary to what it first claimed,[41] the group revealed the security flaw to Gawker Media before AT&T had been notified,[40] and also exposed the data of 114,000 iPad users, including those of celebrities, the government and the military"
That's a pretty silly security breach, but it's still a real security breach. Not comparable to scraping twitter
ai hype is exhaustingly stupid.
I'd go right for /b/. Perfect.
Train your AI on that ...
Respectfully speaking, fully in accordance with HN terms, conditions and guidelines.
He already did.
Twitter (I will not call it X, because that's just stupid), is free to attempt to change their Terms of Service, policies, etc, but we do not have to accept it or agree with it or be resigned to it. Also, it should not be retroactively applied to past content, and it should be an opt-in consent -- but that is pie-in-the-sky wishing at this point given the garbage heap Musk, and others, has made of Twitter.
Also, it's pretty obvious that everyone was training on Twitter data already before they cracked down on scraping. It is, after all, a public forum.
Realistically, a company is only obligated to pay the fine if they have a physical presence within the EU, and care to keep that presence.
For some companies, the calculus says pay the fines and cooperate with EU laws.
But for most companies, they can and do safely ignore GDPR and other EU laws. EU laws do not apply outside the EU... despite what many Europeans want to believe.
You’re right but when did you meet Europeans who said that a company selling in India has to respect GDPR? You only need to apply EU regulations if you serve EU citizens, or you get either fined or blocked
I will note that we've narrowed the claim from "you're free to delete your content" to "If you live in some countries you're very likely to be able to delete your content", which I agree is probably true.
Europeans seem to believe things like GDPR apply to the entire world. They don't.
If your company has no physical presence within the EU - ignore EU laws as much as you want. There is nothing they can do about it.
Keep in mind that search engines and other third parties may still retain copies of your public information, like your profile information and public Tweets, even after you have deleted the information from our services or deactivated your account.
And I'm not even sure it means anything even for users in the EU. If Twitter doesn't have any offices/subsidiaries/bank accounts in the EU, then even if the EU fined them, I'm not sure how that would ever be enforced?
Looking online I can't find any recent information.
Basically, given the way Musk has been ignoring other regulations and/or not paying for things, I'm wondering if he even cares about GDPR. And if he doesn't care and shuts down any legal European presence, then does it matter?
(Of course if the Ireland office is still active and receiving lots of European advertiser revenue, then of course the GDPR has teeth.)
Shutting that office down and safely ignoring the GDPR is probably a valid concession for a US-based business built around data collection.
It's not like EU users will stop or be blocked from using X anyway.
Years ago, "dark triad incarnate" was checked by lack of nearly as established a position* and a longer timeline over which to grow and consolidate wealth and power (which does require time and attention).
He's past 50 now, and started transitioning into his own "late Putin phase" (substitute your own 'favorite megalomaniac' at will) quite aggressively in the past 5 years (especially). Now the game is using that wealth and power for a kind of ultimate "spoiled child fantasy camp".
Regardless of how directly any of them channel the childishness of the archetype, the traits are always there - "I'm special", "your (parent-style) 'rules' don't apply to me"**, "I will get my way", etc. It's the whole point, and the motivation that people who don't think this way miss. The motivation that makes the behavior make at least some sense.
Musk is one of the real extreme examples in terms of how transparent the behavior is - whenever he does something that seems hard to explain, ask yourself how the situation might look "in a sandbox". Seriously. This may sound like typical rhetoric, but I'm serious: try it. Twitter is a perfect example - "if I can't have it my way, then I'll make sure no one can have it" ...
... and, just like in the analogy, there are layers of goals. I.e., it's also good if while we (may ultimately) destroy "the sandbox", we can use it to harm those we don't like who've been playing in it. Either directly (e.g., firing employees of Twitter), or in various indirect ways (reporting "troublesome users" to their authoritarian governments [when applicable], etc.).
* Specifically, still needing something from others here and there - most recently and likely the final example: funding for Twitter deal
** People who think this way can't help 'telegraphing' - it's one way they identify members of their own flock, in part. "Nanny state", "snowflake", etc.
"Snowflake" has got to be a personal favorite. Every time someone uses that one, I know I'm going to need a WHINE break after a few sentences... https://youtu.be/tl4VD8uvgec?si=H2MadAVDduLfolfS&t=1m17s
You cannot expect people to retroactively consent to a thing the existence of which they were not even aware of when they gave consent for some limited OTHER use.
Besides, it's just rude to do things for which there was no consent in general, no matter whether it is AI or anything else. No consent is no consent.
Further, you can't expect people who have a life in general to sit around all day just waiting to figure out where they have to delete accounts before they are misused.
Bazillions of websites nowadays want you to create an account, the normal behavior is that users just abandon accounts which they don't need anymore. Nobody has the time to delete all of them.
Your profile says you were "Senior Director of Monetization at Reddit". Considering all the outrage that company has caused with its users as well (redesign cough cough - just one example out of a truckload), and that the outrage-causing things largely seemed to be aimed at monetization, perhaps you should do some soul seeking to figure out whether your values are aligned with common societal morals.
Or in other words: How many more people will you make angry until you realize that maybe you're the baddie?
Ask yourself whether that would be a place worth living in, or rather a hellscape.
https://youtu.be/mH3La3RJdNA?si=qIojw1j5NkH0pBCj
... cause that seems like a pretty kickin' party to me.
Oh, wait, there's nothing legal about some of what goes on at those concerts. So, not even that level of fun. Gotcha.
So, surely, you can't change ToS retroactively and expect that any of it applies?
By adding that they can train AI they are trying to get out of a future lawsuit that may happen if the courts require consent for training (current anyone can train on anything).
Your right to sue them for using the data they hold might be lost if you continue to use after the term change.
If they asked for your car and you refused they could stop service but they can't take your car.
They can't change payment terms from the past and sue for them. But if they change the tos to say it costs more now your next bill will go up. If they say they can use your data now that they hold and you have an active account they could take that as an acceptance that past/future can be used to train ai.
A better example might be a right given. For a year you could download photos for AI training. Today they forbid that for all future and past posted photos. Anything downloaded before the date can be legally used to train.
A poor, or even no, moral compass.
I'd lean as far out of the windows to speculate that what has often been said about people in positions of power applies here:
Those positions attract people who are completely unable to perceive empathy, and who act solely out of the desire for power and narcissism.
It's a shame, you cause so much harm for society - and you probably are unable to even perceive the harm you're causing because your brain is just not wired to be capable of empathy.
If you want to do the world a favor, go read up what a psychopath is, and by that I do NOT mean to insult you, but rather the actual medical term "psychopath".
Ask yourself whether it applies to you, and learn to protect society from yourself if it does.
Well for a lot of us humans, ownership is not a stopper for conversations about ethics in technology.
> insufferable whiner
I don't think this is a proportionate response to a convo about ethics in tech.
The ethical discussion is kind of the entire point. We all know what the law is.
Source?
> he really has been screwing up bad, and in ways that are indefensible by anyone with common sense.
Examples?
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful. You may not owe $CelebrityBillionaire better but you owe this community better if you're participating in it.
If you want to keep your thoughts private, then maybe don’t post them publicly?
If WhatsApp or another private messaging app started doing this, I’d be right there with the people calling that absolutely unacceptable.
But I’m not surprised at all that Twitter is doing this, and I don’t know how anybody even remotely tech savvy could be.
This seems so critical to a functioning society that one would have thought it would have been considered in the Constitution. Oh well!
But by agreeing to the terms of use, Twitter retains certain rights over what you post on the platform.
If that is not acceptable, then don’t use Twitter. If you think your thoughts are too valuable for Twitter to use, then write them into a book or blog or some other venue where your intellectual property can be protected.
I do not think there are any violations of intellectual property law given that there is surely a waiver of ownership of posts in the TOS.
I, of course, do not have the kind of free time required to do something like engage with Twitter, and accordingly I have no account, cannot post, and have not agreed to the TOS.
I think you have misconstrued my post, but that’s ok.
>Senior Director of Monetization at Reddit
You of all people should know that "free to" and "should" and consent by default is perfectly legal. And yet, it's also gross and slimy. So I guess I'm not at all surprised to find out you're a monetization person.
Gross.
the data landscape is ever-evolving, and what was acceptable or even conceivable years ago may not be the same today.
companies should not only be transparent but also dynamically update users on how their data is being used and offer an option to opt-out.
ignoring this not only impacts individual users but also has broader societal implications.
consent fatigue is real; expecting users to keep track and delete their accounts across numerous platforms is neither practical nor ethical.
also, cancelling your account or laboriously deleting all of your content doesn't necessarily guarantee that all your data will be deleted on the backend... did you think your comment through at all?
Corporations have more power in our government than individuals. They have a better understanding and coordination. Acting like a public forum owned by a company is immune from criticism because it's a private company is sweeping so much under the rug.
I imagine you wouldn't be where you are in life if you didn't believe such things, though.
> We are a separate company from X Corp, but will work closely with X (Twitter), Tesla, and other companies to make progress towards our mission
Even they call it Twitter!
[0] https://techcrunch.com/2018/03/23/elon-musk-deletes-own-spac...
LOL, it's the latest craze to change company names. When I see their "new" logo somehow my mind immediately associates it (correctly) with the X11 logo. Facebook another one that decided to change its name for something that maybe turns out to be biggest money burn a company has ever done. Maybe tomorrow we will wake up with Pear instead of Apple, who knows. Now that I mentioned FB, what's the current status of the so called Metaverse? Are we there yet? Or are they still furiously pouring millions and millions and getting nothing out of it?
Like "the artists formerly known as" Prince, Kanye, Snoop Dogg, etc. There's basically no getting away from the old branding because it has to be included with the new branding so one knows what we're even talking about.
As far as rebrands go, X just seems dumb. The more an article/news segment talk about X, it feels like an unfilled mad libs made it to air. Or it feels like they're talking about something general, like when X Company does Y thing.
If they attempt this will open them to lots of lawsuits
Open source can liberate us from this, but we need someone to build really good and competitive alternatives to Twitter, Zoom et al.
I started Qbix to do it. LA Weekly just published this piece about my company and what it’s doing differently: https://news.ycombinator.com/item?id=37353229
Capitalism is characterized by PRIVATE ownership of the “means of production”. That’s the term used in the 19th century, but today we could point to the technological infrastructure which enables each new user to engage with a network.
“Ownership” means exercising exclusive control over this, and excluding others from using (even a copy of) it.
Musk controls Twitter. Zuck controls Facebook. Durov controls Telegram. Moxie controls Signal. And so on. This is centralized control by people who won’t give you their back-end software. They’ll at best let you have your own custom client for a while, until they don’t (Reddit).
But in the meantime they’ll spy on you everywhere so they can mine your data and try to extract profits for shareholders. It’s called surveillance capitalism: https://en.m.wikipedia.org/wiki/Surveillance_capitalism
Cory Doctorow recently wrote about the “enshittification” that happens as the end result of all this private ownership. “I built it — I own it!” Well, if you believe that, you shouldn’t complain when a privately owned company does something, not even when they deplatform you. What you should complain about is the lack of open source alternatives.
Does Linus own Linux?
Does TimBL own the Web?
Does Rasmus Lerdorf own PHP?
Does Vitalik own Ethereum?
Just because one specific company in an ecosystem is privately owned does not mean the network infrastructure is centrally controlled by a few people.
In fact our company has experimebted with ways to reward contributors properly:
https://qbix.com/blog/2016/11/17/properly-valuing-contributi...
Wordpress, Drupal, Magento, Linux etc. can be hosted anywhere. It is a free market. By contrast, Twitter and Facebook (oh sorry, X and Meta) are digital feudalism!
https://qbix.com/blog/2021/01/15/open-source-communities/
We also are working on utility tokens that, unlike shares, entitle people only to services in that free market, and not to expect rents to be extracted forever. If Qbix or Automattic extracts too much rents from their open source ecosystem, or doesn’t do the best hosting in, say, Hawaii, then a competitor can arise and compete with them, locally or globally.
In fact, Qbix can be used to host social networks in areas with bad internet, including rural villages, cruise ships and planes. They can help young people of all sexes be educated in rural areas with bad internet. Can the same be said of Google or Facebook? NO! Their capitalist ideas always involve sending the signals back to their own server farms. Whether it’s Project Loon (google) or the solar-powered drones (facebook), what they don’t offer is local villages to simply load their own forked copy of their backend software, and owe them nothing!
We do. We give the source code away and help hosting companies install it. We are working on creating an entire decentralized ecosystem where we don’t have centralized control … so if host locally, you NEVER have to worry about us training our AI models on your data, or any of the other thousands of things to ebtray your trust. It’s YOUR choice who will run your infrastructure — and it could be your friend on a local computer and connecting your town over a mesh network:
It's a more vibrant, open, and honest community than ever, despite the organized and coordinated (I wonder by who) advertiser boycott. If anything, a lot of garbage has been removed from the heap.
People have been plenty clear about why they find him disgusting, if you care to look.
https://news.ycombinator.com/item?id=37352719
The Register has started calling it Xitter. I like that!
Quite a few years ago I remember working in Paypal's X API. Part of me wondered if I misremembered this, but no ... there are still references to it online. Maybe Musk named it. He wanted to name the entire company X, right?
https://www.paypalobjects.com/webstatic/en_US/developer/docs...
Those items along with your posts and social network will be used for advertising, monetization, tracking, selling services to government agencies, and training their machine learning bots.
Not for your benefit, but for X's uses and benefits. You are the product and X wants to sell your info and attention so Elon can earn his money back.
Of course they’re doing ML training on user data! Every major tech company has been doing ML training on user data for over a decade at this point! Twitter used to have official tutorials on how to do ML on data you pulled from their APIs!
Seriously, how the f** do you all think their moderation algorithms worked before? Unicorn farts? Some guys in a room that could read tweets really fast?
The only thing that’s changed is that they feel the need to put it in their terms of service.
And no, this isn’t a defense of Twitter or Musk. If you actually give a shit about this and don’t just want to perform slacktivism on social media, get off ALL of these platforms. Migrate away from Gmail. Close your Facebook. Use paid services from trusted privacy-focused vendors.
https://www.popularmechanics.com/technology/robots/a43126181...
Truly I do not understand why they are not even attempting to advantage of their competition actively driving people away. And they have history of Google+ to reflect on! But they're making the same mistake.
I spend a few days on there after getting the invite and didn't enjoy the experience. Every other post was making fun of Twitter/Musk and for some reason there's a lot of furry porn there.
An open platform and open protocol makes it harder to prohibit AI bots from ingesting your published thoughts than when on a private, centralised service.
Mastodon, the fediverse, Bluesky, are really enabling AI learning more than prohibiting it.
¹ well, BSKY is really still just a single server.