Or rather the quality of the training data?
Or rather the quality of the training data?
Calling these models open source is like calling a binary open source because you can download it.
Which in this day and age isn't far from where were at.
Synthetic data can be more diverse if you sample carefully with seeded concepts, and it can be more complex than average web text. You can even diff against a garden variety Mistral or LLaMA and only collect knowledge and skills they don't already have. I call this approach "Machine Study", where AI makes its own training data by studying its corpus and learning from other models.
Mistral models are one example, they never released pre training data and there are many fine tunes.
Is anyone else just assuming at this point that virtually everyone is using the pirated materials in The Pile like Books3?
https://arstechnica.com/tech-policy/2023/02/report-musk-had-...
https://mashable.com/article/twitter-releases-algorithm-show...
And personally I never used Twitter much, but I certainly did not follow Elon Musk when I did - yet I had to see lot's of his posts in my feed. Surely just coincidence.
No, and that's not what the article says either. They were just tracking how well his tweets were doing versus others. They were not favoring Elon.
Yeah, and adjusting it, so he comes out best. That was Musks demand, as the other article shows, that is linked inside, after a Biden tweet performed better than Musk:
https://mashable.com/article/elon-musk-super-bowl-joe-biden-...
They officially boost people, who pay a little bit. Elon payed a lot.
And the source is clearly not the production source and never where in this shape - otherwise why sue someone, who open sourced it?
"But, the release of this source code also comes days after Twitter forced Github to take down other parts of Twitter's source code that was allegedly posted by a former employee without the company's permission. So, clearly, there's still plenty of Twitter that Musk still doesn't want us to see."
Also, you probably missed that:
"Zoë Schiffer of Platformer reported that Twitter actually removed part of the source code that affected the reach of Musk's and other user's tweets before releasing the algorithm to the public."
Which is consistent with quite some other statements, also from Twitter itself and the fact, that the source has not been updated in 8 months.
See also this HN comment and discussion about it:
https://news.ycombinator.com/item?id=35391854
"But the underlying policies and models are almost entirely missing (there are a couple valuable components in [1]). Without those, we can't evaluate the behavior and possible effects of "the algorithm.""
So changes in power users stats would also result in audience balancing?
Most likely the code was used for analytics and for tracking balance; Elon was a pain in the ass and asked to have custom analytics for his account and devs eventually added him as an audience to be able to get analytics about him easily. A bit dirty but it works.
Most likely the balancing code is somewhere else and it affects only republican / democrats.
https://github.com/twitter/the-algorithm
So clearly they aren't running it in production.
Also they didn't open source the list of people who are being artificially boosted e.g. Elon.
Are you sure or is it the literal opposite and you’re just speculating?