This week, xAI will open source Grok
twitter.com
twitter.com
The engineers likely learned of this news via the tweet.
Did Elon Musk promise this?
They also invited to contribute to that repo... but none of the serious PRs got ever merged as far as I can tell. Basically, they were never serious about doing this.
If e.g. Amazon open sources some part of its software infrastructure should they also open source the data it uses or their configuration files?
What do you mean? There exists only one binding definition of open source
and either some product does satisfy it, or it doesn't. As far as I am aware
> https://github.com/twitter/the-algorithm
does satisfy the open source definition, so your sarcasm looks demagogical to me, but I am very willing to learn something new.
Insert Obama awarding himself meme. Who said that this is the "only binding definition"?
Is grok really noteworthy or is it just a nothing burger?
Why shouldn't we treat LLM weights like LLM creators treat ebooks and open source code? Namely, that it is not subject to copyright?
To say that the Llama training process bypasses the copyright of all the training data creators, and yet the output is copyrighted by Facebook, seems a uniquely pro-corporation stance.
Get inspired and training on prev work, creating something new.
You’re absolutely right. It’s very one sided at the moment.
If we follow their ebook usage practice, it’s not even required that they declare it to be open source. Just need someone to publish their copyrighted work online [0] without their agreement and then - per their rules - it’s totally acceptable to download and use those weights with abandon.
Maybe it could be called “weights3”
[0] I’m not actually suggesting anyone should do this.
Take Wikipedia's content, licensed under Creative Commons - by who? Donald Duck? Then when Pikachu and Tony Stark edit the article it becomes a derived work?
> Creative Commons licenses give everyone from individual creators to large institutions a standardized way to grant the public permission to use their creative work under copyright law.
>....so long as attribution is given to the creator.
Who is the creator I must attribute to?
I don't think any of WP is CC? Without at least a full name and claim of authorship I cant satisfy the requirements of the license? Or can I? Then if I can satisfy attribution I will have to disclose who I am in order to allow further sharing.
When Scratch[0] took off lots of kids re-uploaded things made by others replacing the description with "I MADE THIS"
I'd say we, the grown ups of this world should know we've messed up when kids mock our ways.
[0] - https://scratch.mit.edu
There is a very big difference between knowingly downloading and using illegally distributed copyrighted works vs scraping the internet in general.
And if we can’t have copyright any more then we need to work out how to allow authors to make a living, (and musicians, and artists, and indie software developers in fact…)
I agree it’s less clear cut about content that has been willingly posted to the internet but that’s not really what I’m most concerned about.
I'm sorry, while true you've made to much of a heart warming story from it. It is an ongoing conflict between sharing and not sharing books, published papers, video, audio and perhaps patents should also be part of the scope. On both sides we have both small and large efforts that range from deserving to not deserving our sympathy.
The main beneficiary of not sharing the content of books are the publishers. For the most part they have proven not to care about authors. Much like the recording industry. They will not stop pushing for more and more control if it benefits them.
They really want (and have) my government adopting/creating/preserving/copying(lol?) laws that cant realistically be implemented. They wont be satisfied even if they can get a scheme like the TV license circus with random assholes searching peoples homes looking for a radio or TV (while even the police has no such rights) You already cant play music in public places without paying various kinds of protection money.
They are already scanning your uploads in various places looking for anything that vaguely resembles something else. When they think they've found it you will be punished. They don't care what life will be like after losing your proverbial google account over a false positive. Oh and Google has to pay for it which means you ultimately have to pay for your own investigation and persecution.
People got enormous fines for tiny offenses. There are efforts to filter out websites at the ISP level. Bittorrent is portrayed as a tool for pirates while it is simply a much superior sharing technology.
Many enormous data centers had to be build just so that we can use inferior means of distribution. You ultimately have to pay for that. You got asymmetric internet connections because hey, you don't need to be uploading anything now do you? We are retooling the entire civilization to protect Harry Potter, Shakin that Ass and Plan 9 from outer space and you get to pay for it.
The industries want to sell new works. That agenda also opposes the distribution of existing works. We have a rich history of book burning so that the old may make room for the new.
Personally, the most worrying part is the desire/agenda to breed a population of illiterate consumers who can barely tie their own shoes but should some how run a democracy.
It should be that if anyone shows an ever so slight interest in a topic we ram all the relevant books, published papers, patents, documentaries and tools in their hands and shout: HERE, READ THIS, WATCH THIS AND HERE IS YOUR FISHING ROD.
This is worth twice the military budget. We can find a way to pay authors. It doesn't seem a very hard problem. I'm not sure there really is a need but plenty of people want this so lets make it.
> There is a very big difference between knowingly downloading and using illegally distributed copyrighted works vs scraping the internet in general.
Not really, you cant look inside peoples head. If I buy something knowing it was stolen or pretending not to know it is still a crime.
I guess Gro(q|k) is the new Spar(k|c)[1][2][3][4][5][6][7][8]
----
EDIT: I originally thought xAI was referring to XAI [9]
[3] https://en.wikipedia.org/wiki/SPARC
[4] https://en.wikipedia.org/wiki/SPARC_(tokamak)
[5] https://en.wikipedia.org/wiki/Adobe_Express
[7] https://en.wikipedia.org/wiki/Spark
[8] https://en.wikipedia.org/wiki/SPARC_(disambiguation)
[9] https://www.darpa.mil/program/explainable-artificial-intelli...
It's like that glorious week in 2018 when we got full self driving.
It might be a great achievement if it hadn’t been hyped up well beyond what has been released
And the backend structure with their AI model as well.
I'm not talking about marketing shit
https://www.theverge.com/2023/5/25/23737972/tesla-whistleblo...
Instead what you have is Fool Self Driving as Tesla knows that it isn't fully autonomous yet, but still market it as such.
Intentionally misleading and irresponsible.
George W. Bush
You can indeed buy Full Self Driving™ (FSD), but even then your Tesla is not capable of self driving, fully (eg, there are many scenarios where a human is still required)