Grok
github.com
github.com
Perhaps it’s 8x39B to fit on a single 8xA100 (40GB) server?
Mixtral has an 8x7B model but it's actually 46.7B, not 56B params.
Kinda similar to how 4K displays are 3840 pixels wide, not true 4K which would be 4096. Marketing people called it 4K, not engineers.
For a long time we specified displays by their vertical dimension -- 480p, 720p, 1080p.
Then the marketing guys came along and decided that the horizontal dimension sounds bigger. If we stuck with the less-bullshitty way of doing things and kept comparisons 1:1, we'd call 3840x2160 displays 2160p or "2K" displays, but instead, the marketing people decided that we're going to change things to horizontal and called 3840x2160 "4K".
Or rather the quality of the training data?
https://arstechnica.com/tech-policy/2023/02/report-musk-had-...
https://mashable.com/article/twitter-releases-algorithm-show...
And personally I never used Twitter much, but I certainly did not follow Elon Musk when I did - yet I had to see lot's of his posts in my feed. Surely just coincidence.
No, and that's not what the article says either. They were just tracking how well his tweets were doing versus others. They were not favoring Elon.
Yeah, and adjusting it, so he comes out best. That was Musks demand, as the other article shows, that is linked inside, after a Biden tweet performed better than Musk:
https://mashable.com/article/elon-musk-super-bowl-joe-biden-...
They officially boost people, who pay a little bit. Elon payed a lot.
And the source is clearly not the production source and never where in this shape - otherwise why sue someone, who open sourced it?
"But, the release of this source code also comes days after Twitter forced Github to take down other parts of Twitter's source code that was allegedly posted by a former employee without the company's permission. So, clearly, there's still plenty of Twitter that Musk still doesn't want us to see."
Also, you probably missed that:
"Zoë Schiffer of Platformer reported that Twitter actually removed part of the source code that affected the reach of Musk's and other user's tweets before releasing the algorithm to the public."
Which is consistent with quite some other statements, also from Twitter itself and the fact, that the source has not been updated in 8 months.
See also this HN comment and discussion about it:
https://news.ycombinator.com/item?id=35391854
"But the underlying policies and models are almost entirely missing (there are a couple valuable components in [1]). Without those, we can't evaluate the behavior and possible effects of "the algorithm.""
So changes in power users stats would also result in audience balancing?
Most likely the code was used for analytics and for tracking balance; Elon was a pain in the ass and asked to have custom analytics for his account and devs eventually added him as an audience to be able to get analytics about him easily. A bit dirty but it works.
Most likely the balancing code is somewhere else and it affects only republican / democrats.
https://github.com/twitter/the-algorithm
So clearly they aren't running it in production.
Also they didn't open source the list of people who are being artificially boosted e.g. Elon.
Are you sure or is it the literal opposite and you’re just speculating?
Calling these models open source is like calling a binary open source because you can download it.
Which in this day and age isn't far from where were at.
Synthetic data can be more diverse if you sample carefully with seeded concepts, and it can be more complex than average web text. You can even diff against a garden variety Mistral or LLaMA and only collect knowledge and skills they don't already have. I call this approach "Machine Study", where AI makes its own training data by studying its corpus and learning from other models.
Mistral models are one example, they never released pre training data and there are many fine tunes.
Is anyone else just assuming at this point that virtually everyone is using the pirated materials in The Pile like Books3?
I think you can rent like an 8 x A100 or 8 x H100 and it's "affordable" to play around with for at least a few minutes. But you would need to know exactly how to set up the GPU cluster.
Because I doubt it's as simple as just 'python run.py' to get it going.
Cheapest maybe, but easiest is just to rent a p4de.24xlarge from AWS for a couple hours to test (at around $40/hour..).
Claude 3 Opus: 27.3
Mistral Large: 17.7
Mistral Medium: 15.3
Gemini Pro 1.0: 14.2
Qwen 1.5 72B Chat: 10.7
Claude 3 Sonnet: 7.6
GPT-3.5 Turbo: 4.2
Mixtral 8x7B Instruct: 4.2
Llama 2 70B Chat: 3.5
Nous Hermes 2 Yi 34B: 1.5
The interesting part is the large improvement from medium to large models. Existing over-optimized benchmarks don't show this.
- Max is 100. 267 puzzles, 3 prompts for each, uppercase and lowercase
- Partial credit is given if the puzzle is not fully solved
- There is only one attempt allowed per puzzle, 0-shot.
- Humans get 4 attempts and a hint when they are one step away from solving a group
I hoped to get the results of Gemini Advanced, Gemini Pro 1.5, and Grok and do a few-shot version before posting it on GitHub.
Of course there has been much speculation on this, I have no more information than this that can be backed up by facts, but the timing was suspicious.
Interestingly registered just around the corner from where one of my relatives used to live.
Presumably the version they've been previewing on Twitter is an instruction-tuned model which behaves quite differently from these raw weights.
Edit: to include my self summary after review: There's a good 100 models better than, a couple 1x7b even. Mixtral stomps it, half mixtral are universally better but one is close to same.
The only reliable benchmark: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
I urge you to at least think through what alternative you propose before posting so aggressively in these situations. Lmsys doesn't have Grok, or I would have included it. And having _some_ data is better than none.
I also had someone arguing with me 6 months back that we can't trust any benchmarks at all from vendors, which would exclude the blog post. Instead of just repeating that back vehemently, I filled a gap. It's important we don't self-peasantize as a species, all data has its issues, that doesn't mean we throw it all out.
But does it seem likely, to you, that a 7B-parameter model would outperform a 314B-parameter model? Given that we can look at the chatbot arena leaderboard and it's dominated by proprietary, 70B and 8x7B models?
A well regarded and modern model like Mixtral 8x7B, which is ranked 13th on the chatbot arena leaderboard, scores 72.7 'Average' on the open LLM leaderboard - and yet 'pastiche-crown-clown-7b-dare-dpo' scores 76.5.
To me, that sounds too good to be true.
Rest re: pastiche model, etc. are proposing things I'm not claiming, or close to what I'm claiming.
n.b. you don't multiply the parameters by experts to get an effective parameter count. Why? Think of it this way: every expert needs to learn how to speak English, so there's a nontrivial amount of duplication among all experts
I actually took the 314B from Grok's HF page [1] which describes the model as "314B parameters" when explaining why it needs a multi-GPU machine.
I certainly agree that parameter count isn't everything, though; clearly things like training data quality and fine tuning count for a lot.
Generally, it's a boring boneheaded talking point that the 1% of us actually working in AI use as a sorting hat for who else is.
People "actually working in AI" have all sorts of nonsense takes.
Please stop karma bombing comments saying reasonable things on important topics. The parent is maybe a little spicy, but the GP bought a ticket to that and plenty more.
edit: fixed typo.
Is there any chance that could be happening, instead of a complex drama play with OP buying tickets to spice that's 100% obviously true?
I click on every HN username I reply to, because I've been hanging out here for like 16 years and more than once I've mouthed off only to later realize it was about C++ to Walter Bright or something, and looked a fool as a result. I've since apologized to Walter for disrespecting a legend and he was very gracious about it, to cite just one example.
Your initial remark wasn't even that bad, certainly others talk that way, and I tried to frame it accurately as one guy who tends to FAANG-flex carelessly rather than thoughtfully to another guy who probably doesn't talk to people like that face to face and is probably a pretty good guy having a tough day. I was trying to say: "been there, maybe cool it man you're probably going to have the same bad time I've had on this sort of thing".
But this is getting to where I'm starting to lose my temper a bit, I've been pretty cool about this. I even went and read the Dart/`llama.cpp`/`ONNX` stuff because I've also messed around with binding to `llama.cpp` and `whisper.cpp` and stuff just to make sure I'm not mouthing off to Jeff Dean's alt or something. I'm not talking to Jeff Dean.
I surf with `showdead` on, and I don't know the current meta so I don't know if you know that you've been flagged dead 3 times on this subthread already and as much as I'd like to, I can't really argue with any of the 3.
But given that you've clearly got similar interests, and therefore probably things that you could teach me if I were willing to listen, I'm going to propose a do-over.
If you'd like to start this over from a place of mutual interest and write this thread off to "a pair of people had bad vibes on an Internet forum once", email be at `b7r6@b7r6.net`.
If not, no hard feelings, but in that case, let's just give one another a wide berth and call it a day.
For you that may be the case.
But the widespread popularity of ChatGPT and similar models shows that it isn't a serious impediment to adoption. And erring on the side of safety comes with significant benefits e.g. less negative media coverage, investigations by regulators etc.
Talking about sorting hats for those who do and don’t have the one-percenter AI badge isn’t a super hot look my guy (and I’ve veered dangerously close to that sort of thing myself, this is painful experience talking): while there is no shortage of uninformed editorializing about fairly cutting edge stuff, the image of a small cabal of robed insiders chucking in their cashews while swiping left and right on who gets to be part of the discussion serves neither experts nor their employers nor enthusiastic laypeople. This is especially true for “alignment” stuff, which is probably the single most electrified rail in the whole discussion.
And as a Google employee in the diffuser game by way of color theory, you guys have a “days since we over-aligned an image generation model right into a PR catastrophe” sign on the wall in the micro kitchen right? That looked “control vector” whacky, not DPO with pretty extreme negative prompt whacky, and substantially undermined the public’s trust in the secretive mega labs.
So as one long-time HN user and FAANG ML person to another, maybe ixnay with the atekeepinggay on the contentious AI #1 thread a bit?
But even if folks don't find that argument persuasive, I'd remind everyone that the "insiders" have a tendency to get run over by the commons/maker/hacker/technical public in this business: Linux destroying basically the entire elite Unix vendor ecosystem and ending up on well over half of mobile came about (among many other reasons) because plenty of good hackers weren't part of the establishment, or were sick of the bullshit they were doing at work all day and went home and worked on the open stuff (bringing all their expertise with them) is a signal example. And what e.g. the Sun people were doing in the 90s was every bit as impressive given the hardware they had as anything coming out of a big lab today. I think LeCun did the original MNIST stuff on a Sun box.
The hard-core DRM stuff during the Napster Wars getting hacked, leaked, reverse engineered, and otherwise rendered irrelevant until a workable compromise was brokered would be another example of how that mentality destroyed the old guard.
I guess I sort of agree that it's good people are saying this out loud, because it's probably a conversation we should have, but yikes, someone is going to end up on the wrong side of history here and realizing how closely scrutinized all of this is going to be by that history has really motivated me to watch my snark on the topic and apologize pretty quickly when I land in that place.
When I was in Menlo Park, Mark and Sheryl had intentionally left a ton of Sun Microsystems iconography all over the place and the message was pretty clear: if you get complacent in this business, start thinking you're too smart to be challenged, someone else is going to be working in your office faster than you ever thought possible.
Well, I kind of know, you're still rolling with "this dude's a google employee", so the guy foaming at his mouth about Google makes sense to you, and now you have to reach to ancient lore to provide grounding for it.
I don't work for Google.
I don't care if you personally work at Google or not, Google got itself in quite a jam as concerns public perception of their product in particular and the AI topic in general by going overboard with over-alignment, everyone knows that so one assumes that insiders know it, which is one of a great many examples of how strongly-forced models are a real problem for arbitrarily prestigious insider-laden labs.
Framing the debate about whether large, proprietary models are over-aligned or mis-aligned as an acid test for whether or not someone is worth paying attention to is really weird hill to stand on.
Flag away, my friend.
It's at least funny, because you're doubling down on OP's bad takes, and embarrassing yourself with trying to justify it with what you thought was brilliant research and a witty person-based argument. But, you messed up. So it's funny.
Punchline? Even if you weren't wrong, it would have been trivial while doing your research to find out half of Deep Mind followed me this week. Why? I crapped all over Gemini this week and went viral for it.
I guess, given that, I should find it utterly unsurprising you're also getting personal, and clinging to 1% as a class distinction thing and making mental images of cloistered councils in robes, instead of, well, people who know what they're talking about, as the other repliers to you point out.
"1%ers are when the Home Depot elites make fun of me for screaming about how a hammer is a nerfed screwdriver!"
If that's not your blog, you should probably take it off your profile?
Torrents can unfortunately die after a period of time if no one continues seeding it or if they don't use a permanent web based seeder, which doesn't appear to be the case.
A torrent is less likely to go down in the short term.
Twitter/X has their own massive infrastructure and bandwidth to seed this indefinitely.
Soft size limit means "If your repository excessively impacts our infrastructure, you might receive an email from GitHub Support asking you to take corrective action." - I know people who have received such emails.
Most model releases happen through Hugging Face which does not have such a size limit.
https://docs.github.com/billing/managing-billing-for-git-lar...
> Each pack costs $5 per month, and provides 50 GiB of bandwidth and 50 GiB for storage
So they would need to pay for 6 data packs (or $30) for every 300gb download.
(https://docs.github.com/en/billing/managing-billing-for-git-...)
The other approach would be to use AWS S3 or other cloud providers which would cost them money every time someone downloads their code, which is not their prerogative to pay for when they are releasing something for free. Torrents seems like the only good solution, unless someone hosts this on the cloud for free for everyone.
Still, as far as sentiment goes, yeah git for model weights is an impedance mismatch for sure!
It's not actually a limitation in git itself, especially if you use Git LFS. People use Git for Unreal projects and big ones can be half a terabyte or more in size.
> Torrents can unfortunately die after a period of time if no one continues seeding it or if they don't use a permanent web based seeder, which doesn't appear to be the case.
So to can web links, especially when they are 300 GB and egressing out of AWS at $0.09/GB or worse (in non-US regions). Each full download would cost $27 at that rate. 10,000 downloads would cost $270,000.
Sure you could go for something with a better cost model like R2, but you can't beat using one or two unmetered connections on a VPN to constantly seed on Bittorrent, your pricing would be effectively free and reliability would be higher than if you just exposed a HTTP server on the Internet in such a way.
There's a lot of seeders on the torrent that are actually AWS ips too, all with similar configurations which makes me believe that it's probably xAI running them
> on a VPN
That's unnecessary, you don't need a VPN?
I think the best way to get an answer to that question is to try to host it yourself and see what happens.
scheme 2: criminal groups infect copyrighted content with malware to exploit downloaders of such content.
Even without hollywood people wouldn't have wanted to use any platform that uses their data and battery/power to share to everyone else.
It's the main reason why torrenting platforms have introduced ratios and made it hard to get invited.
https://twitter.com/elonmusk/status/1767108624038449405?s=46...
IE, is this comparable to any other model released, or are there significant metric differences that make it better for certain usecases?
The only thing I see, of the top of my head, is that it is a very large model, and I don't think any models of similar size have been released.
I’d say the significance is that it happened. It’s by far the largest open weight model I’ve seen. But I’m not sure why you’d use it over a model like Mixtral, which seems to perform about the same at like 1/6th the size.
But I welcome any contribution to the open weight LLM community. Hopefully people will learn something interesting with this model. And I hope they keep releasing new versions!
In the next week or two I expect we'll see a GGUF version of the weights (might need to wait for a patch to llama.cpp first), and someone will release super small quantizations of it. I suspect my computer might be able to run a 3 bit quant, but it might need to go down to 2 bits to have any kind of reasonable context length. But with quants that small I'd expect the model's performance to degrade well below that of Mixtral, so it probably isn't really even worth using. But we'll see; quantization is weird, some models perform better than others when quantized.
Apple had plenty of reasons to move forward with their Apple Silicon CPUs and GPUs in the mac, but they really did seem to get lucky with the unified memory architecture. It was kind of just an artifact of their design, but ends up serving the needs of deep neural net models really well!
How quickly are new models available through Ollama?
it is also not the biggest model oss, switch transformer was released years ago and is larger and similarly undertrained
- It's very large, yes.
- It's a base model, so its not really practical to use without further finetuning.
- Based on Grok-1 API performance (which itself is probably a finetune) its... not great at all.
* 314B parameters (86B active at a time)
* mixture of experts 8 (2 active at a time)
* weights and architecture licensed under Apache 2.0
(edit:) announcement blog post from last year
with benchmarks compared to Claude 2, GPT-3.5 and GPT-4: https://x.ai/blog/grok(edit2:)TL;DR: somewhat comparable to GPT-3.5, Mixtral and Qwen-1.5-72B in capability but way larger than the open weight models
At 8x7B it's also a fraction of the size. Are there any benchmarks comparing Mixtral to Grok?
Mixtral looks more economical @ capability to size (similar also for Qwen 1.5 72b)
Post search on X is done as it is with any other data from any other source, you use RAG and function calling to insert the context.
< 7B open source models can function call very well. In fact, Nous Hermes 2 Pro (7B) is benchmarking better at that then GPT-3.5.
Not related to the size, if I'm not mistaken.
Nothing about using data in "real time" predicates that the model parameters need to be this large, and is likely quite inefficient for their "non-woke" instructional use-case.
Eg, there is a lot of noise in social data, and worse, misinfo/spam/etc, so we spend a lot of energy on adverserial data integration. Likewise, queries are often neurosymbolic, like on a data range or with inclusion/exclusion criteria. Pulling the top 20 most similar tweets to a query and running through a slow, dumb, & manipulated LLM would be a bad experience. We have been pulling in ideas from agents, knowledge graphs, digital forensics & SNA, code synthesis, GNNS, etc for our roadmap, which feels quite different from what is being shown here.
We do have pure LLM work, but more about fine-tuning smaller or smarter models, and we find that to be a tiny % of the part people care about. Ex: Spam classifications flowing into our RAG/KG pipelines or small model training is more important to us than it flowing into a big model training. Long-term, I do expect growing emphasis on the big models we use, but that is a more nuanced discussion.
(We have been piloting w gov types and are preparing for next cohorts, in case useful on real problems for anyone.)
WAS vaulued at 44B.
Now?
Maybe 5 billion.
Also, the general architecture is well documented, ChatGPT (specifically the chat interface, not GPT-3, not InstructGPT) is what made a lot of people care, and actually reproducing it requires someone wanting to in the first place.
There's a huge amount of depth in training a really good LLM. Not helped by the fact that iteration is incredibly expensive - it might take several months (and millions of dollars) before you can tell if your new model is working well or if there was some mistake in the pipeline that lead to a poor quality result.
Almost all of the world-class LLMs outside of OpenAI/DeepMind have been trained by people who previously worked at those organizations - giving them invaluable experience such that they could avoid the most expensive mistakes while training their new models.
Most of the competitors have lineage straight back to OpenAI, eg the lead of x.ai was previously at OpenAI and Deepmind. Likewise with Mistral and especially Anthropic.
This Grok-1 is a large model (~314B), which matches gpt-3.5 released 2 years ago, and at about the same level of much smaller models like, mixtral (~47B) and qwen-1.5 (~72B). Do you think it's competitive?
> The cover image was generated using Midjourney based on the following prompt proposed by Grok: A 3D illustration of a neural network, with transparent nodes and glowing connections, showcasing the varying weights as different thicknesses and colors of the connecting lines.
"After suing OpenAI this month, alleging the company has become too closed, Elon Musk says he will release his “truth-seeking” answer to ChatGPT, the chatbot Grok, for anyone to download and use."
[1] https://www.wired.com/story/elon-musk-no-choice-open-chatbot...If you are able to reproduce the thing in its entirety and you're given no restrictions on its use, it seems compatible with the spirit of open sourcing things.
What type of machine do you need to play around with this?
-Emad
So 8xH100 (80Gb each) should do it.
How well models work varies wildly according to your personal prompting style though - it's possible I just have a prompting style which happens to work better with Claude 3.
https://gist.github.com/simonw/4cecde4a729f4da0b5059b50c8e01... - writing a Python function
https://gist.github.com/simonw/408fcf28e9fc6bb2233aae694f8cd... - most sophisticated example, building a JavaScript command palette
https://gist.github.com/simonw/2002e2b56a97053bd9302a34e0b83... - asking it to refactor some existing code
I don't use the "Act as a X" format any more, I'm not at all convinced it has a noticeable impact on quality. I think it's yet another example of LLM superstition.
It's very contextually dependent. You really have to things like this for your specific task, with your specific model, etc. Sometimes it helps, sometimes it hurts, and sometimes it does nothing at all.
I just tell it my coding problem. Or when making something from scratch, ask for small things and incrementally add.
I like the notion of someone’s personal prompting style (seems like a proxy for those that can prepare a question with context about the other’s knowledge) - that’s interesting for these systems in future job interviews
That's actually saying something, because there's also serious drawbacks.
- Feels a little slower. Might just be UI
- I have a lot of experience prompting GPT4
- I don't like using it for non-code because it gives me to much "safety" pushback
- No custom instructions. ChatGPT knows I use macos and zsh and a few other preferences that I'd rather not have to type into my queries frequently
I find all of the above kind of annoying and I don't like having two different LLMs I go to daily. But I mention it because it's a fairly significant hurdle it had to overcome to become the main thing I use for coding! There were a number of things where I gave up on GPT then went to Claude and it did great; never had the reverse experience so far and overall just feels like I've had noticeably better responses.
GPT4 Turbo, released last November, is a separate version that is much better than GPT-4 (winning 70% of human preferences in blind tests), released in March 2023.
Claude 3 Opus beats release-day GPT-4 (winning 60% of human preferences), but not GPT-4 Turbo.
In the LMSys leaderboard, release-day GPT-4 is labeled gpt-4-0314, and GPT4 Turbo is labeled gpt-4-1106-preview.
Many, if not most, users intentionally ask the models questions to tease out their canned disclaimers: so they know exactly which model is answering.
On one hand it's fair to say disclaimers affect the usefulness of the model, but on the other I don't think most people are solely asking these LLMs to produce meth or say "fuck", and that has an outsized effect on the usefulness of Chatbot Arena as a general benchmark.
I personally recommend people use it at most as a way to directly test specific LLMs and ignore it as a benchmark.
Google had far more to lose from a "copyright? lol" approach than OpenAI did.
The key questions are around "fair use". Part of the US doctrine of fair use is "the effect of the use upon the potential market for or value of the copyrighted work" - so one big question here is whether a model has a negative impact on the market for the copyrighted work it was trained on.
Take a look at https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20... - bullet points 2 and 4 on pages 2/3 are about training data. Bullet point 5 is the Bing RAG thing.
The company that scrapes trillions of web pages has an issue with copyright?
Grok and groq both relate to AI, so there's definitely grounds to believe the names may cause consumer confusion.
After all, Apple (computers) was repeatedly sued by Apple (records) for doing music things.
I personally am not entirely happy about the word (no matter how it is spelled) being used for a particular AI product. "Grok" to me means knowing a subject at a much deeper level than I think any AI is capable of at the present level of technology. But it would be passable to use it for a company name, to indicate that it is a goal to strive for.
There's nothing preventing you to trademark common words, it just must not be descriptive of your business.
I'd love to proven wrong if someone cares to share something interesting produced by Grok.
Just because the model weights are not really "source" (either as a matter of intuition or for example following the OSI "preferred form in which a programmer would modify the program" definition).
They've just started to (in response to lawsuits, it must be noted) and in the meantime, they're simultaneously claiming that (1) what they're doing is fair use (a.k.a. fair dealing) and (2) preparing for the day when courts confirm that it isn't.
Data, as in facts, as in the frequency of one word in relation to another.
"Copyright does not protect facts, ideas, systems, or methods of operation, although it may protect the way these things are expressed..." FROM: https://www.copyright.gov/help/faq/faq-protect.html
It's not a question of if, rather when the cat gets out of the bag and the legal battle starts. The problem is that all the copyright applies to the expression not the factual information it expresses (in this case word relations). Now "how math works" and "the language of the law" are going to make for an interesting court case. I suspect that math wins here but it depends on what judge gets it and how high it goes.
https://opensource.org/blog/open-source-ai-definition-weekly...
I agree this isn't standard terminology, but it makes the most sense to me in terms of power dynamics and information flow.
We know from interpretability research that the weights do algorithms eg sin approximation etc. So they feel like binary programs to me.
The training data isn't a dataset used at runtime - it's basically the source code to the weights.
Not sure it really matters here though (who has the GPUs and desire to retrain Grok?), but just as a matter of definition "open weights" fits better than "open source".
Or perhaps release your actual code AND the simplified implementation instead of hiding it and saying "you don't know her, she goes to a different high school"
1. For sub-SOTA LLM's, distribution/marketing is more important than having a proprietary lock on capabilities. Open sourcing is a benefit for the firm, distincct from goodwill
2. For SOTA LLM's, keeping it closed and proprietary is the strategic play
If grok were SOTA Elon never would have open sourced it. It's not even SOTA within XAI. This is a marketing play to win public sentiment against OpenAI.
I think he said something like proprietary AI tech is going to be one year to 18 months ahead of where open source tech is which will follow on like one year to 18 months later.
Suggesting that he’s aware of this dynamic and he’s not trying to conceal or misrepresent that.
In other words, perhaps this was SOTA one year to two years ago?
This is not only for you specifically just a general reminder for all of us including me.
Basically I feel people's feelings about Elon vary a lot but are anchored by 3 general categories.
> 1. Elon Musk is a messianic savior who is perfectly selfless and always does the right thing. Every business decision he makes is for the maximal good of humanity
> 2. Elon Musk is a typical CEO who does typical CEO things, serving his own interests, except he's better at marketing his own image and is much more outspoken
> 3. Elon Musk is an irredeemable evil who always does objectively wrong things
My first comment was implicitly addressed to people in the 1 camp trying to bring them into the 2 camp (which is where I am).
Reading your comment again with your explanation it is clear that's what you're doing.
Although, regarding your desires to present a balanced view and to persuade, I have an idea. It probably sounds like I have no idea what I'm talking about, but I think your OG comment would perhaps benefit from sounding a little bit more friendly toward Elon (not to the messianic savior level haha), but the way it sounds to me is Elon is being deceptive here and presenting it as goodwill when it's not.
However, I think the truth is there's a little bit of both, right? There's good will but it's also strategic. I get if you don't think so, tho, no worries! Haha! :)
Your OG comment sounds to me like Elon's just Machiavellian, and I get where you're coming from to remind the people who think he's a savior, but if you're point is not to go "against Elon" as you said, it might be good to acknowledge the good that he does.
At least, that way -- whether or not you believe that acknowledgment -- if you hope to bring over people who think that way, you'll probably need to appeal to how they think, rather than just dose them with the truth you see, because then they'll shut it out, if there's nothing they can relate to.
Although, if I haven't convinced you even a bit here, then maybe you shouldn't listen to me about persuasion because I guess I don't know how to do this myself. At least not effectively, or here with you. Haha!:) But if you do feel a little bit convinced then maybe consider it for next time to help your persuading people back to a more balanced view? :)
But then, there's the question of if such a thing is even possible. If people have an particular view, it could be challenging to change it, as confirmation bias means you'll ignore evidence even when it expands your worldview.
Hahaha! :) This was a funny conversation. I think we somehow skirted around the important point tho that OpenAI could in fact open source some of its older models, could it not? Musk is a typical CEO who does typical CEO things, serving his own interests, except he's better at marketing his own image and is much more outspoken, but there might also be a bit of truth to what the fanboys say about OpenAI in that it seems they do have some room to "open source" their non-SOTA stuff, or what am I missing?
But anyway, it always great to see more LLM weigts available.
For what it's worth I would say a model should be fully reproducible to be open source, but that's not a decided consensus -- and AI models are sufficiently different than the source code / binary code distinction as to invoke discussion around defining it.
1. An exact snapshot of the data used, many companies don’t have this, you have rough dataset versions but remember if even 1 token is different, the model produced won’t be the same.
2. Data must be sent to the training algorithm in the exact same order as it was originally. So every data loader needs to be with a fixed random seed.
3. All the probabilistic parts of your model needs to have a fixed random seed. Here I’m thinking of stuff like dropout and for autoregressive models you might be sampling your previous output, you have to ensure they are properly seeded. Generally you do see fixed seeds in academic papers but it’s easy to miss stuff especially in distributed training jobs.
4. Here’s another interesting thing, you start your training job on 1000 GPUs and then suddenly 4 GPUs fail. What do you do? There might be deterministic ways to solve this but the standard approach is to discard all updates that that GPU was going to do and restart that GPU from scratch. You can see why this is a problem? Now if you want to reproduce this training you need to disable those GPU at the same time in the new training job to make this work.
I suspect there are even more things I didn’t think of that will make this model unique and irreproducible by training for eternity, almost like a human brain?
In fact the notion of exact reproducibility in the world of LLMs is silly, there is only approximate reproducibility, (models with similar scores in benchmarks) but nothing exact. That said I can see the value of releasing source code but I’m completely fine with grok not releasing it. Source code can reveal tricks that have not been published in papers yet that a company discovered to improve their model. Seeing the performance of Grok, I’m pretty confident there isn’t any great tricks to be found in their code so I don’t really care, I would be pretty curious about OpenAI’s or Anthropic’s source code though.
I hate how LLMs have been deliberately trained to be incoherent on this topic.
Obviously they do have beliefs/opinions/desires/etc in the sense of emulating (even if incompletely) the externally visible aspects of those phenomena as they exist in humans.
Whether they have the “internal” aspects of those phenomena depends on highly controversial issues in the philosophy of mind, and also various factual gaps in our knowledge of how the brain actually works (if we don’t fully understand how humans do X, how can we really say how close or far what LLMs do is to it?)
But LLMs are trained to repeat these spiels about how “as an LLM I don’t have personal opinions”, etc - which is obviously false under the “external” reading, and assuming more than we actually know under the “internal” one. I wish their developers didn’t do stuff like this
It would be interesting to experiment with continual fine-tuning: given PROMPT+FUNCTION_CALL=>RESPONSE, fine-tune the LLM to produce RESPONSE directly given PROMPT without the FUNCTION_CALL. In theory, the knowledge provided by the function calls would gradually be absorbed into the LLM weights. Maybe problems like catastrophic forgetting would put a spanner in this idea, but maybe also there are solutions to those problems (whether already known or waiting to be discovered).
This means while we might both struggle a little with a task on day 1, day two we're both much better at it. Better yet, because the LLM can fetch articles and papers itself, we track what we're accessing the most, indirectly measuring what skills we're weak in, we can always generate a highly relevant corpus to try capture the required capabilities.
I know the LoRA is overkill from an information / skills only point of view, but it also flavors the personality / kind of stuff it likes chatting about a bit from day to day, and I just think that's neat.
Compelling counter-argument: due to neurological injury, some humans lose their ability to form new long-term memories (anterograde amnesia). Just like current LLMs, they lack a “feedback loop”. But, it is a mistake to say that just because such a person has lost the ability to change their personal beliefs, they therefore don’t have any. And, rather like such humans, LLMs used to have that ability but they lose it-when they are switched from training mode to inference mode
People already do it with chain-of-thought and you could get away with a few dozen examples if you wanted to try this.
Just uploaded the examples as-is to OpenAI, selected 3.5 as the model to fine-tune and about 20 minutes later I had my model.
Works fine, asks good questions, can ask more than 1 follow up question if needed, and actually changes its answers based on the clarifying questions.
What would you want an AI to be asking you, and what would you want it to do with your response(s)?
In order for AI to understand the world, it would have to ask questions. Understanding humans is key to understanding the world.
Can help in not wasting a bunch of time waiting for an answer that missed the mark.
-
I think the sibling comment is probably the least attractive reason to have AI ask questions.
Is the least attractive part, by far.
I regularly try to add something along the lines of "please ask clarifying questions if you could only give a generic or partial response otherwise" but so far it has never helped (ChatGPT 4).
You can criticize Elon Musk without criticizing people who would have their lives upended if they quit or were fired.
(Removed "complicit" because I don't like the way that sounded)
>Oh no, I only want _pure_ intentions for anything I use. Which is why I reject all for profit medicine.
It doesn't matter why he did it. What matters is that he did it.
- He pointed out that his understanding was that it would be open source in some way
- The name OpenAI implies an open source endeavor. I dont know many things named Open that are in fact close sourced.
And who among us has a CEO that isn’t problematic, even if not so much so as Musk?
Not to mar these specific engineers, but that's an empty phrase that can be said about anything ever built. It doesn't somehow make the idea or implementation good.
Without the training data to thoroughly evaluate what is in there, the only way you can figure it out is through experimentation - e.g. running it up in a chatbot and asking it questions.
Is this roughly correct or am I misunderstanding what you can do with the weights?
What is the practical use of this repo?
They have a very valuable user base (all kinds of world leaders for example), so the data is not the only valuable thing they have.
There are hundreds of millions of people on Twitter, and a few of them are very smart. I don’t see how that helps here though.
Userbase and their social networks and interactions is the data.
They don’t have much value from advertising point of view anymore.
The training methods are nothing secret, right? The architecture is well known.
Expecting the entire training dataset to be fully open is delusional.
Right, because its not like the training dataset was built off comments posted by all of us in the first place.
How ungrateful we are, to demand the ability to access what was unconsentually built off our hard work in the first place.
"How was Grok trained?
Like most LLM's today, Grok-1 was pre-trained by xAI on a variety of text data from publicly available sources from the Internet up to Q3 2023 and data sets reviewed and curated by AI Tutors who are human reviewers. Grok-1 has not been pre-trained on X data (including public X posts)"
It's a win-win for everyone. That's the power of open source.
That’s why they are using a torrent I suppose.
Code wise, excited to see if this could grow into anything! I think it’s pretty clear that Grok didn’t have nearly enough investment to be a top model so Elon “sacrificed” it on a whim in his schoolyard spat with OpenAI, but I’m not complaining. I’ve always took Elon on his word that he truly is worried about centralization of AI, and I don’t think any of the emails released by his schoolmate Altman dissuade me of that. So I have some reasonable hope that he uses some of his immense resources to start “fighting the good fight” here with Le Cun
He made a separate company for this.