Ironic.
Ironic.
https://arxiv.org/abs/1910.10683
This included full model weights along with a detailed description of the dataset, training process, and ablations that led them to that architecture. T5 was state-of-the-art on many benchmarks when it was released, but it was of course quickly eclipsed by GPT-3.
It was common practice from Google (BERT, T5), Meta (BART), OpenAI (GPT1, GPT2) and others to release full training details and model weights. Following GPT-3, it became much more common for labs to not release full details or model weights.
Not at all. When you're the underdog, it makes perfect sense to be open because you can profit from the work of the community and gain market share. Only after establishing some kind of dominance or monopoly it makes sense (profit wise) to switch to closed technology.
OpenAI was open, but is now the leader and closed up. Meta and Google need to play catch up, so they are open.
When is the last time they released something in the open?
That is purely the language of commerce. OpenAI was supposed to be a public benefit organisation, but it acts like a garden variety evil corp.
Even garden variety evil corps spend decades benefitting society with good products and services before they become big and greedy, but OpenAI skipped all that and just cut to the chase. It saw an opening with the insane hype around ChatGPT and just grabbed all it could as fast as it could.
I have a special contempt for OpenAI on that basis.
So open sourcing simple models brings PR and possibility of biasing OSS towards your own models.
For those interested in some of the recent MoE work going on, some groups have been doing their own MoE adaptations, like this one, Sparsetral - this is pretty exciting as it's basically an MoE LoRA implementation that runs a 16x7B at 9.4B total parameters (the original paper introduced a model, Camelidae-8x34B, that ran at 38B total parameters, 35B activated parameters). For those interested, best to start here for discussion and links: https://www.reddit.com/r/LocalLLaMA/comments/1ajwijf/model_r...
The question is not about Google but about OpenAI.
Initially? It fueled dethroning MSFT and help gain marketshare for Chrome. On a go-forward basis it allows Google to project massive weight in standards. In extension to its use with Chrome, Chrome is a significant knob for ad revenue that they utilize to help meet expectations. That knob only exists because of its market share.
Isn’t there a whole anti-trust case going on around this?
[0] https://www.nytimes.com/interactive/2023/10/24/business/goog...
Google stood on the shoulders of others to get out a browser that drives 80% of their desktop ad revenue.
How does that not affect GOOG?
https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/
Today big corp A will open up a little to court the developers, and tomorrow when it gains dominance it will close up, and corp B open up a little.
who were easily bought off.
When they were first talking about this, lots of people ignored this by saying "let's just keep the AI in a box", and even last year it was "what's so hard about an off switch?".
The problem with any model you can just download and run is that some complete idiot will do that and just give the AI agency they shouldn't have. Fortunately, for now the models are more of a threat to their users than anyone else — lawyers who use it to do lawyering without checking the results losing their law licence, etc.
But that doesn't mean open models are not a threat to other people besides their users, as all the artists complaining about losing work due to Stable Diffusion, the law enforcement people concerned about illegal porn, election interference specialists worried about propaganda, and anyone trying to use a search engine, and that research lab that found a huge number of novel nerve agent candidates whose precursors aren't all listed as dual use, will all tell you for different reasons.
Models have access to users, users have access to dangerous stuff. Seems like we are already vulnerable.
The AI splits a task in two parts, and gets two people to execute each part without knowing the effect. This was a scenario in one of Asimov's robot novels, but the roles were reversed.
AI models exposed to public at large is a huge security hole. We got to live with the consequences, no turning back now.
It's important there are companies publishing models(running locally). If some stop and others are born, it's ok. The worst thing that could happen is having AI only in the cloud.
Seems like anyone who is releasing open weight models today could close it up any day, but at least while competition is hot among wealthy companies, we're going to have a lot of nice things.
That barrier is the first basic moat; hundreds of millions of dollars needed to train a better model. Eliminating tons of companies and reducing it to a handful.
The second moat is the ownership of the tons of data to train the models on.
The third is the hardware and data centers setup to create the model in a reasonable amount of time faster than others.
Put together all three and you have Meta, Google, Apple and Microsoft.
The last is the silicon product. Nvidia which has >80pc of the entire GPU market and being the #1 AI shovel maker for both inference and training.
The funny part is that the real answer is: Some random French company is running circles around them all.
I mean who the hell just drops a torrent magnet link onto twitter for the best state of the art LLM base model for its size class, and with a completely open license. No corporate grandstanding, no benchmark overpromises, no theatrics. That was unfathomably based of Mistral.