GPT-4 details leaked?
threadreaderapp.com
threadreaderapp.com
With the original source being: https://www.semianalysis.com/p/gpt-4-architecture-infrastruc...
The twitter guy seems to just be paraphrasing the actual blog post? That's presumably why the tweets are now deleted.
---
The fact that they're using MoE was news to me and very interesting. I'd love to know more details about how they got that to work. Variations in that implementation would explain the fluctuations in the quality of output that people have observed.
I'm still waiting for the release of their vision model which is mentioned here but we still know little about, sans a few demos a few months ago.
"MoE" in the context of artificial intelligence typically stands for "Mixture of Experts". This is a machine learning technique that is based on the idea of dividing a problem into sub-problems, solving each sub-problem with a specialized "expert" (or model), and then combining their outputs.
Then another "routing model" decides which model is most suitable for the given user prompt.
Given they use relatively few experts, each one is likely similarly capable to the others on many tasks. I assume this make deployment easier and is a "more conservative" less risky approach. Even if the wrong model is chosen by the router, answers should still tend to be somewhat acceptable, for instance.
> The fact that they're using MoE was news to me and very interesting.
Maybe adds some legitimacy to the claim.
It goes to show how LLMs are nothing like AGI. I think combining it with a calculator is just a bandaid. A useful bandaid, but its not going to be able to do science ever.
GPT-4 isn't "nothing like AGI" any more than its dense equivalent would be.
what a colossal asshole
The whole point of accurate titles is that you'll get less votes on uninteresting content.
1. Training took 21 yottaflops. When was the last time you saw the yotta- prefix for anything?
2. The training cost of GPT-4 is now only 1/3 of what it was about a year ago. It is absolutely staggering how quickly the price of training an LLM is dropping, which is great news for open source. The google memo was right about the lack of a moat.
That really doesn't change anything at all. The more training large models gets cheaper, the more large corporations are able to train larger models than everyone else.
Suppose the gross price of rice was $0.001 a kg. That's dirt cheap! Yet, if I had a million dollars and you had a thousand dollars, I could still buy a thousand times more rice than you.
Sure - they might be 'good enough' to build a business on. But if a competitor builds their business on top of a more accurate model, their product will work better, and they will win the market.
Yes, when i can run GPT4 in my closet, OpenAI will have GPT7 or w/e - but it doesn't change the fact that i have something useful running in my closed network and that opens up all kinds of data integration that i'm unwilling to ship to OpenAI. In that day i'll probably still use GPT7, but i'll _also_ have GPT4 running in my closet and integrating with a ton of things on my local network.
Like Betamax?
I think a better frame is, if rice got so absolutely cheap to make that anybody could spin up a bag of rice on a demand, anybody whose business model was based on selling rice sacks would be in trouble, especially if their specialty was selling rice in bulk instead of, eg, mom-and-pop restaurants selling cooked rice with flavors and a focus on customer experience.
(Not sure the metaphor is a good fit for AI. Maybe OpenAI comes up with GPT-5 and makes something so powerful that by the time OSS projects get to GPT-4 level nobody cares. But if GPT-5 is only incrementally better than GPT-4, then yeah, they have no moat.)
Right now the bottleneck is "how big a model can you fit on an H100 TPU". It's possible that in a few years, when bigger cards come out and/or we get better at compressing models, we'll get even better models just by increasing the scale.
And all that rice would be useless since you could only eat one cup a day.
The richest person in the world and someone who is solidly middle class both use the exact same iPhone. After a point more dollars doesn't necessarily mean better or more useful technology. If training "good enough" models becomes cheap enough to be achievable by small-time developers then OpenAI/Google/Anthropic etc. will definitely lose some of their edge in the space.
And?
...the market for rice will totally collapse because it would cost more to transport it than the farmer would make by selling it. Feel free to substitute "rice" for whatever commodity which becomes "too cheap to meter".
The "invisible hand" has a tendency to bitchslap people who don't have an even modest understanding of economic principles.
"Chinchilla showed that we need to be using 11× more data during training than that used for GPT-3 and similar models. This means that we need to source, clean, and filter to around 33TB of text data for a 1T-parameter model." https://lifearchitect.ai/chinchilla/
GPT4 has been trained on images exactly for this reason (it might not have been worth it separately from multi-modality, but together these two advantages seem decisive).
...and billions would be lifted out of poverty, and world hunger would be solved. The rice metaphor doesn't quite apply here.
If the price of GPU training continues to drop at the present rate, then it would be possible to train a GPT-4 level LLM on a $3000 card in 10 years. The ability to run inference on it would come way sooner.
One day some new startup will train on all of libgen and torrent networks, but it will be very hard to prove. You'll keep getting these gaps up in questionable morality and legality, and even openai will complain about playing fair
https://www.theverge.com/2023/7/9/23788741/sarah-silverman-o...
Skim read it, mark out some grammar errors, assign it a grade based on the quality of the opening and closing paragraphs.
Blogs upon blogs full of worthless pap that is there for SEO reasons have existed for like a decade already.
Reddit, twitter, etc.. raising prices is going to make it more expensive.
5 months on, and nobody has yet beaten their result quality. I think there is a moat.
Also, I think for many usecases, smarter is better. If a few cents can buy a more accurate answer, then it is always worth paying those few cents. So, while more hardware and more data can train a bigger better model, then that is the moat.
though google may have something up its sleeve with the corpus of google books! I have been wondering if openAI secretly pulled in scihub or zlibrary to neutralize that potential advantage.
Yes, and great news for shills, bad actors, agitators, trolls, foreign intel, and propagandists. I'm impressed by the tech but terrified because for once I cannot conceive of what this means for the future. My guess is that this kills the open web and laws get passed which bury it.
In other words: the speculation was likely right, I'll propose a specific mechanism explaining it, but then still insult the people bringing it up and keep gaslighting them.
Calling the belief "that the new GPT-4 quality had been deteriorated" a "conspiracy theory" goes beyond claiming the belief itself is wrong - it's also claiming that holding this belief implies significantly compromised reasoning skills. That is, it's just a drive-by insult.
For instance - MoE yes, but 16 experts at 111B parameters? Doesn't make sense. GPT 3 had 175B parameters. I doubt they would go less on base models from now on. The number that makes more sense is ~220B parameters per model and 8 expert models. That is the same inference cost in total.
The 13T tokens of training data seems pulled from thin air.
i'd love to know whats going on in that team.
The secret sauce and moat lies in data though. I have heard rumour that they have paid competitive coders to write and annotate code with information like complexity for them.
>(Today, the pre-training could be done with ~8,192 H100 in ~55 days for $21.5 million at $2 per H100 hour.)
Why flex both system size and training time to arbitrary numbers?
>For example, MoE is incredibly difficult to deal with on inference because not every part of the model is utilized on every token generation. This means parts may sit dormant when other parts are being used. When serving users, this really hurts utilization rates.
Utilization of what? Memory? If you're that worried about inference utilization, then why not just fire up a non-MOE model?
Here's what the post said about MQA:
>Because of that only 1 head is needed and memory capacity can be significantly reduced for the KV cache
This is close but wrong. You only need one Key and Value (KV) head, but you still have the same amount of query heads.
My guess is that this is all a relatively knowledgeable person, using formulas laid out by the 2020 scaling paper and making a fantasy system (with the correct math), based on that.
Put differently, I could probably fake my way through a similar post and be an equal level of close but definitely wrong because I'm way out of my league. That vibe makes me very suspicious.
Having multiple query heads does not affect the cache size, which is the limiting factor in MHA decoding for both memory capacity and bandwidth reasons.
Emphasis mine, source here [0]
[0] https://arxiv.org/pdf/2305.13245.pdf
FWIW the original MQA paper is called One Write head is all you need.
Here's the quote from that referencing multiple heads [1]
>We propose a variant called multi-query attention, where the keys and values are shared across all of the different attention "heads", greatly reducing the size of these tensors and hence the memory bandwidth requirements of incremental decoding. We verify experimentally that the resulting models can indeed be much faster to decode, and incur only minor quality degradation from the baseline.
I haven't registered for Twitter since it started and I'd rather not now (though I probably will if it's the only way to get leaked gpt4 training details)
Also, I'm dubious about this unsubstantiated claim. The biggest past innovation (training with human feedback) actually shrunk the size of a model. Compare Bloom-366B with falcon-40B (much better). I would be mildly surprised if it turned out Gpt4 has 1.8T parameters. (even if it's a composite model as they say)
The article says they use 16 experts 111B each. So the best thing to assume is probably that each of these experts is basically a fine tuned version of the same initial model for some problem domain.
If someone legitimate put together a crowd funding effort, I would donate a non-insignificant amount to train an open model. Has it been tried before?
Given that the price since the original training effort has already dropped to ~$20 million, and that (a) the fundraising will take time, and (b) improvements are being made every day with regard to resource usage, you could probably get away with aiming for a much lower number.
Pulling a number out of my arse, I'd guess that training a comparable model will only cost $1-5 million in 12 months time, with the hardest part of doing so once you have the funds being acquiring the training data.
$65 million sounds pretty high though.
Worth noting, though, that it isn't just the computing budget that's missing here - it is also (and perhaps even more importantly) the high quality data to actually train the model.
People don't need to own A100s, they just need to be willing to be part of a distributed supercomputer by running a background app that downloads chunks of data, processes them, and sends the result back. The utility comes from having enough people participate (which worked quite well for SETI@home, but helping find "signals from outer space" is a little bit more interesting than "helping train an LLM")
HuggingGPT works similar to this. It automatically chooses, downloads and runs the right "expert" model from HuggingFace https://arxiv.org/abs/2303.17580
Even if bits and pieces of the book text are distributed across the internet and you end up picking up portions of the book, you still read the book.
It is extremely sad but ChatGPT will be taken down by the end of this year and replaced by a highly neutered model next year.
But I think that unless GPT starts reciting large parts outside of the context of learning/education/research, reciting smaller snippets would fall into "fair use" and not be illegal.
You can't steal a book, photocopy some pages, then claim the photocopied pages are fair use.
Especially with ChatGPT you can probe the model by asking certain questions about the material at hand to see if it has seen the entire book.
Also you don’t have to be able to recite the book verbatim for it to have been in your training set. The snippets I am referring to are on the side of the training data
It's, however, very interesting to see if they fund efforts to massively (re)start books digitalisation.
Nevermind, fuckers, actually it's just to take your jobs and make a few VCs richer. We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us.
https://github.com/ggerganov/llama.cpp
https://github.com/openlm-research/open_llama
https://huggingface.co/TheBloke/open-llama-7b-open-instruct-...
https://huggingface.co/TheBloke/open-llama-13b-open-instruct...
You can use the above without paying OpenAI. You don't even need a GPU. There are no license issues like with the facebook llama.
Just to be clear, there's no science being kept secret because there is no science being done. OpenAI's is a feat of engineering, borne aloft by a huge budget supporting a large team whose expertise lies in tuning neural net systems, and not in doing science.
Machine learning, as it is practiced today, is not science. There is no scientific theory behind it and there is no scientific method applied. There are no scientific questions asked, or attempted to be answered. There is no new knowledge produced other than how to tune systems to beat benchmarks. The standard machine learning paper is a bunch of text and arcane-looking formulae around a glorified leaderboard: a little table with competing systems on one side and arbitrarily chosen benchmark datasets on the other side; and all our results in bold so everyone knows we're winning. That's as much doing science as is racing cool-looking sports cars.
I judge things as "what you can do", not "what can you predict". The only demonstration of knowledge and understanding is being able to do something. Not predict. Not "scientific method" and ridiculous "peer review" (actually peer pressure), not blind trials and not rigorous statistical analysis.
In the end of the day, you either manage to do something or you don't. So much of so called science had lost all contact with reality because our judgement of success isn't successfully doing something, it is successfully jumping through "scientific" hoops. Look at string theory and social sciences. The scientific process, instead of being a tool, became the purpose. It became a stamp of validity to seek. A stamp of validity with gatekeepers in the academia, in the peer review process, in the media coverage afterwards all the way to social media censorship and "fact checkers".
What used to be the frontier of creative people became a stagnant beurocratic machine worshipped like a new religion. The side of the heretics burned at the stake became the ones crying out heresy.
Enjoy your new brand of science. I'll stick to the older brand of heretics and mad men which did whatever it was the prevailing orthodoxy told them to avoid doing and thinking, and I'll remind you that the only real reason those are remembered is because they did something useful, not because of the social traditions they adhered to or the rigorous scientific standards they followed.
Because science gives you the tools to know that you're wrong. If you're a good scientist, you will be wrong _all the time_. That's how science advances: one mistake at a time. But you can't make mistakes if all you ever do is doing stuff with computers, like beating all the benchmarks, because that is a meaningless result judged by its own, self-chosen, measure of success that can never fail; and so can never inform.
Peer review is also not science, but it has been great to catch errors in my papers. Not in conferences, mind. I try to stay away from conferences. Everybody flocks to conferences because of quick turnaround, instant gratification. Journals have the good reviewers who can take their time understanding your work and helping you find where you've gone wrong. "Reject with encouragement to resubmit" is the best review result I ever got.
>> I judge things as "what you can do", not "what can you predict".
The goal of science is not to make predictions, but to understand how the world works, and why. Put that into instrumentalism's pipe and smoke it.
The hard sciences would like a word.
PS: I would be happy to connect, you can find my socials in my bio.
While people praise the scientific method, the majority of achievements we attribute to “science” are not derived from guess and check grad students doing their thing.
The standard paper is not very good. And a lot of the reason is that our scientific model is largely designed to find tiny effect sizes. Which is fine for like… medical stuff. But the juice is in large effect sizes. Stuff so obvious you don’t really need the stats to evaluate it. It doesn’t really matter if you followed the scientific method or not when you discover penicillin. It just works, clearly. And you can demonstrate it working again.
But today, academia is a victim of organizational and political capture, making it less competitive for talent.
These days I only believe large effects- like smoking causes heart disease
I’d argue it’s pedantic to assume all science has to be completely novel or revolutionary. More neurons good, is a perfectly reasonable reason to experiment. I think you’re being a bit “snooty” in setting such a high bar to gatekeep science…
Total horseshit. There are tons of scientific papers on ML published. In fact it is MORE like traditional science than typical CS, because it is trying to reverse engineer how something we encountered in the real world works. We know NNs do amazing things, and we don't fully yet understand how.
Predictability means that if your hypothesis is correct then you'd be able to formulate other improbable (and ideally currently untestable) predictions from it. And falsifiability means that if these predictions fail to occur then your initial hypothesis was also almost certainly wrong. So for instance Newton's hypothesis was that gravity was driven by a mathematical relationship between the mass/distance of two bodies. It was good science because it lead to the shocking ability to be able to dramatically simplify orbital dynamics, and create a complete predictive system of these bodies. It was even used to mathematically discover a completely unknown planet - Neptune.
Incidentally, his theory would also be able to be shown to be false if any of these unexpected predictions ended up being false. And that's actually exactly what happened. Observation of Mercury's orbit about the Sun showed it was off by ~1/3600th of one degree per century, relative to what was expected. And it's from there that people knew there was a mistake, which would only be explained later by Einstein hundreds of years later, who hypothesized a system with far more absurd predictions... and so the story continues.
What OpenAI is doing doesn't even resemble science at all.
Nobody has really explained why it works.
I don't think you would call it "science" for a bunch of single cellular organisms to cooperatively evolve a multicellular one. Similarly you wouldn't call it science when humans create digital lifeforms that require actual science to be done to understand how they work.
One thing is identifying (age of copper; age of bronze) the best ways of smelting ore to obtain the metal through trial and error, another is to try and understand the nature of materials.
My interest has always been in using machine learning as a tool to help understand some underlying phenomena, not in trying to push to the top of the leaderboard on benchmark datasets. I think this kind of attitude isn't uncommon in academic labs focused on doing real science, though the methods used aren't necessarily state of the art, and ML typically plays only a supporting role.
Curiously, I'm going to present at a conference later today and will give a toy example showing why those leaderboards are not necessarily reliable for distinguishing between the quality of different models.
I like the analogy attempt but there's a lot of science behind Formula 1
Reminder that such situation - science being out of focus, below the attempt to obtain practical results "empirically" - is not necessary but contingent.
The effort towards "science" is just postponed and/or relegated to other researchers in the same field.
And anyway, it is not that the effort towards "explanation" is completely absent. The situation is that many are working on the prospect of achieving big results through bets.
Is a scientist doing a linear regression not doing science?
Anyway, alphafold, while not particularly "scientific", did answer one scientific question: it is possible to predict the structure of most proteins thru a combination of limited structural and extensive sequence information, combined with a sophisticated (and "non-scientific") algorithm. That was an open question for some time and their results convinced the community that their methods were right. What's amazing is that while it's entirely nonscientific, the results have been absolutely blockbuster in the scientific field. And even better, the only reason Alphafold was able to show this is because there was a well-defined protein structure leaderboard.
That said, most of machine learning is not just engineering. If anything, there is little engineering done at all.
but do not forget the apparent point in the original poster's is the lacking approach below the implementations of above said techniques.
The techniques you listed in a substantial way expanded our knowledge ("oh, we can do this ... and solve more problems") while raising little more questions.
The said implementations on topic give results but raise many more unanswered questions.
(This post is unfortunately an "immature" reply - just a provisional tentative point. The matter of "ML algorithms ∩ science" requires more time and concentration and consultation of the texts of the Great Ones.)
For some reason people tend to consider a field that is more "formal" (like pure math, some CS concepts like lambda calculus) more science, even though historically formal systems came very late, and practically very few systems can be described that way.
I really wonder whether people who regurgitate "machine learning isn't science" think theory of evolution is science or not.
If you prefer to use an "instruct" model à la ChatGPT (i.e. that does not need few-shot learning to output good results) you can use something like this: https://huggingface.co/TheBloke/Wizard-Vicuna-30B-Uncensored... The interesting thing with these Uncensored models is that they don't constantly answer that they cannot help you (which is what ChatGPT and GPT-4 are doing more and more).
The open replacements for LLaMA have yet to reach 30B, let alone 65B.
that's great to hear. that political correctness in gpt is annoying.
Furthermore, Not many people discuss the significance of proper output sampling. I myself used to just test open source models with the greedy decoding only. Who knows if they wouldn't even beat (not at all)OpenAI with some clever output sampling scheme.
Apparently "Hugging Face" have some internal swift code that works (but it has not been released). I'm keen to see how it performs on a maxed out Mac Studio (with all that unified memory available).
I actually wrote about getting an LLM chatbot up and running a while ago: https://blog.kronis.dev/tutorials/self-hosting-an-ai-llm-cha...
It's good that the technology and models are both available for free, and you don't even need a GPU for it. However, there are still large memory requirements (if you want output that makes sense) and using CPU does result in somewhat slower performance.
There are async use cases where it can make sense, but for something like autocomplete or other near real time situations we're not there yet. Nor is the quality of the models comparable to some of the commercial offerings, at least not yet.
So I don't have it in me to blame anyone who forks over the money to a SaaS platform instead of getting a good GPU/RAM and hosting stuff themselves.
Here's hoping the more open options keep getting better and more competitive, though!
But I can easily imagine more conventional forms of entertainment, as well. Like a game of D&D that's narrated by the AI, or a text based adventure set in the Mass Effect universe, Lord of the Rings, Warhammer or any other fandom, really. Maybe like those old Choose Your Own Adventure games.
I think some companies are also experimenting with characters in video games that get their dialogue from these models - where the developers give the character a persona, provide information about events in the world and let players interact with them, like the Detective Origins demo. Of course, due to the slightly unpredictable nature of these models and their hardware requirements, no idea how viable this will be.
Despite the size of the regular online porn industry, it still gets strangled by payment processors. There is plenty of appetite out there for restricting porn in various ways.
I mean the guy who created GPT-4 literally demanded a ban of any system more powerful than GPT-4.
What I've seen from the horse's mouth is more like:
"""There are several other areas I mentioned in my written testimony where I believe that companies like ours can partner with governments, including ensuring that the most powerful AI models adhere to a set of safety requirements, facilitating processes to develop and update safety measures, and examining opportunities for global coordination."""
Now, sure, I could probably take a malevolent moustache twirling villain monologue and use chatGPT to turn it into a bland and prosaic statement like that, or even something utterly milquetoast, but if I try to imagine Altman playing 5D chess with the aid of his own AGI, there's the much easier solution of… just not telling anyone you have GPT-4 in the first place while using its power to manipulate etc.
"First, it is vital that AI companies–especially those working on the most powerful models–adhere to an appropriate set of safety requirements, including internal and external testing prior to release and publication of evaluation results. To ensure this, the U.S. government should consider a combination of licensing or registration requirements for development and release of AI models above a crucial threshold of capabilities, alongside incentives for full compliance with these requirement"
https://www.judiciary.senate.gov/imo/media/doc/2023-05-16%20...
It’s strange because their corporate structure is highly unique. They generate fixed returns on investments, so investors can’t make that much money. Also, the profit-generating division is wholly owned and controlled by the nonprofit. I don’t understand why Elon and others have so many problems with this.
If your choice is between $100M + doing earnest work, or $1B+ and debasing our free and competitive society, and you take the latter... what does that say about your collective character as an organization and as a set of people?
Honestly even reddit and teenagers on TikTok have a more accurate view on OpenAI vs. local LLMs than HN.
...
Here's a round up of open source projects focused on allowing you to run your own model's locally ('AI'), they all take slightly different approaches although under the hood many use the same models.
https://lnkd.in/exKqJZm8 A gradio web UI for running Large Language Models like LLaMA, llama.cpp, GPT-J, Pythia, OPT, and GALACTICA. Its goal is to become the AUTOMATIC1111/stable-diffusion-webui of text generation.
https://lnkd.in/etVCmZHB With OpenLLM, you can run inference with any open-source large-language models, deploy to the cloud or on-premises, and build powerful AI apps. State-of-the-art LLMs: built-in supports a wide range of open-source LLMs and model runtime, including StableLM, Falcon, Dolly, Flan-T5, ChatGLM, StarCoder and more.
https://lnkd.in/e7-NKGzJ LocalAI is a drop-in replacement REST API that's compatible with OpenAI API specifications for local inferencing. It allows you to run LLMs (and not only) locally or on-prem with consumer grade hardware, supporting multiple model families that are compatible with the ggml format. Does not require GPU.
https://lnkd.in/ef_Sa9AN Multi-platform desktop app to download and run Large Language Models(LLM) locally in your computer
https://lnkd.in/e288q-Wb A desktop app for local, private, secured AI experimentation. Included out-of-the box are: A known-good model API and a model downloader, with descriptions such as recommended hardware specs, model license, blake3/sha256 hashes etc... A simple note-taking app, with inference config PER note. The note and its config are output into plain text .mdx format A model inference streaming server (/completion endpoint, similar to OpenAI)
https://lnkd.in/eycRJn6b Transcribe and translate audio offline on your personal computer. Powered by OpenAI's Whisper.
https://lnkd.in/eUrtE3uQ The easiest way to install and use Stable Diffusion on your computer. Does not require technical knowledge, does not require pre-installed software. 1-click install, powerful features, friendly community.
Besides, what does your post add to the discussion, and why is it the top posting?
Create your local LLM, use it, tell other people about how you did it exactly, and be happy. But why the heck do you need to fight a company in that space?
Wehre have the times gone when someone motivated to do something nice just went ahead and did it, instead of running in circles and telling everyone else what they should NOT do.
OpenLLaMa uses a dataset which does not seem to have gotten propper commercial licensing for the training data. There is potential licensing issues because the copyright situation is not well defended.
You're right though, that's arguably still up for debate, but I think the precedent of transformative work is pretty well attested.
Because condensation is literally part of the definition of derivative, and basically the weights are a condensed form of the input data. It's some sort of lossy compression, when looking at it from the right point of view.
Summarization and translation are also clearly derivative.
The definition of transformative I found:
- add something new (context of other books I guess, this one might pass)
- with a further purpose or different character (further purpose clearly yes)
- do not substitute for the original use of the work (this one I find difficult. In the case of books, probably. In the case of github, it aims to replace quite some aspects of it) [Rumors that start to become lawsuits]
Some speculations are:
- LibGen (4M+ books)
- Sci-Hub (80M+ papers)
- All of GitHub
This is the most funny, but in the end sad aspect. If ChatGPT was indeed trained on pirated content and is able to be(come) such a powerful tool, then the copyright laws should have been abolished yesterday. If ChatGPT was not trained on all these resources out there, then think how much powerful a tool it would be if it were trained, then copyright laws are actively stifling advancement and should have been abolished yesterday.In short, we do not have the maturity to handle AI; we are playing with fire. Every person who contributes anything to AI development is responsible for the disasters that AI will bring, and all AI research should be destroyed.
GPT4 costs are ridiculously cheap for the value you get out of it. Any other company wouldn’t even release it to the public like they’ve done
Or at least Science is allowed to progress and get funded only as long as it serves the interest of Capitalism.
This is essentially the Capitalist Credo, expressed in practical vs theoretical terms.
Not really, Capitalism is free-trade between two parties.
The more government involvement you have the more it moves towards socialism or communism where the government controls trade.
At least this was the historical meaning. These days Capitalism is being redefined to mean private (non-government) Communism. That is, power concentrated in the hands of a few.
But if you don't have the time for that, just read the intro paragraph.
i've heard this called super-capitalism, yeah.
1. I don't think this is the right place for this kind of content, perhaps find your way back to Twitter or Reddit
2. Have you contributed funds to OpenAI? If not, where did your sense of entitlement come from?
3. What makes you think that any of what OpenAI has produced and provided would be available without funding? I assume the answer to 1 above will be no, so, how do you expect them to build without funds?
and 4. What's stopping you from creating what you thought OpenAI should be? Feel free. Nobody stopping you.
Is your argument that we need to have given a company money before we can be opposed to their unethical business practices? We're talking about a company that wantonly disregarded copyright, broke the DMCA billions of times (proving that law is for completely destroying individuals who want to personally enjoy some media they probably otherwise wouldn't buy, not for corporations who want to hoard the wealth collectively made by all of society), charges its users even when its product is not delivered due to their own server errors, and then goes in front of congress to try to put up a regulatory moat to make competition illegal. Everything they do is based on "the law applies to you, not us." And you're saying that we're entitled for being outraged by this?
> 4. What's stopping you from creating what you thought OpenAI should be? Feel free. Nobody stopping you.
That's literally what Sam Altman is trying to get congress to do. This will literally be illegal if we don't fight back. This is literally what the poster you're arguing against is saying is happening. How many more "literally"s do I need here!?
If Elon hadn't pulled the rug out from under them after they refused his forceful takeover*, they wouldn't have had to go to Microsoft and they'd still be open.
* a takeover which he predicated on the claim that OpenAI was "doomed to fail"
But secondly, even if they were a private company, it's dishonest and reprehensible to claim to congress that you want to "protect the public" when you really only give a shit about protecting your moat, I'm not happy about that either.
I'm also tired of tech oligarchs general tomfuckery in all our daily lives, as I suspect many more people here are. OpenAI is just particularly egregious about it.
I also think it's my civic duty to let other developers know that OpenAI does not have, by any stretch of the imagination, a stranglehold on this technology or any secret sauce. That's why they're lying and sweating in front of congress.
They didn't start the private company until the person who promised them $1B reneged. That person reneged because they tried to forcefully takeover and were rebuked.
They were running out of money and forced to raise funds in a very for-profit way or fold.
-
Edit: Rate limited because the hivemind has decided that it's unacceptable to insult their leader
People keep acting like if it wasn't for OpenAI we'd be in some LLM utopia. The reality is some other big tech giant would reach the current SOTA first and we'd be in the same predicament except with a company with 1000x more machinery to do the things you're complaining about.
It's ridiculous the sense of entitlement some people must have to keep insisting that OpenAI should have crawled into a cave and died because some megalomaniac threw a tantrum.
A non-profit and charity are two different things.
While a charity is a form of non-profit, it has to follow certain rules to qualify as one. Their profits must go towards the charity.
A non-profit is a company that is set up not to make a profit. It is allowed to make a profit if it does. This is what OpenAI was.
They switched to a "capped" for-profit model so that they could get more funding. It also allowed their employees to invest in the company, and openAI gave equity to their employees.
There was no lies or abuse. Where did you get that information from?
https://hackernoon.com/how-openai-transitioned-from-a-nonpro...
Responsible for an attempt to arrive at the construction of production facilities for a good that seems to be in dire scarcity in today's world: intelligence.
If somebody comes and implements "artificial morons", that is actually out of the root that made the field of research necessary.
> Mixture of Expert Tradeoffs: There are multiple MoE tradeoffs taken: For example, MoE is incredibly difficult to deal with on inference because not every part of the model is utilized on every token generation.
Are these experts able to communicate among them in one query? How do they get selected? How do they know who to pass information to?
Would I be able to influence the selection of experts by how I create my questions? For example to ensure that a question about code gets passed directly to an expert in code? I feel silly asking this question, but I honestly have no idea how to interpret this.
I obviously don't know how GPT-4 do it (or if it even does it) but think of partitioning your network into a couple of very isolated sub-graphs (the "experts"), and add another learnable network between the input tokens and the experts, that learns to route tokens to 1 or more expert sub-graphs. Then the gain is that you can potentially ignore running the unused sub-graphs completely for that token, and you can distribute them on other GPUs as except for the input and output they are independent of each other.
It all depends on the problem, data, and if the gradient descent optimizer can find a way to actually partition the problem usefully using the router and "experts".
Interesting to think about in comparison to the challenges today around parallelizing 'commodity' GPUs. Scare quotes because he A100 and H100 are pretty impressive machines in and of themselves.
Whether or not this specific theory is true something along these lines seems like the most likely explanation for the quality degradation that many have noticed; where OpenAI's claims about not changing the model are both technically true and conpletely misleading.
- https://www.washingtonpost.com/technology/interactive/2023/a...
- https://pile.eleuther.ai/ (data hosted by https://the-eye.eu/, where it's not too hard to find pirated, copyrighted books, e.g. https://the-eye.eu/public/Books/cdn.preterhuman.net/texts/li...)
What does this mean?
There's no magic here.
[1] That's probably twenty or so orgs right now, which will blow away OpenAI's moat and margins.
Hahahaha, the truth of anyone who has worked with quanty types running Python code at scale on a cluster
You must be thinking of some other post, or you're just making stuff up.
- parameters
- layers
- "Mixture Of Experts"
- tokens
That's about as far as I made it
Layers: In AI, layers are like stacked building blocks within a neural network, which is the fundamental structure of many AI models. Each layer performs different computations, transforming the input data as it passes through them. Think of a layer as a specific task or filter that the AI model can utilize to understand and process information. The deeper the neural network, the more layers it has, allowing for more complex patterns and representations to be learned.
"Mixture Of Experts": An "MoE" is an approach in AI that combines multiple specialized AI models, known as "experts," to work together on a task. Each expert focuses on a particular subset or aspect of the problem, leveraging their expertise to contribute to the final result. It's like having a team of experts who specialize in different areas collaborating to provide the best solution. By dividing the task and letting each expert handle their niche, the AI model can achieve better overall performance.
Tokens: In the context of AI and language models, tokens are chunks of text that are used as input or output. They can be individual words, characters, or even subwords, depending on how the language model is designed. For example, in the sentence "I love cats," the tokens would be "I," "love," and "cats." Tokens help the AI model understand and process language by breaking it down into manageable units. They allow the model to learn patterns, context, and relationships between words to generate meaningful responses or predictions.
It seems that much like running a co-op or commune the barrier is whether enough people care to take up the fraction of effort and cost it takes to join such an arrangement.
Even if the social winds did miraculously change, LLMs will never be decentralized for the sole reason that once they're good enough to operate a gun and servo motors every government on the planet is going to lock that shit up fast (assuming massive relentless cyber campaigns somehow don't trigger that response first).
I will admit to using it all the time for simple programming tasks, and it occasionally does them correctly. Often it comes close enough that I can fix them. (Interestingly in most of these cases I can’t talk it into fixing itself. It kinda gets into wrong-approach ruts), and sometimes it’s horribly wrong (like here).
I find the horribly wrong cases funny.
I'd guess that takes it out of the top 0.1%, but not the top 1%.
The dunking makes sense to me—“AI is taking our Jobs” is a real concern, so pointing out how bad ChatGPT is at coding on an arguably coding social network is one way to control the narrative.
2. I looked in The Art of Computer Programming for a quantum computer algorithm to square a number and it didn't have one. If it's a CS textbook it obviously isn't a very good one. In fact it's amazing that anyone thinks Knuth is worth anything at all. I immediately threw my copies in the recycling.
In fact you can divide everything into the set of things which know how to square a number on a quantum computer (let's call that the set of valuable things) and everything else. Everything else can be discarded.