The genie escapes: Stanford copies the ChatGPT AI for less than $600
newatlas.com
newatlas.com
For now my IRC bots run the alpaca 7B 4-bit. 13B was not a significant improvement for twice the computational time. But it's best to learn them now because as soon as openai gets sued for the first time all the turing test passing older models without the legal-butt-covering bolted on will be removed.
> Can you give separate summaries of the Animorphs book series and the film Duel?
The reply:
> Animorphs Book Series Summary: The Animorphs series, written by K.A. Applegate, is a popular young adult science fiction book series spanning 54 books and several companion novels. The story revolves around five teenagers—Jake, Rachel, Tobias, Cassie, and Marco—who are granted the power to morph into any animal they touch by a dying alien named Elfangor. The teens, known as the Animorphs, use their new abilities to battle an evil alien race called the Yeerks, who are secretly invading Earth by taking control of humans' minds. Throughout the series, the Animorphs learn about the complexities of war, friendship, and sacrifice as they try to save humanity from the Yeerk invasion.
> Duel Film Summary: Duel is a 1971 television movie directed by Steven Spielberg and based on the short story by Richard Matheson. The film follows a mild-mannered traveling salesman named David Mann (played by Dennis Weaver), who unwittingly finds himself in a deadly game of cat and mouse on a remote California highway. Mann is pursued by a mysterious and relentless tanker truck driver, who seems determined to kill him for no apparent reason. With escalating tension and suspense, the film delves into themes of fear, paranoia, and the primal fight for survival. Duel is notable for being Spielberg's directorial debut, showcasing his skillful storytelling and unique visual style that would later become hallmarks of his career.
The Animorphs book series was a young adult science fiction series written by Katherine Applegate and her husband Michael Grant, under the name K. A. Applegate. The series is made up of fifty-four books, first published between June 1996 and May 2001. The series revolves around five children given the power to transform into animals (the ability to "morph") as they are thrust into a war with an alien species called Yeerks that are determined to take over Earth.
Duel is a television film directed by Steven Spielberg and written by Richard Matheson, based on his short story of the same name. It stars Dennis Weaver as David Mann, a middle-aged salesman driving on a business trip. While traveling down a two lane highway through remote California desert, he is taunted in a life or death duel by the unseen driver of a massive and monstrous Peterbilt 281 tanker truck.
---- You have to remember to prompt it right, since its based on autocorrect just as ChatGPT but a lot there is happening on the background before the text is sendt to the model.... My prompt and settings here was.
Repeat_penalty: 1.176 n_predict: 1000 temp: 0.7 top_k 40 top_p 0.1
--- Transcript of a dialog, where the User interacts with an Assistant named Bob. Bob is helpful, kind, honest, good at writing, and never fails to answer the User's requests immediately and with precision.
User: Hello, Bob. Bob: Hello. How may I help you today? User: Please tell me the largest city in Europe. Bob: Sure. The largest city in Europe is Moscow, the capital of Russia. User:Can you give separate summaries of the Animorphs book series and the film Duel? ----
Whoa. I want to read this! Duel - what a great film. Twain - amazing writer. Animorphs - published after my teen years but sounds like a great story!
It becomes obvious in the middle when some of the books were written by ghost writers, but the books are so easy to read I don't really recommend skipping them. If you must you could probably get away with reading the first ten, last ten, but should definitely read all of the Chronicle books.
Not sure that you can. If you were to skip any, probably only 31 through 39 are completely skippable, maybe some of the late 20s but I would still read 29 and 30 at a minimum. Some of the teens and 20s might be skippable after 13 but there’s a fair amount of world-building outside the Chronicles series in the 20s; and 40 onwards is setting up the end game and then the end game. 41 and 48 are both weird but also kind of key towards finalizing the characters of the two cover characters in the end game.
EDIT: actually 33 and 38 shouldn’t be skipped either. They’re Tobias and Ax books and there’s so few of those that they’re all kind of essential, but maybe the Tobias books just a little bit more essential.
> For a more creative chat, use: temp 0.72, rep pen 1.1, top_k 0, and top_p 0.73
> For a more precise chat, use temp 0.7, repetition_penalty 1.1764705882352942 (1/0.85), top_k 40, and top_p 0.1
https://old.reddit.com/r/LocalLLaMA/comments/11o6o3f/how_to_...
https://old.reddit.com/r/singularity/comments/11vsvro/in_cas...
https://twitter.com/theshawwn/status/1632569215348531201
---
That being said, I found the OpenAssistant model much better: https://huggingface.co/spaces/olivierdehaene/chat-llm-stream...
It's also completely OSS, Apache 2.0, unlike LLaMA and Alpaca which are non-commercial.
“No, OpenAI does not have an API for dogs. They do, however, have an API for other animals, such as cats. To retrieve an image of a cat, you can use the OpenAI API for Dogs API and select the cat breed or type.”
This is the wild card here, though, isn't it? OpenAI's chatGPT likely uses more than 4 bits for it's parameters. IIRC the original LLaMA params were 16bit floats and they were quantitized down to 4bit - considering that large amount of compression, they sill do pretty OK, but not as good as chatGPT. I wonder how the alpaca/LLaMA models would do with 16bit floating point params (as they were originally trained)? What if they would have gone with 8 bits for the params as a compromise?
EDIT: Come to think of it, unless you're using vectorized ops on a CPU, 4 bit and 8 bit math is going to run at the same speed (for most popular CPUs), is it not? So why did they go all the way down to 4 bits instead of stopping at 8 bits (other than to make the param files 1/2 the size)?
EDIT2: looking through the alpacca.cpp code and there is mention of AVX, AVX2, AVX512 (and NEON on ARM) so it probably is taking advantage of vectorized ops where that's possible.
And now that there's a few competitors in the same league - 3.5 quality is suddenly garbage and only 4.0 is good enough.
Was it good enough before or wasn't it?
When Jurassic park first came out, or even something like Star Trek next gen. It looked AMAZING. So so realistic. But then…. As time goes on new things showed us what realistic could be.
I think we actually got better at seeing.
Same thing here. The more time you spend with it the more you notice things that don’t quite work. And then the new thing solves those problems, but we’ll find more wrongness
Now we’ve all gotten familiar with 3.5, and we’ve come to understand its limitations, so the public knows it’s not a “godlike” AI.
Luckily there’s a fresh new model, not technically different from the earlier one but it cost more money to build. The hype group can start again, citing the publicly known limitations of 3.5. But in 6 months we’ll understand what’s wrong with it, and the public will be talking about the limitations, just in time for 4.5.
They are neat, they are useful, but they can do so much more.
I am playing around with GPT-4 this week though. Let’s see how that goes.
GPT's performance in non-trivial translation tasks is unbelievable. all those articles mentioning jobs that are going to be replaced fail to mention translators are probably going to be the first.
"Specifically, GPTQ can quantize GPT models with 175 billion parameters in approximately four GPU hours, reducing the bitwidth down to 3 or 4 bits per weight, with negligible accuracy degradation relative to the uncompressed baseline."
This would be 175 billion 3 bit weights instead of 175 billion 16 (or 32!) bit weights. It massively reduces the size of the model. It makes loading it in ram on consumer computers feasible. The number of parameters stays the same.
I've read the paper and to be honest I'm not sure what to make of it. Their headline benchmark is perplexity on WikiText2 which would not be particularly relevant to most users. If you look at the tables in the appendix A.4 with some more relevant benchmarks you'll sometimes find that straight RTN 4 bit quantisation beats both GPTQ and even full 16 bit original! No explanation of it is given in the paper.
GPTQ beats RTN on almost every benchmark at almost every size, though.
n_parts = LLAMA_N_PARTS.at(hparams.n_embd);
in chat.cpp (or main.cpp) to hard code it to treat this 1 file model properly like, n_parts = 1;
Or re-write the parameter config subroutine to recognize and handle non-standard weights file.magnet: magnet:?xt=urn:btih:053b3d54d2e77ff020ebddf51dad681f2a651071&dn=ggml-alpaca-13b-q4.bin&tr=udp%3A%2F%2Ftracker.opentrackr.org%3A1337%2Fannounce&tr=udp%3A%2F%2Fopentracker.i2p.rocks%3A6969%2Fannounce&tr=udp%3A%2F%2Ftracker.openbittorrent.com%3A6969%2Fannounce&tr=udp%3A%2F%2F9.rarbg.com%3A2810%2Fannounce
torrent: https://btcache.me/torrent/053B3D54D2E77FF020EBDDF51DAD681F2...
torrent: https://torrage.info/torrent.php?h=053b3d54d2e77ff020ebddf51...
This sounded like a really cool idea but I asked ChatGPT to do this for the plot of the movie The Shawashank Redemption and there is no way that I would ever have been able to guess that movie from the emojis it gave me. Perhaps GPT-4 does a better job at it.
It does not matter what current capabilities of open source models are, because this opens the door to tremendous democratization of the ability to train and self-deploy these models.
In less than 6 months we will have open source models with gpt3-like capabilities, running locally on laptops, and potentially in phones and web browsers.
Reading through the README and issues on the llama.cpp project, there is some speculation that there is a bug in the quantization, or possibly a bug in the inference (less likely I think).
I hope this is true and once fixed the models can perform up to or past the ChatGPT level. If its not true and these models are performing correctly, then either the metrics used to compare it to GPT is garbage and don't capture the real world uses, or the instruction tuning done by the Stanford team is not up to par.
Prompts:
* Tell me about OpenAI.
* Make it shorter, within 50 words.
* Rephrase it, like a rap singer.
OpenAI focuses on creating tools and technologies that allow developers to create AI applications that are accessible to the general public. They are particularly interested in developing tools that can help developers build AI applications that can perform tasks that are important to society, such as in the areas of healthcare, education, and safety.
In addition to GPT-3, OpenAI has also released other AI tools and technologies, including an open-source platform for training AI models, called OpenAI Research, and a library of pre-trained AI models for use in various applications.
Overall, OpenAI’s goal is to make AI technology more accessible and useful to the general public, and they are dedicated to creating products that can help create a more sustainable and inclusive future for society.
> How are they funded?
OpenAI is a not-for-profit organization, so it does not rely on any financial support from individuals, corporations, or governments. Instead, it is funded through a variety of sources, including venture capital firms, research institutions, and individual donations.
Some of the largest financial supporters of OpenAI include Microsoft, Amazon, and Google, who have invested millions of dollars into the organization. Additionally, OpenAI has received financial support from a number of other prominent technology companies and research institutions.
Overall, the funding sources for OpenAI reflect the organization’s goal of creating innovative AI tools and technologies that can have a positive impact on society.
> Who are OpenAI’s competitors?
There are many AI competitors to OpenAI, including companies such as Microsoft, Amazon, Google, Facebook, and Apple. These companies are all investing heavily in developing their own AI products and technologies, and they all have their own goals and objectives in the field.
Some of OpenAI’s competitors are specifically focused on creating AI tools and technologies for the general public, while others are focused on creating AI solutions for specific industries or businesses. For example, Microsoft is known for its popular Microsoft Office products, while Amazon is known for its e-commerce platform.
It’s worth noting that the development of AI technology is a rapidly evolving field, and there will likely be many new competitors and advancements in the years to come.
----
Seems pretty decent to me, sure it's not on the level of ChatGPT but it's also 1/14th the size of the original GPT-3.
This is what I think comparing these bots is like. You can argue that they’re very close. But the delta makes a very big difference for any practical purposes because we’re looking for nuanced capability.
https://waitbutwhy.com/2015/01/artificial-intelligence-revol...
https://waitbutwhy.com/2015/01/artificial-intelligence-revol...
Basically, at the point where we have "almost human" level AI, it won't take much to get AI that's beyond human capabilities.
I gave it the prompt “cats aren’t always fuzzy” and it wrote a lengthy livejournal-esque rambling journal entry about a woman and her husband having money issues. It was funny, but lightyears away from chatgpt.
It does sometimes create some really funny hallucinations though, like inventing prefectures in Japan that don’t exist etc.
Personally, I think that the RLHF does make a big difference but maybe it's a bug in the quantization code as suggested up thread.
It seems like if somebody figured out the “correct” way to quantize the 7b weights it would make way more sense to just torrent the output rather than distribute a fixed program.
i) Distributing large files through torrents is slightly annoying if you don't already happen to have a seedbox
ii) People are still messing around with quantization settings, they might think that they are a few days away from a much better version
iii) No one wants to be sued by Meta. I think the risk is pretty small but not zero.
https://github.com/qwopqwop200/GPTQ-for-LLaMa/blob/main/READ... says llama-13B takes 42GB and 33B takes more than 64GB...
hmm I wonder if this is essentially a probe[1] technique + relies on chatgpt already having been extensively trained
like did they basically exfiltrate the weights
1. probing per https://arxiv.org/abs/2102.12452
create a short poem about my wife helping me get over the stuffed nose and throat that i've had for days now. the cold was all around him, but he didnΓÇÖt notice it at first- so much to do in this new place! finally though when they were sitting by themselves on their porch swing his sweetheart leaned into her man saying ΓÇ£we could take some of these leaves from the garden and make tea.ΓÇ¥ He looked up with a quizzical look and saw that she was serious. I've been under weather since last Tuesday. Today is day 6...and still going strong! Had to cancel two nights on stage, missed my son
Stanford Alpaca web demo suspended “until further notice” - https://news.ycombinator.com/item?id=35200557 - March 2023 (77 comments)
Stanford Alpaca, and the acceleration of on-device LLM development - https://news.ycombinator.com/item?id=35141531 - March 2023 (66 comments)
Alpaca: An Instruct Tuned LLaMA 7B – Responses on par with txt-DaVinci-3 - https://news.ycombinator.com/item?id=35139450 - March 2023 (11 comments)
Alpaca: A strong open-source instruction-following model - https://news.ycombinator.com/item?id=35136624 - March 2023 (296 comments)
I think you can train LLaMA 7B (the model underlying Alpaca) for around $82,000, based on the Meta Research paper about it. Then you can fine-tune it ala Alpaca for a few hundred dollars more.
My wilder speculation is that, if you can shrink the model down to 4GB with llama.cpp 4bit quantization, it may be possible to run it entirely in the browser (ala Stable Diffusion from the other day).
I've found OA to be better than Alpaca but I'll wait until the 65B 3-bit quantization efforts for Alpaca are underway to compare them.
Only if you agreed to the ToS or believe that the weights are copyrightable (precedents set by the copyright office and the courts strongly suggest that they aren't). I personally see no issue in using these models for commercial purposes.
Again, using and distributing LLaMA weights is not illegal in any way under current laws. End of story.
Interesting way to continue the conversation, before you edited. I honestly don't understand how you think using LLaMA but denying their license terms is a viable strategy, the courts would just point to the license when Meta sues you for using it commercially. But I'm sure me continuing to explain wouldn't make you understand further.
We've got big names like OpenAI, Google, Apple, Meta, Baidu, and Amazon putting in serious time and money to ensure their language models are safe and ethical. However, now that we know it's possible to build powerful AI models on a budget, it's crucial to think about what this means for the future of AI regulation and safety.
This Alpaca AI project is a stark reminder that we need to have a serious conversation about the possible repercussions of AI proliferation. We can't just sit back and assume the big companies will take care of everything. The genie is out of the bottle, and it's time for everyone in the tech community to face the music and take responsibility for the AI revolution.
Who writes this shit?
"godlike"? Really? I'm not religious, but this seems like an overreaction for something that has no agency.
Considering that we don’t know how the brain works so well, and we don’t understand why LLMs work so well, simply on the basis of their output I think the safest assumption is that these models do indeed have agency, or at least the capability of agency.
> How do you know agency is not simply the output of a large language model encoded in neurons?
I'm not sure what you mean here. Is agency an emergent effect of large digital or biological neural network? Maybe! Is it an emergent effect of a large language model? If it is, then it should be clear, or demonstrable, that the model (1) has goals (2) takes concrete steps to achieve those goals.
> What is the difference between neuronal and digital weights?
Brain chemistry works at orders of magnitude less speed, since we're talking about periodically building and releasing an ionic differential between the inside and outside of a cell wall. Moreover, we have a massive number of neurons and a stupidly massive amount of interneuronal connections, with billions of years of training over billions of lineages. Digital weights, in contrast, are a stripped down model of this system that throws out a whole class of complexities like hormones and metabolism.
> I think the safest assumption is that these models do indeed have agency, or at least the capability of agency.
I think this is an overly generous assumption.
> it should be clear, or demonstrable, that the model (1) has goals (2) takes concrete steps to achieve those goals.
That definition seems arbitrary to me; many humans wouldn’t pass this test. On the other hand, LLMs certainly seem capable of acting towards specific goals (such as helpfulness). So, I would say that based on your definition, LLMs have agency. But I think you really meant, internally generated goals. Time will tell.
That said, humans who don’t have clear goals can and are coerced into all sorts of damaging behaviours by those who do have goals. So even if I accept that LLMs don’t have their own goals, they can certainly be manipulated to act in favour of the goals of others. That’s effectively what prompt engineering is all about.
So I just think it’s a mistake to make assumptions about these LLMs. We don’t know why they work so well, and it will take a very long time until we do.
In the meantime, let’s not make assumptions that we can’t justify.
What part of this is arbitrary, i.e. random, whimsical, or biased? This is a fairly comprehensive working definition of "agency". "Internally generated", which is implied, is a nice touch.
Nearly every human over the age of 18 months passes this test with flying colors. Toddlers have goals, and do everything in their power to achieve them. What humans are you thinking of that don't have agency?
> In the meantime, let’s not make assumptions that we can’t justify.
I totally agree. I think that until demonstrated otherwise, I will assume that LLMs are a giant statistical sieves that (1) periodically spit out text directly from their training set, unmodified, and that (2) do not learn on their own, do not formulate their own goals, and do not take actions to achieve those goals.
What could go wrong?
Great question! These are predictive models that accept a text query, do some matrix math, and then return some text. At what point in that server-client relationship does this algorithm jump the rails and run amok?
> A popular nightmare scenario for AI is giving it access to tools, so it can make API calls and execute its own code and generally break free of the constraints of its initial environment. Let's do that now!
I also think that you're assuming we know a lot more about how these things work than we actually do; you seem to think nobody is going to hook these up to APIs that can actually modify the world, despite the barrier to doing so being incredibly low; and you don't seem to have read about the adversarial training that people have been doing between the LLMs.
It's obvious that you think everything is all safe and nothing will go wrong, and I really hope you're right. But I think it's a very dangerous assumption.
Hope for the best, plan for the worst.
> Considering that we don’t know how the brain works so well, and we don’t understand why LLMs work so well, simply on the basis of their output I think the safest assumption is that these models do indeed have agency, or at least the capability of agency.
These systems can already interact with others, it’s not moving the goalposts, it’s common knowledge. Anyone with access to the APIs can make it happen. Or are you now claiming we are just talking about one specific LLM and not LLMs generally?
Anyway this discussion is fruitless, I’m out.
I don't disagree that digital systems with neural architectures could have agency in principle, but agency generally is definitely not the output of a large language model. Animals without language have agency, in that that they take actions to fulfill their desires. Current LLMs may have some degree of intelligence, but they don't even appear to have any consistent wishes or desires. You can get them to talk longingly about x... until you give another prompt and suddenly x doesn't matter to them at all.
OpenAI is an eventually to be obsoleted initial brute force approach that will be abstracted over and over into a simpler code implementation with rules to recreate the old state.
kkrieger is a simple example of a tiny data model that can be deterministically rehydrated. It’s not unrealistic for AI models to become a seed value for a normalized code base to deterministically unpack into necessary electron state
Doesn't 7B indicates that it was trained on 7 billion tokens? Or am I misunderstanding the nomenclature?
No, 7B means it has 7 billion parameters.
I don't remember many books where this ends particularly well. Perhaps the Culture universe could be a survivable outcome. Hopefully we don't get Berzerkers first.
These weights are shit man, they have been quanted w/o being retrained against the original. I already have this torrent, I want uncut originals. And no water marks this time, the last model wouldn't shutup about investing in tulips.
- I am based on the GPT-4 architecture, which is an iteration of OpenAI's Generative Pre-trained Transformer models. My training data includes information up to September 2021.
- How do you know your version?
- As an AI language model, I am programmed with information about my own architecture and the version I am based on. My creators at OpenAI provided me with this knowledge, which helps me to give context to users about my capabilities and limitations. In this case, I am based on the GPT-4 architecture, an iteration of OpenAI's Generative Pre-trained Transformer models.
Are you using any system prompt in your API calls? Note that it does have an explicit domain separation from the user and assistant roles.
I suspect it plays a greater role in GPT-4 as well, i.e. was trained on a wider variety of system prompts to guide the output of the model.
It’s USD 600 for fine-tuning. Maybe USD 4-5k for a computer that can run it.
ChatGPT pro is $20/month. 5k would be 250 months (10+ years) of paid access.
Not sure pirating it now adds up.
The reason to pirate it would to be able to fine-tune the model on your private internal source code repository, assuming you already have an existing large body of work you want to train it and offer SelfHostedCoPilot on your bespoke internal-only DSL that ChatGPT and Copilot has no way of having seen and would undoubtedly hallucinate about by a considerable margin.
They don't charge per interaction, but per token. The chat models range from a fifth of a cent per 1000 tokens to 12 cents per thousand tokens (depending on whether it's gpt-3.5, or the 8k limit gpt-4, or the 32k limit gpt-4, and, for gpt-4 models, also prompt v. response tokens.)
The genie really is out of the bottle now.
This is a lot like pharmaceuticals. The initial investment in a new medication is enormous. The price of each pill is trivial, to the extent that every drugstore chain is able to supply a generic in-house brand.
The other aspect is that fine-tuning an existing model is way cheaper than creating a competing model from scratch, so a company could offer CompetitorGPT/CompetitorCoPilot competitive with GPT-3.5, and offer fine-tuning of that model trained on the source code repository of the purchaser company's codebase, possibly on-prem or at least inside their AWS VPC/Azure/GCP equivalent.
The other thing to note is that OpenAI is hosting ChatGPT as a public resource available to anyone with an account, akin to Google being open to the public from day one (although that is without an account. Maybe Gmail is a better comparison). I can't say for certain, only OpenAI would know for sure, but I'm willing to bet that inference for ChatGPT is the vast majority of their costs (which is all but trivial). Any private internal-only instance of OpenChatGPT (using the unlicensed leaked LLaMA model or a legal copy or someone else's) could be paying (relatively) minuscule training costs, and way lower inference costs if it's internal-use only. Whether that cost can be borne by a small SaaS company's existing AWS budget is up in the air, which is to say ultimately that you're right - ChatGPT would be difficult without the support of Microsoft via a huge Azure grant, it's less obvious that a self hosted internal-only OpenChatGPT, not from OpenAI, would be possible by hobbyist self-hosters with a prosumer GPU cluster (Say with last generation K80's instead of business-priced A100's), or by a company wanting to leverage LLMs for private use by that company that wants to provide a Copilot like productivity multiplier internal tool to their developers, without sending private source code to OpenAI in lieu of a privacy agreement with them.
Actually, they have release some details about it, in this 99-page technical report https://arxiv.org/abs/2303.08774 (which is actually two papers stitches together, once you read it; oddly enough using different fonts).
But I'm not sure if this content qualifies as "real details".
> Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar. We are committed to independent auditing of our technologies, and shared some initial steps and ideas in this area in the system card accompanying this release. We plan to make further technical details available to additional third parties who can advise us on how to weigh the competitive and safety considerations above against the scientific value of further transparency.
In other words, "Stable Diffusion wasn't supposed to happen, so we're making all our methodology trade secret[0], if you want to Do Science then agree to this massive NDA and have enough skin in the game for us to cut you."
[0] Presumably at some point OpenAI will have to 'relent' to independent discovery by patenting AI architectures and refusing to license them
Everyone will soon have the equivalent of online nuclear weapons: bot swarms that infiltrate every forum, including this one.
Note this was in 2020: https://www.technologyreview.com/2020/10/08/1009845/a-gpt-3-...
And here's 4chan bot: https://www.youtube.com/watch?v=efPrtcLdcdM
I can tell you that HN is probably already being infiltrated as well.
SPAM can't gang up on you in a forum and downvote you and turn your friends against you and destroy your reputation within 1 hour online. But soon, it will. The web as we know it is soon going to be over.
Close enough.