Hello Dolly: Democratizing the magic of ChatGPT with open models
databricks.com
databricks.com
I just see a zip download that AFAIK also doesn't contain the weights/checkpoints. I find this a bit odd, the contents of the zip (from the gdrive preview) look like they should be in a git repo, and I assume they download the model from somewhere? GDrive usually has rate limits which I'm concerned about.
If anyone from databricks reads this - are there plans to publish this on a git repo somewhere, as well as the weights/checkpoints?
EDIT: Oh I just noticed
> Contact us at hello-dolly@databricks.com if you would like to get access to the trained weights.
This... seems odd for a article titled "Democratizing the magic of ChatGPT with open models"?
Perhaps Databricks suspected another big announcement coming soon and wanted to get this announcement out?
Alpaca was made to fine-tune LLaMa, however they also released their dataset they used to do this, and it looks like Dolly is this dataset applied to GPT-J, and does not use LLaMa itself.
> This fine-tunes the [GPT-J 6B](https://huggingface.co/EleutherAI/gpt-j-6B) model on the [Alpaca](https://huggingface.co/datasets/tatsu-lab/alpaca) dataset using a Databricks notebook.
> Please note that while GPT-J 6B is Apache 2.0 licensed, the Alpaca dataset is licensed under Creative Commons NonCommercial (CC BY-NC 4.0).
...so, this cannot be used for commercial purposes
The legal relation between models and training data sets seems murky; of course, with the build tooling, you can also substitute in another instruction-following training set if you want to avoid licensing issues with the Alpaca set, whereas if you aren't concerbed with them, you can just blaze ahead.
Can't they also release the fine-tuned weights as non-commercial as well?
No, because you as a human looking at "art" over your lifetime and learning from it is not "fair use" of the copyright, it's no-use at all. This is the crux of every argument for both for language models and AI Art models, that their tools are learning how to draw, learning what styles and characteristics of input art correspond the most with words, and creating art with that knowledge just like any other human, not simply collaging together different pieces of art.
or you can raise $30,000,000 right now and worry about the copyright infringement lawsuit in 2026 or never.
The implication being that you're only "democratizing" something if people can make money off of it?
But maybe "democratize" is starting to mean something similar to "open". All the good words.
https://github.com/databrickslabs/dolly
Sorry it took us a day to get the external repo setup.
Was the Alpaca dataset being licensed as non-commercial only the reason you aren't releasing the weights? Is it possible to just release them under the same license?
Working on a model without this issue. Certainly our goal is totally open models anyone can use for anything.
I've been a bit jaded by the "open/democratizing ai" stuff and then having companies stiff us at actually making it open - but not wanting to be the first to litigate these new types of issues ml brings is very understandable.
Question - Would you consider benchmarking a single 4090 for your training? While training in a few hours with 8x A100's is impressive, myself and I think others are curious how that translates to consumer hardware. IMO running/fine-tuning on consumer hardware is the ultimate endgame for all ai models.
I guess: "You Either Die A Hero, Or You Live Long Enough To See Yourself Become The Villain"?
OpenAI renounced being open source. Don't let the name fool you.
AI being tuned to be "safe" by an exceedingly small set of humans is the thing we should be afraid of. It's the effective altruism effect: if you bombard people enough with "safety" and "alignment" speak, they will look past the fact that you're mainly interested in being a monopoly. My bigger conspiracy theory is that Bill Gates getting behind "AI alignment" is a calculated move to get people to look past Microsoft's unilateral involvement.
We know what their immediate goals are: to make as much money as possible. The only question is what their longer-term goals are.
The employees are usually they are "in on the message". All you have to do is get your employees to believe the message. This is easy when the company is growing exponentially. CEO's word is gospel.
There is probably a lot of political pressure on OpenAI to be as closed as possible. Remember the US government has banned Nvidia from exporting A100/H100 to China/Russia. Those are the same chips OpenAI uses for both training and inference.
We started seeing this in our testing. OpenAI's Curie model is responding very well to our fine-tuning experiments for chatbot-style interface. I am trying to keep us focused on quality of training data rather than obsessing over raw network size. Davinci (and derivatives) might turn out to be overkill for our use cases.
It feels like there should be an xkcd for this.
Tangentially, it seems like most of the results for both searches were autogenerated with TTS programs. I wonder if our pronunciations will shift towards TTS mistakes over time. Probably not, these videos only have a few thousand views, but neat if true.
This is like when Khan Academy came out and there was a guy online saying it's a terrible brand because it sounds like Con Academy which it doesn't in my dialect.
Took a while to get it.
It's the KH sound that doesn't really exist in English hence many get it wrong.
But I do like the hang for whimsical naming schemes in that field. First sesame street characters, now apparently everything sheep...
No they're not.
It's clear the model already "knows" the format of a Tweet (short length, attention-grabbing, contains hashtags). The model also knows stuff about language models (word2vec, tokenization), and can include entities from the question in its response (Dolly, Databricks). Yet, it just doesn't put these pieces together in the right way without the Q&A training.
Edit: For kicks, I asked GPT-4 this question: https://imgur.com/a/sM4uyBn
Nice job DataBricks, nice numbers too. Looking forward to more improvements.
> Contact us at hello-dolly@databricks.com if you would like to get access to the trained weights.
Like giving away a website template without the demo content, it's perfectly normal.
Someone lay out the reason they should package the weights with this, when they're allowing you to apply your own ?
This repo isn't what you think it is.
Can you guess what their reply would be?
"You may not [...] except as permitted through the API, use any automated or programmatic method to extract data or output from the Services, including scraping, web harvesting, or web data extraction;"
P.S. this comment does not reflect my personal values. But I would rather someone with values try it almost like a white hat pen test.
The widespread adoption of local AI won't obsolete a well-priced AI API. I feel like we learned that lesson pretty thoroughly in the SaaS era.
Unless I am misunderstanding (?), this seems like an overgeneralized lesson. There are many key differences between these situations that make such a connection unlikely. Could you explain your reasoning?
1. OpenAI, the organization, is not equivalent to its chat offering.
2. Saying "the" real risk isn't persuasive. Let's examine many risks before claiming one is the most significant. Also, "real" is this usage often a throwaway (i.e. unneeded) word, in editor speak.
3. Let's talk about OpenAI's "business model" (though such discussions are tricky).
3A. Originally, OpenAI wasn't trying to "hold onto" AI advancements. It claimed to be a broadly funded way to explore fundamental questions of artificial intelligence in a non-commercial, ethical way.
3B. Of course, the above claim was largely aspirational, because it wasn't baked into their DNA in way that could survive the surrounding temptations for more funding, glory, and resources.
3C. Even with their more commercialized model of the last several years, it seems their business model feels like (a) fundraise in exchange for (b) (claimed) collective good open source, tools and shared research.
3D. OpenAI feels to me more and more like a commercial research lab; there does seem to be a lot of commercial partnering with their funding organizations (e.g. Microsoft).
4. I doubt the leadership there views the current ChatGPT models as unchanging. I expect there is a considerable revenue stream around the space. OpenAI is well positioned to play the game several steps ahead of others.
I would frame the broader question this way: for many years, there has been a hunger for this deeper AI research, due not only to (i) the expertise and resources required, but also (ii) to this hope that there is an organization that can maybe keep it within human or ethical bounds.
Unfortunately, this amorphous hope doesn't seem to be matching the actual organizational incentives nor dynamics. It is also unclear how much demand the public in free market will have for nobler research.
My position on these kinds of things is simple: follow the money. If we want an accountable public interest, AI research laboratory it's going to have to be designed, funded, and overseen very differently.
They want to stay top of mind.
Think about CocaCola, anyone can make a drink just as good. But it’s almost impossible to build their brand and distribution from scratch.
Note: I work at Databricks and am familiar with this project but didn't work on it.
Also Meta's licensing here https://github.com/facebookresearch/llama/blob/main/LICENSE
Can't be sure what that license actually reffers to, the language model or just the tooling in the Git Repo.
I agree its a minefield, but with Meta I would eer on the side of caution.
Companies routinely ban users for ToS violations. Just look at any thread about Google on here to see people complaining about it.
[1]: https://www.ftc.gov/advice-guidance/competition-guidance/gui...
Google might be banning people for enforceable violations of their ToS but imagine the uproar if they banned a Bing engineer for using Google search to find solutions for some Bing problem (which is similar to the problem here). The upside for Google or OpenAI would be somewhat limited but the downside is almost boundless.
With these things, it is usually the other way around.
If you are a small fish, no one will care. But if you are big enough, that money could be extracted from you, then they will come. A big org just has better lawers and negotiating power, but they really cannot ignore the law. Especially not, if there is a competitor with money to sue.
So if you are small and want to become big, better be cautious on the legal ground you are walking.
A. it's an output gained via following the letter of the law (TOS).
B. TOS only applies directly to people who've accepted the TOS, unless alpaca's license/TOS ALSO forwards the same criterion as it's source at openai, then derivatives wouldn't apply.
It's like if an app developer on IOS violated a TOS, and apple tried to go after everybody who ever used the app, they didn't agree directly to the TOS, only the developer did.
It's no surprise really though, from what I see they recognised some way to monitize and rolled back their commitment.
But this Dolly doesn't depend on Llama (unless I'M missing something), so you don't have to use it.
Besides: How would anyone ever know which model generated the output you are serving? AFAIK there is no fingerprint in any model’s output. And even if there was, it would probably be destroyed by fine tuning “over it”.
It seems like there easily could be. What if some of the data they trained it on didn't exist anywhere else except in the training set, and was put there specifically for this purpose? For instance they could have taught it a few poems that don't exist anywhere else. If you can coax the LLM of unknown origin into reciting those poems back to you, you know where it came from.
There's precedent for "whatever you can get away with" in tech companies, but establishing a culture of that at the start of this new big change could end up undesirable for most people.
For example, it could relieve demand for more legal and sustainable ways, until it's too late. (Look at the history of digital entertainment media piracy and DRM and legislation, for example. Or look at the history of software piracy, where some big companies seem to actually want their product to be pirated, partly because it builds a bigger moat against competitors, and they can legally strongarm some of those pirates later.)
This uses a fully open source (liberally licensed) model and we also open sourced (liberally licensed) our own training code. However, the uptraining dataset of ~50,000 samples was generated with OpenAI's text-davinci-003 model, and depending on how one interprets their terms, commercial use of the resulting model may violate the OpenAI terms of use. For that reason we are advising only noncommercial use of this model for now.
The next step here is to create a set of uptraining samples that is 100% open. Stay tuned.
I wonder how small can these models get? From 175B to 6B with comparable performance is huge, but can it go lower?
Will also commend the authors for not falling into the “LLMs can’t perform without 200B params!” fallacy. For anyone reading, 6B params is enough to train on a 3090. A PC rig for training or running inference with this would put you back maybe 4k$.
The end game here is likely getting the model to perform well in millions of parameters on specific tasks. Most business uses of ChatGPT are pretty closed domain tasks, it wouldn’t be a huge step to distill this model on a specific task and get it down to 150-350M params (which is roughly BART size and can run on AWS Lambda even).
Not many humans would even get this answer correct.
I am impressed.
Download GPT-J-6B from Eleuther
Download Alpaca Fine Tuning Code + Alpaca Examples
Train for 6 hours or so.
Get vaguely good RLHF model
In short, yes. More data, higher quality data, more epochs on the data. That is the name of the game.
> Although we may observe an emergent ability to occur at a certain scale, it is possible that the ability could be later achieved at a smaller scale—in other words, model scale is not the singular factor for unlocking an emergent ability. As the science of training large language models progresses, certain abilities may be unlocked for smaller models with new architectures, higher-quality data, or improved training procedures. For example, there are 14 BIG-Bench tasks5 for which LaMDA 137B and GPT-3 175B models perform at near-random, but PaLM 62B in fact achieves above-random performance, despite having fewer model parameters and training FLOPs.
So it's not obvious that it should be so straightforward.
If you have a use case or a bunch of disposable income then go with the “bitter” one.
Isn't there? I'm certainly not sure, based on the results published over the last weeks and months.
The giant GPT-{3.5,4} models show that if you make the model big enough and throw enough data at it you can produce an AI capable of conversing on basically any topic, in dozens of languages. There are plenty of different takes on how near-human its abilities are on specific tasks, but it's worth stepping back and appreciating how super-human the breadth of this knowledge is.
But it's also not clear if a mega-model is anything close to the most efficient way of storing knowledge. After all, you don't need to memorize every fact in Wikipedia if you know how to effectively search it.
And we're currently seeing a daily explosion in these capabilities. Today's flavor is interfacing with Wolfram, but we've also seen web searches, python coding, etc. That, I think, it the real superpower that comes out of this: you or I can answer a question by "doing a web search" or "query a database" or "use wolfram" or "develop a python program that finds the answer" However, an AI could do tasks like this just by "thinking" about it. Maybe it would be as natural as we find blinking.
That to me is the real breakthrough in stuff like Alpaca -- start with a mega-model and prompt it with something like: "After this paragraph, you are going to be speaking to a AI model similar to yourself but much more primitive. Its task will involve interfacing with English speakers, so converse with it only in that language. It has access to the same {X,Y,Z} APIs you have so any time it has trouble answering a question, prefer to give hints about how it could find the answer using those APIs rather than providing the answer directly yourself. Only give an answer directly if it repeatedly fails to be able to answer it by using an API. I've provided a large set of standardized tests used by humans at this URL -- start by asking it questions intended for a preschool-aged child. Each time it is able to answer new questions at a given level correctly 99% of the time increase the material's level until it is able to achieve that score on a test designed for a Computer Science PhD candidate"
How large would the "student" model have to be to succeed at this deep but narrower task? I think the answer right now "we have no idea". However if the model has the advantage that it can rely on external knowledge and tools from the start (and is rewarded by the "teacher" for doing just that) I bet it'll be a lot smaller than these mega-models. Sure, you wouldn't be able to disconnect the "student-AI" from its APIs and expect it to converse with you in Hungarian about the history of yacht design, but that might not be a capability it needs to have.
My personal hunch is that we're going to find these "AI-taught specialist AI, with API access" models will be a lot smaller than most people are expecting. That's the moment when things REALLY change: instead of pairing a human with a mega-model AI, if specialized models are cheap someone can say "spin up 100K expert-programmer AIs and have them supervized by 5K expert-manager AIs and have them build XYZ"
Or if you need it to work on an existing task you'd specialize further -- you'd go to your AI vendor and say "I'd like to license the weights for your expert-programmer model, but first have it read these 200 books I consider important to my problem domain and then show it every commit ever made by a human to my git repo and every design document I have"
Yes, but what if you do? Imagine your hyper-specialzied API-heavy model takes 10x less resources to answer a question (or at least a question relevant to the task at hand) Won't it be more powerful to have a model that can run 10 times as fast (or run 10 instances in parallel)?
What if the ratio turns out to be 100x or 1000x?
So I agree that the cutting edge of "best possible AGI" might mean building the largest models we can train on massive clusters of computers and then run on high-end hardware. My hunch, though, is that models that can be run on cheap hardware and then "swarmed" on a problem space will be even more powerful in what they can perform in aggregate.
Again, it's just my hunch but right now I think everybody's predictions are hunches.
I'll actually go one bit further: even for a linear task that can't be "swarmed" in the same way, it could be that cheaper-per-token models could even do better on linear problem-solving tasks. Existing models already have the ability to use randomness to give more "creative", if less reliable, answers. This is inherently parallelizable though -- in fact Bard seems to be exposing this in its UI in the form of multiple "drafts". So what if you just ran 100 copies of your cheap-AI against a problem and then had one cheap-AI (or maybe a medium-AI) judge the results?
Or at the risk of a getting too anthropomorphic about it: imagine you as a human are writing a program and you get stuck on a tricky bit -- you know that the problem should be solvable but you've never doing anything similar and don't know what algorithm to start with. Suppose then you could tell your brain "Temporarily fork off 100 copies of yourself. 10 of them go do a literature review of every CS paper you can find related to this topic. 10 of you search for open source programs that might have a similar need and try to determine how their code does it. The other 80 of you just stare off into the middle distance and try to think of a creative solution. In two human-seconds write a summary of your best idea and exit. I'll then read them all and see if I/we are closer to understanding what to do next"
For us, this type of mental process is so alien we can't even imagine what it would feel like to be able to do. It might come completely natural to an AI, though.
yeah you're onto something. models good enough to sustain a conversation where I bring my own data as a primer are probably more useful that models that have a frozen knowledge of everything. the killer feature of gpt-4 is the 32k token size, which allows unprecedented amount of input to be fed into the knowledge graph and queried.
how long until IBM, Tesla and Oracle announce Me-Too LLMs?