DeepSpeed Chat: Easy, fast and affordable RLHF training of ChatGPT-like models
github.com
github.com
Not that reproducing GPT-4 is going to be easy with this, but it'll definitely get rid of some major hurdles. I read a report about the difficulties HuggingFace had with producing their Bloom model, and a lot of it was the sort of straight forward systems engineering that goes into tooling like this.
Is the Bloom model considered a failure by the community? If you read the introduction it was supposed to include improvements over GPT3, but it performs much worse, I guess because of lower quality training data? I wonder what sort of company would have high enough quality data that they could use this project to fine tune a public model to the point where it would be better in some scenario than plain old GPT4 would be. Especially when you can just inject extra info in to the GPT4 prompt, like phind does for example. What even is the use of fine tuning given GPT 4 exists?
What nothing else can provide is something that will reason intelligently with the data you give it, with similar or better quality to paying for something like MTurk, for much cheaper and nearly instant delivery. That reasoning ability comes from the model size and training data quality, and in real applications using CoT, LangChain etc a lot of it comes from the context length. 8k is better than anything else I've tried at real use cases, and I very much want to try 32k because that opens up a lot of space to do new things (e.g. dump in a textbook on the domain you want the model to reason about). I want even longer context lengths than that too, but we'll have to see how it develops. From what I understand context length/block size is a pretty pure relationship to the amount of compute and memory they're willing to devote during training. RWKV's architectural changes may shake that up a bit, we'll see when Stability releases it.
your point?
In my mind, MSFT spent that money to acquire a head start on getting LLM-style capabilities into MS’s profitable product portfolio. This is money well spent:
1. MSFT can and will make money on these capabilities.
2. If MSFT didn’t do this, they would take the substantial risk of someone else pulling it off and attacking their moat.
I can’t really imagine today’s Google pulling this off with Google Docs. Adobe doesn’t target MS’s market directly enough to be an immediate risk. Apple doesn’t seem interested in competing with MS. Meta is doing its own thing in the corner. But someone really could attack MS with something amazing and make short-term sales, which could turn into a long term loss for MS. (Salesforce? They don’t seem able make things that normal people want to use.). But MS is now ahead of the curve, and they didn’t really spend that much money to get there.
Keep in mind that LibreOffice is not vastly less capable than Office 365, and Teams is not exactly an insurmountably strong piece of technology.
The LLM space is moving fast. OpenAI may stay on top for a long time, or it may not. But I expect Microsoft’s use of LLMs to be valuable for MS and likely market-leading in the office AI space for quite some time regardless of what happens to OpenAI.
personally I think in 10 years people will joke about the machine generated boilerplate in the same way they joke about Clippy today
The idea is to get the people who arent willing to pay hooked on what they offer. Once you are used to a system you will probably want the same thing at your workplace, where they can charge a a prumium. Same thing was done with windows in asia.
This seems like more evidence that under the "commoditize your complement" framework, all intellectual property is the complement, and the only thing actually worth selling for Microsoft is subscriptions and server time.
Databricks wants to have people use their compute to use other LLMs and Microsoft wants that compute to be Azure.
> With just one click, you can train, generate and serve a 1.3 billion parameter ChatGPT model within 1.36 hours on a single consumer-grade NVIDIA A6000 GPU with 48GB memory. On a single DGX node with 8 NVIDIA A100-40G GPUs, DeepSpeed-Chat enables training for a 13 billion parameter ChatGPT model in 13.6 hours. On multi-GPU multi-node systems (cloud scenarios),i.e., 8 DGX nodes with 8 NVIDIA A100 GPUs/node, DeepSpeed-Chat can train a 66 billion parameter ChatGPT model under 9 hours. Finally, it enables 15X faster training over the existing RLHF systems
> The following are some of the open-source examples that are powered by DeepSpeed: Databricks Dolly, LMFlow, CarperAI-TRLX, Huggingface-PEFT
(disclaimer: MSFT/GH employee, not affiliated with this project)
I wouldn't call an A6000 "consumer-grade" -- it's about $5000 and 99% of consumers that have graphics cards wouldn't have that.
Top of the line consumer grade GPU would be a Nvidia RTX 4090/3090 with 24GB VRAM.
You can link two 3090s with nvlink to increase bandwidth.
Which totally begs the question, Microsoft AI researchers get two hour coffee breaks?
—edit—
The way they word it is confusing but, yeah, fine tune the model on your two hour coffee break.
You can also bootstrap RLHF training data from the gpt4 api. Vicuna is probably the best public model created with gpt4 data available as of today.
On the other hand, if you're being cheeky, I bet there's a way to datamine from websites like ShareGPT and profit off shared ChatGPT <> User interactions.
I know there was a summary but the point is that ChatGPT just really accelerates a LOT of bulk work we were used to having to do manually.
It's an amazing time to be alive!
It doesn't increase it's knowledge of the world or increase it's capabilities.
EDIT: Is the idea that the critic model learns via the PPO process and gives a value estimate to prefixes of the responses?
I am hoping that someone makes a very simple Jupyter notebook where I can enter my RLHF file and select a few other settings and just run (on AWS or Azure; willing to pay per fine-tuned model say $100-$500 for cloud credits + notebook access).
Aka, the article under discussion.
(Though some people use other words for the F part.)
Haven’t felt this way since I realized about 50% of my social network meant “In My Honest Opinion” and the other 50% meant “In My Humble Opinion”