Stable LM 3B: Bringing Sustainable, High-Performance LMs to Smart Devices
stability.ai
stability.ai
I don't know how to fine tune an LLM. Does anyone have good resources on how to do this?
They refuse to even acknowledge why those people might be annoyed their work was used without their consent or compensation to put them out of work.
The HF toolchain is pretty mature and most llm finetuning projects are a wrapper around HF models, HF Trainer and some config templates.
An LLM by default would be trained like in the example below, but it would take a lot of VRAM and time.
https://huggingface.co/docs/transformers/training
That’s where things like PEFT LoRa, gptq and accelerate contribute to make your training faster/require less VRAM so you could do it on a consumer GPU with 16-24GB.
For example: https://huggingface.co/docs/peft/quicktour
Then for the tips and tricks, either reddit as suggested by a sibling comment, or Discord communities. Huggingface, EleutherAI and LAION discord servers are all great and have super helpful, friendly and knowledgeable people.
> Login or Sign up to review the conditions and access this model content.
How does that work with just calling `tokenizer = AutoTokenizer.from_pretrained("stabilityai/stablelm-3b-4e1t")`
looking at the 3b results (here https://github.com/Stability-AI/StableLM#stablelm-alpha-v2 ?), it looks like Mistral (which outperforms Llama-2 13b) is far more powerful
There are improved versions coming but this is the best 3b model and Mistral is the best 7b model.
You can check out real world performances on devices here: https://llm.mlc.ai/
Do you have plans to train/release a fine-tuned 3b chat version or other variants?
How so? It's on the general order that seems prevalent (and significant) with LLMs? 3, 7, 15 billion?
Not really. A 3B model quantized to 4 bits should run in any reasonable smartphone (using around 2GB of memory).
My phone can run a 7B parameter model at 12 tokens per second, which is probably faster than most humans are comfortable reading, and definitely faster than a virtual assistant would speak.
Out of curiosity, I tested a 3B parameter model, and it runs at about 21 tokens per second on my phone.
Many of these use cases are possible today with either specialized models, or are old school and rule-based. Being able to have an LLM apply soft judgment on a device that generates so much contextual information, and completely locally/privately, is bound to make smartphones an entirely new kind of device.
7 hours * 3600sec/hr * 21token/sec = 530,000 tokens per night on this hardware, assuming no thermal throttling. (I don’t have data to say what the sustained rate would be, throttling could happen.)
There are other reasons to want a higher throughput. To perform retrieval or for a chain-of-thought approach, you typically need to run several prompts per user prompt, effectively impacting user-perceived performance of your LLM based solution.
3b is also small enough to fit in a wasm runtime for browser based text local text generation.