532 karma · joined February 2, 2015
Last time we only released the quantized GGUFs. Only llama.cpp users could use it (+ Ollama, but without vision).
Now, we released the unquantized checkpoints, so anyone can quantize themselves and use in their favorite tools, including Ollama with vision, MLX, LM Studio, etc. MLX folks also found that the model worked decently with 3 bits compared to naive 3-bit, so by releasing the unquantized checkpoints we allow further experimentation and research.
TL;DR. One was a release in a specific format/tool, we followed-up with a full release of artifacts that enable the community to do much more.
- Encoder based models have much faster inference (are auto-regressive) and are smaller. They are great for applications where speed and efficiency are key. - Most embedding models are BERT-based (see MTEB leaderboard). So widely used for retrieval. - They are also used to filter data for pre-training decoder models. The Llama 3 authors used a quality classifier (DistilRoberta) to generate quality scores for documents. Something similar is done for FineWeb Edu
- Base model: Mixtral 8x22B, 8 experts, 141B total params, 35B activated params
- Fine-tuned with ORPO, a new alignment algorithm with no SFT step (hence much faster than DPO/PPO)
- Trained with 7K open data instances -> high-quality, synthetic, multi-turn
- Apache 2
Everything is open:
- Final Model: https://huggingface.co/HuggingFaceH4/zephyr-orpo-141b-A35b-v...
- Base Model: https://huggingface.co/mistral-community/Mixtral-8x22B-v0.1
- Fine-tune data: https://huggingface.co/datasets/argilla/distilabel-capybara-...
- Recipe/code to train the model: https://huggingface.co/datasets/argilla/distilabel-capybara-...
- Open-source inference engine: https://github.com/huggingface/text-generation-inference
- Open-source UI code https://github.com/huggingface/chat-ui
Have fun!
- 2 million patches
- 1068x1068 pixel patches
- 2.5 trillion pixels
Read more in https://huggingface.co/posts/aliFrancis/293058125194160
I really appreciate you taking the time to provide all this feedback. This feedback + additional resources are extremely useful.
I agree that the subtitle is not as accurate as it could be. I'll revisit it! As for content updates, I've been doing some additional updates in the last days based on feedback (e.g. more info about tokenization and the token embeddings). Although diving in some of your suggestions is likely out of scope for this article, I in particular agree that expanding the attention mechanism content (e.g. the analogy with databases or explaining what is dot product) would increase the quality of the article. I will look into expanding this!
I also think a more rigorous, separate mathematical exploration into attention mechanisms and recent advancements would be a great tool for the ecosystem.
Once again, thank you for all the amazing feedback!
That said, the whole neural network will be sensible to large values, so it won't be fixed by a numerically stable softmax. The normalization is a key aspect for the network to work.
Let's say you want to predict if you'll pass an exam based on how many hours you studied (x1) and how many exercises you did (x2). A neuron will learn a weight for each variable (w1 and w2). If the model learns w1=0.5 and w2=1, the model will provide more importance to the # of exercises.
So if you study for 10 hours and only do 2 exercises, the model will do x1w1 + x2w2=10x0.5 + 2x1 = 7. The neuron then outputs that. This is a bit (but not much) simplified - we also have a bias term and an activation to process the output.
Congrats! We built our first neuron together! Have thousands of these neurons in connected layers, and you suddenly have a deep neural network. Have billions or trillions of them, you have an LLM :)
MoEs are especially useful for much faster pre-training. During inference, the model will be fast but still require a very high amount of VRAM. MoEs don't do great in fine-tuning but recent work shows promising instruction-tuning results. There's also quite a bit of ongoing work around MoEs quantization.
In general, MoEs are interesting for high throughput cases with high number of machines, so this is not so so exciting for a local setup, but the recent work in quantization makes it more appealing.
- Large language models: LLaMA, LLaMA v2, Falcon, Phi-v1.5, StarCoder. - Quantized models with the llama.cpp approach: LLaMA, T5, Phi-v1.5 - Image generation: Stable Diffusion (v1.5, v2.1, and XL), Wuerstchen. - Computer Vision: DINOv2, yolo-v3, yolo-v8, Segment-Anything Model. - Text-to-speech: Whisper.
Thanks to using pure Rust, it's easy to directly rpedict in the browser using WASM. See these examples using Yolo, Whisper, SAM, T5 and Llama 2 https://huggingface.co/collections/radames/candle-wasm-examp...
- Multilingual: Generates speech in 13 different languages
- Voice cloning with 3 seconds of audio
- Cross language voice cloning as well
- 24khz quality
- Blog: https://coqui.ai/blog/tts/open_xtts
- Demo: https://huggingface.co/spaces/coqui/xtts
- Model: https://huggingface.co/coqui/XTTS-v1
- Trained on 3.5 trillion tokens
- 7 million GPU hours
- Quality on par with PaLM 2, outperforming Llama 2 and GPT
-3.5 across benchmarks
- 4-bit and 8-bit show little degradation
https://github.com/huggingface/safetensors https://huggingface.co/blog/safetensors-security-audit
This allows using LoRA to fine-tune SDXL with DreamBooth on a T4 GPU, which is provided by free in Google Colab.
As an example, you can run GPT-4 on your Mac or Pixel! On M1 Pro, you can achieve 16 tokens/sec.