HNHacker News
TopNewBestAskShowJobs

osanseviero

532 karma · joined February 2, 2015

submissionscomments
osanseviero··on DiffusionGemma: 4x Faster Text Generation
Hi! What implementation are you using? Right now VLLM is the one recommended. llama.cpp is in an early draft
osanseviero··on Gemma 3n preview: Mobile-first AI
Hi! The model is 8B if you also load the vision and audio components. We just used the text model in LMArena.
osanseviero··on Gemma 3 QAT Models: Bringing AI to Consumer GPUs
Hi! Omar from the Gemma team here.

Last time we only released the quantized GGUFs. Only llama.cpp users could use it (+ Ollama, but without vision).

Now, we released the unquantized checkpoints, so anyone can quantize themselves and use in their favorite tools, including Ollama with vision, MLX, LM Studio, etc. MLX folks also found that the model worked decently with 3 bits compared to naive 3-bit, so by releasing the unquantized checkpoints we allow further experimentation and research.

TL;DR. One was a release in a specific format/tool, we followed-up with a full release of artifacts that enable the community to do much more.

osanseviero··on Gemma 3 Technical Report [pdf]
Please make sure to update to the latest llama.cpp version
osanseviero··on Plasticlist Report – Data on plastic chemicals in Bay Area foods
Nat Friedman leads the project. He was GitHub's CEO, among many other things. He funds many interesting ambitious projects, such as the Vesuvius Challenge (https://scrollprize.org/)
osanseviero··on BERTs Are Generative In-Context Learners
Yes, they are still used

- Encoder based models have much faster inference (are auto-regressive) and are smaller. They are great for applications where speed and efficiency are key. - Most embedding models are BERT-based (see MTEB leaderboard). So widely used for retrieval. - They are also used to filter data for pre-training decoder models. The Llama 3 authors used a quality classifier (DistilRoberta) to generate quality scores for documents. Something similar is done for FineWeb Edu

osanseviero··on Open source AI is the path forward
Yes, there are a few dozen full open source models (license, code, data, models)
osanseviero··on DevRel at HuggingFace
Hi all! I'm Omar from Hugging Face. Happy to answer any questions you might have about Hugging Face in general, llamas, and open ML!
osanseviero··on Zephyr 141B, a Mixtral 8x22B fine-tune, is now available in Hugging Chat
Zephyr 141B is a Mixtral 8x22B fine-tune. Here are some interesting details

- Base model: Mixtral 8x22B, 8 experts, 141B total params, 35B activated params

- Fine-tuned with ORPO, a new alignment algorithm with no SFT step (hence much faster than DPO/PPO)

- Trained with 7K open data instances -> high-quality, synthetic, multi-turn

- Apache 2

Everything is open:

- Final Model: https://huggingface.co/HuggingFaceH4/zephyr-orpo-141b-A35b-v...

- Base Model: https://huggingface.co/mistral-community/Mixtral-8x22B-v0.1

- Fine-tune data: https://huggingface.co/datasets/argilla/distilabel-capybara-...

- Recipe/code to train the model: https://huggingface.co/datasets/argilla/distilabel-capybara-...

- Open-source inference engine: https://github.com/huggingface/text-generation-inference

- Open-source UI code https://github.com/huggingface/chat-ui

Have fun!

osanseviero··on Grok
The model is also at https://huggingface.co/xai-org
osanseviero··on European Space Agency open-sources largest Sentinel-2 dataset
The dataset has

- 2 million patches

- 1068x1068 pixel patches

- 2.5 trillion pixels

Read more in https://huggingface.co/posts/aliFrancis/293058125194160

osanseviero··on Understand how transformers work by demystifying the math behind them
Hey @godelski! Author of the blog post here.

I really appreciate you taking the time to provide all this feedback. This feedback + additional resources are extremely useful.

I agree that the subtitle is not as accurate as it could be. I'll revisit it! As for content updates, I've been doing some additional updates in the last days based on feedback (e.g. more info about tokenization and the token embeddings). Although diving in some of your suggestions is likely out of scope for this article, I in particular agree that expanding the attention mechanism content (e.g. the analogy with databases or explaining what is dot product) would increase the quality of the article. I will look into expanding this!

I also think a more rigorous, separate mathematical exploration into attention mechanisms and recent advancements would be a great tool for the ecosystem.

Once again, thank you for all the amazing feedback!

osanseviero··on Understand how transformers work by demystifying the math behind them
Yes, you're correct. I tried to connect a common training problem (gradient explosion and vanishing gradient) with the issue of softmax being sensitive to large values. I agree it's misleading/inaccurate, so will rewrite that part.

That said, the whole neural network will be sensible to large values, so it won't be fixed by a numerically stable softmax. The normalization is a key aspect for the network to work.

osanseviero··on Understand how transformers work by demystifying the math behind them
Just math, and not even that fancy.

Let's say you want to predict if you'll pass an exam based on how many hours you studied (x1) and how many exercises you did (x2). A neuron will learn a weight for each variable (w1 and w2). If the model learns w1=0.5 and w2=1, the model will provide more importance to the # of exercises.

So if you study for 10 hours and only do 2 exercises, the model will do x1w1 + x2w2=10x0.5 + 2x1 = 7. The neuron then outputs that. This is a bit (but not much) simplified - we also have a bias term and an activation to process the output.

Congrats! We built our first neuron together! Have thousands of these neurons in connected layers, and you suddenly have a deep neural network. Have billions or trillions of them, you have an LLM :)

osanseviero··on DeciLM-7B: The Fastest and Most Accurate 7B-Parameter LLM to Date
Base model, not fine-tuned
osanseviero··on Mixtral of experts
This blog post might be interesting - https://huggingface.co/blog/moe

MoEs are especially useful for much faster pre-training. During inference, the model will be fast but still require a very high amount of VRAM. MoEs don't do great in fine-tuning but recent work shows promising instruction-tuning results. There's also quite a bit of ongoing work around MoEs quantization.

In general, MoEs are interesting for high throughput cases with high number of machines, so this is not so so exciting for a local setup, but the recent work in quantization makes it more appealing.

osanseviero··on Purple Llama: Towards open trust and safety in generative AI
Model at https://huggingface.co/meta-llama/LlamaGuard-7b Run in free Google Colab https://colab.research.google.com/drive/16s0tlCSEDtczjPzdIK3...
osanseviero··on Kyutai – non-profit AI lab dedicated to open science with 300M euro of support
Kyutai, non-profit AI lab, is announced with 300M€ of support by Xavier Niel, Eric Schmidt, and Rodolphe Saadé, with a strong team from Meta, Google DeepMind, INRIA, and more, with experience in LLMs (Llama), Audio (MusicGen, AudioLM), and infra tooling (Candle).
osanseviero··on Update on the Candle ML Framework
We've first announced Candle, a minimalist ML framework in Rust 6 weeks ago. Since then, we've focused on adding various recent models and improved the framework so as to support the necessary features in an efficient way. You can checkout a gallery of the examples, supported models include:

- Large language models: LLaMA, LLaMA v2, Falcon, Phi-v1.5, StarCoder. - Quantized models with the llama.cpp approach: LLaMA, T5, Phi-v1.5 - Image generation: Stable Diffusion (v1.5, v2.1, and XL), Wuerstchen. - Computer Vision: DINOv2, yolo-v3, yolo-v8, Segment-Anything Model. - Text-to-speech: Whisper.

Thanks to using pure Rust, it's easy to directly rpedict in the browser using WASM. See these examples using Yolo, Whisper, SAM, T5 and Llama 2 https://huggingface.co/collections/radames/candle-wasm-examp...

osanseviero··on Data accidentally exposed by Microsoft AI researchers
You should check out safetensors. They are used widely in diffusion models and LLMs https://huggingface.co/blog/safetensors-security-audit
osanseviero··on Data accidentally exposed by Microsoft AI researchers
The safetensors format was created exactly for this - safe model serialization

https://huggingface.co/blog/safetensors-security-audit

osanseviero··on XTTS: Open-source Foundation TTS model
XTTS, the production-quality TTS model from Coqui, is released

- Multilingual: Generates speech in 13 different languages

- Voice cloning with 3 seconds of audio

- Cross language voice cloning as well

- 24khz quality

- Blog: https://coqui.ai/blog/tts/open_xtts

- Demo: https://huggingface.co/spaces/coqui/xtts

- Model: https://huggingface.co/coqui/XTTS-v1

osanseviero··on Open-source web Gaussian Viewer
Code: https://github.com/dylanebert/gaussian-viewer Demo: https://huggingface.co/spaces/dylanebert/gaussian-viewer
osanseviero··on Falcon 180B
- 180B parameters

- Trained on 3.5 trillion tokens

- 7 million GPU hours

- Quality on par with PaLM 2, outperforming Llama 2 and GPT

-3.5 across benchmarks

- 4-bit and 8-bit show little degradation

osanseviero··on Hugging Face raises $235M from investors including Salesforce and Nvidia
FYI the file format, safetensors, was proposed, developed and maintained by HF, and involved people from groups such as Eleuther and Stability for external security audits.

https://github.com/huggingface/safetensors https://huggingface.co/blog/safetensors-security-audit

osanseviero··on Idefics: Open Access 60B multimodal model
IDEFICS is an open access visual language model. It's an open source reproduction of Flamingo and similar to GPT-4, it can receive images and text inputs to generate text outputs. The 9B and 60B versions are released
osanseviero··on SDXL Dreambooth fine-tuning with LoRA on free Colab
Using the diffusers library and combining multiple optimizations - Gradient checkpointing - Mixed precision training - 8-bit Adam - Gradient accumulation

This allows using LoRA to fine-tune SDXL with DreamBooth on a T4 GPU, which is provided by free in Google Colab.

osanseviero··on Replit's new Code LLM: Open Source, 77% smaller than Codex, trained in 1 week
You can also check the NLP with transformers book
osanseviero··on Bloomz.cpp: Run multilingual BLOOM model with C++
bloomz.cpp allows running inference of BLOOM-like models in pure C/C++ (inspired by llama.cpp). It supports all models that can be loaded in transformers for BLOOM.

As an example, you can run GPT-4 on your Mac or Pixel! On M1 Pro, you can achieve 16 tokens/sec.

osanseviero··on SantaCoder: A new 1.1B code model for generation and infilling
The demo is up again!
Page 1 of 2Next →