Bagel: Open-source unified multimodal model
bagel-ai.org
bagel-ai.org
It works following readme instructions at least on Ubuntu, on my RTX 3090 GPU with 24 gigs of memory, just barely. Have to close most other windows and lower screen resolution to be able to load the model. Then it generates or edits images in 2-3 minutes. I only have this one GPU and am using Chrome to use the browser interface on the same machine.
The original release won't run on this hardware, but the compressed one is supposed to give identical results.
If you want to try only voice, Try unmute.sh by Kyutai which will be eventually open-sourced
I'm still interested in what keyword I could use to search for the latest research in voice models.
It could just be using speech to text (e.g. Whisper) on your input, and then using its text model on the text of your words. Or has OpenAI said that they aren't doing this?
> Advanced voice uses natively multimodal models, such as GPT-4o, which means that it directly “hears” and generates audio, providing for more natural, real-time conversations that pick up on non-verbal cues, such as the speed you’re talking, and can respond with emotion.
From https://help.openai.com/en/articles/8400625-voice-mode-faq
Has anyone here experimented with fine-tuning this for domain-specific applications?
ETA: oof, and it's still getting hands wrong. (Editing demo #12)
It seems to be 7b, but like with other new architectures expect to not be able to run it quantizised.
So quick napkin math can give you the VRAM usage for loading the model. 7b can be ~14GB full, 7GB in fp8 and ~3.5GB in 4bit (AWQ, int4, q4_k_m, etc). But that's just to load the model in VRAM. You also need some available VRAM to run inference, and there are a lot of things to consider there too. You need to be able to run a forward pass on the required context, you can keep a kv cache to speed up inference, you can do multiple sessions in parallel, and so on.
Context length is important to take into account because images take a lot of tokens. So what you could do with a 7b LLM at full precision on a 16GB VRAM GPU might not be possible with a VLM, because the context of your query might not fit into the remaining 2GB.
The direct parm conversion math tends to be much less reliable than one would expect once quants are involved.
e.g.
7B @ Q8 = 7.1gb [0]
30B @ Q8 = 34.6gb [1]
btw you can also roughly estimate expected output speed too if you know the device memory throughput. Noting that this doesn't work for MoEs
Also recently discovered that in CPU mode llama.cpp does memory mapping. For some models it loads less than a quarter into memory.
If you wanna call it Bagel, just call it Bagel. No need to make up a justification.