Llama.cpp Now Supports Qwen2-VL (Vision Language Model)
github.com
github.com
Qwen2-VL is a decent vision model. You can try it out online here: https://huggingface.co/spaces/GanymedeNil/Qwen2-VL-7B - I got great results from it for OCR against handwritten text: https://simonwillison.net/2024/Sep/4/qwen2-vl/
Qwen2.5-Coder-32B is an excellent (I'd say even GPT-4 class) model at generating code which I can run on a 64GB M2 MacBook Pro: https://simonwillison.net/2024/Nov/12/qwen25-coder/
QwQ is the Qwen team's exploration of the o1-style of model that has built in chain-of-thought. It's absolutely fascinating, partly because if you ask it a question in English it will often think in Chinese before spitting out an answer in English. My notes on that one here: https://simonwillison.net/2024/Nov/27/qwq/
Most of the Qwen models are Apache 2 licensed, which makes them more open than many of the other open weights models (Llama etc).
(Unsurprisingly they all get quite stubborn if you ask them about topics like Tiananmen Square)
Works with +700 year old books w some tweaks. took like $400 to train. can't share more because i don't know more.
My guess is it “thinks” in non-eligible tokens
https://openai.com/index/learning-to-reason-with-llms/#chain...
I have seen the summaries dip into completely random languages like Thai, so it might switch between languages occasionally.
Has anyone made a political censorship eval yet?
Someone open sourced it with langchain
I wonder how the abliterated variants respond to this query.
The funniest was asking for an ascii graphics depiction of a minecraft watch recipe, and I was actually feeling quite sorry for it, 'wait that can't be right' 'let me try' 'still not right' round and round it went, at least a few pages at which point it decided to try the second recipe I'd asked about to see if that helped with the first.
I didnt know about the other models, 'coder' is downloading now, and fingers crossed it fits in 32GB and knows a bit about Zig.
It sounds like you got the vision one running locally on your M2, nice. I'm running Asahi Linux and not tried anything AI/SD/graphical orientated yet. But nice that you got some SVG out of coder, I never thought of using a coding model in that way.
I noticed that too, but I haven't seen it think in numbers in Chinese like most bilingual chinese speakers prefers. Or at least I haven't been able to trigger it.
It's also great to see qwq's open chain of thought baked in a OSS LLM so you can see it reason with itself in real-time, it's the kind of secret sauce that proprietary LLMs like o1 would prefer to keep hidden to try build a moat around.
We've got a lot to thank Meta and Qwen for in continually releasing improving high quality OSS models which also encourages others to follow. High quality OSS models are the best thing keeping the cost of LLMs down, you can get unbelievable value on OpenRouter with qwen2.5-coder:32b at $0.08/$0.18 M/tok qwq:32 available at $0.15/$0.60 M/tok which is more than 18x cheaper than Anthropic's latest budget Haiku 3.5 model at $0.80/$4 M/tok (4x price hike over Haiku 3.0).
I collected notes on the lowest cost hosted LLMs from the major vendors when I wrote up Amazon Nova last week: https://simonwillison.net/2024/Dec/4/amazon-nova/
Nova Micro is $0.035/$0.14 and Google's Gemini 1.5 Flash 8B is $0.0375/$0.15 - just beating those OpenRouter prices, but it may well be that the Qwen models provide better results.
Also worth shouting out you can get Meta's latest llama-3.3:70b (comparable to llama3.1:405b but must faster and cheaper) within GroqCloud's free quotas running at an impressive 276 tok/s.
Right now, with the availability of open weights for cutting-edge models, it feels like this wave of technological advance is pleasantly decentralised however. I can download and run a model and tinker with things which at least feel like the seeds of such a future, where I _might_ be able to build things with my own interests at heart.
But what happens if these models stop being shared, and how likely is that? Reading about the vast quantities of compute deployed to train them, replicating the successes of the main players with a community of volunteers just seems an order of magnitude less achievable than traditional OSS efforts like Linux. This wave feels so tied to massive scale for its success, what do we do if big-tech stop handing out models?
But I don't see why Meta and Alibaba would stop releasing their best models as OSS, since they benefit from the tooling, optimizations and software ecosystems being developed around their OSS models and don't benefit from a future where the best AI models are centralized behind the big tech Corps. As long as their core business remain profitable I don't expect them to stop improving and sharing their OSS models.
Having a personal robot would be great, but they have to invent a fully offline real positronic brain before I will consider allowing one in my house.
Fully open source might be too much to hope for, but that would obviously be the ideal. If it is closed source it definitely should be offline. I can have another, carefully sandboxed, AI in my computer that can help out with tasks that require online access. No need for the two types to be built into the same device.
Prediction 2: The ecosystem around open source models will grow to be much larger, richer, and deeper than closed source models.
If these are true, then OpenAI and Anthropic are in a precarious place. They basically burned a lot of capital to show the open source second movers what to build.
thinking that the Chinese government might have built in a back door gives me a little pause though
Letting an LLM run arbitrary commands in your main user account seems risky even without worrying about conspiracies.
I was thinking "here's an IP address and ssh key" would be what to phone home with, and that could be encrypted/hidden pretty well, but any network access should be pretty suspicious right away.