WizardLM 2
wizardlm.github.io
wizardlm.github.io
We're still crunching the 8x22B model to get it ready, and the 70B model isn't yet available.
I've tried downloading it twice now.
Here's one example of the same prompt against the three models: https://twitter.com/blixt/status/1780191143747101072
(Running on an M3 Max 128GB it's quite fast – the times shown are including initial loading time which does take a while)
* 8x22B (based on Mistral 8x22B)
* 70B (based on llama2?)
* 7B (based on Mistral-7B-v0.1)
and the 70B model is not available yet on huggingface
https://huggingface.co/collections/microsoft/wizardlm-661d40...
Does anyone know what is going on? Were there any announcements on its being unpublished?
But Claude Sonnet is not open source and can't be run on your own hardware like Wizard.
It looks like the people who focasted that AI models would need to keep growing to improve their performance where completely misguided. I wonder if in 3 to 4 years we'll end up with models with less than 4B parameters we comparable performance as today's State of the Art.
[1]: https://paperswithcode.com/sota/multi-task-language-understa...
A 4B has limited capacity to encode knowledge and algorithmic circuits. It's also too small to learn programs whose execution exceed the sizes of circuits it can encode. There is a hard cap on how much we can squeeze out of small models. What we need is better consumer hardware, so we don't have to hope for miracles.
Another hope is that the 1.58 bit/ternary quantization aware training of model scaling pans out. That'd be another axis of inefficiency beyond just parameter count.
That's exactly what I mean by misguided, they were focusing on the wrong metric.
> A 4B has limited capacity to encode knowledge and algorithmic circuits. It's also too small to learn programs whose execution exceed the sizes of circuits it can encode. There is a hard cap on how much we can squeeze out of small models.
There must be some kind of cap indeed, but at this point we have no idea where it is, especially now that there are emerging infrastructures that aren't just transformers popping out.
> What we need is better consumer hardware, so we don't have to hope for miracles.
Hardware improvements are going to be a big factor, but there's no reason to think the software part (+ training data) isn't going to continue improving a lot.
7B: 32k 8x22B: 65k
So like that.
I usually use Mistral 7B-Q4 with Ollama's REST interface. I am curious about and looking forward to slightly adjust the prompts I use for wizardlm2 and see how it does.
If openai cares enough to stop people scraping responses, then we’ll just crowdsource them like open assistant did with their dataset.
I think the TOS should have no effect on the license and copyright of the resulting model, since outputs of AI Models cannot be copyrighted (at least in the US).
If so, what’s the easiest way?
Running WizardLM-2 70B or lower WizardLM-2 7B is much more feasible however.
Look into Ollama: