Vicuna v1.5 series, featuring 4K and 16K context, based on Llama 2
twitter.com
twitter.com
my favourite leaderboard is this one as it compares open source and closed source:
I use wizardlm myself right now and have used vicuna in the past. Do you mean there's a model trained on the datasets of both wizardlm and vicuna?
Vicuna + Wizard - Censoring:
https://huggingface.co/TheBloke/Wizard-Vicuna-30B-Uncensored...
Seems to be based upon this: https://github.com/melodysdreamj/WizardVicunaLM
It puts LLaMA2 Chat 70B at 92.66% and GPT-4 at 95.28%, only a ~3% difference.
The base Llama-Chat models use something called "ghost attention" (they describe it in their paper). No clue how it works, but the result is, the model sticks to the system prompt extremely well. If you tell it that it's Marvin the Paranoid Android in the system prompt, it will stick 100% to that.
In llama-derived models like Vicuna, if you tell it to act as Marvin the Paranoid Android, eventually the regular "assistant" voice starts to bleed through, after only a few chat turns.
Doesn't sound like a big deal, but in cases where you have strict rules you want the model to follow, then the base llama-2-chat models are far better than any derived ones that do not implement ghost attention.
But during inference, there's no trick. The system message remains once at the top.
So would need to make sure comparing apples to apples.
Would be nice to know how close they are to the paid ones.
Edit: Found one here:
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
It doesn't include a comparison to the paid ones, but it's a great place to look at different ones and see which ones you should experiment with and sink time into.
Have you tried any of the high-scored ones?
That said, if you’re looking for these to be a similar quality to OpenAI’s ChatGPT 4, they’re not even remotely close.
I figured as much, but it's incredibly hard to find any comparison that's kept up to date.
What I was after was a local javascript LLM coder that i could train on a local codebase, but even that was fruitless except for some unmaintained "promptr" project.
I guess OpenAI embeddings is the only option.
As another poster mentioned though, it's nowhere near the level of GPT-4. It's close enough to GPT 3.5 though, you should try it out!
i.e., do these benchmarks taken together capture anything resembling general intelligence?
Bad programming is just bad programming, no need to pull what language is being used into it. Just happens to be JavaScript because that's what browsers natively support. But if Python (or whatever) was used instead, the very same programmer would have done the very same mistake.
Counterpoint: It's not what programming languages do, it's what they shepherd you to
https://nibblestew.blogspot.com/2020/03/its-not-what-program...
You're right it isn't the language but since js is the only front end language it gets called out.
Just so we are clear, I love JavaScript and write it everyday but I avoid the “modern” eco system like it is the plague.
The sorting is actually fine. On my Mac the sort completes in less than 1ms.
The performance issue is because of all of the hooks of the svelte framework trying to do their thing in the middle of the sort, and then another frame-sizing script trying to do its thing in the middle of the sort. This is an endemic problem in Angular apps as well.
Not javascript's fault, or even the sort algos fault.
Its interesting to see how fast these models evolve.