Best 7B LLM on leaderboards made by an amateur following a medium tutorial
huggingface.co
huggingface.co
It's a bug, not a feature, that the stock leaderboard ends up with endless fine tunes, and as you point out and is demonstrated by the article, its more about something else than about quality
That said, we do have public 120b models now that genuinely feel better than the original gpt-3.5.
The holy grail remains beating gpt-4 (or even gpt-4-turbo). This seems to be out of reach on consumer hardware at least...
-- and as long as we're asserting andecdotes freely, I work in the field and have a couple years in before ChatGPT -- it most certainly is not a well-kept secret or a secret or true or anything else other than standard post-millenial self-peasantization.
"Outright lie" is kinder toward the average reader via being more succinct, but usually causes explosive reactions because people take a while to come to terms with their ad-hoc knowledge via consuming commentary is fundamentally flawed, if ever.
gpt-4-turbo also has an extremely annoying quirk where instead of producing a complete response, it tends to respond with "you can do it like this: blah blah ...; fill in the blanks as needed", where the blank is literally the most important part of the response. It can sometimes take 3-4 rounds to get it to actually do what it's supposed to do.
But it does produce useless output much faster indeed.
It's frustrating for both of us, I assume.
I'm tired of people fact-free asserting it got worse because don't you know other people saw it got worse? And it did something bad the other day.
You're tired of the thing not doing the thing and you have observed it no longer does the thing. And you certainly shouldn't need to retain past prompts just to prove it.
https://en.wikipedia.org/wiki/No_free_lunch_theorem
In the speed/quality trade-off sense, there have /often/ been free lunches in many areas of computer science, where algorithmic improvements let us solve problems orders of magnitude faster. We don't fully understand what further improvements will be available for LLMs.
“medium” should be capitalised in the title, as it refers to the blogging platform.
LLM Leaderboards and benchmarks do not show the complete picture for YOUR specific usecase.
Mixtral does about average at summarization. Not enough details are presented in the summary. Passable for non-decisionbound workflows. Using quantized models injects occasional spelling errors, so fall back on Mistral unquantized. Skip all the Mistral finetunes. Only choose Instruct models.
There are no good local code comprehension/generation models (and that includes CodeLlama, DeepSeek, Starcoder, WizardCoder, Phind). You will spend as much time guiding/rerolling/correcting output as coding it yourself, assuming your task is nontrivial and exceeds 100+ LoC. Code completion via continue.dev using local models did not yield good results, either.
No local models are good at agentic tasks (e.g., via AutoGen). You might have better luck rolling your own with function-calling via NexusRaven2 than using Agent frameworks with local LLMs.
Claude-Instant edges out gpt-3.5-turbo for longform summarization with its 100k context. It is what I use for https://HackYourNews.com because chunked/rolling summarization is noisy.
For every non-hobby task, I am switching to gpt-4-*, choosing between base, 32k, and turbo depending on speed/correctness/cost/length tradeoffs. They are not perfect, but there is no competition.
Yes it is usecase specific but thats also why its crucial for real world trial results to be shared so that theres a better understanding thats qualitatively accurate.
P.S: hackyournews.com is great!!
What I AM surprised about is that it is not clear what CultriX did that was better than what a ton of others have done.
Any clues?
This tweet is still very true: https://twitter.com/karpathy/status/1737544497016578453
https://en.wikipedia.org/wiki/Goodhart%27s_law
Nevertheless, scoring so well on this benchmark is an accomplishment, though I'm not in a position to evaluate how significant it is.
https://www.reddit.com/r/LocalLLaMA/comments/18xbevs/open_ll...
Many top models are overfitting to top leaderboards rather than be actually useful.
Under this scenario, the ones that achieve the top performance are the closest relation to the poison model?
an important point to keep in mind is that at inference, the lora adaptors are made to be merged into the base model so they don't affect inference speed. (you need to explicitly do it though, if you train your own adaptor)
That and they have great SEO, you basically can't avoid them.