DeciLM-7B: The Fastest and Most Accurate 7B-Parameter LLM to Date
deci.ai
deci.ai
"groundbreaking", "outshines its competitors", "remarkable", "pivotal transformation"
Totally unnecessary bravado.
“As a large language model, I can explain. I and my relative were trained at the bestest universities, only in superlatives about humans who are optimum et vetustissimum octagenarium, i.e. the “bestest of all time! Could win nathan’s hot dog-eating contest while making the Jersey Turnpike Marathon unfair and telling you about how it used to cost a nickel at the ferry. Everyone else stinks stonks!””
I'll never respect marketing like this because it's standing on the shoulders of thousands of others peoples work.
Everyone is currently talking about it because they got a massive investment and people started posting links. However, no one was talking about them 2 weeks ago.
There's plainly 7B models that surpass it on https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
Am I missing something?
I keep seeing them pass each other on different benchmarks and leaderboards and I can't help but imagine that so many of them are only good at benchmarks and not much else.
I haven't gotten a chance to play with this stuff yet so I have no basis to go off of, but I'd like to hear from anyone who's actually been impressed by any of these small models for general purpose tasks.
But they're very nice for making PoCs on complex systems since they're near free to run.
The finetunes of 0.1 are already extremely impressive at general tasks.
https://huggingface.co/TheBloke/Mistral-7B-Instruct-v0.2-GGU...
GPT-4 turbo has a 128k token context, which might be good enough for your book.
You might also use another strategy: write a summary of what your book is about and the title of each chapter, then pass the summary plus the chapter title to the LLM and it will generate for you. This would allow you to go beyond the 100k word limit.
I think it's a new foundation model, but the announcement doesn't make that clear to me.
"DeciLM-7B’s superior performance is rooted in its strategic implementation of variable Grouped Query Attention (GQA), a significant enhancement over traditional Multi-Query Attention (MQA) and standard GQA."
The literal opposite of Mistral's "no marketing" just a torrent with the data.
Off course it may be actual new work but colour me very sceptical with this extreme confidence and marketingspeak.
> The banana is still in the kitchen, as it was placed on the plate before it was moved to the living room.
He's a little confused, but he's got the spirit.
- AI regulation is partly based on the number of parameters
- It's generally accepted that bigger models perform better, so counterexamples to that notion are valuable
- Smaller models are needed for local, offline processing with current hardware
For me I care about different companies reaching mistral level so that mistral or whoever is in top has to release the model weights else competitors will.
This is the worst form of defence.
I still use Mistral as was clear from the last post. This model is not that good to make me switch. As I said I just want Mistral or the top model to remain open weight and so I want competition.