Falcon 40B LLM (which beats Llama) now Apache 2.0
twitter.com
twitter.com
Inference is very slow right now but it works!
https://huggingface.co/tiiuae/falcon-40b
I'll also add that the fact something is in C++ doesn't mean it will run on arm or that it can be compiled in it.
The fundamental limit for hardware acceleration are number of gates you can squeeze on a die, right now. (Or, alternatively. memory bandwidth)
But if you're willing to spend $1500 on two used RTX 3090, it's the sweet spot in terms of the ability to run large models right now. Everything beyond that is much more expensive.
I thought that 4090s were "nerfed" and nvlink support removed - https://www.windowscentral.com/hardware/computers-desktops/n...
I feel like the best benchmark atm is the orig gpt-4 version.
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
https://www.morningstar.com/news/business-wire/2023052900504...
I've never seen this kind of "strategy" with an ML model before. Maybe I'm seeing something that isn't there...
It could be a question of not being used to see blatantly commercial advertising in places we're used to being about software. Feels like we're moving more towards the bro-ification of AI.
Seriously though, I think it's mostly the opposite. How much advertising did Georgi Gerganov do for ggml / llama.cpp and it's super popular. Maybe other people are just being more subtle, but I feel generally merit stands out on it's own and advertising is a poor substitute.
https://towardsdatascience.com/attention-is-all-you-need-dis...
What PR noise they make is secondary to that in my mind.
Cynics logic is always a race to the moral bottom
Another comment mentions llama may get an open license, and there are other emerging alternatives. In six months there will be lots of options. I would not spend my time building anything around a model that started in such a sketchy way.
It would be interesting to hear about why they decided to change their license and what their plans are for the future.
I admit Falcon has really pissed my off because they pretended they had an open source model when they really released a sleazy freemium thing.
Well now Falcon is open source. Given they are giving away something that was very expensive to train, I am grateful.
But all that's void now that they've gone Apache.
Will they keep maintaining their code and improving the models and continue releasing everything under Apache 2.0?
I'm just saying I feel like we got a glimpse of what they're about and it wasn't pretty, so why build around them when there are lots of options.
Look at Stable Diffusion 1.5 and LLaMA: They are thriving, but the original implementations are ancient history, and Meta/StabilityAI/RunawayML have done precisely nothing. And to be blunt, their legality is very ugly, which already makes Falcon more attractive.
Does "disparage" have a settled meaning in law? Is a parody disparaging?
https://github.com/CompVis/stable-diffusion/blob/main/LICENS...
I mean, that’s true of SD 1.5 in the sense that what the original creators have done since is new versions (SD 2.0, 2.1, and currently SDXL, which is apparent another SD2-architecture model, and DeepFloyd.) 2.1 has also seen some community uptake, and XL likely will once it is released unless there’s something inhibiting that. DF seems to be slowed by different architecture and high resource cost, but I’ve seen posts about people integrating the DeepFloyd early stage models with other models from the SD ecosystem for the last stage upscaling and final rendering, so I wouldn’t be surprised to see it integrated in some of the community UIs as both an integrated workflow and with access to the individual models for mix-and-match workflows.
Deepfloyd is niche.
SDXL is indeed interesting, especially if its happy with 4/8 bit quant... we will see about that.
Nevertheless StabilityAI seems kinda disconnected from all the innovations going on in the community compared to, say, huggingface.
I agree that it's a shame that StabilityAI seem to struggle so much to actually leverage their community (ideally with much more open development)... One could say they're a little too "full of themselves" and think they know better than everyone else.
Maybe I am cynical, but I dont see the incentive for Meta to contribute an open model.
There's no advantage here. Meta just spent $10 million on releasing fun chaos into the world and increasing their recruiting power.
ChatGPT was the fastest growing app in history, leaders at Meta (The ones who do M&A, strategy, etc) probably raised an eyebrow. They don't really give a crap about some stupid talking chatbot, but OpenAI getting smart and building a Social Network around millions of brand new users could be an existential problem for them. When Lecun wanted to OSS it they were probably like, sure, we can kill a few birds with one stone. If LLMs are a commodity that stops OpenAI and Google before they even get off the ground.
I believe it's more to make sure that others also continue to share their research.
Or also a general genuine good mindset of people involved in those groups.
They aren't looking to create a technological edge themselves, they want to remove the edge that OpenAI has so that they can win using their user count/brand recognition/etc.
From what I understand, even to run locally you/your team needs to be able to afford a machine with a 4090. These are super expensive in some countries.
I played around with the smaller Llama/Alpaca models and it wasn't really viable to build anything with.
Not really seeing a use-case for fine-tuning either compared to just few-shot prompting.
Can someone fill me in on what I'm missing? It feels like I'm out of the loop
So... not exactly a serious use-case. But it's what I'm using, and now I'm saving 10s of dollars on inferencing costs per month!
[0] https://github.com/go-skynet/LocalAI
I'm also using this to improve acceleration - https://cloudmarketplace.oracle.com/marketplace/en_US/adf.ta...
This use-case is alright for a toy I guess - which is the extent that I was originally expecting these things to be useful for.
> Are they any good?
Yep, free tier allows you to spec up to 24gb of RAM without paying, which is cool. The bottleneck is really the disk speed, but that's not an issue with mmaped models. There's enough cached memory that it loads instantly, so it's good-ish for this use case.
> Is the free-tier time limited?
No, but there are a lot of strings attached:
- The cores are vCPUs, not dedi (duh)
- You can't create new instances when demand is high (unless you add a credit card)
- Technically Oracle reserves the right to shut down the instance if demand gets really high (although I haven't heard any stories about this personally)
Proceed with caution. It's still a great place to start before you shell out $1/hr for dedi GPU rackspace.
1. Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks
https://huggingface.co/papers/2305.14201
2. Gorilla: Large Language Model Connected with Massive APIs
https://arxiv.org/abs/2305.15334
Consider also these 2 papers supporting the feasibility of fine-tuning:
3. LIMA: Less Is More for Alignment [showing that a very small number of high quality examples is sufficient to align a base model]
https://arxiv.org/abs/2305.11206
4. QLoRA: Efficient Finetuning of Quantized LLMs [showing that LLMs can now be fine-tuned quickly on consumer-grade GPUs]
https://arxiv.org/abs/2305.14314
—-
Adding up these developments (all of which occurred during the span of one week), I don’t see how huge, slow, general-purpose models maintain their relevance in the long term, when a lean, domain-focused model is right there within reach of every application developer.
If you want a specific kind of interaction with the model then you could take up 1/3rd of the 2048 token context window with few-shot or you could simply finetune it with QLoRA for a few hours on a consumer GPU and then get to use the full 2048 context with the finetuned model.
That's because we're only half a year into LLMs becoming mainstream. Give it 3-4 years. The advancements in bringing down model size, optimizations, and newer GPUs, SoCs from Nvidia, AMD, Apple, Intel, Qualcomm, etc will make it so that top LLMs will run on a highend laptop/desktop.
All advances in this direction do indicate that it will be easier and easier for more people to do things with it.
This doesn't need to work for everyone.
A 4090 costs today 2k, the 3090 with also 24gb costs today 1k and costed 2k.
GP Core count is much lower than than the 4090 but it still does 275 int8 TOPS for only $2k
I spent some time today exploring HuggingFace's Inference API but if the model is sufficiently large (> 10gb), HF requires you to use their commercial offerings.
Some of which are quite affordable ($80 per month). Larger ones can be like 2000 a month which is still ok to prototyping phase. You're basically paying for aws/gcp infrastructure.
I quite liked the UX of it, very intuitive. My trouble was finding a model that executes out-of-the-box tho. All of the GPT ones crash on startup.
Used DDR3 ECC would be roughly half that.
Falcon 40B is probably too much for it, but apparently there's similar cheap hardware that could work.
Someone else might be better able to confirm the pricing, but in any case you don't need to purchase the hardware.
I still use GPTQ for 30B, but even CPU generates quickly enough at q5_1 on modern hardware.
Here somebody quantized it down to 29929.56MB .
Sure, it might be a lot slower, but that's a lot better than "I give up, go buy $20K worth of hardware"
For example Guanaco-33B generates ~10 token per second running fully from VRAM of my 3090, the ~1 token/second running from DDR4 RAM of my Ryzen. I would imagine it would do like a token per minute from NVM SSD.
I had previously assumed it was safety concerns, since I don't see what stops someone from finetuning away all guard rails.
There no reason, particularly, to believe either that it does, or it would.
For openai to scramble and try to “catch up” with a competitor and make such a massive change in strategy would require someone to be offering an equivalent service (hosted inference) that was either orders of magnitude cheaper than their offering and just as good, or significantly better than it. Or legal compulsion.
This is none of those things. They won’t care.
E.g. 10 layers x 2048 tokens x 1024 embedding model using full attention.
Falcon seemed good till I read the license fine print about pre approvals and what not. This seems to fix that
There will be an ASIC as soon as serious money is being made from LLMs, most use cases atm seem to be in prototype/toy stage, but I imagine we'll start seeing that change.