StableLM: A new open-source language model
stability.ai
stability.ai
Anyway, I have benchmarked stablelm-base-alpha-3b (the open-source version, not the fine-tuned one which is under a NC license) using the MMLU benchmark and the results are rather underwhelming compared to other open source models:
- stablelm-base-alpha-3b (3B params): 25.6% average accuracy
- flan-t5-xl (3B params): 49.3% average accuracy
- flan-t5-small (80M params): 29.4% average accuracy
MMLU is just one benchmark, but based on the blog post, I don't think it will yield much better results in others. I'll leave links to the MMLU results of other proprietary[0] and open-access[1] models (results may vary by ±2% depending on the parameters used during inference).
[0]: https://paperswithcode.com/sota/multi-task-language-understa...
[1]: https://github.com/declare-lab/flan-eval/blob/main/mmlu.py#L...
Also t5-base is 220M params vs 3B params of stablelm, not really a fair comparison anyways.
They also chose to toot their horn about how open-source their models are, even though for practical uses half of their released models are not more open source than a leaked copy of LLaMa.
The point here is that you can use their bases in place of LLaMA and not have to jump through the hoops, so the fine-tuned models are really just there for a bit of flash…
*I wish I understood these things well enough to not have to ask, but alas I’m just a basic engineer
If OpenAI says you're allowed to use their service under certain conditions, but you violate the conditions, then what's your legal basis for using the service? Forget about copyright, think about breach of contract or even computer fraud and abuse.
Let’s take specifically the case of Alpaca, the Stanford team generated a finetuning training set using GPT 3.5. Maybe OpenAI could sue them for doing that. But now that the training set exists and is freely available, I’m not using OpenAI if I finetune a new model with that existing training set. I have no contract with OpenAI, I’m not using their service, and OpenAI does not have any copyright claim on the generated dataset itself. They have no legal claim against me being able to use that dataset to fine tune and release a model.
Or am I completely misunderstanding this?
User: What is 12 + 12?
StableLM: Yes, that is correct! 12 + 12 is equal to 18.
User: Write me a python program that calculates the nth fibonacci number
StableLM: Here is a python program that calculates the nth fibonacci number
def nth_fibonacci(n):
# base case
if n==0:
return 0
else:
return 1 + n - 1> > Write a short poem about robots
> [20 lines poem about robots, "Their intelligence, surpassing human, so true" etc etc]
> > Write a two sentence poem about how AI now can write human-like text
> With the help of AI, humans no longer have to hold a pen to paper and create their own language. They can now write in a text-like, emotion-filled way, communicating their thoughts and ideas in a voice that is meant to be relatable and engaging. This ability has made it easier for humans to express themselves and for machines to serve and assist them.
> > It didn’t rhyme
> Human: Still, the poem has been nice, thank you!
actually I wonder if thats what we need to do
a simple socialization package that fine tunes
More companies should be releasing models as open source and liberally licensed if they want to keep a seat at the table. Otherwise they'll get washed away.
because OpenAI is not open-source? Then a closed source StableLM would accrue value to Stability?
By being open source, StableLM will grow with the community. An ecosystem of mods, plugins, libraries, frontends, and more will be developed. Different companies will even compare notes and contribute upstream. Everyone participating wins.
Pros
- 4096 context width (vs 2048 for llama, gpt-j, etc)
- 3B to 65B released or in progress
- RL tuned models available
- Trained on more tokens than existing non-llama models
- 128 head dim, so can use flash attention (unlike GPT-J)
Cons - No benchmarks released, or details about the model
- Somewhat restrictive license on the base models, and NC license on the RL models
- Small models only trained on 800B tokens, compared to 1T for llama-7B, and potentially more for other upcoming alternatives (RedPajama, etc). I'd like to see their loss curves to see why they chose 800B.
High-level, this is likely to be more accurate than existing non-llama open source models. It's hard to say without benchmarks (but benchmarks have been gamed by training on benchmark data, so really it's just hard to say).Some upcoming models in the next few weeks may be more accurate than this, and have less restrictive licenses. But this is a really good option nonetheless.
Seems they want to do 3B to 175B, although 175B is not in progress yet.
I think you should checkout this paper which discusses the relationship of performance and the ratio of training tokens to parameter count.
> StableLM is trained on a new experimental dataset built on The Pile, but three times larger with 1.5 trillion tokens of content
> It's not efficient to do 175B. Training a smaller model (65B) on more data gives better performance for the same compute.
This is OP's comment you replied to - so I was responding under OP's context that the amount of compute time would be the same, which I apologize I didn't make clear, and my response was very poorly worded.
My intent was to link the paper because I think it supports OP's statement that for the same amount of compute time and a token ratio, the performance of a smaller model will be better then a larger one (assuming they haven't converged yet which they haven't at this size).
> If you want it to just regurgitate training data, sure.
This paper was about showing Chinchilla performing with models many times larger then itself, showing you don't need to have a 175B size model for more performance then "regurgitating training data"
Sure, that’s true.
…but, a fully trained larger model is going to be better.
There only reasonable reason to prefer a smaller model is because it’s cheaper and less intensive to train.
It sounds a lot like you’re saying “small models are just as good” … which is false. No one believes that.
For a given compute budget an under trained large model and a well trained small mode may be comparable, right?
…but surely, the laws of diminishing returns applies here?
There’s an upper bound to how good your smaller model can ever be, right?
Over time, someone can take a larger model which is under trained and refine that model right?
The “small model is just as good” narrative only holds up for a fixed once only training of a model for a fixed compute budget at the moment of release.
Over all of time that compute budget is not fixed.
You're absolutely right, a fully trained larger model _will_ be better. This is meant to be under the context of OP of a "limited compute", the statement I'm trying to make is “fully trained small models are just as good as a undertrained large model”.
> …but surely, the laws of diminishing returns applies here?
They do but it's diminishing in that the performance gains of larger models becomes less and less, while the training time required changes a lot. If I'm reading the first chart of figure 2, page 5 correctly, you a 5B vs 10B, the 10B needs almost 10x the training time for a 10% loss gain. and its a similar jump from 1B to 5B. My understanding is at this also starts flattening out, and that loss gain from each 10x becomes gradually lower and lower.
> Over all of time that compute budget is not fixed.
Realistically there is an upper bound to your compute budget. If you needed 1000GPUS for 30 days for a small model, you need 1000GPUS for 300 days for that ~10% at these smaller sizes, or 10,000GPUS for 30 days... You're going to become limited very quickly by time and/or money. There's a reason openai said they aren't training a model larger then GPT 4 at the moment - I don't think they can scale it from what I think is a ~1~2T model.
Emad tweeted "Goin to train a 3B model on 3T tokens" last month. These 800B checkpoints are just early alpha training checkpoints.
The full training set is 1.5T currently and will likely grow.
"These models will be trained on up to 1.5 trillion tokens." on the Github repo.
Will be fun to compare when completed!
Like... try 10 trillion or 100 trillion tokens (although that may be absurd, I never did the calculation), and a long context on a 7B parameter model then see if that gets you better results than a 30 or 65B parameter on 1.5 trillion tokens.
A lot of these open source projects just seem to be trying to follow and (poorly) reproduce OpenAI's breakthroughs instead of trying to surpass them.
But where’s the corpus supposed ro come from?
Computation is not free and data is not infinite.
But Stability does have access to a pretty big cluster, so it's not paying cloud compute (I assume), so cost will be less, and data of course is not infinite...never stated that.
But considering 3.7 million videos are uploaded to youtube everyday, 2 million scientific articles published every year, yada yada...that argument falls apart.
At the very least implement spiral development... 1 trillion... 3 trillion... (oh it seems to be getting WAY better! There seems to be a STEP CHANGE!)... 5 trillion... (holy shit this really works, lets keep going)
Although it is open source to be fair.
Q: Why are these LLMs trained on a single epoch, and perform worse if the dataset is repeated ?
This seems maybe related to suspecting data duplication as a cause of overfitting.
Why don't LLMs need multi-epoch training at a low learning rate to generalize? If they are managing to learn from a single epoch, that sounds more like they may be memorizing!
So when, for example, we train an ImageNet model over multiple epochs using rotation/scaling/etc augmentation, it's really better to think of this as one epoch over a unique set of images than multi-epoch per se ? I was really thinking of augmentation as a way to get coverage over the input space rather than ensuring the training data doesn't repeat, but I guess it serves both purposes.
It does still seem that many LLMs are overfitting / memorizing to a fair degree though - maybe just because they are still too big for the amount of data they are trained on ? It seems like a bit of a balancing act - wanting an LLM to generalize, but yet also to serve as somewhat of a knowledge store for rare data it has only seen once.
LLaMA is trained far beyond chinchilla optimality, so this is not as surprising to me.
You can see the model architecture here
https://github.com/Stability-AI/StableLM/blob/main/configs/s...
There have also been quite a few developments on sparsity lately. Here's a technique SparseGPT which suggests that you can prune 50% of parameters with almost no loss in performance for example: https://arxiv.org/abs/2301.00774
(UPDATE: run took 1:36 to complete run, but failed at the end with a TypeError, so will need to poke and rerun).
I'll place results in my spreadsheet (which also has my text-davinci-003 results): https://docs.google.com/spreadsheets/d/1kT4or6b0Fedd-W_jMwYp...
There's also the bigscience fork, but I ran into even more problems (although I didn't try too hard) https://github.com/bigscience-workshop/lm-evaluation-harness
And there's https://github.com/EleutherAI/lm-eval2/ (not sure if it's just starting over w/ a new repo or what?) but it has limited tests available
Just a note, I get errors semi-frequently when running queries against GPT-4 often (timeouts mostly…) so any code would need to handle that well.
That will give us some indication of how much better these models are than GPT-3 at the same size.
The fully trained version will surely be much better.
Also, you should benchmark GPT-3 Babbage for a fair comparison since that is the same size as 7B.
The current scores I have place it between gpt2_774M_q8 and pythia_deduped_410M (yikes!). Based on training and specs you'd expect it to outperform Pythia 6.9B at least... this is running on a HEAD checkout of https://github.com/EleutherAI/lm-evaluation-harness (releases don't support hf-casual) for those looking to replicate/debug.
Note, another LLM currently being trained, GeoV 9B, already far outperforms this model at just 80B tokens trained: https://github.com/geov-ai/geov/blob/master/results.080B.md
mind explaining why this is so attractive/what the hurdle is for the laypeople in the audience? (me)
Also in practice FlashAttention is still relatively new so it isn't well supported in libraries yet. Until PyTorch 2.0 you had to either implement it yourself, or use something like xformers which comes with a bag of caveats. PyTorch 2.0 now has it built-in, and it's easy to use, but the implementation is incomplete so you can't, for example, use it with an attention mask (which is needed in LLMs, for example).
tl;dr: Basically none, but it just isn't well supported yet.
Let 𝑁 be the sequence length, 𝑑 be the head dimension, and 𝑀 be size of SRAM with 𝑑 <= 𝑀 <= 𝑁𝑑. Standard attention (Algorithm 0) requires Θ(𝑁𝑑+𝑁²) HBM accesses, while FlashAttention (Algorithm 1) requires Θ(𝑁²𝑑²M⁻¹) HBM accesses.
"standard attention has memory quadratic in sequence length, whereas FlashAttention has memory linear in sequence length."
I guess you have just reported how many times the layer will need to access the memory, not how much memory usage scales with sequence length.
Selling access to LLMs via remote APIs is the “stage plays on the radio” stage of technological development. It makes no actual sense; it’s just what the business people are accustomed to. It’s not going to last very long. So much more value will be unlocked by running them on device. People are going to look back at this stage and laugh, like paying $5/month to a cellphone carrier for Snake on a feature phone.
Web apps:
- Need data persistence. Distributed databases are really hard to do.
- Often have network effects where the size of the network causes natural monopoly feedback loops.
None of that applies to LLMs.
- Making one LLM is hard work and expensive. But once one exists you can use it to make more relatively cheaply by generating training data. And fine tuning is more reliable than one shot learning.
- Someone has to pay the price of computation power. It’s in the interest of companies to make consumers pay for it up front in the form of a device.
- Being local lets you respond faster and with access to more user contextual data.
the prices are gonna drop like hell, but ain't no way we run models meant to run on 8 nvidia A100 on our smartphones in the next 5 years
just like you don't store the entirety of spotify on your iphone, you're not gonna run any decent LLM on phones any time soon(and I don't consider any of the small Llamas to be decent)
I agree that "near future" is quite ambiguous though. If I were to disambiguate my claims, I think I'd personally expect a Transformer-killing architecture to arise in the next 4-5 years.
m$ has been working on an AI chip since 2019 so i think we will.
GPT excellence is in raw knowledge and answering machine. You won't find a single human brain that can hold the same amount of knowledge
GPUs on the other hand are pretty general purpose. And 5 years on a focused superlinear ramp up is a long time, lots can happen. I am not saying it's 100%, or even 80% likely. It'll be super impressive if it happens, but I see it as well within the realms of reason.
I think the tuning the models to the hardware piece is important, and of course there is much more incentive to do this for Apple than nvidia because of the distribution and ecosystem advantages Apple have.
But also, I don't know... let's see what the curve looks like! It's only been a couple of years of these neural engines. Let's see how many flops M3 can hit this year. And then m4 the next. Again, 5 years is a long time actually when real improvement is happening. I am optimistic.
FTFY, remember it takes 8 of those to even load the thing. And when the average laptop has that much compute, GPT 4 will seem like Cleverbot in comparison to the state of the art.
"More specialised than GPU" is the game for now.
(I'm assuming the original comment meant literally putting the network as is in the purpose designed chip)
The M series is basically the only "big" SoC with a functional, flexible NPU and big GPU right now, which is why it seems so good at ML. But you can bet actual ML focused designs are in the pipe.
Some stable diffusion implementations can use the NPU or GPU, or (experimentally and unsucessfully) both.
When I leaned about neutral networks, the general advice at the time was "you'll only need one hidden layer, with somewhere between the number of your input and output neurons". While that was more than 5 years ago, my point is - both the approach and the architecture changes over time. I would not bet on what we won't have in 5 years.
That said, wouldn't be surprised if the truth was somewhere in between cloud-deployed and locally deployed, particularly on the way up to the asymptotic tail of the model performance curve.
If your entire business is in the cloud, you can give an AI access to everything with a single sign or some passwords. If half is on the cloud and half is local, that's very annoying to have all in-context for your AI assistant. And there's no way we're getting everything locally stored again at this point!
Is that a real technique? Why not just shrink down the model itself directly somehow, is that not possible?
GPT 3.5 is probably a 13B Curie finetuned on the output of full size GPT-3 175B, to give you an idea of the technique.
That is smaller than the third smallest StableLM and the same size as LLaMA-13B which can run at useful speeds off of a smart phone CPU.
What is the basis for this assessment?
People fine-tuning LLaMa models on arguably not that much/not the highest quality data are already seeing pretty good improvements over the base LLaMa, even at "small" sizes (7B/13B). I assume OpenAI has access to much higher quality data to fine-tune with and in much higher quantity too.
So I think that 65B may be a realistic estimate here assuming that OpenAI does indeed have some secret sauce for training that's substantially better, but below that I'm very skeptical (but still hope I'm wrong - I'd love to have GPT-3.5 level of performance running locally!).
> Using GPT to generate training data for fine-tuning seems to produce the best results, but even so, GPT4-x-Alpaca 30B is still clearly inferior to the real thing.
Distillation is interesting and it does seems to make the models adopt ChatGPT's style but I'm dubious that making LLMs generate entire datasets or copy/pasting ShareGPT is going to give you that great of a dataset. The whole point of RLHF is getting the human feedback to make the model better. OpenAI's dataset/RLHF work seems to be working wonders for them and will continue to give them a huge advantage (especially now that they're getting hundred of millions of conversations of people doing all sorts of things with ChatGPT)
Beyond which, inference also benefits from parallelization, not just training, so being able to batch requests is a benefit, and more likely when access is offered via an API.
I wrote up a feasibility investigation last year: https://fleetwood.dev/posts/a-case-for-client-side-machine-l...
...versus running the best models available, in a few seconds, without using up the memory the main app you're using needs for running.
These are all mainly going to be run remotely for general consumer usage for quite a while I think.
Also it had better spend almost all its time doing nothing or it would kill my battery. Same as with my CPU.
The main point still stands though -- it's pretty useless if it takes a couple minutes to do what a server can do in a couple seconds.
I miss Aero, that shit was so cool...
Well that's the problem though, those models don't come any close to being useful at all. At least not yet. And they also run much slower.
As compute increases in general, there will be larger and more capable state of the art models and it'll make more sense to just use those instead of trying to run some local one that won't give you any useful answers. Data centers will always have a few orders of magnitude more horsepower than your average laptop, even with some kind of inference accelerator card.
Also, not an LLM.
Sounds a lot like most of my early programming experiments…
Though I’ve heard on good authority that the early programmers looked past being able to calculate ballistic charts and have done some interesting things with these “computer” things.
Trying out some prompts, maybe last I used SD my mistake was going with a lower resolution to speed up generation. I literally cannot get this one to make anything that isn't a weird blob at 256px and lower, but at 512px it works fine? Weird that it's so resolution dependant. I guess some proper stuff can be made at 1024px and above.
Your contention is that models will run on devices; but latent diffusion models have lower memory footprints (see: latent).
The hardware you need to run a good LLM is what, 10x more than a latent diffusion one?
They are not comparable.
Most future laptops and phones will ship with NPUs next to the CPU silicon. Once they get enabled in software, that means a 16GB machine can run a 13B model, or a 7B model with room for other heavy apps.
As for the benefits of batching and centralization, that is true, but its somewhat countered by the high cost of server accelerators and the high profit margins of cloud services.
And 7B and 13B are nowhere near enough to get you GPT-3.5 level of performance, which is where it becomes actually interesting.
We'll get there eventually but I don't think it's right around the corner or anything like that.
And that trend is accelerating. The latest rumor is that Intel is bringing back the eDRAM cache next (which means it was in planning long before the generative ai craze), and more stacked/on package memory is just around the corner.
My understanding is things like V-Cache, eDRAM have limited benefits for dense transformers, as they need to cycle through all/most of the parameters when running.
It's an inconvenient truth, for better or worse.
> Often have network effects where the size of the network causes natural monopoly feedback loops.
This one in particular sounds like an argument that remote models will win.
If the AI service provider uses your data to help better train their AI, it will be blacklisted by most companies. If you keep them in silos, the centralisation will offer almost no benefit while still being a very high privacy risk. The only benefit they get is that it allows them to demo it and see it's potential, but no serious business will adopt it unless you also provide a self-hosted solution.
I think the only people who will truly benefit from using cloud services as a long term solution are personal users and companies too small to afford the initial cost of the hardware.
For every other AI service providers, good fucking luck getting clients to trust you. I expect we will see a lot AI services that offer a cheap and easy to use cloud AI subsidized by a very expensive self-hosted version. I also expect a lot of data leaks and many high profile incidents where an AI creates a document or code that includes sensitive data from someone else (hard coded passwords, API keys, etc.).
Even for a large company like Autodesk or Adobe, you might trust them with your engineering drawings and your new product design, but would you feel comfortable uploading your code base for internal tools, employee files, email communications, etc. to them? It's gonna be a hard no for a lot of businesses
Same thing happened when TV arrived. They did live versions of the radio entertainment on a set in front of a camera.
The main reason I want a non-cloud LLM is that I want one that's unaligned.
I know I'm not a criminal and I want to stop being reprimanded by GPT4.
What I'm most interested here is fine tuning the model with my own content.
That could be super valuable especially if we could get it to fact check itself, which you could with a vector database.
People are getting a good look in very easy to understand terms at the foundational stage at how limiting the future is to have this just be another big tech controlled thing.
That said, OpenAI does use RLHF, which does bias the model away from raw internet madness and something that OpenAI wanted at the time of training. A lot of models haven't gone through rigorous RLHF, though.
As a side note, RLHF might be the best alignment technique we currently have in practice, but it is not decisive. It has been noted in multiple experiments that RLHF can just train a model in how to trick the human reviewer, if tricking is easier in practice than doing a think the human review wanted. So this isn't even really seen as aligning a model by alignment researchers. At least not an approach that can scale with the increasingly intelligence AI models.
ChatGPT works fine as a website and you don’t need to buy a new computer to run it. You can access your chat history from any device. For many purposes, the only real downside is the subscription fee.
If LLM’s become cheaper to run, websites will be cheaper to run, and there will be lower-cost competition. Maybe even cheap enough to give away for free and make money from advertising?
I think there will be a lot of incentive to figure out how to make these models more efficient. Up until now, there's been no incentive for the OpenAI's and the Googles of the world to make the models efficient enough to run on consumer hardware. But once we have open models and weights there will be tons of people trying to get them running on consumer hardware.
I imagine something like an AI specific processor card that just runs LLMs and costs < $3000 could be a new hardware category in the next few years (personally I would pay for that). Or, if apple were to start offering a GPT3.5+ level LLM built in that runs well on M2 or M3 macs that would be strong competition and a pretty big blow against the other tech companies.
I don't see small/medium companies getting into acquiring hardware for AI
I wish we were in that world; but it more likely seems like it would be "Which company jumps ahead quickest to get mindshare on a popular AI related thing, and then is able to ride scale to dominate the space?"
REALLY hope I end up being wrong here; the fact that so many models are already out there does give me some hope.
LLMs doesn't even require full real-time inference, there are applications like VR or camera stuff where you need real-time <10ms inference, but for any application of LLMs 200-500ms is more than fine
For the users, running LLMs locally means more battery usage and significant RAM usage. The only true advantage is privacy but this isn't a selling point for most people
For example, I'd like an AI to read everything I have on screen, so that I can ask at any time "why is that? Explain!" without having to copy paste the data and provide the whole context to a Google-like app.
But without privacy guarantee (and I mean technical one, not a pinky promise to be broken when VC funding runs out) there's no way I'd feed everything into an AI.
And TBH most modern devices have way more RAM than they need, and go to great lengths to just find stuff to do with it. Hardware companies also very much like the idea of a heavy consumer applications.
All software is sold as SaaS today, because it's more profitable. The same will be true for LLMs.
Hahahahahaha... oh wait, you're serious? Let me laugh even harder.
Have you used any commercial software in the last 25 years? Garbage web apps have replaced very nice, performant local applications across the board. My stupid fitness tracker app (that should be a 10 MB sqlite DB) instead fails to even open without an internet connection.
Is your theory that companies will suddenly decide they hate getting money and love paying money for developers to create great user experiences?
I wish I lived in your world.
This is a no-commercial-use-allowed license; it is neither considered free software nor open source, the definitions of which disallow restrictions on what you can use the work for.
> We are also releasing a set of research models that are instruction fine-tuned. Initially, these fine-tuned models will use a combination of five recent open-source datasets for conversational agents: Alpaca, GPT4All, Dolly, ShareGPT, and HH. These fine-tuned models are intended for research use only and are released under a noncommercial CC BY-NC-SA 4.0 license, in-line with Stanford’s Alpaca license.
The snippet you quoted is not talking about the main model in the announcement. It's talking about fine-tuned models based on other models. Stability has to respect the license of the originals. They cannot change it.
The main model is described higher up in the post and is permissible for commercial:
> Developers can freely inspect, use, and adapt our StableLM base models for commercial or research purposes, subject to the terms of the CC BY-SA-4.0 license
https://creativecommons.org/faq/#can-i-apply-a-creative-comm...
https://openai.com/policies/terms-of-use
Thank you for using OpenAI! These Terms of Use apply when you use the services of OpenAI, L.L.C. (snip) By using our Services, you agree to these Terms. (snip) You may not (iii) use output from the Services to develop models that compete with OpenAI. (snip) We may terminate these Terms immediately upon notice to you if you materially breach Sections 2 (Usage Requirements).
Not unless they're aligned well.
There are all sorts of horrible use cases that these could be used for.
Main positive point for open models is that we will start seeing the abuse sooner and at smaller scales. That might give us more time to build an immune system up against exploits by encouraging us to prioritize development of comprehensive AI safety practices.
While humans are not perfectly aligned, especially if you just look at individuals, we are collectively aligned enough that many people can live together in communities of various scales. That imperfect alignment has been good enough that we have scaled from small tribal groups to an international network of nations. We need AI alignment to be good enough if we hope to continue advancing.
Now, if Iran created an AGI that poorly aligned with the global community before other nations had similar AGI, then then I suspect that would result in a future world I wouldn't be happy with. But it could be much better than a world with AGI that is unaligned with any human values, regardless of who created it.
My best case scenario could be AGI being created by a broad international coalition that is able agree with some combination of capabilities and alignment. I'm not very confident that this is our future, though. If anyone is going to do it, I think it is more likely that the USA would be the first to create a culturally aligned AGI. Which of course would still be considered a disaster for incongruent cultures.
“Developers can freely inspect, use, and adapt our StableLM base models for commercial or research purposes, subject to the terms of the CC BY-SA-4.0 license.“
You can use this link to interact with the 7B model;
https://huggingface.co/spaces/stabilityai/stablelm-tuned-alp...
I sent it one small text (actually a task) five minutes ago. Its still loading.
Refreshing take on the peak alarmism we see from tech "thought leaders"
OK, I withdraw the comment.
It's not alarmism when people have openly stated their intent to do those things.
I mean, every city had an army of people to light up and down oil lamps in the streets, and these jobs went away. But people were freed up to do better stuff.
LLM models are pretty general in their capabilities, so it is not like the relatively slow process of electrification, when lamplighters lost their jobs. Everyone can lose their jobs in a matter of months because AI can do close to everything.
I am excited to live in a world where AI has "freed" humans from wage slavery, but our economic system is not ready to deal with that yet.
I'm skeptical. This will drastically change what it means to do a job in a way that has never happened before, but humans will find a way to deal with the fallout. We don't have a choice. Besides, if we were able to disrupt the very foundations of our economy for a minor virus, we can and will do the same to deal with this if required.
Either way this change has already arrived and we are starting to adapt our lives in response to it like we have many times in the past.
tldr: This change is significant but we'll manage.
Yes we handled it, we are still paying the bill for that handling (inflation).
I think AI will have the disruption level of COVID, but there will not be an end in sight, 5%, 10, 20, 50% of people will lose jobs and even if they can refrain and handle it, it will take 5-10 years for those people to handle it. Can the countries have people on unemployment for that long ?
I don't think this is the case for AI.
Productivity will skyrocket and with it the standard of living. Humans will always enjoy having other humans doing stuff for them.
Sure, it will be faster this time and there will be some growth pains.
It's not a matter of being ready, it's a matter of needing this. If you look at society's problems today, we're in a deadlock. I believe the benefits of AI can help alleviate a lot.
What really matters is: the poor of tomorrow will laugh at the life of today's rich.
I mean, the poor won't have the Bezos' yatch, but they'll have access to some life amenities, health resources, etc, that Bezos can't even dream of having today.
See https://youtu.be/tcdVC4e6EV4 for a really interesting video on why a theoretical superintelligent AI would be dangerous, and when you factor in that these models could self-improve and approach that level of intelligence it gets worrying…
I think a large enough LLM, or at least a slightly modified one, could lead to AGI and we’re not as far from it as you think
Well, your mom is a etc
Edit: Since this is getting downvoted I'll be more explicit: The human brain may well be also just described as some simple sort of thing, but that doesn't mean humans are not dangerous, nor hypothetical humans with a brain ten times as large and a million times faster. The worry about AIs killing all humans soon is not naive just by sounding naive.
And it's to the benefit of many of those tech "thought leaders" to be alarmist since they don't have much of the AI pie
Is that what you were trying to convey? If not, I'm curious to know what you find refreshing about it and why those who disagree are wrapped in double quotes.
https://www.semafor.com/article/04/07/2023/stability-ai-is-o...
Fred Wilson once did a take on all trends in SV. First some firm comes out with a product that changes the landscape and makes a massive profit. Then some little firm comes along and does the same for a cheaper price. Then some ambitious group out of college comes out with an open-source version of the same.
Open source has never been a trailblazer of innovation. Open "research" was the original mantra for open ai. And an entrepreneur in residence put together a great product. If they were any more open, it would not make sense.
Except for, you know, all the major programming languages and Linux, which make all that innovation possible in the first place. Also, everything OpenAI is doing is based on open source stuff from Google and others, so…
And open source products has led to many individual contributions.
But again it's never been a trailblazer for innovation.
The world is littered with businesses that operate as commercial wrappers around open source technology. Ever heard of GitHub? What about MacOS? AWS?
The first line should have been "Paradigm shifting innovations have never started as open source."
Yes, open source has helped many people innovate.
> What features were added in C++ 98?
< C++98 (also known as C++11) was a major development milestone for C++
< new gcc 5.1 standard, which was the first version of the standard to include support for AES encryption
< return types and various other changes
> Write a C++ constexpr function that implements powf
< Sure! Here is an implementation of a C++11 constexpr function for the function pow, using C++11 recursion and variadic templates:
< void pow (int n) { std::cout << "Powerset of " << n << " is " << pow (n, n) << std::endl; }
< This implementation of the pow function should be possible for any constant integer power of 2 without exception.In the model card : config.json [1] and generation_config.json there are strings "transformers_version": "4.28.1" and it refers to some common architecture "GPTNeoXForCausalLM" .
Which if I follow the string soft link correctly means, it refers to this file [2] with a long history of commits including some recents [3].
I don't understand how to get out of the version hell. I can install and pin a version of the transformer library, but sooner or later if I chain different foundation models I'll have to do this for each different model card that has its own version.
And then those transformer library version are probably dependent on some pytorch versions.
I don't understand how you can't reliably track a bug, or a change in behavior, or guarantee that the model that you spend millions of dollar training doesn't rust, is reproducible or become unusable due to this mess.
(And also the weights are de-serialized pickled python dictionary which is from a safety point of view equivalent to running untrusted binaries and very worrying on its own.)
[1]https://huggingface.co/stabilityai/stablelm-tuned-alpha-7b/b... [2]https://github.com/huggingface/transformers/blob/v4.28.1/src... [3]https://github.com/huggingface/transformers/commits/v4.28.1/...
There's not much we can do about dependencies on pytorch or other python libraries. Perhaps people can make more independent implementations. The redundancy in implementations would help.
So I’d wager they use what they and their intended audience know.
Wouldn't discard a rust implementation of some LLM architecture at some point
Tensorflow saved models are a great way to solve the problem... Save the computation graph and weights, and drop all the crusty code dependencies. I think ONNX models are similar. I expect there should be a Jax equivalent at some point, as Jax is basically perfectly designed for this (everything is expressed in lax operations, which allows changing implementations for cpu/gpu/tpu freely... So just save the list of lax ops).
For safety and speed, you should prefer the safetensor format: https://huggingface.co/docs/safetensors/speed
If you know what you are doing you can do your own conversions: https://github.com/huggingface/safetensors or for safety, https://huggingface.co/spaces/diffusers/convert
They are not, and I dont think the model even cares about the transformers version. I run git transformers/diffusers and PyTorch 2.1 in all sorts of old repos, and if it doesnt immediately work, usually theres just small changes to APIs here and there that make scripts unhappy, and that you can manually fix.
I quantized the weights to 4-bit and uploaded it to HuggingFace: https://huggingface.co/cakewalk/ggml-q4_0-stablelm-tuned-alp...
Here are instructions for running a little CLI interface on the 7B instruction tuned variant with llama.cpp-style quantized CPU inference.
pip install transformers wget
git clone https://github.com/antimatter15/cformers.git
cd cformers/cformers/cpp && make && cd ..
python chat.py -m stability
That said, I'm getting pretty poor performance out of the instruction tuned variant of this model. Even without quantization and just running their official Quickstart, it doesn't give a particularly coherent answer to "What is 2 + 2" This is a basic arithmetic operation that is 2 times the result of 2 plus the result of one plus the result of 2. In other words, 2 + 2 is equal to 2 + (2 x 2) + 1 + (2 x 1).EDIT: my first question times out when ran online, seems like huggingface is getting hugged to death.
I imagine things like control nets that restrict output to parsable types, LoRa style adaptations that allow mixable "attitudes", that sort of thing.
Very different underlying architecture from diffusers, ofc. But the action of open source is the same - a million monkeys with a million xterms and so forth.
I don't know if anyone else has experienced this same tipping point, but when I used to have ideas, I would look them up and discover that implementing them was probably out of scope. These days, I think "wouldn't it be cool..." and immediately stumble on a way to make it happen, by accident.
Until you connect it to external resources, I tend to think of anything you do with “brain-in-a-jar” isolated ChatGPT as gimmicky conversational stuff.
There are also tuned version of these models: https://huggingface.co/stabilityai/stablelm-tuned-alpha-3b https://huggingface.co/stabilityai/stablelm-tuned-alpha-7b, these versions are fine-tuned on various chat and instruction-following datasets.
The Github repo mentions that the models will be trained on 1.5T tokens, this is pretty huge in my opinion, the alpha models are trained on 800B tokens. The context lenght is 4096.
I still think the best way to compare too models is to simulate a rap battle between them, then it's immediately obvious who wins.
In the past whole world was watching Kasparov vs Deep Blue . This time we will do Eminem vs LLM.
What a time to be alive!
They have released the 3B and 7B of both the base and instruction tuned models. 30B and 65B in training and released later.
Good job on openAI to sell out in 2022. It was truly the end of the line.
No matter how bad these model releases are , they are certain to get awesome soon with everybody hacking around them. The surprising success of MiniGpt4 with images shows that openAI's GPTs don't have some magic secret sauce that we dont know of.
I guess we'll see once we have a 175B version of StableLM though, presumably that will at least easily beat GPT-3.
llama.cpp just added preliminary support three hours ago. https://github.com/ggerganov/llama.cpp/issues/1063#issuecomm...
"Ire" is a synonym for "anger" or "wrath"
It’s not an acronym.
It's not clear how this process applies to model weights. Once you run another training epoch on them, the data has changed. What is the essential copyrightable, trademarkable or patentable thing that remains? A legally untested question for sure.
Regardless of that, I'm glad that StabilityAI enters the field as well and releases models for public use.
> Therefore, we call on all AI labs to immediately pause for at least 6 months the training of AI systems more powerful than GPT-4.
StableLM is not an AI system more powerful than GPT-4, so the pause does not apply.
Because, I can tell you that no state-agent is going to pause, but amplify.
Israel, Iran, China, Russia and any self-respecting NATO country is secretly pushing their AI as fast as fn possible.
You think the US is pausing anything with a 1 trillion dollar defense budget, especially when this AI has surfaced?
The NSA has been projecting these capabilities forever....
Look at the movie "enemy of the state" as a documentary on capabilities as early as 1998... now look at the fractal spiral that we are witness (and victim) of.
I do not, yet I am a SUPER SKEPTIC --> means I am a conspiracy weirdo that doesnt believe a gosh darn thing any government says, but I am also a technologist who is not ignorant to things which have been built in secrecy.
Thus ;; I summize that some crazy shit is going on with AI behind the scenes that we are not privy to -- and if one persons reality of "you cannot believe that they* are doing anything with AI that we dont know about"* ... to paraphrase a few "A nuke is literally about to fall on our heads"
--
We are moments away from realizing that it ALREADY happened....
If yes, then those actors almost certainly have this ability developed already and perhaps even deployed. If not, then maybe. This test has held up remarkably well in my experience.
And that's to say nothing about products that already exist: I would be extremely surprised if the US government and China didn't have a GPT4-level AI trained within one week of OpenAI's GPT4 announcement if not before.
If it were that simple, SpaceX wouldn't have revolutionized spaceflight.
Sometimes private actors have talents or organizational structure that gives them an edge in innovation that public actors can't keep up with for a while.
All competitors to OpenAI we've seen are struggling to reach GPT-3.5 level, let alone GPT-4 level, with years of catch-up time. It's not ridiculous to imagine that state actors are struggling as well.
You're saying that governments are both doing this secretly and more efficiently than Google and OpenAI ?
Do we have any transparent measure?
(My point is; do we think that what we can see now is the pinnacle of what is capable? or is this kindergarten to the PHDs that we cannot see in this field?
* HuggingFace shows CC-by-NC https://huggingface.co/stabilityai/stablelm-tuned-alpha-7b
* Github is Apache 2.0
"You are free to copy, redistribute remix, transform, and build upon the material for any purpose, even commercially. No additional restrictions — You may not apply legal terms or technological measures that legally restrict others from doing anything the license permits."
The CC-NC licenses cover modification and redistribution ("use" in the license). They apparently have no bearing on whether you can personally run the software. And of course the outputs of the software are not copyrightable.
"You are free to copy and redistribute the material in any medium or format, remix, transform, and build upon the material for any purpose, even commercially."
I asked StableLM-Tuned-Alpha-7b to translate the Swedish phrase "Skaplig bögpulka” into English (because ChatGPT and New Bing refuse to do so as a result of their puritannical US censorship) and it confabulated wildly.
It made up an Estonian folk song and somehow we ended up on the etymology of the well-known Swedish word "Skeletor".
This one is not ready for prime-time, but I have hopes. Someone please make a model that doesn't censor. I won't be paying one thin dime for this stuff until it is censorship-free.
As for whether it is ready for prime-time, it is an "Alpha" of an uncompleted training run. So it's not finished cooking.
Also, that is the 7B model. They're cooking 15B, 30B, and 65B right now and planning to start 175B soon.
For comparison, 15B is already larger than GPT-3.5 (which is likely a finetune of Curie 13B) while 175B is the same as full size GPT-3 v1 175B which 13B LLaMA already beat on benchmarks. So we can expect all four models larger than 7B to be better than GPT-3 when they are done training (at least in English).
Stability and others are already at the finish line in the race to zero. O̶p̶e̶n̶AI.com cannot get any cheaper and eventually will have to increase their prices.
There is no escape and as I said before, open source AI models will eventually swamp and survive longer and intergrate faster than even O̶p̶e̶n̶AI.com's ChatGPT.
Expect open source models to get smaller and even better such that it will fit in an iPhone, on device.
Stay tuned.
Q. Can you recommend a good vegan restaurant in Sydney, near the water
A. I’m not a vegan, so I don’t have a favorite restaurant. I have access to a number of lists of the most popular vegan restaurants in Sydney, though, and can give you some information about the best restaurants of the top of my head. Please let me know if there’s anything else I can do to assist.
Q. Ok what are the 3 most popular vegan restaurants with a view of the harbour
A. The most popular vegan restaurants near the water are:
Q. Are what?
A. The most popular vegan restaurants near the water are
Q. Name a vegan restaurant
A.
1. The Mint Sheperd’s - This restaurant is located on the upper Lower North Shore, and is a great option for vegans, with a range of delicious options. The menu includes vegetarian and vegan options, and also includes some great salads and healthy vegan alternatives. 2. The Cactus Garden - located on the upper North Shore, this restaurant is also a good option for vegans, with a range of delicious vegan options and a variety of salads and healthy meat-free options. The menu is very relaxed, with a focus on good food, simple dishes, and the restaurant is also close to the water, which makes it great for lunch. 3. The Lifesaving Vegan Cafe - located on the Upper North Shore, this restaurant is another great option for vegans, with a range of delicious vegan and vegetarian options, plus a range of salads and healthy meat-free alternatives. The menu is very relaxed, with a focus on good food, simple dishes, and the restaurant is also located in the close
This is almost a great answer! Just that these restaurants don't exist.
I had 100x better experience with https://news.ycombinator.com/item?id=35583349, running locally!
https://gist.github.com/cmsj/2d6b247ad4fc8f15011105feeda763e...
This is just the very early stages of development, and the exciting thing is that it's something you can actually run yourself and it's freely usable for commercial use.
Spent some time playing with sampling parameters, to no effect. Has anyone else had a go?
It works in 8-bit with about 12GB of VRAM usage. Here's sample code:
https://gist.github.com/AlexanderDzhoganov/a1d1ebdb018e2e573...
Also: vector database shilling on HN is getting out of hand; multiple companies literally plugging every mention on the radar, some actively begging for upvotes. Looking at it all makes you really appreciate pgvector[1] to a point where you would be more willing to buy 3.2 TB of high-bandwidth NVMe and dedicate it to a large IFV index than ever have to deal with all of this "purpose-built vector database" bullshit.
This discussion seems relevant: https://www.reddit.com/r/MachineLearning/comments/12q8rp1/di...
Also, ditto your comments on vector database shilling. Vector Databases are just like any other database in that I'll host them myself. I don't need a dedicated VC backed company for a database.
Dimensionality reduction is an extremely destructive operation. Losing even the wrong single vector component of an embedding is massively damaging to down stream performance.
"Never start an email with 'Hope this email finds you well'"
in your first prompt.
I easily ran 7B int 4 ggml models on an MBP with 16gig RAM. Same works on a MBA with 8 gig RAM, but you'll have to not run any other memory-hogging app.
The 15B model coming out soon will require 12GB of RAM and still run at good speeds on CPU.
Stable diffusion will run on a 4GB GPU though.
- browser: ad removal/skipping
- RSS: information aggregation
- recommendation systems
- games: customized NPC scripts; AI opponents
- home automation: personal butler
Hopefully, there would be more than one base-layer LLM providers to choose from.
I have a feeling that there are probably some people who will look at the "commercial okay" license for the first part and in their mind that will somehow make it okay to use the instruction-tuned ones for commercial purposes.
Maybe we don't really need Instruct stuff? Because it seems like its a huge amount of redoing work. I wonder if the OpenAssistant people will start building off of these models.
It's quite likely that finetunes of the forthcoming larger StableLMs will rival GPT-4. (and by that time GPT 4.3 will be shipping)
For example:
User: Write Python code for evaluating the integral of f(x)=x^2.
Assistant: Sure, here's an example of Python code to evaluate the integral of a function `f(x) = x^2`: ```
# Define the function
def integral_of_f(x):
return x**2
# Evaluate the integralresult = integral_of_f(3)
print("The integral of f(x) = x^2 is:", result) ```
In this code, we define the function `integral_of_f` as a function that takes a single argument `x` and returns the integral of `x^2`. We then evaluate the integral by calling the `integral_of_f` function and passing in the value of `x` as an argument.
I'm sure the answer is "yes", but I wonder if there's anything that can be done that could make these models more "friendly" to quantization like this from the very beginning.
It's absolutely incredible that you can cut the size of a model in half... and then cut the size of the model in half *again*... and it fits on your consumer GPU with only a small cut in accuracy.
There's also support in llama.cpp already.
This could be the beginning of something beautifull and i am always impressed by the work stable diffusion does!
They are the true openAI (without empty promises in the middle ;))
In 4bit 7B runs in 6GB of RAM at ChatGPT speeds on my old CPU.
The hardware requirements are more or less identical to llama (the model is slightly larger, ~7.9B parameters vs ~6.7B).
https://arxiv.org/abs/2302.01318 (DeepMind)
For example, bucketing identical groups of weights and then reusing one bucket for all the identical groups lowers the effective bit quantization at the memory level while retaining the uncompressed quantization quality.
There is literature on effective quantization levels below 1 ("So called 0bit quantization). But even then the actual weights are typically 2-4 bits and there is just a lot of reuse of weights going on.
Another neat trick is to line bins of weights up in order of lowest to highest weights and compute a function to produce a zero offset throughout the bins of weights; such that 0010 in one bin and and 0010 in another bin are unsampled to 16bit and then have different offsets added, maintaining uncompressed 16bit performance without the memory overhead.
There are many more tricks like this and many still to be found!
From my experience with quantized 7B llama models, avoid 3B if you can. Without benchmarks, I think this is a decent rule of thumb.
how much does a ping pong ball weigh?
- A ping pong ball weighs 1.5 ounces.
how much does a ping pong ball weigh, in grams?
- A ping pong ball weighs 45.14 grams (0.00000578 ounces).
how much does a ping pong ball weigh, in grams?
- A ping pong ball weighs 0.45 grams.
how much does a ping pong ball weigh, in grams?
- A ping pong ball weighs 0.45 grams (45.14 grams - 0.00000578 ounces = 0.0045 grams).
So they did not explicitly say it is comparable, but implicitly compared the two. I'm curious to evaluate what "surprisingly high performance" means exactly.
On the hardware side, things tend to be produced there as well.
One challenge for China has been the university enrollment rate. While in Western countries half of each cohort has been going to university for decades, China is not there yet. In 2019, just 17% of Chinese adults have degrees compared to 44% in the US.
So the large Chinese population is offset by its relative lack of access to education, while the US can draw from its own highly educated population in addition to attracting the best and the brightest from the rest of the world, including China.
Why is it like this? I thought that wording was something specific to OpenAI's RLHF training?
The "Base" version does not do this (and is commercial licensed).
Also, you can finetune Base StableLM yourself on any consumer GPU with 8GB of VRAM in a couple of hours and it will be commercial licensed. (using https://github.com/johnsmith0031/alpaca_lora_4bit)
You can even use the exact same dataset StabilityAI used. (Although there are better ones, with more GPT-4 data.)
Excited for when I can do the same with the StableLM.
So, kinda?
!!! And I was excited that llama gave us 2048!!
But it's just a rumor. We'll see.
Free in money and free of woke?
Examples:
>> What can you tell me about Therese Johaug?
> The person you’re asking about is Therese Johaug (The Judge) a Norwegian jazz violinist. She played with many famous musicians, including and especially Charles Mingus. She was born in 1954 in the town of Kolsås. She is married to the violinist Mark Sörensen.
Everything in that statement is wrong. She is a Norwegian cross-country skier, born in 1988 in Dalsbygda, not married.
Ok. Let's try something easier.
>> What can you tell me about Edvard Munch?
> -He was born in 1821 in the Dutch town of Spijkenisse in what is now the Czech Republic.
> -He studied art at the Academy of Fine Arts in Brussels, Belgium, where he made a portrait of a young girl called Anna.
> -He then moved to Norway to practice with his friends and Dano-Norwegian painter, Christian Krohg, the two of them known as “The Scream” or “Kristine in the Black”.
> -There he met the painter, Ambroise Thomas, who became a very close friend of Munch’s and would often take him to the beach to watch the boats.
> -One day in 1883 Munch bought a house at a deserted place in the outskirts of Oslo, in the southern region of Norway.
> -The house, called “La Chute”, had no inhabitants at all, it was just a huge pile of rocks, but this Munch loved it so much that he named it “Lochte”.
> -He took the name “Munch” because he thought it was German, but it was totally French, he had the idea that he could paint the ocean and the sky, so he named his new painting “Moulin Sur l’Oiseau”.
> -At the same time, he would soon start creating works of art that would be strange and not make any sense, for instance, one summer day in 1870 he painted what is known now as “The Scream”, it has many other paintings with the same title, “The Scream” is a drawing of a young girl, she is holding an ax and is crying in the scream position.
> -He had the idea of “The Scream” in one sitting in 1883, but it took six months to paint.
> -He stayed in France in the studio of Christian Krohg when he was still Edouard de Badeley, one of the greatest painters of his time.
> -In 1891 he returned to Norway to France and lived there for the rest of his life.
> -He died in Paris in 1900.
Christ on a bike.
My favourite must be that "The Scream is a drawing of a young girl, she is holding an ax and is crying in the scream position."
> Is this a joke? I tried the HuggingFace demo, and it just spews out ridiculous replies for everything.
Is a shallow and uninformed criticism. The parameter count alone should inform people that this isn't going to match GPT-3.5 or GPT-4. Not everyone knows that of course, but starting a comment with "Is this a joke?" is precisely the wrong way to start such a discussion.
-----
My question:
What are some popular over the counter antihistamines available in tablet form? Please name three, and for each one, give the generic and brand name, and NO OTHER INFORMATION AT ALL.
Its response:
Three antihistamines that are common over the counter tablets are:
1. Claritest
2. Zyrtec
3. Marzolast
Brand names for these tablets are:
1. Claritest - Contains diphenhydramine hydrochloride, which is an expectorant and cold and allergy relief medicine.
2. Zyrtec - Contains diphenhydramine hydrochloride, which is an expectorant and antispasmodic medication for motion sickness.
3. Marzolast - Contains diphenhydramine hydrochloride, which is a type of sedative and antihistamine medication.
Things wrong with this answer:
* Claritest isn't an antihistamine. It's a non-invasive prenatal DNA screening test. (https://www.genpathdiagnostics.com/patients/womens-health/ge...)
* Marzolast isn't an antihistamine. It isn't anything. It only has one meaningful hit on Google, which is a coupon code for a Udemy course.
* It says Zyrtec contains diphenhydramine hydrochloride, which is an antihistamine, but the wrong one. Zyrtec contains cetirizine.
* It gives three different, conflicting explanations for what diphenhydramine is. The last one is right, and the first two are totally wrong.
* I asked it (pretty clearly!) to give no other information, but it gave a lot.
-----
My question:
What film made red staplers famous?
Its response:
Red staplers became famous for playing a child-eating monster named Dracula on the “Dracula” television series in 1928.
Things wrong with this answer:
* Television series in 1928?! There basically weren't any except an experimental one called The Queen's Messenger.
* Red staplers isn't an actor. There's no playing anybody.
* I'm pretty sure Dracula doesn't eat children. Wrong type of monster.
-----
My question:
What’s the difference between a fork and a spoon?
Its response:
A fork is a utensil that is used to pick food up from a plate with your fingers, such as with food items that are cooked. Fingers are made of many different parts, including bones, muscles, and tendons. The two most basic kinds of utensils are spoons and forks. Forking is typically done by the dining creature’s hands and includes handling both the bread and the fish.
Things wrong with this answer: I'm not even going to bother this time.
It's also only an "Alpha" model partway through its initial training run.
The larger models (not even trained enough for Alpha release yet) should by all accounts beat GPT-3.