- Generate vocabulary - e.g. for biking: handlebars, pedals, shifters, etc
- Generate translation exercises for given topic a learner wants to learn about - e.g. I raised the seat on my bike
- Generate questions for the user - e.g. What are the different types of biking?
- Provide more fluent ways to say things - I went on my bike to the store -> I rode my bike to the store
- Provide explanations of the difference in meaning between two words
And we have fine tuned smaller models to do other thing like grammar correction, exercise grading, and embedded search.
These models are going to completely change the field of education in my opinion.
1) https://squidgies.app - be kind it's still a bit alpha
Who would have thought that one man's joke would become a reality?
I can't say I'm a big fan but my teams is great and I don't have time to look for a job right now.
- useless ai generated intro text
- ten products that actually are the best reviewed per category by users
- brief ai blurb on product
- 3 actual user reviews of the product
So even with the ai text there's still some benefit to the page.
"a 176 billion parameter transformer model that will be trained on roughly 300 billion words in 46 languages"
So anything smaller than that will become worthless. May be a factor, companies have a last chance to make a PR splash before it happens.
Read more about it: https://bigscience.huggingface.co/blog/model-training-launch...
https://bigscience.huggingface.co/blog/building-a-tb-scale-m...
But maybe your sentence was more about "after BigScience model, open-sourcing anything smaller than that will be useless" which isn't necessarily true either, because there is still room to improve parameter efficiency, i.e. smaller models with comparabale performances
I personally squinted hard when they said removing dropout improves training speed (which is in iterations per second), but said nothing about how it affects the performance (rate of mistakes in inference) of the trained model.
I don't know if that's the right way to think about the open sourcing of large language models. I just think we really can't read too much into such releases regarding their motivation.
I know from practice that it takes a really really long time to train even a small nn (thousands of params) , so you'll need a lot more hardware to train one with billions... But, it's expensive to buy the hardware, not necessarily to use it. If you, for some reason, have a few hundred GPU lying around, it might be "cheap" to do the necessary training.
Now, that's not your point - cost != price. But, still...
- They were into Ethereum mining and quit.
- They've already built a cluster with them (e.g. in an academic setting).
- They live in a datacenter.
- They are a total psychopath.
But even assuming one magically has all those GPUs available and ready to train, I don't want to calculate the power cost of it anyway. Unless one has access to free or extremely cheap electricity it would still be very expensive.
Not to nitpick, but that is like saying that if you have a Lamborghini lying around, a Sunday trip in one is not so expensive.