Stanford Alpaca, and the acceleration of on-device LLM development
simonwillison.net
simonwillison.net
> Alpaca shows that you can apply fine-tuning with a feasible sized set of examples (52,000) and cost ($600) such that even the smallest of the LLaMA models—the 7B one, which can compress down to a 4GB file with 4-bit quantization—provides results that compare well to cutting edge text-davinci-003 in initial human evaluation.
this is the most exciting thing. the cost of finetuning is rapidly coming down which means everyone will be able to train their own models for their usecases.
Looking for the contrarians on HN: what is being left unsaid here that people like myself and Simon might be getting too optimistic about? what are the known downsides that people in academia already know about?
Generally speaking, research is not usually done with consumer usage in mind, so what this is, and Dreambooth etc. for Stable Diffusion was, is that gap between researcher software and accessible software being bridged.
The past week has felt like a wake-up call to enthusiasts. Running models locally has been available for a while (even small, fairly coherent ones), and the majority of "improvements" recently have come from implementing the leaked LLaMa model.
The results from 7B are an improvement on what we had a year ago, but not by much. We're learning that there's room to optimize these models, but also that size matters. ChatGPT and 7B are both great at bullshitting, but you can feel the difference in model size during regular conversation. Adding insult to injury, it will almost always be faster to query an API for AI results than it will be to run it locally.
Analysis: Things are moving at a clip right now, but people expecting competitive LLMs running locally on their smartphone will be disappointed for quite a while. As the technology improves, it's also safe to assume that we'll find ways to scale model intelligence with greater resources, and the status quo will look much different than it does today.
There is no company on the planet that would benefit from providing people local means to run LLMs. As a result only hacks and leaks will be how individuals can manage to run LLMs outside of heavily monitored remote API calls.
Facebook is not a major player in the Llm field, the technological advantage of openai is to large, BUT they can reduce the expected gains of their competition by providing less powerful alternatives for free.
(Though to be clear I do hope Facebook does release more models because I want to play with them)
Or they’ll do a home network device like implement it like a HomeKit hub. So the plugged in resource has more computational power and feeds it to iphones when nearby. Disconnected, the iphone uses a much more rudimentary one or falls back to Siri or a cloud service.
Only two things are for certain; you'll have to pay for it, and it will not be cheap.
I work in biology with some interface to wildlife managers/'end users' of our biological insights. I was hoping we could use these models as advanced chatbots so that our end users could ask biological questions from our data.
However, even the replies to the most basic biological questions in LLAMA/ChatGPT/OpenChatKit are hallucinated, even when setting temperature to 0. We simply can't trust any of these answers; the bot needs to be as truthful as humans in its replies, but none of them are so far. So what good are these models then?
Take a look at this chat log from Bing. Note that the user only sees the stuff that is marked #message and #suggestions.
https://www.make-safe-ai.com/is-bing-chat-safe/Prompts_Conve...
Then take a look at https://langchain.readthedocs.io/en/latest/, specifically the Document Loaders and the Indexes sections to begin with.
Edit: I just tried it with a single task of my own (that I've successfully used with ChatGPT and Bing) and it flubbed it horribly, so this model at least is noticeably inferior to the SOTA, which is not surprising given how small it is.
I am optimistic for 30B or 66B to catch up with OpenAI, but 7B is unlikely to have the same quality.
https://i.imgur.com/V4lzLz7.png
It does quite well at simple reasoning. More complex stuff, it does struggle sometimes, but I am impressed by the output and the fact that it even runs on my computer.
I used the oobabooga repository on Ubuntu 22.
it's a subtle distinction, but i think it shapes and reflects how you view ai as a tool for humans or as a replacement.
The second catch is that you would get much higher quality out of the 65b model, but would need to lay out a few thousand for the hardware.
The third catch is that you need the fine tuning data, but that seems easier than ever to create out of more capable LMMs.
The base config is $2500 to begin with. So not thousands plural, it is one grand more than the base config and a lot of devs/pros go for the 32GB anyhow.
Apple laptops also have atypically high resale value, I usually sell mine after a couple of years for 50%+ of what I paid for it. By the by, Apple have a two week return policy. ;)
I think only your first catch applies. But if you can come up with the right mvp even that might fall away.
You could spend $600 and build your own PC that runs LLaMA 65b.
First of all it is $4,700 on Newegg and secondly it is out of stock. This also leaves me with the rest of the computer to put together which is not $0. And finally it would absolutely not be 10x faster for this use case than a 64GB MBP.
The advantage of going with Apple is that I can pick one up immediately, use it as a laptop (soon with Asahi linux when it matriculates fully) and if need be easily sell it. You might not see the value in that, that's alright, people have different use cases and needs.
I'll add the possibility of network-based service downtime to the list.
But this performance will still page behind OpenAI for a bit but within a month or a few months we should have something almost comparable it’s just frustrating that Stanford is blocking the code and model release when it’s just a matter of time for a full copy to come out.
Alpaca: A strong open-source instruction-following model - https://news.ycombinator.com/item?id=35136624
Also recent and related:
Large language models are having their Stable Diffusion moment - https://news.ycombinator.com/item?id=35111646 - March 2023 (355 comments)
> Second, the instruction data is based OpenAI’s text-davinci-003, whose terms of use prohibit developing models that compete with OpenAI.
Wow, that seems really sketchy on the part of OpenAI. Even considering their overall lack of openness, this clause feels particularly egregious.
Someone should ask the researchers, either via email or via github pull request, it shouldn't even be that hard to do.
And in the alpaca-data file there's this:
{
"instruction": "Perform the following arithmetic operation.",
"input": "(5+7)\*3.",
"output": "60."
},
Do we need human QA on the training data? It's contaminating the set! I wonder how incorrect data like this will be filtered out in the future. Unless we're ok with that kind of new math.{“Problem”: “Perform the following arithmetic operation." "input": "(5+7)3”} ‘’’ I got the answer: ‘’’ Yes, a program can be written to perform the arithmetic operation given in the input. Here is a Python implementation:
input_str = "(5+7)3" # input string result = eval(input_str) # using eval function to evaluate the expression print(result) # output: 36 The eval() function in Python evaluates the expression given as a string and returns the result. In this case, the input string "(5+7)*3" is evaluated to 36. The program then prints the result using the print() function. ’’’
DaVinci is also not instruct tuned, but it has an instruct tuned variant like the other OpenAI text completion models.
The article is otherwise good.