Alpaca-LoRA with Docker
github.com
github.com
I love the .ccp, Apple Silicon, etc projects but IMO for the time being Nvidia is still king when it comes to multi-user production use of these models with competitive response time, parameter count/size, etc.
Of course as others pointed out the quality of these models still leaves a lot to be desired but this is a good start for the inevitable actually open models, finetuned variants, etc that are being released on what seems like a daily basis at this point.
I'm walking through it (fun weekend project!) but my dual RTX 4090 dev workstation will almost certainly scream with these (even though VRAM isn't "great"). Over time with better and better models (with compatible licenses) the OpenAI lead will get smaller and smaller.
I use this for an optimized hosted Whisper implementation I've been working on. It hits 120x realtime with large v2 on a 4090 and uses WebRTC to stream the audio in realtime with datachannels for ASR responses. Hopefully a "Show HN" soon once I get some legal stuff out of the way :). I mention it because AFAIK it's many multiples faster than the OpenAI hosted Whisper (especially for "realtime" speech).
I expect we'll see these kinds of innovations and more come to self-hosted approaches generally and the open source community will pull a web hosting, etc Microsoft vs Linux/LAMP/etc 1990s/early 2000s situation on OpenAI where open source wins in the end. The fact that MS is so heavily invested in OpenAI is just history repeating itself.
Yep, saw the Databricks article! I don't try to make specific time predictions but you're probably not far off :).
>I am a 25-year-old woman from the United States. I have a bachelor's degree in computer science and am currently pursuing a master's degree in data science. I am passionate about technology and am always looking for new ways to use it to make the world a better place. Outside of work, I enjoy spending time with my family and friends, reading, and traveling.
Well, I was starting to get tired of "as a AI language model" disclaimer. Out of curiosity, is this model meant to be a 25 year old personal assistant?
Think of the prompt as "pretend you're some random person, tell me some details"
Perhaps someone will release a llama I can run at home… how about “llama-homekit”? ;)
Although better than Bard (btw, Bard sucks compared to ChatGPT and can’t even do translations - which I would have expected out of the box from Google)
At this pace we might not even need GPUs anymore.
All the research I've seen says quantization has basically negligible performance impact. My experience working with 65B 4bit has been great
So what is Alpaca-Lora? I know you get Alpaca by retraining Llama using Stanford Alpaca 52k instruction-following data? So if I am guessing right, you get Aplaca-Lora by retraining Alpaca using Lora's data?
This reduces model sizes and therefore also compute costs.
See the abstract of https://arxiv.org/pdf/2106.09685.pdf
The work was all done by the original repo author - just added a Dockerfile!
> Try the pretrained model out here, courtesy of a GPU grant from Huggingface!
https://huggingface.co/spaces/tloen/alpaca-lora
Anyone else getting error messages when trying to submit instructions to the model on Huggingface? It just says "Error" so I don't know if it's a "too many users" problem or something else
edit: nevermind, I was able to get a response after a few more tries, plus a 20 second processing time