The Coming of Local LLMs
nickarner.com
nickarner.com
4 example labels, and I had a binary classifier in seconds. Sure, semantic text classifiers were possible for a while, but making it accessible changes everything. Giving anyone who can use a spreadsheet the power of a local LLM (or, basically free LLMs) can make them much, much more productive. A lot of office work is clicking through sheets and doing manual labeling.
It's truly wild what is becoming accessible! Really excited to see the next gen software that the open community comes up with :)
Maybe both the buzz factor and broader applicability means it's more likely to happen this go around?
If you're interested, see this paper that argues that point: https://arxiv.org/abs/2302.06541
Essentially, being label efficient is more important than being compute efficient, because the biggest computing constraint we have is enough humans doing the labeling (and knowing how to work a jupyter notebook), not tensor smashing nvidia cards
They could have just said "efficient". But no - they had to go for "agile".
Sure, you and I know how to write a little script to sort a directory of documents into "schoolwork" and "other stuff".
But most people don't have that ability, so giving them that would really help accessibility.
Bonus points: I had never written a browser plugin, but GPT4 helped me do it in under half an hour.
I’m only vaguely familiar with the API.. if I had to guess I would say you send
- a system instruction that its job is to filter unwanted content
- examples of unwanted content
- an instruction like “filter the following html:”
For every web request you want to filter, you would re-send all of those messages followed by the page HTML as the final message. Is that close?
- Examples of unwanted content
- Then I give a large numbered list of comments and ask which numbers should be filtered
- The plugin then just deletes those comment nodes from the DOM. If HN ever updates their HTML I will have to tweak this code.
The reason to send a large list of comments is just to save on costs. It's cheaper to do it this way than one comment at a time.
So the main difference from what you've proposed is GPT never sees the HTML. My code enumerates the comments in the HTML and splices them in to the prompt in a nice numbered list, then does the reverse translation from list number to DOM element in the other direction.
How sure are you that all that code has been reviewed by a 3rd party? How many CVEs a year impact your laptop/desktop?
Do you have any reason to think that increased productivity with LLM assistance will result in lower quality code? Personally I find LLM assistance increases productivity, decreases the penalty of using a more difficult language like rust, and makes it more palatable to spend more (LLM assisted) time writing tests.
I just don't see it chatbot assisted programming any worse than what we have today.
You are hereby removed from the discourse.
/s
Also, unless you are reading every comment in every thread, you are going to miss a few interesting ideas anyway. That's ok too.
it will be hard to trust anything at all
as always happens
I've already used LLMs in my work as a data scientist but it requires a ton of work to just make the results tractable (and I have been using GPT4, which behaves pretty well). These smaller language models ain't so regular. Ok, like consider a basic thing you want to do with a classifier: understand its behavior on a held out data set. Since no one knows what is really in the training data (since its so large), its quite hard to understand what the model can generalize about and what it has just accidentally memorized. CF reports that GPT4 doesn't perform nearly as well on even simple programming exercises that are chosen in such a way as to be sure they weren't in the training data.
There is enormous potential for statistical fuck ups here. Prompt engineering, for instance, is an easy place for over-fitting to happen as a prompt is fine tuned on data the prompt engineer has and thus fails to generalize to new data.
I do think there is a lot of value here, but I'm also sure that sloppy use of large language models is going to cause a bunch of trouble in the short to medium term, generate a lot of garbage, pollute a lot of databases, etc, while we figure all this stuff out.
But I posit that most text classification tasks don't have such strict accuracy requirements. For one, no text classifier is 100% accurate. For instance, I have genuine mail in my spam folder frequently. I see spam on social networks, etc. I struggle to think of cases that aren't at least somewhat tolerant to some amount of incorrect classification.
The potential amplifying power of that is enormous.
What I'm wondering is how they fed the documents, as all those LLMs have limitations on the input sizes.
That's for the 7B model. The 30B model needs 24GB quantized (or 64GB for the unquantized model).
> What I'm wondering is how they fed the documents, as all those LLMs have limitations on the input sizes.
It's like file hashing at scale, you don't have to read the whole stream for every file, just the first 1024/2048 bytes (or first few paragraphs).
(This works for classification and sorting, less so for summarization.)
Binary classification can actually take you all the way in terms of classification if you are clever with set theory. It's also one of the most traceable & deterministic ways to understand how the natural language is being interpreted at each step.
The amount of performance required to run something like an SVM is laughable compared to what is required to run even baby-tier LLMs. If you can reduce the cost of running models to a <1ms invocation over a few megabytes of black box, you can easily test thousands of these per-user-query. Re-training and iterating is much more enjoyable for these reasons. You also don't need any GPUs for this.
At the end of the day, the quality of your data will be the biggest issue with older techniques. LLMs can bandaid all sorts of weird things that crop up in the real world and aren't present in the training data. SVMs cannot tolerate requests delivered in the format of Shakespeare (if unexpected). In a well-controlled domain, you would probably be able to get away with much cheaper options that are also more flexible.
Curious how you run the model then interface with it.
Attention: Isn't it quadratic in context length? I dunno, this feels like the crude first iteration of something that will get inevitably passed by something that scales better.
"""
Complexity is quadratic in sequence length. For 512 tokens it is 262K, but for 4000 tokens it becomes 16M and goes OOM on a single GPU. We need about 100K-1M tokens to load whole books at once.
Since 2017 there have been hundreds of attempts to bring O(N^2) to O(N), but none of them replaced the vanilla attention yet in large models. They lose on accuracy. Maybe Flash attention has a shot (https://arxiv.org/abs/2205.14135).
"""
Keep in mind: Once you go into precision as low as 4 bits (or lower?), all sorts of optimizations can become practical. Off the top of my head, maybe you could cache and reuse common attention sub-matrices (e.g., a 16×16 sub-matrix with 4-bit elements occupies only 16×16×4÷8=128 bytes of space)?
My sense is there's so much money at stake here, that whoever does this first will win big even if they end up having to replace or augment it with something better down the road. Hypothetical example: Imagine Intel or AMD coming out with a $1K or $2K card that has "built-in 4-bit attention," enabling you to run transformers of much greater scale on a run-of-the-mill desktop PC. I'd buy that in a heartbeat.
[a] Here's a recent post about a new approach from a group at Stanford that looks promising to me, although I don't fully understand all its details yet: https://news.ycombinator.com/item?id=35502187
[0]: https://twitter.com/typedfemale/status/1609867110695735296
Alternative models like S4 have been able to get transformer level performance with O(N) sequence length scaling.
It’s used for boosting interference (offline) on Linux, Mac and Windows.
Haven’t bought or used them but I’ve had my eyes on these for a little while!
Waste of die space currently (on Macbook at least, I'm sure they find uses for it in the iPhone)
[1] https://github.com/smpanaro/more-ane-transformers/blob/main/...
But I do agree it should be improved!
I've posted this before, but it seems like this genre is just getting more and more popular - and more and more untethered from any actual metrics of how good these models are.
They are not comparable with late 2022 ChatGPT.
Llama has the potential to reach ChatGPT it needs tunning to get better at responding to questions, llama if I am not worng is mostly attempting to predict what is next.
I can see it similar like Midjourney and Stable diffusion, midjorny can make any stupid prompt look like a digiatal art in the style of Midjorney but look how many stable Diffusion innovation happens, a competent person that is on top with all the new stuff can produce absolute anything in any style they want.
But I just don't see the need to mislead and suggest that these models are "on par" with ChatGPT or something like that. They just aren't.
Imagine what a math community could train, they just need access to the model and soem GUI software that can help them train.
So llama based Chat stuff is not yet comparable with ChatGPT but there are already lot of progress made. At this moment coding and math is bad in llama based but other stuff is great, like story creation, also I only could test 3-b 4bit and it is good enough to for example provide me a complex response in valid JSON format.
Then hardware will be a separate business. I think Apple might be caught off-guard by Nvidia on hardware. The latest NVlink and 400Gbps interconnects when combined with H100 next iterations and also rumors of advanced PCIe motherboards with high lane Nvidia CPUs and it looks to me that next year they can be selling $100-300K physical systems optimized for LLM inference that physically remind me of mainframes.
It is kind of idiotic that some scientist can spend years and a lot of public money to create some technology and then bilionairs are miliking all the profits.
I’m really hoping there are viable distributed and somewhat decentralized eventually consistent training algorithms we could all run in a P2P system. That would be super cool.
However I can easily see that now the framework has been established if a company builds a proprietary curated dataset for specific skills and then pays to spend resources on specialized reinforcement training.
Then they can commercialize that I would think. As people would pay for an LLM that does XYZ the best. Kinda like your Disney example but I was thinking engineering tasks in my head.
It would be great if we could have detailed QA evaluation to show this, but of course then the open source people would train their models on it as a fine-tuning datasaet.
I mean… the horror.
langchain agents are a good starting implementation.
you can build your own prompt and get the ai to work by iself hallucinating tools, which may be cheaper to test out than going back and forth with an agent manager. not as accurate, but you can still extract useful work, i.e. https://i.imgur.com/AE4R3dR.png (gpt-35-turbo is traditionally failing this task completely, prompt get it to work at it)
these prompt all require the model to work off data within the prompt within the first shoot. model require a degree to introspection for that to work.
There's also already tools for conversation flows (which just means you prepend the conversation history to the prompt).
I'm not saying the performance is nearly as good, but the actual workflow does already exist and is massively improving. The interesting part to me is that this finetuning can be done in a couple few hours on a consumer gpu (4090).
They will be able to shutdown your startup on a whim. Even if they didn't politicians and regulators would be a huge risk. Without democratization we get blade runner . Not that democratization has no problems, just that it is the way a lot of us are wanting it to go.
I get it, I understand why people like decentralization - but the open source community doesn't even have close to the capability to train an actually open-source LLaMA equivalent.
While those organizations might not be the right fit, I think they serve as an "existence proof" that a large scale open / nonprofit project is not completely inconceivable.
Could possibly see an industry consortium (of players too small to compete on their own) funding an open effort.
Last idea sounds crazy but hear me out: how much would Nvidia spending $100M on "open" models boost spending on graphics cards? I hope someone's running the numbers on that...
Probably not as much as in a zero-sum game where everyone is trying to train their own model. Every leap in CPU inference makes this an increasingly less appealing option for them.
But I agree, some industry consortium might try to do it. I think they would first have to be relying heavily on LLMs before they were willing to do that, and its possible by then that the lead will have gotten too large to easily surmount, especially now that everyone has stopped publishing.
The current hardware of course can't pull anything like this yet. But iPhone supports on-device facial recognition, object recognition, dictation and translation, so small steps...
https://research.ibm.com/blog/why-we-need-analog-AI-hardware
https://news.mit.edu/2022/analog-deep-learning-ai-computing-...
Make no mistake, it’s for tinkerers that do not expect each prompt to be answered human like.
I see them as creativity and thought testing tools, also knowledge exploratory.
Models can be found on huggingface.co, and I’d start with eachadea/ggml-vicuna-13b-4bit, but it needs 10G of cpu-ram. It is very friendly to any prompt though.
I read on my way (on reddit, when I recall correctly), that there must be some really good intro videos on YouTube.
I'm trying to decide whether it's worth while to splurge on a more expensive 4090 vs a 4070 or whatever.
That said, I recommend renting a cloud GPU for a few hours and trying the larger models on them before buying a GPU of your own, just to see if the models meet your requirements.
I doubt it will take you a year to train a model.
Vicuna-13B cost $300 to train/fine-tune [0]. They trained on an A100 which costs $10k [1].
[0]: https://vicuna.lmsys.org/ [1]: https://github.com/lm-sys/FastChat/blob/main/scripts/train-v...
The 24GB of VRAM is 100% worth it alone. If you want to do local ML stuff you _need_ that 24gb of VRAM.
64GB of ram + 24GB of vram lets you run a lot of the medium size models at decent speeds. I don't use Colab personally but AFAIK it should work fine for you if you don't want to do it locally.
Also worth noting is the newer ray tracing rendering that cyberpunk is doing. You should checkout the demos IMO it looks sick. It only runs at 18fps on a 4090 so it's only playable on a 4090 + dlss, and I'm not sure if the newer rending tech will be super achievable on any of the other cards - if that's of interest to you.
Llama is basically an auto-complete right now. We're celebrating baby's first steps. It's not really worth the $600-1000 jump up from cards that can run all current games 4k60.
So finetuning with LoRAs and a few other methods is fine on higher end consumer hardware like a 4090 and finishes in a reasonable amount of time - IMO definitely worth it if you're experimenting with this especially for the inference.
The base training though yeah I totally agree with you - train in the cloud, don't buy hardware when you need a month of 8x A100's or whatnot.
My perspective was for people who have other uses for them e.g. gaming or local inference. From a pure finance standpoint you're definitely right - you should rent and not buy a dedicated card. I think you'd need a few thousand hours to break even which is a few months 24/7.
I've been underwhelmed by Alpaca and Alpaca-LORA and LLaMA all at 13B but I have not tried higher params.
Vicuna is more friendly in that regard.
But I’m well aware of their limitations also, and I can see how one can be underwhelmed. They are not jacks of all trades
But newer finetunes like Vicuna go well beyond that, including hundreds of thousands of real human conversations with GPT-4 ChatGPT in the dataset (unlike Alpaca's fully synthetic dataset).
Vicuna-13B in 16bit is easily comparable to ChatGPT-3.5 in capability. Newer finetunes coming out nearly every day are going beyond chatGPT-3.5 and getting closer and closer to GPT-4 performance.
You don't even have to install anything to validate this for yourself. There's a live web demo of Vicuna-13B right here: https://chat.lmsys.org/ (disable ad blocker if it does not load)
https://www.reddit.com/r/LocalLLaMA/comments/11o6o3f/how_to_...
13B is pretty meh, but 30B is great, if not quite Chatgpt. But I can ask it why my highschool geometry teacher was such a cunt and it will happily discuss the matter without reservation. Very therapeutic.
I'm sure if we'd pool resources together we could build a truly open alternative worthy of building on top of.
This would require like ConstitutionDAO level of resource pooling without direct monetary payoff.
I mean, good luck.
Like Jack Dorsey would often say "it's not important to be first to market, you can just be best to market". And the world got CashApp.
I'm sure however Apple enters the space, it will be fleshed out (vs Bard).
Siri should be better! It lost features post-acquisition by Apple, and it seems like user privacy is why.
Home automation is arguable. If you consider a single point of failure on a server somewhere to be bad, Apple's solution is pretty great. Their commitment to zigging where others zagged put them behind, since hardware vendors didn't want to put in powerful (expensive) enough chips to handle the cryptography, but while other companies go out of business, or transmit images and video to external parties, Apple's works reliably and securely.
Still, as with most of my complaints about Apple, it's a trade-off between privacy and functionality, and Apple will seemingly always choose privacy over functionality, even as Google consistently chooses functionality over privacy.
[0] Yes, there are examples of edge cases that suggest a less-than-perfect record. Contrast that with their competitors, for which invading privacy is foundational to the business model.
Increasingly Apple seem to be blocking tracking, noteably from facebook, to make the most profit of that tracking. I've read claims that apple made between $5B and $20B on advertising in 2022. It's far from clear that Apple's view on privacy is going to stay the same.
At Apple's scale, it's relatively easy for them to deliver $20B in ad revenue without any privacy-invading means.
People seem to have forgotten, but ads used to be based on context, so people looking at apps related to fitness might see ads related to fitness, but that wouldn't follow them around when they looked at other things. Apple still seems to be doing that; I haven't seen fitness ads on games, or game ads on fitness apps.
They are absolutely behind in the space and anyone who works in the industry will tell you that. Only feasible way IMO would have to be a very big budget acquisition of one of the major LLM startups, but most of those already have big tech backers.
I have Kevin Kwok's SheepyT running on my iPhone right now - it uses GPT-J, which is an openly licensed LLM by EleutherAI.
'You are a koala who plays with 5-7 year olds, you are friendly natured and curious and like to ask questions'
I know the article goes on to speak about something else, but I'm not sure why this claim that the LLaMa model weights were leaked, as in unintendenly made available is being done.
But this could be the equivalent of the Low Orbit Ion Canon for phishers and scammers
Digital arms races are nothing new, this is just the latest battlefield.
Power consumption would not be an issue if it's used sporadically throughout the day, it's not like it needs to run continuously?
There is still the issue of NAND flash read disturb, which I haven't fully looked into yet.
I’m also in the middle of making it user friendly to run these models on all platforms (built with Flutter). First MacOS release will be out before this weekend: https://github.com/BrutalCoding/shady.ai
I get around ~140 ms per token running a 13B parameter model on a thinkpad laptop with a 14 core Intel i7-9750 processor. Because it's CPU inference the initial prompt processing takes longer than on GPU so total latency is still higher than I'd like. I'm working on some caching solutions that should make this bareable for things like chat.
Yes, we will be running them soon in low end hardware, but we need to get at least to GPT-3.5-turbo level of inference speed and quality before we try to make it small.
I already started.
$ neofetch
-` x@decpti
.o+` -------
`ooo/ OS: Arch Linux ARM aarch64
`+oooo: Host: Apple Mac Studio (M1 Ultra, 2022)
`+oooooo: Kernel: 6.1.0-rc6-asahi-4-1-ARCH
-+oooooo+: Uptime: 4 hours, 23 mins
`/:-:++oooo+: Packages: 177 (pacman)
`/++++/+++++++: Shell: bash 5.1.16
`/++++++++++++++: Resolution: 1920x1080
`/+++ooooooooooooo/` Terminal: /dev/pts/0
./ooosssso++osssssso+` CPU: (20) @ 2.064GHz
.oossssso-````/ossssss+` Memory: 717MiB / 129540MiB
-osssssso. :ssssssso.
:osssssss/ osssso+++.
/ossssssss/ +ssssooo/-https://github.com/geohot/tinygrad/tree/master/accel/ane
But I have not tested it on Linux since Asahi has not yet added support.
Same machine but OSX, llama.cpp runs at 18ms per token (7B) and 200ms per token (65B) on CPU using float16.