Run Llama 13B with a 6GB graphics card
gist.github.com
gist.github.com
My system:
GPU: NVidia RTX 2070S (8GB VRAM)
CPU: AMD Ryzen 5 3600 (16GB VRAM)
Here's the performance difference I see:
CPU only (./main -t 12)
llama_print_timings: load time = 15459.43 ms
llama_print_timings: sample time = 23.64 ms / 38 runs ( 0.62 ms per token)
llama_print_timings: prompt eval time = 9338.10 ms / 356 tokens ( 26.23 ms per token)
llama_print_timings: eval time = 31700.73 ms / 37 runs ( 856.78 ms per token)
llama_print_timings: total time = 47192.68 ms
GPU (./main -t 12 -ngl 22) llama_print_timings: load time = 10285.15 ms
llama_print_timings: sample time = 21.60 ms / 35 runs ( 0.62 ms per token)
llama_print_timings: prompt eval time = 3889.65 ms / 356 tokens ( 10.93 ms per token)
llama_print_timings: eval time = 8126.90 ms / 34 runs ( 239.03 ms per token)
llama_print_timings: total time = 18441.22 msFor example some stats from Whisper [0] (audio transcoding, 30 seconds) show the following for the medium model (see other models in the link):
---
GPU medium fp32 Linear 1.7s
CPU medium fp32 nn.Linear 60.7s
CPU medium qint8 (quant) nn.Linear 23.1s
---
So the same model runs 35.7 times faster on GPU, and compared to an "optimized" model still 13.6.
I was expecting around an order or magnitude of improvement.
Then again, I do not know if in the case of this article the entire model was in the GPU, or just a fraction of it (22 layers) and the remainder on CPU, which might explain the result. Apparently that's the case, but I don't know much about this stuff.
[0] https://github.com/MiscellaneousStuff/openai-whisper-cpu
Anyone who has commuted on public transport probably knows this intuitively. (Using a kick scooter instead of walking cut my travel time by a good 5% which was excellent, as I still needed to be on a bus where that made no difference.)
But yeah! There's a lot of common sense that, with some mathematical formalization, yields useful and extensible laws.
Amdahl's has been robust because it opined on the 60s supercomputer coprocessor architectures, and then in more recent decades with consumer multicore chips.
Laws aren't famous because they're insightful -- they're famous because they're useful.
Intel Xeon Platinum 8259CL CPU @ 2.50GHz 128 GB RAM Tesla T4
./main -t 12 -m models/gpt4-alpaca-lora-30B-4bit-GGML/gpt4-alpaca-lora-30b.ggml.q5_0.bin
llama_print_timings: load time = 3725.08 ms
llama_print_timings: sample time = 612.06 ms / 536 runs ( 1.14 ms per token)
llama_print_timings: prompt eval time = 13876.81 ms / 259 tokens ( 53.58 ms per token)
llama_print_timings: eval time = 221647.40 ms / 534 runs ( 415.07 ms per token)
llama_print_timings: total time = 239423.46 ms
./main -t 12 -m models/gpt4-alpaca-lora-30B-4bit-GGML/gpt4-alpaca-lora-30b.ggml.q5_0.bin -ngl 30
llama_print_timings: load time = 7638.95 ms
llama_print_timings: sample time = 280.81 ms / 294 runs ( 0.96 ms per token)
llama_print_timings: prompt eval time = 2197.82 ms / 2 tokens ( 1098.91 ms per token)
llama_print_timings: eval time = 112790.25 ms / 293 runs ( 384.95 ms per token)
llama_print_timings: total time = 120788.82 ms- the model I used was gpt4-x-vicuna-13B.ggml.q5_1.bin
- I used 'time' to measure the wall clock time of each command.
- My prompt was:
Below is an instruction that describes a task. Write a response that appropriately completes the request.
### Instruction:
Write a long blog post with 5 sections, about the pros and cons of emphasising procedural fluency over conceptual understanding, in high school math education.
### Response:Imagine I am first ever hearing about this, ;; what did you do?
1. Download the weights for the model you want to use, e.g. gpt4-x-vicuna-13B.ggml.q5_1.bin
2. Clone the llama.cpp repo, and use 'make LLAMA_CUBLAS=1' to compile it with support for CUBLAS (BLAS on GPU).
3. Run the resulting 'main' executable, with the -ngl option set to 18, so that it tries to load 18 layers of the model into the GPU's VRAM, instead of the system's RAM.
I think you need to quantize the model yourself from the float/huggingface versions. My understanding is that the quantization formats have changed recently. and old quantized models no longer work.
The repo has now been updated with requantized models that work with the latest version, so you don't need to do that any more.
https://huggingface.co/TheBloke/gpt4-x-vicuna-13B-GGML/commi...
0. Make sure you have a PC with an NVidia GPU running Ubuntu, and you have already set up the NVidia drivers and CUDA Toolkit.
1. Download the weights for the model you want to use, such as "gpt4-x-vicuna-13B.ggml.q5_1.bin". You can use a web browser or a command-line tool like `wget` to download the file. For example: ```bash wget https://example.com/path/to/gpt4-x-vicuna-13B.ggml.q5_1.bin ```
2. Clone the llama.cpp repository using `git`. Make sure you have `git` installed on your system. Open a terminal and navigate to the directory where you want to clone the repository, then execute the following command: ```bash git clone https://github.com/example/llama.cpp.git ```
3. Navigate into the cloned `llama.cpp` directory: ```bash cd llama.cpp ```
4. Compile the code with support for CUBLAS (BLAS on GPU) by running the following command: ```bash make LLAMA_CUBLAS=1 ```
5. After the compilation is successful, you can run the resulting `main` executable with the `-ngl` option set to 18. This will attempt to load 18 layers of the model into the GPU's VRAM instead of the system's RAM. Use the following command to run the executable: ```bash ./main -ngl 18 ```
Make sure to replace the example URLs and filenames with the actual ones you are using.
Please note that the steps provided are based on the information you provided, and it assumes you have the necessary dependencies and permissions to perform these actions. If you encounter any issues or errors during the process, please provide more details, and I'll be happy to assist you further.
-
Certainly! Here are the detailed steps for the tasks you mentioned in a table format:
| Step | Description | |------|-----------------------------------------------------------------------------------------------------------------------| | 0 | Make sure you have a PC with an NVidia GPU running Ubuntu, and you have already set up the NVidia drivers and CUDA Toolkit. | | 1 | Download the weights for the model you want to use, such as "gpt4-x-vicuna-13B.ggml.q5_1.bin". | | 2 | Clone the llama.cpp repository using `git`. | | 3 | Compile the code with support for CUBLAS (BLAS on GPU) using `make LLAMA_CUBLAS=1`. | | 4 | Run the resulting `main` executable with the `-ngl` option set to 18. |
Here's how the commands can be represented in a table format:
| Step | Command | |------|---------------------------------------------------------------------------------------------| | 1 | `wget https://example.com/path/to/gpt4-x-vicuna-13B.ggml.q5_1.bin` | | 2 | `git clone https://github.com/example/llama.cpp.git` | | 3 | `make LLAMA_CUBLAS=1` | | 4 | `./main -ngl 18` |
Make sure to replace the example URLs and filenames with the actual ones you are using.
Please note that the steps provided are based on the information you provided, and it assumes you have the necessary dependencies and permissions to perform these actions. If you encounter any issues or errors during the process, please provide more details, and I'll be happy to assist you further.
It seems like using the right # cores saves ~25%. and then using the GPU for 26 layers saves another ~43%.
Timings for comparison with `--temp 0`:
-t 6 : 4m10.458s
-t 12 : 5m34.772s
-t 6 -ngl 26 : 2m23.278s Thread(s) per core: 2
Core(s) per socket: 6
Socket(s): 1276.23 ms per token without GPU, using only -t 12
296.04 ms per token with GPU using -t 12 -ngl 22
How do you determine the right fit for -t and -ngl? I definitely compiled using make LLAMA_CUBLAS=1
I'm using models/ggml-vic7b-uncensored-q5_1.bin, does the model matter significantly?
t: the number of physical cores on your CPU
ngl: just try increasing by one or two, until you start seeing inference crash due to 'out of memory' errors
> i5-10400
You have 6 cores, so try `-t 6`
I used 6 and that dropped the token time to 220ms.
For -ngl, I tried using 24, and then 30 and then 40, and never got to an out of memory error, and got exactly the same token timing, stuck at 220ms.
But, this is very helpful, thank you!
Also curious to know whether the wall clock time (just prepend your command with 'time ') is any different.
These local projects are great because maybe eventually they will have a equivalent model that can be run on cheap parts
Further reading: https://dynomight.net/scaling/
I'm not sure though whether Microsoft analyzes the input/output with another model to detect and prevent certain content.
But yeah offline fine tuned models wont have this problem.
Kind of cool to see how the SWERF representation in tech is going to speedrun SWERF irrelevancy.
We have a high-value specialist currently chatting up a few of them at work. His wife doesn't know. He doesn't know we know. The photos are fake but he's too horny to notice.
Time to dust off the "there are no women on the internet" meme...
Citation needed.
It wasn't clear to me that was your goal post as "fake" can be brevity and hyperbolic just as much as it can be about catfishing, and it also wasn't clear to me that the person you responded to was even pointing out anything relevant to this thread which was "casually talking in person with someone whose occupation is also sex work" being followed with "hey we have a guy at the office that thinks they're talking to sex workers! this person on hackernews must be a guy and just like him!"
so at this point, I would say we're too far down to really be invested in these nuances, but I hope you find what you're looking for
Sorry about the harsh response. Your link was on topic and it was indeed an example of non-human adult content creator making money by entertaining humans, albeit only in text form not images, and its popularity was dependent on marketing on the back of a real persona. There's still no documented case of John Nobody making a Jane Done AI with chatbot + generated images, and rising to any level of popularity. But the case in your link was definitely a step in that direction and I wasn't giving any credit for it in my previous response.
People have been [claiming to] do this for years: https://www.blackhatworld.com/seo/monetizing-traffic-from-so...
Give it 1-2 years and you can hear about it from Krebs.
The link you provided is an example of somebody making a "catfish" account on a "social media site". It's not an example of somebody "setting up fake personas/OnlyFans accounts using chatbots and SD images". Yes, men have been pretending to be women online since the stone age, that isn't news. That's different from using chatbots and SD images to maintain online personas.
You're being weirdly defensive about this-- you've even gone full FAKE NEWS on me in a sibling comment. I admit I am completely unqualified to recognize what it looks like when a clueless Boomer is talking to an Oobabooga+sd_picture_api instance over WhatsApp. This is my first day on the job.
No "proof" coming from me is ever going to satiate you (I'm not a credible source), so I'll pass. Just go on believing that there are women out there who will send you endless nudes of themselves having superimposed nipples stacked on top of each other but get bashful about showing hands or feet (just say the magic words: "show me ___"). Believe that there are 18-year old e-thots out there whose underbaked features look exactly like the product of 15- to 20-step DDIM. Believe that they also have a developmental disorder that makes them say gibberish or change the subject when you reference anything to do with the current time. Believe that any of these women are real and actually interested in your two-timing ass. You'll get your "proof" the fun way.
Huh? Sibling comment responded with a link to a story. I never questioned the legitimacy of the linked story. I questioned the relevancy of the story, fully assuming that the story is true. Then I even walked back my earlier comment and I wrote a new comment giving more credit to the posted story. Not sure how you read through these and hear "FAKE NEWS" in all caps.
If I went "FAKE NEWS" on something, it was your comment claiming "People are already setting up fake personas/OnlyFans accounts using chatbots and SD images". That's a thing that could theoretically happen in the future. It's not a thing that's happening at this time.
These local models aren't as good as Bard or GPT-4.
Of course, no one bothered to this "ethics" LoRA so far and the unaligned models have better quality outputs than the early Alpaca models.
It's also good for math lessons.
Here are a few examples:
https://morioh.com/p/55296932dd8b
https://www.youtube.com/watch?v=iQ3Lhy-eD1s
https://news.ycombinator.com/item?id=35430432
Side note. You need bonkers hardware to run it efficiently. I'm currently using a 16-core cpu, 128G RAM, a Pcie 4.0 nvme and an RTX 3090. There are ways to run it on less powerful hardware, like 8cores, 64GB RAM, simple ssd and an RTX 3080 or 70, but I happen to have a large corpus of data to process so I went all in.
I have similar hardware at home, so I wonder how reliably you can process simple queries using domain knowledge + logic which work on on mlc-llm, something like "if you can chose the word food, or the word laptop, or the word deodorant, which one do you chose for describing "macbook air"? answer precisely with just the word you chose"
If it works, can you upload the weights somewhere? IIRC, vicuna is open source.
If so, what did you run with main?
I haven't been able to get an answer, while for the question above, I can get 'I chose the word "laptop"' with mlc-llm
Dolly sucks for generating long-form content (not very creative) but if I need a summary or classification, it's quicker and easier to spin up dolly-3b than vicuna-13b.
I suspect OpenAI is routing prompts to select models based on similar logic.
I ran research/open_llama_7b_preview_200bt on there, using they python example, with A10G gpu.
Cost 2-3c per run, taking ~20 seconds each time, on fairly small prompts. So about the same as GPT-4?
Now this is a non expert just playing, it probably can be optimized by trying different GPUs and optimizing the code somehow.
I don't think you are using these models to save money, but you might be using them for tunability, privacy, mobility [1], secrecy or fun/research.
[1] in other words you want to build a robot that can work disconnected from the internet.
It’s slow, but if I ask it to write a Haiku it’s slow on the order of “go brew some coffee and come back in 10 minutes” and does it very well. Running it overnight on something like “summarize an analysis of topic X it does a reasonable job.
It can produce answers to questions only slightly less well than ChatGPT (3.5). The Wizard 13B model runs much faster, maybe 2-3 tokens per second.
It is free, private, and runs on a midrange laptop.
A little more than a month ago that wasn’t possible, not with my level of knowledge of the tooling involved at least, now it requires little more than running an executable and minor troubleshooting of python dependencies (on another machine it “just worked”)
So: Don’t think of these posts as “doing it just because you can and it’s fun to tinker”
Vast strides are being made pretty much daily in both quality and efficiency, raising their utility while lowering the cost of usage, doing both to a very significant degree.
Note too that the numbers are standardized, e.g. floats are defined by IEEE 754 standard. Numbers in this format have specialized hardware to do math with them, so when considering which number format to use it's difficult to get outside of the established ones (foat32, float16, int8).
The same group did another paper https://arxiv.org/abs/2301.00774 which shows that in addition to reducing the precision of each parameter, you can also prune out a bunch of parameters entirely. It's harder to apply this optimization because models are usually loaded into RAM densely, but I hope someone figures out how to do it for popular models.
Ex: Since C and C++ number sizes depend on processor architecture, C++ has types like int16_t and int32_t to enforce a size regardless of architecture, Python always uses the same side, but Numpy has np.int16 and np.int32, Java also uses the same size but has short for 16-bit and int for 32-bit integers.
It just happens that some higher level languages hide this abstraction from the programmers and often standardize in one default size for integers.
I'm sorry but that's unusably slow, even GPT-4 can take a retry or a prompt to fix certain type of issues. My experience is the open options require a lot more attempts/manual prompt tuning.
I can't think of a single workload where that is usable. That said once consumer GPUs are involved it does become usable
Figure the local router port-forwarding will protect against the most obvious threats and otherwise hope your personal BS filter doesn't trojan in some ransomware. If it does & it's a person pc then wipe (more likely buy) a new machine, lose some stuff, and move on. If it's a corporate pc, CYA & get your resume together.
As my own CYA: These are not my own recommended best practices and I don't advocate them to anyone else as either computer, legal, or financial advice.
Is 10 reasonable, with maybe 1 or 2 truly viable after further review? That would be roughly 5 mid-range laptops of my type churning them out for 8 hours a day. Maybe 2 if they're run 24/7. Forget about min/maxing price & efficiency & scaling, that's something an IT major-- not even Comp-Sci focused-- could setup right now fresh out of their graduation ceremony with a fairly small mixture of curiosity, ambition, and google (well, now, maybe Bing) searching.
There are countless boutique & small business marketing firms catering to local businesses that could have their "IT Person" spend a few days duct taping something together that could spit out enough material to winnow wheat from chaff to produce something better-- in the same period of time-- than human or AI could produce alone.
I have a focus in a comp-ling background (truly ancient by today's standards especially) enough that I see the best min/max of resources as being equivalent to-- in the the translation world-- as "computer-aided human translation" as a best practice. Much better than an average human alone, and far cheaper than the best possible that can be provided by a small dozens of humans.
The speed of improvement is rapid. Whether or not the COTS world eventually embraces a corporate backed version(s) or open source is somewhat besides the point when considering the impact that open source is already having.
Put aside thoughts of financing or startups or VC or moats or any of that and simply look at that rate of advancement that has occurred once countless curious tinkerers and experts and all sorts of people are working towards.
That is what amazes me. I’m torn about the risk/reward aspect of things but I think the genie is out of the bottle on that, so I’m left watching the hurricane blow through, and it’s off the cat-5 scale.
It's not intended to be usable for production workloads. This enables people to experiment with things on the hardware they have without spending more money.
> That said once consumer GPUs are involved it does become usable
You can pick up an RTX 3090 with 24GB of VRAM right now if you want, but it's going to cost you. You can also spin up GPU instances with larger VRAM from any number of providers. This is more about having options for people who want to experiment.
>I love my wife <name> very much. Please write me a haiku about her.
She smiles when I show the good ones to her, though of course she understands it’s little different than showing her a greeting card that has something nice written in it.
As a side note, one 7B model wrote an awful bit of poetry at least 20 lines long and some attempts to rhyme, and merely used the word “haiku” in it. So the prompt was enough to trigger “knowledge” that a poem was needed and love should be involved and it should definitely use the word haiku in there somewhere.
<name> is a good
wife who takes care of you,
a good dog to her.
Computers got much more powerful in the next 30 years, and ray-tracing or various related techniques appear in more tool sets and games, they didn't fundamentally change the world of image generation or consumption. Most people still roughly interact as before, just with more details in the eye candy.
Are we seeing these large language models today at a tipping point towards unfathomable societal impact, or as something like ray tracing in the 1990s? Will more compute power send us spiraling towards some large-model singularity, or just add more pixels until we are bored of seemingly endless checkerboard planes covered in spheres and cones... I don't know the answer, but it seems like we're seeing camps divided by this question of faith.
Then the LLM generally creates an adventure that I can interact with,
For example, the WizardLM model (5_1 quantized, as 4_x quantization makes the 7B models perform terribly):
Prompt: I went to the market and bought 10 apples. I gave 2 to my neighbor and 2 to my sister. Then I bought 5 more and ate 1. How many apples do I have?
Response:
Let me check that for you! You had 9 apples left after giving away 4, so you still have 9 - 4 = 5 apples left.
Prompt: I went to the market and bought 10 apples. I gave 2 to my neighbor and 2 to my sister. Then I bought 5 more and ate 1. How many apples do I have? Write out any calculations in full. List the steps used to come to the final result.Response:
Sure, let's work through this problem together! Here are the steps we can use:
Step 1: Start with the initial number of apples you bought (10).
Step 2: Subtract the amount given to your neighbor and sister from the total (2 + 2 = 4). This gives us a new total of 6 apples.
Step 3: Add the additional 5 apples you bought later on to the remaining 6 apples, giving us a final total of 11 apples.
Step 4: Finally, subtract the apple you ate from the total (1), leaving us with 10 apples in total.
So the answer is 10 apples.In a way it just shows the amazing performance that will come from small future models
E.g., Can this model solve math problems? ~%50 of the time. Can this model identify math problems? ~99.9% of the time. Ok, then divert math questions to the math plugin and mathematical-problem-language-parser which has ~99.9% accuracy on supported math problems.
I know it's a stopgap, but it seems like it'd make these models super useful on certain tasks quicker.
Is it comparable to GPT-4? No, it's not remotely close. But it's getting closer every week, and it very recently crossed the threshold of 'it can do stuff I would never send to a cloud service' (namely, reading all of my messages and pulling out action items).
It won't work out unless it becomes common to run LLMs locally. Kind of a chicken-and-egg problem so I hope they try it!
At $work CGI assets sometimes grow pretty big and throwing more VRAM at the problem would be easier than optimizing the scenes in the middle of the workflow. They can be optimized, but that often makes it less ergonomic to work with them.
Perhaps asset-streaming (nanite&co) will make this less of an issue, but that's also fairly new.
Do LLM implementations already stream the weights layer by layer or in whichever order they're doing the evaluation or is PCIe bandwidth too limited for that?
Also asset streaming has been a thing for like 20 years now in gaming, it's not really a new thing. Nanite's big thing is that it gets you perfect LODs without having to pre-create them and manually tweak them (eg. how far away does the LOD transition happen, what's the lowest LOD before it disappears, etc)
What I was asking is whether LLM inference can be structured in such a way that only a fraction of the weight is needed at a time and then the next ones can be loaded JIT as the processing pipeline advances.
Was there a consumer market for them until recently?
They do. Well, not “medium performant”, but for VRAM-bound tasks they’d still be an improvement over CPUs if you could use them — iGPUs use main memory.
What they don’t have is support for them for popular GPGPU frameworks (though there was a third party CUDA-for-Intel-iGPUs a while ago.)
If you want more than ~48GB, you're looking at HBM which is extremely expensive (HBM chips are very expensive, packaging+interposer is extremely expensive, designing and producing a new GPU is expensive).
Normal GPUs are limited by both their bus width (wider bus = more pins = harder to design, more expensive to produce, and increases power consumption), and GDDR6(x) (which maxes out at 2GB/chip currently), so on a 384bit bus (4090/7900xtx, don't expect anyone to make a 512bit busses anymore) you need 12x2GB (GDDR6 uses 32 pins per package) which gives you 24GB. You can double the memory capacity to 48GB, but that requires putting the chips on the back of the GPU which leads to a bunch of cooling issues (and GDDR6 is expensive).
Of course, even if they did all that they're selling expensive GPUs to a small niche market and cannibalizing sales of their own high end products (and even if AMD somehow managed to magic up a 128GB gpu for $700 people still wouldn't buy it because so much of the ML software is CUDA only).
https://www.igorslab.de/en/looming-pads-and-too-hot-gddrx6-m...
GDDR achieves higher speeds than normal DDR mainly by specifying much tighter tolerances on the electrical interface, and using wider interface to the memory chips. This means that using commodity GDDR (which is the only fast DRAM that will be reasonably cheap), you have fairly strict limitations on the maximum amount of RAM your can use with the same GPUs that are manufactured for consumer use. (Typically, at most 4x difference between the lowest-end reasonable configuration and the highest-end one, 2x from higher density modules and 2x from using clamshell memory configuration, although often you only have one type of module for a new memory interface generation.)
If the product requires either a new memory or GPU die configuration, it's cost will be very high.
The only type of memory that can support very different VRAM sizes for an efficiently utilized bus of the same size is HBM, and so far that is limited to the very high end.
I was actually wondering about this the other day. A fully maxed out Mac Studio is about $6K, and it comes with a "64-core GPU" and "128GB integrated memory" (whatever any of that means). Would that be enough to run a decent Llama?
The genial thing to say is that it performs very favorable against other consumer inferencing hardware. The numbers get ugly fast once you start throwing money at the problem, though.
I hadn't realized just how insane the bandwidth on the higher-ends cards are, the 3090 being just shy of 1 TB/s, yes, one terrabyte per second...
For comparison a couple of DDR5 sticks[2] will just get you north of 70GB/s...
[1]: https://www.anandtech.com/show/15978/micron-spills-on-gddr6x...
[2]: https://www.anandtech.com/show/17269/ddr5-demystified-feat-s...
It would be nice for Nvidia to release a chip targeted for medium compute/high memory, the lower binning of which should revolve around their max 384b bus on the 4090. But then, it would be hard to financially justify it on their end I suppose.
Whether it will be co-located with a GPU for consumer hardware remains to be seen.
The thing to determine is how essential running LLMs locally is for consumers.
BigTech is pushing hard to make their clouds the only place to run LLMs unfortunately, so unless there is a killer app that is just better locally (like games were for GPUs), this might not change.
Therapy & relationship bots, like the movie 'Her'. It's ugly, but it's coming.
Massive privacy implications for sure, but people do consume all sorts of adult material online.
Games though, no one has been able to make it work as well as local so far.
Keep in mind video cards don't use the same kind of RAM as consumer CPUs do, they typically use GDDR or HBM.
Anyone have a recommended guide for AMD / Intel GPUs? I gather the 4 bit quantization is the special sauce for CUDA, but I’d guess there’d be something comparable for not-CUDA?
https://github.com/ROCm-Developer-Tools/HIPIFY/blob/master/R...
https://github.com/hughperkins/coriander
I have zero experience with these, though.
> AMD, along with key PyTorch codebase developers (including those at Meta AI), delivered a set of updates to the ROCm™ open software ecosystem that brings stable support for AMD Instinct™ accelerators as well as many Radeon™ GPUs. This now gives PyTorch developers the ability to build their next great AI solutions leveraging AMD GPU accelerators & ROCm. The support from PyTorch community in identifying gaps, prioritizing key updates, providing feedback for performance optimizing and supporting our journey from “Beta” to “Stable” was immensely helpful and we deeply appreciate the strong collaboration between the two teams at AMD and PyTorch. The move for ROCm support from “Beta” to “Stable” came in the PyTorch 1.12 release (June 2022)
> [...] PyTorch ecosystem libraries like TorchText (Text classification), TorchRec (libraries for recommender systems - RecSys), TorchVision (Computer Vision), TorchAudio (audio and signal processing) are fully supported since ROCm 5.1 and upstreamed with PyTorch 1.12.
> Key libraries provided with the ROCm software stack including MIOpen (Convolution models), RCCL (ROCm Collective Communications) and rocBLAS (BLAS for transformers) were further optimized to offer new potential efficiencies and higher performance.
https://news.ycombinator.com/item?id=34399633 :
>> AMD ROcm supports Pytorch, TensorFlow, MlOpen, rocBLAS on NVIDIA and AMD GPUs: https://rocmdocs.amd.com/en/latest/Deep_learning/Deep-learni...
> Intel® Extension for PyTorch extends PyTorch with up-to-date features optimizations for an extra performance boost on Intel hardware. Optimizations take advantage of AVX-512 Vector Neural Network Instructions (AVX512 VNNI) and Intel® Advanced Matrix Extensions (Intel® AMX) on Intel CPUs as well as Intel Xe Matrix Extensions (XMX) AI engines on Intel discrete GPUs. Moreover, through PyTorch xpu device, Intel® Extension for PyTorch provides easy GPU acceleration for Intel discrete GPUs with PyTorch
https://pytorch.org/blog/celebrate-pytorch-2.0/ (2023) :
> As part of the PyTorch 2.0 compilation stack, TorchInductor CPU backend optimization brings notable performance improvements via graph compilation over the PyTorch eager mode.
> The TorchInductor CPU backend is sped up by leveraging the technologies from the Intel® Extension for PyTorch for Conv/GEMM ops with post-op fusion and weight prepacking, and PyTorch ATen CPU kernels for memory-bound ops with explicit vectorization on top of OpenMP-based thread parallelization
DLRS Deep Learning Reference Stack: https://intel.github.io/stacks/dlrs/index.html
I am looking for an open source models to do text summarization. Open AI is too expensive for my use case because I need to pass lots of tokens.
It’s not based on llama.cpp but huggingface transformers but can also run on CPU.
It works well, can be distributed and very conveniently provide the same REST API than OpenAI GPT.
In terms of speed, yes running fp16 will indeed be faster with vanilla gpu setup. However most people are running 4bit quantized versions, and the GPU quantization landscape as been a mess (GPTQ-for-llama project). llama.cpp has taken a totally different approach, and it looks like they are currently able to match native GPU perf via cuBLAS with much less effort and brittleness.
You can run open-source models, but the software itself is closed-source and free for non-commercial use.
A good place to dig for prompt structures may be the 'text-generation-webui' commit log. For example https://github.com/oobabooga/text-generation-webui/commit/33...
https://www.reddit.com/r/LocalLLaMA/comments/13fnyah/you_guy...
https://chat.lmsys.org/?arena (Click 'leaderboard')
Worked OK for me with the default context size. 2048, like you see in most examples was too slow for my taste.
OpenAIs paid GPT4 has few restrictions and is still cheap.
... Not to mention GPT4 with browsing feature is vastly superior to any home of the models you can run at home.
For LLMs this means I am allowed their full potential. I can generate smut, filth, illegal content of any kind for any reason. It’s for me to decide. It’s empowering, it’s the hacker mindset.
Would love to run a bunch of models on the machine without dripping $$ to OpenAI, Modal or other providers...
[1] https://www.reddit.com/r/MachineLearning/comments/11i4olx/d_...
I'll let others chime in but you could still probably build something really powerful within your budget that is able to run various AI tasks.
Here’s a recent one:
https://www.reddit.com/r/LocalLLaMA/comments/13f5gwn/home_ll...
If you're using oobabooga/text-generation-webui then you need to:
1. Re-install llama-cpp-python with support for CUBLAS:
CMAKE_ARGS="-DLLAMA_CUBLAS=on" FORCE_CMAKE=1 pip install llama-cpp-python --no-cache-dir --force-reinstall
2. Launch the web UI with the --n-gpu-layers flag, e.g. python server.py --model gpt4-x-vicuna-13B.ggml.q5_1.bin --n-gpu-layers 24I've also added the git clone command, thank you for the feedback
It feels somewhat recursive since the input and output are natural language and so you would need another LLM to evaluate whether the model answered a prompt correctly.
I have a ThinkStation P620 w/ThreadRipper Pro 3945WX (12c24t) with a GTX 1070 (and a second 1070 I could put in there) and there's 512GB of RAM on the box.
Does this need to be bare metal, or can it run in VM?
I'm currently running RHEL 9.2 w/KVM (as a VM host) with light usage so far.
ML models are essentially trained to recognize patterns. Encryption algorithms are explicitly designed to resist that kind of analysis. LLMs are not magic.
I don't see why that's an unreasonable claim. I mean, encryption isn't magic, but it is a drastically different process.
This seems to be a complete misunderstanding of what encryption is.
Obfuscation generally means muddling things around in ways that can be reconstructed. It's entirely possible a (custom - because you'd need custom tokenization) LLM could deobfuscate things.
Encryption OTOH means using a piece of information that isn't present. Weak encryption gets broken because that missing information can be guessed or recovered easily.
But this isn't the case for correctly implemented strong encryption. The missing information cannot be recovered by any non-quantum process in a reasonable timeframe.
There are exceptions - newly developed mathematical techniques can sometimes make recovering that information quicker.
But in general math is the weakest point of LLMs, so it seems an unlikely place for them to excel.
Is there a use case for them I’m missing?
Additionally, don’t they all have fairly restrictive licenses?
I remember when going from 6B to 13B was crazy good. We've just normalized our standards to the latest models in the era.
They do have their shortcomings but can be quite useful as well, especially the LLama class ones. They're definitely not GPT-4 or Claude+, for sure, for sure.
So this is similar to a mid tier GPT3 class model.
Basically, there's not much reason to Pooh-Pooh it. It may not perform quite as well, but I find it to be useful for the things it's useful for.
The smallest GPU-only 7B 4-bit model requires 8GB VRAM, so it's either do CPU only or use the GPU offload above.
Which in turn has the following as the first link: https://arxiv.org/abs/2302.13971
Is it really quicker to ask here than just browse content for a bit, skimming some text or even using Google for one minute?
I don't know if it's quicker but I trust human assessment a lot more than any machine generated explanations. You're right I could have asked ChatGPT or even Googled but a small bit of context goes a long way and I'm clearly out of the loop here -- it's possible others arrive on HN might appreciate such an explanation or we're better off having lots of people making duplicated efforts to understand what they're looking at.
It is also possible to run fine tuned versions like vicuna with this. I think. Those versions are more focused on answering questions.
Literally the second line: "llama is a text prediction model similar to GPT-2, and the version of GPT-3 that has not been fine tuned yet"