3090s are consumer grade and common on gaming rigs. I’m hoping game devs start experimenting with locally deployed Mixtral in their games. e.g. something like CIV but with each leader powered via LLM
3090s are consumer grade and common on gaming rigs. I’m hoping game devs start experimenting with locally deployed Mixtral in their games. e.g. something like CIV but with each leader powered via LLM
An Apple M2 Pro with 32GB of RAM is in the same price range as a gaming PC with a 3090, but its another example of normal people with moderately high performance systems "accidentally" being able to run a GPT-3.5 comparable model.
If you have an Apple meeting these specs and want to play around, LLM Studio is open source and has made it really easy to get started: https://lmstudio.ai/
I hope to see a LOT more hobby hacking as a result of Mixtral and successors.
Llmstudio is, but I suspect that was a typo in their comment. https://github.com/TensorOpsAI/LLMStudio
Is there any speed/performance/quality/context size/etc. advantage to using LLM Studio or any of the other *llama tools that require more setup than downloading and running a single llamafile executable?
I tried using Ollama on my machine (same specs as above) and it told me I needed 49gb RAM minimum.
ollama run dolphin-mixtral:8x7b-v2.5-q3_K_SEDIT: I just checked, it runs great, thanks.
You can buy a whole PC for that, I refuse to believe that a GPU priced that highly is "consumer grade" and "common".
Are there any GPUs that are good for LLMs or other genAI that aren't absurdly priced? Or ones specifically designed for AI rather than gaming graphics?
https://store.steampowered.com/hwsurvey/Steam-Hardware-Softw...
3090 isn't even in the top 30.
Search the page for 3090 and see for yourself, it's on the list twice.
Apple is the one doing the best in terms of making consumer-friendly hardware that can perform AI/ML tasks...but that involves a different problem regarding video games.
In the case of high-end video games, that's unlikely.
The bigger problem is memory capacity and bandwidth, but I suspect folks will eventually figure out some sort of QoS setup to let the system crunch LLMs using otherwise unused/idle resources.
Running LLMs locally to create custom dialogue for games is still years away.
I had think about this, you need a small LLM for a game, you do not need it to know about movies, music bands, history, coding and all the text on the internet. I am thinking we need a small model similar to phi-2 , trained only on basic stuff then you would train it on the game world lore. Then the game would also use some "simpler" graphics( we had good looking game graphics decades ago so I do not think you could be limited to text adventures or 2D graphics, just you need some simpler and optimized graphics, it is always interesting when someone shows his unreal demo that uses most RAM/VRAM for a simple demo level then a giant game like GTA5)
Or even better look at something like phi-2. It's likely to go even lower. I'm sure there are people here who can detail more specifics.
Its 1/24 and we're already there more or less.
What do you mean? Most gamers do have an nvidia GPU.
Edit: unless you talk about mobile gamers, and not PC gamers?
https://store.steampowered.com/hwsurvey/Steam-Hardware-Softw...