Nvidia's Chat with RTX is an AI chatbot that runs locally on your PC
theverge.com
theverge.com
Previously there was TensorRT for Stable Diffusion[1], which provided pretty drastic performance improvements[2] at the cost of customisation. I don't forsee this being as big of a problem with LLMs as they are used "as is" and augmented with RAG or prompting techniques.
[1]: https://github.com/NVIDIA/Stable-Diffusion-WebUI-TensorRT [2]: https://reddit.com/r/StableDiffusion/comments/17bj6ol/hows_y...
https://github.com/NVIDIA/trt-llm-rag-windows https://github.com/NVIDIA/TensorRT-LLM
It's quite a thin wrapper around putting both projects into %LocalAppData%, along with a miniconda environment with the correct dependnancies installed. Also for some reason the LLaMA 13b (24.5GB) and Ministral 7b (13.6GB) but only installed Ministral?
Ministral 7b runs about as accurate as I remeber, but responses are faster than I can read. This seems at the cost of context and variance/temperature - although it's a chat interface the implementation doesn't seem to take into account previous questions or answers. Asking it the same question also gives the same answer.
The RAG (llamaindex) is okay, but a little suspect. The installation comes with a default folder dataset, containing text files of nvidia marketing materials. When I tried asking questions about the files, it often cites the wrong file even if it gave the right answer.
I’ve been working with it for a while and it’s… Rough.
That said it is extremely fast. With TensorRT-LLM and Triton Inference Server with conservation performance settings I get roughly 175 tokens/s on an RTX 4090 with Mistral-Instruct 7B. Following commits, PRs, etc I expect this to increase significantly in the future.
I’m actually working on a project to better package Triton and TensorRT-LLM and make it “name and model and press enter” level usable with support for embeddings models, Whisper, etc.
But the HW requirements state 8GB of VRAM. How do those models fit in that?
If it means that weights for an LLM can be 4 bits well that's just mind boggling.
I was skeptical of it for some time, but it seems to work because individual parameters don’t encode much information. The knowledge is embedded thanks to having a massive number of low bit parameters.
It did not go well. The generation gap was perhaps even starker than it is between real people.
https://forums.overclockers.com.au/threads/chatgpt-vs-dr-sba...
Ok, tell me your problem $name.
"I'm sad."
Do you enjoy being sad?
"No"
Are you sure?
"Yes"
That should solve your problem. Lets move on to discuss about some other things.
Also wow that creative app/sound/etc brings back memories.
I would really want it to have that Dr. Sbaitso voice, though, telling me how to be fitter, happier, more productive.
https://www.pandorabots.com/pandora/talk?botid=b8d616e35e36e...
Like other bogus things like tarot or horoscopes, it's amazing what you can discover when you talk about something, it asks you questions, and what you want or desire eventually floats to the surface. And now people are even more lonely...
>Human: do you like video games
>A.L.I.C.E: Not really, but I like to play the Turing Game.
Who is the target audience for this solution?
Llama.cpp is much slower, and does not have built-in RAG.
TRT-LLM is a finicky deployment grade framework, and TBH having it packaged into a one click install with llama index is very cool. The RAG in particular is beyond what most local LLM UIs do out-of-the-box.
No, it answers questions from the documents you provide. Off the shelf local LLMs don't do this by default. You need a RAG stack on top of it or fine tune with your own content.
> Are LLM tools better or worse than e.g. meilisearch or elasticsearch for searching with snippets over a set of document resources?
> How does search compare to generating things with citations?
pdfGPT: https://github.com/bhaskatripathi/pdfGPT :
> PDF GPT allows you to chat with the contents of your PDF file by using GPT capabilities.
GH "pdfgpt" topic: https://github.com/topics/pdfgpt
knowledge_gpt: https://github.com/mmz-001/knowledge_gpt
From https://news.ycombinator.com/item?id=39112014 : paperai
neuml/paperai: https://github.com/neuml/paperai :
> Semantic search and workflows for medical/scientific papers
RAG: https://news.ycombinator.com/item?id=38370452
Google Desktop (2004-2011): https://en.wikipedia.org/wiki/Google_Desktop :
> Google Desktop was a computer program with desktop search capabilities, created by Google for Linux, Apple Mac OS X, and Microsoft Windows systems. It allowed text searches of a user's email messages, computer files, music, photos, chats, Web pages viewed, and the ability to display "Google Gadgets" on the user's desktop in a Sidebar
GNOME/tracker-miners: https://gitlab.gnome.org/GNOME/tracker-miners
src/miners/fs: https://gitlab.gnome.org/GNOME/tracker-miners/-/tree/master/...
SPARQL + SQLite: https://gitlab.gnome.org/GNOME/tracker-miners/-/blob/master/...
https://news.ycombinator.com/item?id=38355385 : LocalAI, braintrust-proxy; promptfoo, chainforge, mixtral
Absolutely worse, LLM are not made for it at all.
And perhaps they will add more models in the future?
From my point of view the only person who would be likely to use this would be the small slice of people who are willing to purchase an expensive GPU, know enough about LLMs to not want to use CoPilot, but don’t know enough about them to know of the already existing solutions.
It's an Nvidia "product", published and promoted via their usual channels. This is co-sign/official support from Nvidia vs "Here's an obscure name from a dizzying array of indistinguishable implementations pointing to some random open source project website and Github repo where your eyes will glaze over in seconds".
Completely different but wider and significantly less sophisticated audience. The story link is on The Verge and because this is Nvidia it will also get immediately featured in every other tech publication, website, subreddit, forum, twitter account, youtube channel, etc.
This will get more installs and usage in the next 72 hours than the entire Llama/open LLM ecosystem has had in its history.
I suppose my counter point is only that the user base that relies on simplified solutions is largely already addressed with the wide number of cloud offerings from OpenAi, Microsoft, Google, whatever other random company has popped up. Realistically I don’t know if the people who don’t want to use those, but also don’t want to look at GitHub pages is really that wide of an audience.
You could be right though. I could be out of touch with reality on this one, and people will rush to use the latest software packaged by a well known vendor.
There is a wide spectrum of users for which a more white-labelled locally-runnable solution might be exactly what they're looking for. There's much more than just the two camps of "doesn't know what they're doing" and "technically inclined and knows exactly what to do" with LLMs.
Most users don't care whether whether the model is run, online or local. They go to ChatGPT or Bing/Copilot to get answers, as long as they are free. Well, if it becomes a (mandatory) subscription, they are more likely to pay for it rather than figure out how to run a local LLM.
Sounds like you are the one who's not getting the message.
So basically the only people who runs a local LLM are those who are interested enough in this. Any why would brand name matter? What matters is whether a model is good, whether it can run on a specific machine and how fast it is etc, and there are objectives for it. People who run local LLM don't automatically choose Nvidia's product over something just because nvidia is famous.
Have you ever tried to use ChatGPT alone to work with documents? In terms of the free/ready to use product it's very painful. Give it a URL to a PDF (or something) and assuming it can load it (often can't) you can "chat" with it. One document at a time...
This is for the (BIG) world of Nvidia Windows desktop users (most of whom are fanboys who will install anything Nvidia announces that sounds cool) who don't know what an LLM is. They certainly wouldn't know/have the inclination to wander into /r/LocalLLaMA or some place to try to sort through a bunch of random projects with obscure names that are peppered with jargon and references to various models they've also never heard of or know the difference between. Then the next issue is figuring out the RAG aspects, which is an entirely different challenge.
This is a Windows desktop installer that picks one of two models automatically depending on how much VRAM you have, loads them to run on your GPU using one of the fastest engines out there, and then allows you to load your own local content and interact with it in a UI that just pops up after you double-click the installer. It's green and peppered with Nvidia branding everywhere. They love it.
What the Nvidia Windows desktop users will be able to understand is "WOW, look it's using my own GPU for everything according to my process manager. I just made my own ChatGPT and can even chat with my own local documents. Nvidia is amazing!"
> why would brand name matter?
Do you know anything about humans? Brands make a HUGE difference.
> People who run local LLM don't automatically choose Nvidia's product over something just because nvidia is famous.
/r/LocalLLaMA is currently filled with people ranting and raving about this even though it's inferior (other than ease of use and brand halo) to much of the technology that has been discussed there since forever.
Again - humans spend many billions and billions of dollars choosing products that are inferior solely because of the name/brand.
I don't really care how many installs it gets, does it do anything differently or better?
What you're missing here is you're already in this area deep enough to know what ooogoababagababa text-generation-webui is. Let's back out to the "average Windows desktop user who knows they have an Nvidia card" level. Assuming they even know how to find it:
1) Go to https://github.com/oobabooga/text-generation-webui?tab=readm...
2) See a bunch of instructions opening a terminal window and running random batch/powershell scripts. Powershell, etc will likely prompt you with a scary warning. Then you start wondering who ooobabagagagaba is...
3) Assuming you get this far (many users won't even get to step 1) you're greeted with a web interface[0] FILLED to the brim with technical jargon and extremely overwhelming options just to get a model loaded, which is another mind warp because you get to try to select between a bunch of random models with no clear meaning and non-sensical/joke sounding names from someone called "TheBloke". Ok... Oh yeah, what's a "model"? GGUF? GPTQ? AWQ? Exllama? Prompt format? Transformers? Tokens? Temperature? Repeat for dozens of things you're familiar with but are meaningless to them.
Let's say you somehow braved this gauntlet and get this far now you get to chat with it. Ok, what about my local documents? text-generation-webui itself has nothing for that. Repeat this process over the 10 random open source projects from a bunch of names you've never heard of in an attempt to accomplish that.
This is "I saw this thing from Nvidia explode all over media, twitter, youtube, etc. I downloaded it from Nvidia, double-clicked, pointed it at a folder with documents, and it works".
That's the difference and it's very significant.
[0] - https://raw.githubusercontent.com/oobabooga/screenshots/main...
https://nvidia.github.io/TensorRT-LLM/performance.html https://github.com/lapp0/lm-inference-engines/
You think a regular user has any chance?
Codeword for people who have hardware specialized and suitable for AI.
a personal assistant to monitor everything i do on my machine, ingest it and answer question when i need.
it's not there yet (still need to manually input url, etc...) though but it's very much feasible.
Based on what? The CPU is a physical storage device on my PC but it still can phone home and has backdoors.
Is there any reason to think Nvidia isn't collecting my data?
If you're on Windows, what makes you think they are not already?
> what makes you think they are not already
That is the point
is the bash history command creepy?
Is your browsers history command creepy?
But those also don't try to reinterpret what I wrote.
Gaming LLM
Checks out
So, reviewing this...
- they are associating AI with RTX (ray tracing) now (??)
- your RTX card cannot chat with RTX (???)
wat
Although you’d realistically need 5-6 bit quantization to get anything large/usable enough running on a 12GB card. And I think it’s just CUDA then, so you should be able to use 2080 Ti.
It is largely an arbitrary generational limit
> I try to run windows 10 on it
> It doesn't work
> pff, Intel cpu cannot run OS meant for intel CPUs
wat
Jokes aside, nvidia been using RTX branding for products that use Tensor Cores for a long-time now. Limitation due to 1st gen tensor cores not supporting precisions required.
https://github.com/NVIDIA/TensorRT-LLM?tab=readme-ov-file#pr...
https://pcpartpicker.com/products/video-card/#c=552&sort=pri...
Chepest 8GB 30xx is $220
https://pcpartpicker.com/products/video-card/#sort=price&c=5...
Smells like artificial restriction to me. I have a 2080 Ti with 8GB of VRAM that is still perfectly fine for gaming. I play in 3440x1440 and the modern games need DLSS/FSR on quality for nice 60++ - 90 FPS. That is perfectly enough for me and I have not had a game, even UE5 games where I really thought I really NEED a new one. I bet that card is totally capable of running that chatbot.
They do the same with frame generation. There they even require you a 40 series card. That is ridiculous to me as these cards are so fast that you do not even need any frame generation. The slower cards are the ones that would benefit from it most so they just lock it down artificially to boost their sales.
Sure you don't mean 11GB[1]? Or did they make other variants? FWIW I have a 2080 Ti with 11GB, been considering upgrading but thinking I'll wait til 5xxx.
[1]: https://www.techpowerup.com/gpu-specs/geforce-rtx-2080-ti.c3...
My next card will be an AMD one. I like that they are open sourcing most of their stuff and I think they play better with Linux Wine/Proton. FSR 3 also not artificially restricts cards and runs even on the competition. I read today about at open source API that takes CUDA calls and runs them on AMD or everywhere. I am sure there will be some cool open source projects that do all kinds of things if I ever even need them.
It was one of the fastest backends last time I checked (with vLLM and lmdeploy being comparable), but the space moves fast. It uses cuda under the hood, torch is not relevant in this context.
NVIDIA GeForce™ RTX 30 or 40 Series GPU or NVIDIA RTX™ Ampere or Ada Generation GPU with at least 8GB of VRAM"
Since the demo is 35 GB, my first assumption was it's bundling a ~13B parameter model, but if the requirement is 8 GB VRAM, I assume they're either doing quantization on the user's end or offloading part of the model to the CPU.
(I also hope that Windows 11 is a suggested and not a hard requirement)
54% have at least that much VRAM
79% have at least that much RAM
44% are on Windows 11
36% have the required video card
Do you think people got an MI300 lying around?
AMD's GPUs simply weren't meant for GPGPU.
Oh and now the install failed with no error message. Lovely.
When asked "What is ChatWithRTX" the reply was:
Chat with RTX is not mentioned in the given context information. It is
possible that it is a separate event or feature that is related to the
NVIDIA Lounge at LTX 2023 or QuakeCon 2023, where attendees can interact
with a virtual assistant or chatbot powered by NVIDIA's GeForce RTX
technology. However, without more information, it is impossible to provide a
more detailed answer.
Reference files: portal-prelude-rtx-inside-the-game-interview.txt2010-2011, vibes.
Newer devices only barely have the juice.
I think it could run it, albeit slowly.
Is there some local chatbot application like this, for Windows, that isn't hell to set up and that is not censored?
You can use it to directly search any models, download and run 100% locally.
M-series Macs are the simplest, they Just Work. Even faster if you tick the GPU box.
Windows needs the right kind of GPU to get workable speed.
But AI models are already running on tablets (not necessarily on Nvidia hardware) and I expect some phone to ship with them within a year (maybe as a stunt, I guess it would be a few years more before this is practical).