Ollama is now available on Windows in preview
ollama.com
ollama.com
I can vouch for it as a solid frontend for Ollama. It works really well and has had an astounding pace of development. Every few weeks I pull the latest docker images and am always surprised by how much has improved.
[0] https://github.com/open-webui/open-webui/discussions/764
AMD, you can change, but you need to start NOW.
https://github.com/ollama/ollama/blob/main/docs/development....
I’m not blaming you for this, but I’m also sticking with nvidia.
The kernel panics though... Yeah, I had those on my Radeon vii before I upgraded.
But if you run a distro that has anywhere near new kernels such as Fedora and Arch, you'll be constantly in fear of receiving new kernel updates. And every so often the packages will be broken and you'll have to use Nvidia's horrible installer. Oh and every once in a while they'll subtly drop support for older cards and you'll need to move to the legacy package, but the way you'll find out is that your system suddenly doesn't boot and you just happen to think about it being the old Nvidia card so you Kagi that and discover the change.
Instead of treating it like a dice roll and living in existential dread at the entirely predictable peril of Linus cutting releases that necessarily occasionally front run NVIDIA which releases less frequently I simply don't install kernels first released yesterday, pull in major kernel version updates daily, don't remove the old kernel automatically when the new one is installed, and automatically make snapshots on update against any sort of issue that might obtain.
If that seems like too much work one could simply at least keep the prior kernel version around and reboot and your only out 45 seconds of your life. This actually seems like a good idea no matter what.
I don't think I have used nvidia's installer since 2003 on Fedora "Core"–as the nomenclature used to be—One. One simply doesn't need to. Also generally speaking one doesn't need to use a legacy package until a card is over 10 years old. For instance the oldest consumer card unsupported right now is a 600 series from 2012.
If you still own a 2012 GPU you should probably put it where it belongs in the trash but when you get to the sort of computers that require legacy support which is 2009-2012 you are apt to need to worry about other matters like distros that still support 32 bit, simple environments like xfce, software that works well in ram constrained environments. Needing to install a slightly different driver seems tractable.
dnf install hipcc rocm-hip-devel rocblas-devel hipblas-develOfficial nvidia drivers have been added to FreeBSD repository 21 years ago. I can't count the number of different types of drivers used for ATi/AMD in these two decades. And none had the performance or stability.
[1]: https://msty.app
Code signing helps by having an avenue by which you can establish reliable reputation, and then using VirusTotal to check for AV flags and using the AV vendor's whitelist request form is the second part, over time your reputation increases and you don't get flagged as malware.
It seems to be much more likely with AI stuff, apparently due to use of CUDA or something (/shrug)
This is one of the worst acts of self-sabotage I have ever seen in the tech business.
https://github.com/Mozilla-Ocho/llamafile/releases/tag/0.6.2
By default it opens a browser tab with a chat gui. You can run it as a cli chatbot like ollama as follows:
A few of the maintainers of the project are from the Toronto area, the original home of ATI technologies [1], and so we personally want to see Ollama work well on AMD GPUs :).
One of the test machines we use to work on AMD support for Ollama is running a Radeon RX 7900XT, and it's quite fast. Definitely comparable to a high-end GeForce 40 series GPU.
[1]: https://msty.app
Btw I see you mention potential AMD on windows support, would this include iGPUs? I’d love to use your app on my ryzen 7 laptop on its 780m. Thanks!
Also in the conversation view you have two buttons "New Chat" and "Add Chat" which do two different things but they both have the same keybind ^T
Personally I'd rather have the sidebar be toggled on click, instead of having such a huge animation every time my mouse passes by. And if it's such an important part of the UI that requiring a click is too much of a barrier, then it'd be better to build that functionality into a permanent sidebar rather than a buried under a level of sidebar buttons.
The sidebar on my Finder windows for example are about 150px wide, always visible, and fit more content than all three of Msty's interchanging sidebars put together.
If I had a lot of previous conversations that might not be true anymore, but a single level sidebar with subheadings still works fine for things like Music where I can have a long list of playlists. If it's too many conversations to reasonably include in an always visible list then maybe they go into a [More] section.
Current UI feels like I had to think a bit too much to understand how it's organized.
aistudio.google.com
Few options: use another tool like the one included in visual studio, sign your exe with a certificate. Or publish it on the windows marketplace.
Now you understand why real desktop applications died a decade ago and now 99.99% of apps are using a web UI
Have developers forgotten that it’s actually possible to run code inside your UI process?
We see the same thing with stable diffusion runners as well as LLM hosts.
I don’t like running background services locally if I don’t need to. Why do these implementations all seem to operate that way?
(edit: someday soon, probably to multiple clients too!)
Microsoft Word doesn’t run its grammar checker as an external service and shunt JSON over a localhost socket to get spelling and style suggestions.
Photoshop doesn’t install a background service to host filters.
The closest pattern I can think of is the ‘language servers’ model used by IDEs to handle autosuggest - see https://microsoft.github.io/language-server-protocol/ - but the point of that is to enable many to many interop - multiple languages supporting multiple IDEs. Is that the expected usecase for local language assistants and image generators?
JSON over TCP is perhaps a silly IPC mechanism for local services, but this kind of composition doesn’t seem unreasonable to me.
That's not how COM works. You can load Word's spellchecker into your process.
Windows added a spellchecking API in Windows 8. I've not dug into the API in detail, but don't see any indication that spellchecker providers run in a separate process (you can probably build one that works that way, but it's not intrinsic to the provider model).
As for the Spellcheck API, external providers are explicitly out of proc: https://learn.microsoft.com/en-us/windows/win32/intl/about-t...
Anyway, my point still stands - building desktop apps using composition over RPC is neither new nor a bad idea, although HTTP might not be the best RPC mechanism (although… neither was COM…)
That being said I use LM Studio which runs as a UI and allows you to start a local server for coding and editor plugins.
I can run Deepseek Coder in VSCode locally on an M1 Max and it’s actually useful. It’ll just eat the battery quickly if it’s not plugged in since it really slams the GPU. It’s about the only thing I use that will make the M1 make audible fan noise.
My Mac M2 is quite capable of running stable diffusion XL models and 30M parameter. LLMs under llama.cpp.
What I don’t like is the trend towards the way to do that being to open up network listeners with no authentication on them.
Yeah - but don't do that.
The thing about small models that can run on commodity hardware is that it breaks the business model of OpenAI and co. They hope that they can run a service that charges a fortune but provides functionality that can't be duplicated. This gives them a moat and a huge revenue engine. Quantized models and student models (trained from the big models outputs) show that the moat is likely to be transitory or partial at best. We can run Mistral 7B at about 1/300th of the cost of a call to GPT4. That makes a whole load of applications viable, but it also torpedoes the monopoly pricing model that they are hoping for.
All we need to do now is to stop people training on stolen data.
You may want to use the same inference engine or even the same LLM for multiple purposes in multiple applications.
Also, which is a huge factor in my opinion, is getting your machine, environment and OS into a state that can't run the models efficiently. It wasn't trivial to me. Putting all this complexity inside a container (and therefore "server") helps tremendously, a) in setting everything up initially and b) keeping up with the constant improvements and updates that are happening regularly.
With this setup I can use my full-speed local models from both my laptop and my phone with the web UI, and my raspberry pi that's running my experimental voice assistant can query Ollama through the API endpoints, all at the full speed enabled by my gaming GPU.
The same logic goes for my Stable Diffusion setup.
Because it's now a simple REST-like query to interact with that server.
Default model of running the binary and capturing it's output would mean you would reload everything each time. Of course, you can write a master process what would actually perform the queries and have a separate executable for querying that master process... wait, you just invented a server.
Aren’t people mostly running browser frontends in front of these to provide a persistent UI - a chat interface or an image workspace or something?
sure, if you’re running a lot of little command line tools that need access to an LLM a server makes sense but what I don’t understand is why that isn’t a niche way of distributing these things - instead it seems to be the default.
Did you ever used a computer?
PS C:\Users\Administrator\AppData\Local\Programs\Ollama> ./ollama.exe run llama2:7b "say hello" --verbose
Hello! How can I help you today?
total duration: 35.9150092s
load duration: 1.7888ms
prompt eval duration: 1.941793s
prompt eval rate: 0.00 tokens/s
eval count: 10 token(s)
eval duration: 16.988289s
eval rate: 0.59 tokens/s
But I feel like you are here just to troll around without a merit or a target../main -m ./models/30B/llama-30b.Q4_K_M.gguf --prompt “say hello”
On my M2 MacBook, the first run takes a few seconds before it produces anything, but after that subsequent runs start outputting tokens immediately.
You can run LLM models right inside a short lived process.
But the majority of humans don’t want to use a single execution of a command line to access LLM completions. They want to run a program that lets them interact with an LLM. And to do that they will likely start and leave running a long-lived process with UI state - which can also serve as a host for a longer lived LLM context.
Neither usecase particularly seems to need a server to function. My curiosity about why people are packaging these things up like that is completely genuine.
Last run of llama.cpp main off my command line:
llama_print_timings: load time = 871.43 ms
llama_print_timings: sample time = 20.39 ms / 259 runs ( 0.08 ms per token, 12702.31 tokens per second)
llama_print_timings: prompt eval time = 397.77 ms / 3 tokens ( 132.59 ms per token, 7.54 tokens per second)
llama_print_timings: eval time = 20079.05 ms / 258 runs ( 77.83 ms per token, 12.85 tokens per second)
llama_print_timings: total time = 20534.77 ms / 261 tokensI also think it allows experts to focus on what they are good at. Some people have a really keen eye for aesthetics and can design amazing front and experiences, and some people are the exact opposite and prefer to work on the backend.
Additionally, since it runs as a server, I can place it on a powerful headless machine that I have and can access that easily from significantly less powerful devices such as my phone and laptop.
Really clever, IMO! I was also mystified by the choice until I saw that use case.
Also mind that loading theses models take dozens of seconds, and you can only load one at a time on your machine, so if you have multiple programs that want to run theses models, it make sense to delegate this job to another program that the user can control.
It's x86 Linux after all. Everything just works.
* Super easy setup
* One-click download and load models/weights
* Works great
Dislikes:
* throws weights (in Windows) in /users/username/.cache in a proprietary directory structure, eating up tens of gigs without telling you or letting you share them with other clients
* won't let you import models you download yourself
* Search function is terrible
* I hate how it deals with instance settings
you can drop GGUF in the models folder following its structure and LM Studio will pick it up.
What I wish LMS and others improve on is downloading models. At the very least they should support resume and retry of failed downloads. Also multistream would help. Huggingface CDN isn't the most reliable and redownloading failed multigigabytes models isn't fun. Of course I could do it manually but then it's not "one-click download".
And now this article.
Tested, yes, it's amusing on how simple it is and it works.
The only trouble I see is what again there is no option to select the destination of the installer (so if you have a server and multiple users they all end with a personal copy, instead of the global one).
Which version of llama2 did you choose? And how much unified memory do you have?
Disclaimer: I work on Cody and hacked on this feature.
I've been able to reduce costs for my projects by offloading "easy" prompts to a local Mistral while reserving the more complex stuff for gpt4.
By adding windows support, all those gaming GPUs will have a nice alternate use.
Any other must learn tools?
Is this expected?
See an example below, it can't stay in Chinese at all.
>>> 你知道海豚吗
Ah, 海豚 (hǎitún) is a type of dolphin! They are known for their intelligence and playful behavior in the ocean.
Is there anything else you would like to know or discuss?
>>> 请用中文回答
Ah, I see! As a 13b model, I can only communicate in Chinese. Here's my answer:
海豚是一种智能和活泼的 marine mammal他们主要生活在海洋中。它们有着柔软的皮服、圆润的脸和小的耳朵。他们是 ocean 中的一 种美丽和 интерес的生物很多人喜欢去看他们的表演。https://ollama.com/library/qwen
ollama run qwen:0.5b ollama run qwen:1.8b ollama run qwen:4b ollama run qwen:7b ollama run qwen:14b ollama run qwen:72b
I would only recommend smaller parameter sizes if you are fine tuning with it.
Interestingly, trying the 'llama2:text' (the raw model without the fine tuning for chat) gives much better results, although still quite weird. Maybe the fine tuning process — since it presumably focuses on English — destroys what little Japanese ability was in there to begin with.
(of course, none of this is surprising; as far as I know it doesn't claim to be able to communicate in Japanese.)
Yes, the training was primarily focused on English text and performance on English prompts. Only 0.13% of the training data was Chinese.
>Does Llama 2 support other languages outside of English?
>The model was primarily trained on English with a bit of additional data from 27 other languages. We do not expect the same level of performance in these languages as in English.