it looks like it's an explicit option to open a window with AI features and without, so you get a choice to enable the features if you want them
12 karma · joined November 5, 2025
it looks like it's an explicit option to open a window with AI features and without, so you get a choice to enable the features if you want them
Why not google/ddg/bing etc. them? That's on the context menu too, but LLMs seem uniquely suited at some problems like acronyms that are shared across many fields but different meanings, highlighting a sentence turns out the right acronyms very fast where search engines would take several attempts and is what I used to do previously.
The only way to use a tool like this is to give a problem that fits context, evaluate the solution it chugs at you and re-roll it if it wasn't correct. Don't tell a language model to think because it can't and won't. It's a way less efficient way of re-rolling the solution
Overall if you're memory constrained it's probably still worth to try and fiddle around with it if you can get it to work. Speedwise if you got the memory a 5090 can get ~50-100tok/s for a single query with 32B-AWQ and way more if you have something parallel like open-webui
from
```
excerpt of text or code from some application or site
```
What is the meaning of excerpt?
Just doesen't seem to work at a useable level. Coding questions get code that runs, but almost always misses so many things that finding out what it missed and fixing those takes a lot more time than handwriting code.
>Overall it still seems extremely good for its size and I wouldn't expect anything below 30B to behave like that. I mean, it flies with 100 tok/sec even on a 1650 :D
For it's size absolutely, I've not seen 1,5B models that form even sentences right most of the time so this is miles ahead of most small models, not just to the hinted at levels the benchmarks would you have believe
Used their recommended temperature, top_k, top_p and so on settings
You can gain a lot of performance by using optimal quantization techniques for your setup(ix, awq etc), different llamacpp builds do different between each other and very different compared to something like vLLM
As long as it's within terms and conditions of whatever agreement you made for that $20. I can run queries on my own inference setup from remote locations too