I just want to run `<some-command> <model-name>` with some default parameters set and for it to run locally.
I just want to run `<some-command> <model-name>` with some default parameters set and for it to run locally.
It's long, I guess, but not cryptic.
You tell llama server where the model is, which context size to use, what to use for the K/V cache quant, that it should do MTP, tune some MTP parameters, and that's kinda it.
Perfectly logical blocks with all the model-specific weirdness (that does exist!) abstracted away.
You could also just run -m <modelfile> and let llama-server do the right-ish thing. The defaults are probably fine, but not how you squeeze out these exact numbers. I think at least. I've never tried. My hubris stopped me from trying auto configs.
isn't that the hard part? You know the ballpark ideal values for these many parameters since you're a knowledgeable expert but the vast majority of people are just like "I want AI" and have no idea what all the jargon even means.
Otherwise, if you're a programmer setting up a local harness, it only takes like 20-30 minutes to learn what the right parameters are.
It's very model, hardware, and use case dependent which is why a one size fits all solution doesn't work
They could ask your current agent to a) search for this type of content online for the optimal setup for their hardware b) have the current agent/harness spin it up have it verify the config run few experiments.
Sure AI may make mistakes, or won't get the best possible config probably, but it certainly do a good enough setup, this is a task with feedback on whether the server crashed or poor performance easily measured so the agent can do a pretty good job.
Because this is kinda the one new thing that arrived in the technology scene, so getting at least some amount of understanding of its "inner" workings might prove useful in the future.
Beside that, it is also just.. interesting? It's fun tuning the machine to see it improve. For some, anyway.
The pain point they raised is this is too complicated for people who just want to get started, that is not true anymore.
It is certainly fun to fine-tune and setup if you like do something like that, however the need to do it hardly is a barrier for those who don't want complexity as OP imagines.
Lower level API/interfaces should not be a barrier for people if they are apply framing that way. More and more people are thinking agent native so this is not really a issue.
Why this headache inducing lingo tho? What does that even mean, and why should I sign up for your webinar about that?
For Claude, I setting a single config file and then download and run Claude Code CLI. Even easier for the GitHub Copilot CLI.
You could override, obviously.
There are some clients that will index the models and allow you to do that but I'm no expert, I've used OLama studio but it always seems to go weird for me.
Even this command above, it's not clear where op got the model from. So I'm with yah.
For example, op uses : Qwen3.8-27B-IQ4_NL.gguf.. But I cant see where to download it. It's not tagged on hugging face at least..
If I want to download the model myself, it's not clear. I thought it was supposed to behave like a package manager. But even in nuGet I can download a zip of the package.
They shared a lot of links, I'm struggling to find yours. Where did yours come from ?
Who is Unsloth AI? Have they modified the model ? Is this really the source of truth ?
Do you see how steep the barrier for entry is to do anything right ?
unslothai is not a name qwen has ever used. So you're sharing a link to a model that isn't from the owner, while saying it's the owner's. I'm not comfortable with that, and I want AI to be a better tool.
You asked where to find the GGUF files of this model for direct download and I provided it. Almost all useful model files that can be downloaded are hosted on Huggingface.
> They shared a lot of links, I'm struggling to find yours. Where did yours come from ?
I went to Huggingface, went to the Unsloth org, as they tend to be the best, went to the Model page, and went to the "Files and versions" tab.
> Who is Unsloth AI? Have they modified the model ? Is this really the source of truth ?
Unsloth AI is a very popular, highly reputable organization that takes upstream model files, performs some optimization, and provides models in various formats. Apart from speed tweaks, they do not modify the models. They also provide useful benchmarks, copious documentation for local execution, and a Studio application for easy execution and post-training of models.
> Do you see how steep the barrier for entry is to do anything right ?
No. Searching for this information is not difficult. The llama.cpp documentation and guides that Unsloth provide are all you need. Search engines can take you further if you want.
> unslothai is not a name qwen has ever used. So you're sharing a link to a model that isn't from the owner, while saying it's the owner's. I'm not comfortable with that
Qwen also provides models in GGUF format on Huggingface, but they will not be as performant. Even when first-party GGUFs are available, most people will prefer quants from Unsloth or a few other popular optimizer accounts.
> I want AI to be a better tool.
Best of luck. Your attitude and unwillingness to even try and learn on your own when people have tried helping have burned my good will, and this is as far as I'm willing to carry you.
"I went to youtube, gmail, ycombinator, deliveroo, then I went to another site I randomly chose, because they're the best, duh"
Okay.
> Unsloth AI is a very popular, highly reputable organization that takes upstream model files
And How am I supposed to know that arriving to huggingface as a new user? Enlighten me.
> Qwen also provides models in GGUF format on Huggingface
Cool.. Why don't they share em because I genuinely cant find em, I'm dumb.
> Your attitude and unwillingness to even try and learn on your own when people have tried helping have burned my good will
Well, I'm willing, but not from people who I might burn good will. Gracious. You do you.
The path I described is all within HuggingFace.
> And How am I supposed to know that arriving to huggingface as a new user? Enlighten me.
Because I told you, knowing it was the best starting point for newbies.
> Cool.. Why don't they share em because I genuinely cant find em, I'm dumb.
You could go to Qwen's organization page on HuggingFace, it has a search function at the top, but you would be better served sticking with Unsloth.
> Well, I'm willing, but not from people who I might burn good will. Gracious. You do you.
Expecting others to do everything for you is not the same as trying things and asking questions about what you found.
It's just not. Where do I see the model download ?
The fact we're here is a loss.
It's not intuitive. Deal with it, or fix it.
Huggingface, Unsloth, and llama.cpp all have documentation you can follow that will exceed anything I can tell you here. LMstudio, Lemonade, or Ollama might be even easier for you to use. Take my suggestions or don't.
llama-server -m model.gguf
That's itYou don’t need to fine tune all of those parameters to get started.
It’s really easy to ask an LLM to adjust the command line if you can’t be bothered to read the help out. Copy the help output into the LLM and tell it your goal.
> Ollama is confusing and doesn't seem to support Qwen3?
Typing “Ollama qwen3” into Google takes you right to this page:
https://ollama.com/library/qwen3
If even Googling for basic Ollama support is too hard, there might come a point where you have to acknowledge that local LLMs are not for you. None of this is really that hard with some basic Google bootstrap skills or by asking an LLM to help with the command.
> Attention: To be updated for Qwen3
on Qwen's official docs: https://qwen.readthedocs.io/en/latest/run_locally/ollama.htm.... It's not like I just made it up. Of course I searched "ollama qwen3" and saw what you linked, but that doesn't mean it "works". I have other things to do besides to try a bunch of poorly documented and executed tools just to see if it works or not.
I guess the TLDR is that I'm stupid or lazy. Also, everyone is responding about how easy it is, and yet, it's apparently so easy that it's hard to document well.
I remember there was a short story in BYTE Magazine about a similar kind of scenario way back when, I think at least 30 years ago, long before LLMs and AI agents became a reality.
Very occasionally I'll have Claude install something on my dev box, but not without close supervision.
Woa, is that still a thing? You mean like SOCKS5 stuff that you have to manually configure in every application that uses the internet?
I mean maybe I'm just living under a rock but I feel like that's a rather niche situation you got there.
Every big company in the world uses a network proxy. LM Studio, as far as I can tell, cannot be configured to work behind such proxies.
It's becoming more rare, now.
A lot of the universal truths about corporate networks from the early 2000s are no longer true today. Some companies are stuck in their ways though.
The overlap between companies that require someone to use a network proxy and companies that have GPU-equipped machines with enough RAM for LLMs and and that allow people to download and run executables of their choosing has to be small.
I would love to see more data on that because I've seen it constantly. There is more isolation maybe where you can do whatever on 'open' network, but always some kind of proxy/vpn connection for hitting anything sensitive.
The operlap is there.. But I would be worried if it was just flat out taken away from secure managed connections just because of AI.. Again, would love to see the numbers of your assumptions.
Makes it rather weird that LM Studio doesn't support it given how their target market, or well at least for their paid products, is very enterprisey.