HNHacker News
TopNewBestAskShowJobs

parthsareen

43 karma · joined May 27, 2021

submissionscomments
parthsareen··on Qwen 3.8 27B
Hi! From Ollama here - you can run: ollama run qwen3.8 (or if on mac qwen3.8:27b-mlx)
parthsareen··on Claude Code: connect to a local model when your quota runs out
Also recently added ollama launch claude if you want to connect to cloud models from there :)
parthsareen··on Ask HN: A good Model to choose in Ollama to run on Claude Code
Hey! One of the maintainers of Ollama. 8GB of VRAM is a bit tight for coding agents since their prompts are quite large. You could try playing with qwen3 and at least 16k context length to see how it works.
parthsareen··on A guide to local coding models
How much ram are you running with? Qwen3 and gpt-oss:20b punch a good bit above their weight. Personally use it for small agents.
parthsareen··on A guide to local coding models
You're welcome to go through the source: https://github.com/ollama/ollama/
parthsareen··on A guide to local coding models
Desktop app is open-source now.
parthsareen··on Ollama Web Search
Since we shipped web search with gpt-oss in the Ollama app I've personally been using that a lot more especially for research heavy tasks that I can shoot off. Plus with a 5090 or the new macs it's super fast.
parthsareen··on Ollama Web Search
Hi - author of the post. Yes it does! The "build a search agent" example can be used with a local model. I'd recommend trying qwen3 or gpt-oss
parthsareen··on Ollama Web Search
Hey! Author of the blogpost and I also work on Ollama's tool calling. There has been a big push on tool calling over the last year to improve the parsing. What's the issues you're running into with local tool use? What models are you using?
parthsareen··on Sampling and structured outputs in LLMs
That's a great idea. Going to try this next :)
parthsareen··on Sampling and structured outputs in LLMs
Hey! I'm the author of the post. We haven't optimized sampling yet so it's running linearly on the CPU. A lot of SOTA work either does this while the model is running the forward pass or does the masking on the GPU.

The greedy accept is so that the mask doesn't need to be computed. Planning to make this more efficient from either ends.

parthsareen··on Sampling and structured outputs in LLMs
Thank you! Maybe not "perfect" but near-perfect is something we can expect. Models like the Osmosis structure which just structure data inspired some of that thinking (https://ollama.com/Osmosis/Osmosis-Structure-0.6B). Historically, JSON generation has been a latent capability of a model rather than a trained one, but that seems to be changing. gpt-oss was particularly trained for this type of behavior and so the token probabilities are heavily skewed to conform to JSON. Will be interesting to see the next batch of models!
parthsareen··on Sampling and structured outputs in LLMs
Thanks for posting! Didn't expect this to get picked up – it was a bit of a draft haha. Happy to answer questions around structured outputs :)
parthsareen··on Structured Outputs with Ollama
Yes! I have checked guidance out, as well as a few others. Planning to refactor sampling in the near future which would include improving using grammars for sampling as well. Thanks for sharing!
parthsareen··on Structured Outputs with Ollama
The constraints will always be met. It’s the data inside that might be inaccurate. YMMV with smaller models in that sense.
parthsareen··on Structured Outputs with Ollama
Hey! Author of the blog here. The current implementation uses llama.cpp GBNF which has allowed for a quick implementation. The biggest value-add at this time was getting the feature out.

With the newer research - outlines/xgrammar coming out, I hope to be able to update the sampling to support more formats, increase accuracy, and improve performance.

parthsareen··on Structured Outputs with Ollama
Hey! Author of the post and one of the maintainers here. I agree - we (maintainers) got to this late and in general want to encourage more contributions.

Hoping to be more on top of community PRs and get them merged in the coming year.

parthsareen··on Structured Outputs with Ollama
This looks really useful. Thank you!
parthsareen··on Structured Outputs with Ollama
I authored the blog with some other contributors and worked on the feature (PR: https://github.com/ollama/ollama/pull/7900).

The current implementation uses llama.cpp GBNF grammars. The more recent research (Outlines, XGrammar) points to potentially speeding up the sampling process through FSTs and GPU parallelism.

parthsareen··on Structured Outputs with Ollama
We’ve been keeping a close eye on this as well as research is coming out. We’re looking into improving sampling as a whole on both speed and accuracy.

Hopefully with those changes we might also enable general structure generation not only limited to JSON.

parthsareen··on Structured Outputs with Ollama
Hey! Author of the blog post here. Yes you should be able to use any model. Your mileage may vary with the smaller models but asking them to “return x in json” tends to help with accuracy (anecdotally).
parthsareen··on Ask HN: Favorite Blogs by Individuals?
The first few of these are my fav: https://dive.sh/thread/81vD2RhjxF