294 karma · joined April 2, 2012
This is an OpenPGP proof that connects my OpenPGP key to this Hackernews account. For details check out https://keyoxide.org/guides/openpgp-proofs [Verifying my OpenPGP key: openpgp4fpr:E3A878624A8C0A996D1926F2033C1FEBE1ED3881]
- https://github.com/ENTERPILOT/GoModel - https://github.com/maximhq/bifrost - https://github.com/BerriAI/litellm
Would you care to share what makes experiential different?
In this case you are placing your trust in OpenAI and Anthropic. I'm not sure about Anthropic but OpenAI has changed their mission corpus quite a lot from its humble beginnings that it results hard to trust them when they say they don't use your stuff to further train their models. If I'm a Big Corp with enough lawyers to putnup a fight, I would then feel ok with such clause, but being a small guy, who is going to defend me when the truth comes out that they have been training their models with my data? Similar fiasco as with Facebook, who had claimed they didn't sell your data, even though they were.
That's where I'm coming from with all this "trust us, we don't train our models with your data". At least this Chinese company is being upfront about it.
Can you please elaborate what you mean by "critical market"?
Edit: formatting
There are packages for Vulkan, ROCm and CUDA. They all work.
I have had the chance to test the main Chinese models through OpenRouter but the Pay-as-you-go model is expensive compared to a subscription model, but I don't want to marry to a single provider.
Thanks for bringing OpenCode Go to my attention. Your comparison is the research I didn't know I needed, and I will be cancelling my Copilot subscription to replace it with OpenCode Go right away.
> when they take that tone with you.
This makes it sound as if you took it personally?
I have a 9070 XT, which has 16GB of VRAM. My understanding from reading around a bunch of forums is that the smallest quant you want to go with is Q4. Below that, the compression starts hurting the results quite a lot, especially for agentic coding. The model might eventually start missing brackets, quotes, etc.
I tried various AI + VRAM calculators but nothing was as on the point as Huggingface's built-in functionality. You simply sign up and configure in the settings [1] which GPU you have, so that when you visit a model page, you immediately see which of the quants fits in your card.
From the open source models out there, Qwen3.5 is the best right now. unsloth produces nice quants for it and even provides guidelines [2] on how to run them locally.
The 6-bit version of Qwen3.5 9B would fit nicely in your 6700 XT, but at 9B parameters, it probably isn't as smart as you would expect it to run.
Which model have you tried locally? Also, out of curiosity, what is your host configuration?
[1]: https://huggingface.co/settings/local-apps [2]: https://unsloth.ai/docs/models/qwen3.5
I am legitimately curious about the parameters that the person used for running the model locally to get the results they got because I am myself currently experimenting with running models locally myself. You can see I am asking similar questions to others in this same thread and correlate the timestamps.
The size of the quantization you chose also makes a difference.
The GPU driver also plays an important role.
What was your approach? What software did you use to run the models?
This has also been my experience. But isn't the harness sending the instructions on how to invoke a tool? Maybe it is missing the formatting part. What do you think?
I have tried the same approach with Kimi K2.5 and GLM 5 but I keep going back fo Qwen3.
I also have access to Perplexity which is quite decent to be honest, but I prefer to keep everything in Kagi.
1: https://help.kagi.com/kagi/ai/assistant.html#available-llms
Is this fine-tunning process similar to training models? As in, do you need exhaustive resources? Or can this be done (realistically) on a consumer-grade GPU?
I can only speak for myself: it can be daunting for a beginner to figure out which model fits your GPU, as the model size in GB doesn't directly translate to your GPU's VRAM capacity.
There is value in learning what fits and runs on your system, but that's a different discussion.
Does anyone have an idea as to why this would be a feature? don't you want to have a discussion with your agent to iron out the details before moving onto the implementation (build) phase?
In any case, looks cool :)
EDIT 1: Formatting EDIT 2: Thanks everyone for your input. I was not aware of the extensibility model that pi had in mind or that you can also iterate your plan on a PLAN.md file. Very interesting approach. I'll have a look and give it a go.
They paid for the access the same as any other.
If anything, this makes them more legit than Anthropic because they are paying for the content, whereas Anthropic just stole *all* the data they got a hold of. So, in this case the Chinese AI labs stand on higher moral ground LOL.