6,921 karma · joined October 24, 2010
Then I installed helix and I just use it without config.
If you like configuring things take pi, if not omp is pretty much great defaults.
Use the model through a fast and reliable provider such as Fireworks directly, skip OpenRouter.
What harness you are using?
There is no reason to pay API prices for Anthropic and OpenAI if they don't cut them down 100x at least.
So if you use MCP a lot, simplify the params, be more lenient on validation and rework the errors.
It is quite good with shell.
Hope they bring MiMo for tests.
I just had like four big sessions going today, paid about $8 in tokens. I see no reason to pay more, this is more than I need for intelligence.
I fully agree with you.
Google may just want to kill us and Apple don't even let this kind of software exist without massive hurdles...
Yep, and tell that to anybody who got either kicked out of the country or lost their livelihood by commenting against the Palestinian genocide in the social media.
Source: living in Berlin.
It's really crazy. We price per token our customers. If Kimi was about 30% of the price of Opus 5 for the same quality, DeepSeek is 1/10th of a price of Kimi K3. We've come down in price so much that I seriously cannot recommend other models before they reduce their pricing.
And what is really interesting is its programming ability. As I've said in my previous comments, I use agents a lot in my work. Since early Opus days until now I have 7-8 agents working in parallel for different tasks. Rust, design, GEPA, evals, analysis. For a long time Kimi K3 was the best model for this work, and before that GPT 5.5. But I still can't really believe how well DeepSeek works here. I really try to find faults from it, trying to see that it must be doing sloppy work and be worse than the others. But it does not. It finishes every task I give to it. And the cost per task is under a dollar, usually 15-30 cents.
In comparison the same task with Kimi would be 3-15 dollars; sometimes closing to 100. And before that with GPT 5.5 a 800 dollar task was not uncommon if I spent days evaluating models.
Now it's less than a dollar.
For me if the other providers will not drop their prices dramatically in the coming weeks I see no reason to use them. Even with a 200 dollar subscription, paying per token for DeepSeek is better value.
My harness: https://omp.sh/
From the large models Kimi K3 is definitely the one burning the smallest amount of tokens. Even if you pay for the fast version in Fireworks it's third of the price of Opus 5 for the same task.
All this really needs evals, the token prices tell nothing.
You replay all your sessions against your harness, and then store all logs all output, everything to a safe place.
Finally use a blind judge to check everything, and score the output.
Then fix your harness, iterate again until better until you are in a point where it's just the model's weakness. If you get to that, use a bigger model.
- Medium for Gemini, high for Deepseek.
- Things like find information, then understand something about it, then send a slack message or email etc.
- Completion rates somewhere in 80-90%, Deepseek a bit better than Gemini
- Quality evaluated by Fable 5.1 and Astra 6.0 acting as a rubric judge.
Gemini quality would probably be better with high thinking level, but that would be 40% more expensive. And Deepseek is already third the price of Gemini.
Building an agent like this by yourself is really easy. Now, we have Gemini's subscription, OpenAI's ChatGPT subscription and all those, 20 bucks a month right?
What if you can spend that 20 bucks in tokens to do your own. And you pay 15 bucks _a year_ in tokens to run that? And you own the data, you own your code and integrations. It's really easy to do, and these flash models are _more than enough_ for simple agentic tasks.
Hallucinations you can't fix. Gemini is a bit worse there than DeepSeek, but there's not much research on how to fix that. The only one is the CaMeL paper by Google, where you tag every prompt and result and then for every assistant response or tool call you first check where it got that data and error if you notice fabrication. This one is really annoying to implement.
With larger models the fabrication starts when the context grows or if you have too many tools, for flash models it's much earlier. We use the flash models for repetitive agentic tasks, where the prompt defines clearly what to do and how. The whole run is about 4-5 steps typically, and context size stays in the comfort zone.
Where Gemini still wins is non-text input what Deepseek cannot do, yet, and Deepseek Flash has this thing of cheaper models where a failing tool call can derail your agent to a retry loop if you're not careful on instructions in the error message.
If they fix and make the tool calls to work better in non-optimal situations, it's much easier to switch from Gemini without a few weeks of evals and bugfixing.
Same in Denmark.