HNHacker News
TopNewBestAskShowJobs

pimeys

6,921 karma · joined October 24, 2010

https://twin.so https://github.com/pimeys https://social.nauk.io/pimeys
submissionscomments
pimeys··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Work for my company. Research, code, analysis.
pimeys··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
It's like choosing between vim and helix. I started my career with vim in the early 2000's, customized the whole thing and had my config in a version control.

Then I installed helix and I just use it without config.

If you like configuring things take pi, if not omp is pretty much great defaults.

pimeys··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
https://omp.sh/ has amazing defaults and it sips tokens. Works really well with DeepSeek V4.1 Flash.

Use the model through a fast and reliable provider such as Fireworks directly, skip OpenRouter.

pimeys··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
You mean 4.1 Flash which is the first great Deepseek? The one that actually surpasses Opus in my books now.
pimeys··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
How on earth you can do 100 dollars in 2-3 days with DeepSeek? I have 7 agents in omp running 24/7 every day. I use maybe 10-15 dollars a day. A rarely see a session going over 2 dollars. My maximum is maybe 3.5 dollars and that session took three days.

What harness you are using?

pimeys··on [dead]
Wdym? If you have any LLM features in your app, how much you want to charge your customers? 100 euros per month? 1000? Or could you just offer them in the cheapest 10 euro package and don't care about cost that much?

There is no reason to pay API prices for Anthropic and OpenAI if they don't cut them down 100x at least.

pimeys··on MiMo v2.6
Harness. Especially if a tool call error doesn't say what to do next and the model is not RL'd with that tool, a retry storm is common.

So if you use MCP a lot, simplify the params, be more lenient on validation and rework the errors.

It is quite good with shell.

pimeys··on MiMo v2.6
a smoking gun
pimeys··on Xiaomi MiMo v2.6
Fireworks is amazing for DeepSeek, Kimi and GLM.

Hope they bring MiMo for tests.

pimeys··on Xiaomi MiMo v2.6
I've used Kimi K3 for a few months as my main model and DeepSeek 4.1 is as fast and about 10x cheaper.

I just had like four big sessions going today, paid about $8 in tokens. I see no reason to pay more, this is more than I need for intelligence.

pimeys··on MiMo v2.6
Open tech is cool. Speeds up all progress...
pimeys··on MiMo v2.6
I would really enjoy that blog post.
pimeys··on Xiaomi MiMo v2.6
Show them you can burn tokens in seven sessions day and night with comparable results to Opus with less energy and less than 10 dollars a day, per dev.
pimeys··on AI and the Destruction of the Creative Commons
Yes. And we should all form small social LLM servers in our neighborhoods to run Deepseek and avoid paying for American corporations.

I fully agree with you.

pimeys··on Android 17 is the first since 3.x to add new APIs without releasing to the AOSP
Android has free and open artificial pancreas that is still easy to install and keeps us with complex type 1 diabetes alive.

Google may just want to kill us and Apple don't even let this kind of software exist without massive hurdles...

pimeys··on CCC invites all model citizens to 40C3
> One older man out of nowhere loudly proclaimed I was Anti-semetic, because I was wearing a Palestinian keffiyeh.

Yep, and tell that to anybody who got either kicked out of the country or lost their livelihood by commenting against the Palestinian genocide in the social media.

pimeys··on CCC invites all model citizens to 40C3
That has pretty much changed now. I rarely use cash in Berlin. About 2-3 years already...
pimeys··on CCC invites all model citizens to 40C3
And there is a big warning if there is a security camera around. And they get destroyed pretty often.

Source: living in Berlin.

pimeys··on DeepSeek v4.1 Flash Is Now Our Best Hacking Model
My experience is complete opposite from yours. The V4 Flash was already quite good, but V4.1 is really very close to SOTA in my books. I've eval'd these models for weeks against Gemini 3.8, Kimi K3 and Opus 5, and DeepSeek absolutely wins these evals.

It's really crazy. We price per token our customers. If Kimi was about 30% of the price of Opus 5 for the same quality, DeepSeek is 1/10th of a price of Kimi K3. We've come down in price so much that I seriously cannot recommend other models before they reduce their pricing.

And what is really interesting is its programming ability. As I've said in my previous comments, I use agents a lot in my work. Since early Opus days until now I have 7-8 agents working in parallel for different tasks. Rust, design, GEPA, evals, analysis. For a long time Kimi K3 was the best model for this work, and before that GPT 5.5. But I still can't really believe how well DeepSeek works here. I really try to find faults from it, trying to see that it must be doing sloppy work and be worse than the others. But it does not. It finishes every task I give to it. And the cost per task is under a dollar, usually 15-30 cents.

In comparison the same task with Kimi would be 3-15 dollars; sometimes closing to 100. And before that with GPT 5.5 a 800 dollar task was not uncommon if I spent days evaluating models.

Now it's less than a dollar.

For me if the other providers will not drop their prices dramatically in the coming weeks I see no reason to use them. Even with a 200 dollar subscription, paying per token for DeepSeek is better value.

My harness: https://omp.sh/

pimeys··on DeepSeek v4.1 Flash
A colleague of mine has a strategy game to compare language models, 4.1 scores pretty high in this:

https://clankerbattle.com/

pimeys··on DeepSeek v4.1 Flash
It's more common than you think. I work in a startup and we pay API prices too. And we cut a lot of money by switching from Anthropic models to Kimi K3.
pimeys··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
Deepseek also burns a lot of tokens, its output on high is 2x of Gemini on medium. But it's dirt-cheap so it still can be 60-70% cheaper.

From the large models Kimi K3 is definitely the one burning the smallest amount of tokens. Even if you pay for the fast version in Fireworks it's third of the price of Opus 5 for the same task.

All this really needs evals, the token prices tell nothing.

pimeys··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
What you want is a bunch of sessions to replay. Something anonymized if it's not yours, and something that's not depending on state.

You replay all your sessions against your harness, and then store all logs all output, everything to a safe place.

Finally use a blind judge to check everything, and score the output.

Then fix your harness, iterate again until better until you are in a point where it's just the model's weakness. If you get to that, use a bigger model.

pimeys··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
- Which versions: 3.6 vs 3.7 vs. 3.8 for Gemini Flash, and v4 0731 for Deepseek v4 Flash, and GLM 5.3 Flash

- Medium for Gemini, high for Deepseek.

- Things like find information, then understand something about it, then send a slack message or email etc.

- Completion rates somewhere in 80-90%, Deepseek a bit better than Gemini

- Quality evaluated by Fable 5.1 and Astra 6.0 acting as a rubric judge.

Gemini quality would probably be better with high thinking level, but that would be 40% more expensive. And Deepseek is already third the price of Gemini.

pimeys··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
Well, it's much more than that. In general everybody's building agents now. You see these things that can help you to do things like adding things like OCR an appointment from a picture of a hand-written paper and add it to your calendar, search things from the internet, find that email with a PDF and add it to your local paperless instance.

Building an agent like this by yourself is really easy. Now, we have Gemini's subscription, OpenAI's ChatGPT subscription and all those, 20 bucks a month right?

What if you can spend that 20 bucks in tokens to do your own. And you pay 15 bucks _a year_ in tokens to run that? And you own the data, you own your code and integrations. It's really easy to do, and these flash models are _more than enough_ for simple agentic tasks.

pimeys··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
All of these flash models have this. You have to build your harness so that it deals with it. Infinite loops are solved by having an error message that says what to do differently on failure, invalid tool calls are solved by making the tool schema less strict and detect things in the runtime etc.

Hallucinations you can't fix. Gemini is a bit worse there than DeepSeek, but there's not much research on how to fix that. The only one is the CaMeL paper by Google, where you tag every prompt and result and then for every assistant response or tool call you first check where it got that data and error if you notice fabrication. This one is really annoying to implement.

With larger models the fabrication starts when the context grows or if you have too many tools, for flash models it's much earlier. We use the flash models for repetitive agentic tasks, where the prompt defines clearly what to do and how. The whole run is about 4-5 steps typically, and context size stays in the comfort zone.

pimeys··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
Yes. I'm working in the agent industry and my god are we excited on new versions of Chinese flash models. The direct competition is Gemini Flash, and these models are much better on agentic tasks with fraction of the task price compared to Gemini. Things like oh here's a set of simple instructions for you to follow, call these tools, return this report. 20-30% of the price per task. And especially Deepseek Flash produces better quality than Gemini does.

Where Gemini still wins is non-text input what Deepseek cannot do, yet, and Deepseek Flash has this thing of cheaper models where a failing tool call can derail your agent to a retry loop if you're not careful on instructions in the error message.

If they fix and make the tool calls to work better in non-optimal situations, it's much easier to switch from Gemini without a few weeks of evals and bugfixing.

pimeys··on Mistral raises €3B
Yeah, and buying a house is still kind of out of reach with this income if you don't start paying your mortgage whey you're 20something...
pimeys··on Mistral raises €3B
And 180k in Germany, even in Berlin is pretty good. You are living a very good life with that salary.
pimeys··on Mistral raises €3B
They try to build them, but for example in Finland where there's cheap electricity and lot of interest to build them, the people started protesting on rising electricity prices and now the politicians are noticing this.

Same in Denmark.

Page 1 of 34Next →