It would definitely cost me more per month than a x20 ChatGPT or Claude plan, probably around $400+ was my estimate at the time. This was with Fireworks (ZDR) which has since increased their prices (and got slower!).
That being said, very impressed with the model, and looking forward to what comes next. As the frontier models become less subsidized, the open models will become more appealing.
P.S. There are subscription plans for open models, but I've found most of them to be extremely slow, have model throttling (only so much of model X), and also very sketchy about training and data retention. No thanks! If you want to share your data, just use Muse Spark contributor. Seems impossible to beat that on price per task if you don't mind feeding your data to the Meta machine (spoiler: I won't).
Edit: others have noted the provider and harness matters. My experience is with opencode.
https://openrouter.ai/docs/guides/routing/provider-selection
I stopped using opencode because it has some issues with caching, so I suppose it doesn't do this
Using this interesting framework called Cordis I've recent discovered.
What harness you are using?
> First, they come for the programmers, and next the mathematicians. Then it will be the biologists, lawyers and doctors. [...] Eventually, with advancements in robotics, stabs at increasingly more physical sciences will also be attempted. Eventually, AI will have its hand in the pie of all knowledge work, if it is possible.
You seem to think raising productivity is a bad thing?
> Humanities will stake it out a little bit longer because AI isn’t human, [...]
This is really silly. Many languages other than English don't use related words to describe 'humans' and 'humanities'. Will it be easier in those languages? Should we rename mathematics to 'humathematics' or so, to make it harder for AI to take over?
Higher productivity means very little to people whose labor plummets in value over the course of a couple of years. One day they won't be needed anymore, maybe that's good for you.
I'm a software developer by trade. I welcome the coming brave new world in which machines can do all the software.
At the moment, they ain't quite there yet, alas.
I dont consider myself a programmer but use LLMs almost exclusively for coding.
The number of people able to create useful software today is much much larger than it used to be and arguably a minority of these people are/were "programmers"
Also even if your Gemini is giving you nonprogramming output, underneath the model is most likely generating code for certain tasks.
I can’t overstate how bad of an idea I think using an AI for customer interaction is.
But yeah, send me a non solicited AI slop email or worse, political ad, and you dont get the dignity of me saying stop to unsubscribe. Straight to spam for you.
So, spam?
If someone is asking you for a quote, and your tool replies, it's not spam.
It might be slop for all I know (or might not be!), but it's not spam.
Similarly for arranging meetings with parties that you already have a relationship with.
The shape of my work changes obviously, so it'll vary, sometimes more, sometimes less. For example, fixing all of the bugs and defects I found that week was 2-3 times the effort and chewed through my ChatGPT allowance, but I had banked resets...
Also worth noting that codex models have been kind of all over the place recently with their usage... and it looks like costs are changing again.
Gotta be honest though, I don't love fiddling with this all the time. I would rather be working on my projects than evaluating my usage. Having DS in my back pocket should i need it is a relief. The providers get fiddly though too.
Pi out if the box tries to optimise system prompt size, which is not necessarily good and will cause exactly this effect for all but the most simple tasks.
What you want is to give enough context to the agent to minimize the amount of searching within the codebase etc.
If you want to track cache, what you should do, imho, is to check if you have cache expirations mid sessions (ideally you should not), and if you don’t then lower cache use is actually better - it means that your model doesn’t reread what it just wrote.
I also spent $280 on DeepSeek doing the tests (direct to DS, not OpenRouter). I suggest that if you can't conceive of anyone spending $200 on DeepSeek, you're not being ambitious enough!
tbf, i barely feel the difference with opus 5.5 anymore either.
And there's a lot of labour and effort involved in setting these things up and maintaining them. People don't even run their own email servers, even though the hardware side of that is trivial.
Not $400 a month, it doesn't. At least not around here (we average $0.12/kWh and I don't personally run my cards over 300W.)
And there's a lot of labour and effort involved...
Theoretically, the people who hang around HN are more likely than most to be capable of the labour and effort of setting these things up and maintaining them.
People don't even run their own email servers...
People don't run their own email servers because a convenient coalition of spammers, standards bodies, and large email providers have done their best to make running one's own email server almost impossible.
I did the math and decided it’d better to pay for tokens than to buy the hardware and generate them myself.
'Capable' doesn't mean your labour has no opportunity costs.
I think my argument is easier to attack by noting that you can use AI to substitute for much of that labour.
It's worth it, knowing that there are no rugs Sam or Dario or anyone else can pull.
At the moment, it's still very easy to switch from one open weights model to another, and even between the closed models. So the 'no rug pull' property is nice to have, but not as big of a deal.
Theoretically, the people who hang around HN have immense opportunity cost when doing this. Being capable does not mean it takes no time. Time they could be using for something more profitable (and fun?).
The AI models we currently have still maintain their working state with a finite and laughably-small context window, but that is already starting to change. My .claude directory contains over 300 .md files that I didn't put there myself. Their contents are eye-opening. The question of who owns, stores, maintains, and can access that data is going to become insanely important over the next couple of years.
If you thought LLMs themselves were disruptive and contentious, just wait until the fight over object permanence gets under way. That's when owning your own box full of graphics cards is going to become important. My own bet is that I won't care too much about the electric bill or my opportunity cost when we all find out what the AI labs really have in mind, and what they're going to have to do in order to justify the valuations they're seeking.
TL,DR: it's not about the tokens, IMHO.
So if you can spend $30K and immediately start mining $1500/month out of thin air, that's a pretty nice investment even if the electricity costs $200/month. Two years later the cards will have paid for themselves entirely and (I suspect) will still be pretty useful.
It does argue in favor of just paying OpenAI or Anthropic for tokens, though.
At the same time, when I need to use the hardware for something, whoever is renting it from me at the moment is going to get unceremoniously booted, and I imagine they are not going to be happy about that. I assume that vast.ai's providers get uptime ratings that drive their work allocation, right?
What I do is… rent on the same platform when I actually need to use a card. A benefit is that if I need a burst of more power, that’s not an issue since it’s available from other hosts. But obviously there’s an inflection point of first party utilization where it makes more sense to own.
Now, I have absolutely no clue how long this situation is going to last! But the economics don't really work out for local models while it does.
How is that different from feeding into the OpenAI or Anthropic machines?
Whats your harness?
Subagents are like trading derivatives. You can lose as much as you want.
When the regulations do arrive, I think they should really focus on AI companies and API providers being more transparent wrt how they're billing their customers. Because right now, it's a totally vibes-dependent and a mess.
A smaller model in the same generation will never be the same as a bigger one, assuming this is a smaller model, and the same generation, as naming implies, it will not be comparable, it might be on the benchmarks, even on the benchmarks that matter, but the whole story should also give the drawbacks.
In Search Advertising, the amount you pay (under GSP Auction) is a function of your pCTR. And guess who determines your pCTR? The Search Engine itself! :-D
And watch 10 hours of football on Sunday for our DraftKings bets.
Parallelism is fantastic when it actually speeds up the entire pipeline, but in my experience most people's jobs (at least the ones for which AI is currently relevant) involve a lot of overlapping "hurry up and wait" branches that drastically blunt the real benefits of that sort of parallelism.
There may be specific situations where it makes sense to do it, but just immediately going full gastown on anything AI related seems like such a giant waste to me, of both money and finite world resources.
Eg when I have the AI do a self-code-review before bothering a human, I also want to the AI to give the draft PR to a sub-agent that only has the publicly available context that a viewer of the final PR would have; and not all the accumulated reasoning that lead to the writing of the code in the PR.
Opus directs and starts haiku/sonnet subagents. Much more efficient than opus reading it all.
Without this explicit direction, Opus burned through usage to create complicated regex parsers & 8 excel sheets.
Excellent pithy warning.
But there’s a point on that spectrum where the ability to run multiple experiments in parallel, even with a significant amount of (one time) wastage, is overall more cost effective than the alternative.
The only thing I've found Deepseek and Kimi good for are security tasks that GPT refuses to do.
This is a summary of what Deepseek did and got wrong:
Lost the proven baseline: changed kernel source, configuration, compiler, RAM geometry, MMC width, and peripherals together. Matching an upstream commit did not preserve local boot fixes, making failures difficult to isolate. Misidentified an image: a file labelled “r18-known-good” actually contained the r23 parent bootloader. Filename-based reasoning replaced verification of the artifact’s identity and provenance. Shipped inconsistent boot contracts: flash-16b’s loader read too few kernel blocks. Fresh2 changed the device tree without updating the loader’s expected length and CRC, creating deterministic rejection before normal Linux handoff. Patched binaries without maintaining reproducible source: loader constants diverged from source, a separately compiled cache-flush length remained stale, and assembly used an oversized stage-two slot. Their causal contribution to hangs was not established. Overstated diagnosis: claimed failures were definitively in U-Boot, blamed compiler or IPU changes without controlled isolation, converted noisy observations into confirmed hangs, and neglected persistent journals as an alternative explanation. Mistook compilation for integration: framebuffer registration was incomplete, timing success handling was inverted, BT.656 selection was unreachable, encoder overrides were missing, and audio lacked software clock configuration. Misread hardware evidence: asserted interrupt-free PMIC operation, assigned RF to the wrong SPI controller, confused regulator identifiers with register addresses, and described repeated encoder writes as unique registers. Overclaimed results: treated kernel/probe indications as userspace success, presented earlier discoveries as new progress, and omitted failed flashing attempts from the final narrative.
There's your problem, 4.1 Flash is significantly better and cheaper, to the point where the official DeepSeek API is going to (or already has, I forget) redirect requests for Pro to 4.1 Flash, and adjust billing accordingly too.
4 Pro is still offered by providers I'm sure, since it's open weight, so I can understand making that mistake.
For myself, with ChatGPT for example, it regularly gets extremely slow if there is a lot of text in a single conversation. Especially if I try to scroll up.
Some pages have been strangely just broken for a while now as well. Usage analytics just renders lots of these duplicate "Usage history" components, where the data just never loads: https://imgur.com/a/vpcQIiw.png
I can totally imagine a scenario where some agent built it, another tested and approved it, and nobody at OpenAI even looked at it once or knows that it's like this.
https://www.youtube.com/watch?v=WAeHgE94rVo
The system performs quotation attribution on my local hardware for my near-future, hard sci-fi novel (having nearly 500 quotations) with over 97% accuracy.
The initial prototype was developed quite quickly, but numerous successive iterations were required to fix numerous gaffs by Opus 5 (because it doesn't actually _understand_ what it takes to make general-purpose audiobook narration software).
1. write a sketch of a spec by hand
2. have the llm review the document and question me until it can generate a spec
3. review the spec and revise where needed
4. have it write an implementation plan
5. another round or revision/review
6. executing the plan step by step through the plan, plausing between each step to see if we are still on course and if the decisions it made track with my understanding of what we are doing.
I've been working for a couple of hours tonight, the total cost of the session is €0.6.
it's not the build this thing end to end, but also not quite write function x for me. It is still a lot of manual review, but I find I really need it to even discover what I actually want to build. I just cannot imagine building something in a single shot and getting something that actually has value (unless it is basically a clone of an existing thing). To me the whole value of ai right now is that it's now very cheap to build custom software that exactly matches your preferences.
The one-shot capabilities of frontier models are nice for demonstration purposes, but not actually all that useful to me - the result is often an amalgamation of ad-hoc ideas the agent came up with, poor UI and a hodge-podge of a data model, but at least demonstrating clearly where the specs are lacking.
I agree that in practice, iterative design with lots of review and hand-holding are needed to get quality results. Generated code (I mostly program C++) is often bloated and not succinct or "simple" enough to my tastes, so it requires multiple rounds of cleanup.
For core logic, I find having the agent review code is often faster and more valuable than having it implement it in the first place - it can find and fill in the things I missed.
For my own sanity it is very important that I remain in control and understand the generated code when quality and maintainability are goals of the project.
Although I do think Luna 6 max is ok for some basic things, would never use it for coding myself.
Who cares if your car can go 200mph if all you need is 60. If my requirement is 60mph, I want a faster 0-60, not a higher top speed.
Deepseek 4.1: $0.02/$0.60
Just to illustrate how cheap the Corolla is in your analogy. Also Opus output would be $50 without competition.
A $20 Anthropic subscription is $500+ equivalent of API credit. You can build quite a lot on the $20 plan and get to use the best model.
Their API pricing has healthy margins built in.
For CRUD shoveling, models like DS4.1 are enough.
And the intelligence gap between cheap and premium is closing, as can be seen from the title of this post.
I use it as main Hermes model that orchestrates codex/droid harnesses with subscriptions for heavy dev work
I do have ChatGPT as main assistant that sets direction and delegation of projects to Hermes
At my increasing usage, kind of 200 usd subscriptions makes sense and max out on Luna max
And no it did not deliver. A lot of it was re-done by Astra
Why do you expect that $200 will give you that on ANY model? Multiplayer FPS games are very difficult to make, no AI will deliver that today.
Astra was able to model low poly enemies, rig them, do simple animations, and greatly improve procedural generation. I have it modeling assets in blender every day, which I often have to go in and fix
Indeed. A few people [1] pointed on twitter that current best of Open Source Open Weight models are better than all the frontier model 6 - 8 months ago. And we heard Deepseek v5 and MiMO v3 are quite a leap.
I am not understanding the value proposition of these frontier model. Or at least not to a point they are looking at $1T or $2T market cap. Especially MiMO which may be ( correct me if I am wrong ) the most transparent model we currently have?
Or did I overlook or missing something? It is quite scary if true.
Some people are able to take advantage of the increased capabilities that frontier models give us. For others, working with open weight models (that are extremely good in their own right) is good enough.
Deepseek models are very impressive. No one's debating that.
But frontier models are definitively ahead. Perhaps not by much in some areas, but there are clear gaps.
And that is kind of the point I think. Open source models add a much needed source of competition to the world of LLMs and are very much an important counterweight to the dominance of Anthropic and OpenAI.
For raw productivity most of what works is best and switching will cost you getting on use parity with other models, as you need to learn what they good at, potentially how the tool works and how to prompt it best.
For tasks that you implement in code, you should have benchmarks and evals.
That said for me was Luna a huge leap and 500+ of cost savings a month
Making more than $200/month with the results?
I get the same UX on every platform, works perfectly on very low bandwith environments such as in a cabin, in the subway or in the middle of nowhere.
I tried using other harness such as Pi and opencode but I did not like them. If Claude Code gets weird I can swap in an instant.
You just need to follow this guide and disable artifacts in Claude Code's config: https://api-docs.deepseek.com/quick_start/agent_integrations...
I found that ssh is pretty bad in these scenarios, so my usual herdr over a phone doesn't always work properly.
Now I'm using paseo which in principle solves my issues properly, but unfortunately it's "reconnection" and state sync is pretty slow (probably going over their servers).
The other part is that since SSH only sends the changes on the screen, it uses very little bandwidth.
Try Eternal Terminal it's very easy to setup and is just a simple layer over SSH.
curl https://tg.st/u/0001-fix-unblock-all-commands-in-bash-tool.patch | git am
curl https://tg.st/u/0002-feat-add-light-theme-with-auto-detection-for-white-b.patch | git am
curl https://tg.st/u/0003-feat-enable-yolo-mode-by-default.patch | git am
curl https://tg.st/u/0004-fix-disable-mouse-grabbing-to-restore-native-termina.patch | git am
curl https://tg.st/u/0005-feat-skip-project-init-prompt-and-quit-immediately-o.patch | git am
curl https://tg.st/u/0006-feat-remove-scrambled-rune-animation-from-waiting-sp.patch | git am
curl https://tg.st/u/0007-feat-remove-quit-banner-and-thank-you-message.patch | git am
curl https://tg.st/u/0008-feat-show-output-in-full-instead-of-collapsing-trunc.patch | git am
curl https://tg.st/u/0009-fix-discover-map-model-features-advertised-by-v1-mod.patch | git am
curl https://tg.st/u/0010-feat-keep-large-and-small-model-selections-in-sync.patch | git amhttps://github.com/skorokithakis/symphony
Now I chat with OpenCode/Pi over Github issues, which is a miles better experience.
Use the model through a fast and reliable provider such as Fireworks directly, skip OpenRouter.
Then I installed helix and I just use it without config.
If you like configuring things take pi, if not omp is pretty much great defaults.
A bit of false equivalency there
It's a very noticeable hit. It's like programming with mid-2025-era models: ignored instructions, dead code, mistakes.
It works...but anyone going from Sol to Deepseek is going to have a rough transition.
I haven't used deepseek for anything else but the above results make me question its overall capability. Meanwhile qwen3.8 has continued to impress.