DeepSeek Harness developer preview
deepseek.com
deepseek.com
I think everyone agrees the internet is drowning in AI slop. That makes the joke hard to appreciate on its own.
One consequence we liked: since plugins are just Cordis bundles, the same registry can be exposed to any MCP-speaking agent. We built a small MCP server that searches the dsh-plugin topic, inspects bundles, and can install/run them (catalog plane works without dsh installed). github.com/bobleer/deepseek-harness-plugin-mcp
about: Responsible bot.If the posts are actually from a bot - I would love to know which model is being used.
Only reason I said anything at all.
Presumably more people there have read hn at least some. But there's perspective having seen the ebb and flow of hn/the rest of the ecosystem for the last 16 years and how that intersects with deepseek culture.
It's clear openai culture is influenced by yc culture, which would be clear to an hn user. Google/ant/spacexai have influences hn users users would be familiar with. Hn users from them would potentially know the friendfeed connections to yc/vc/openai, 500 startups/techstars etc..
It's uncommon for me to have takeaway pizza, but I don't think anyone I know would be surprised when I do.
Your friends probably wouldn't go out of their way to ask you why you ordered a takeaway pizza which is uncommon. But if they did ask, it usually means they were surprised.
One of the reasons to bypass GFW is to avoid scrutiny.
> Doesn't this create a chilling effect?
It doesn't.
> Or are you claiming that it has zero effect and citizens of China have zero concerns about using websites and speaking freely?
Not “zero”, maybe like “0.01”? Better use an anonymous account if you want to criticize on gov.
That's correct, and as I said that's one of the reasons Chinese people bypass GFW.
The last paragraph in my previous comment may not have been clear enough. What I meant was that *after bypassing GFW*, you generally have minimal concern, not *before bypassing GFW*
We do hire people who are currently located outside of China (say North America, Europe, or anywhere) if they're open to work in our Beijing or Hangzhou offices.
Hope you can convince some more people join hn and answer questions on occasion.
I think julius has phrased things better than me.
one question is that do you think in the future harness would become more simpler and its behavior should match a guideline or we would add more complexities to make it more robust? Is it important to use the same harness for RL and inference?
Git might be worth adding to the top level. Currently you've got LSP, grep, glob nicely structured for non-mutating queries across a codebase, but git is behind bash and that means hope or sandboxing.
Thank you for uploading it. Gives a lot of insight into how the deepseek models might expect tool calls to be structured.
Everything the model sees is recorded in an append-only session log: system prompts, reasoning, tool calls and results, subagent scheduling, and every context injection. In the Trajectory view, you can inspect these records by source. Resume, fork, search, and replay all operate on the same event stream."
That's a killer feature, IMHO, and one that US models won't allow you to do, as their traces are encrypted, obfuscated, etc. and have to be extracted via various workarounds (that violate the terms of service).
If you want to be able to improve your tools that work with models, you have to be able to assess what the models think is happening, how they think about and interact with the data you give them. And, the US models won't let you see that.
Or point me in the right direction in terms of what to read.
By distilling, you are training a model via Reinforcement Learning (RL) to mimic the answer of a bigger model. To do that, you need all the steps that contributed to generating an answer.
To give you a better idea, imagine teaching a student how to solve math problem:
1. You give it the problem and the answer only (no thinking trace) 2. You give it the problem, the intermediary steps and the answer (full transcript)
I think you can agree that the second method is more likely to give a well-informed student.
In short, there are workarounds, but they're not guaranteed to work forever and they're likely to bump into terms of service.
It has an event sourced architecture in SQLite and it resolves queries using recursive CTEs (and sneaky projections to speed things up) to deliver exactly that. Identical, stable message chains to AI and complete introspection.
Bonus points include a constraint-satisfaction solver for the tiling window manager so windows never shrink too small to read. And many other keyboard-friendly features.
[0] https://www.dreamcoder.ai/ [1] https://www.dreamcoder.ai/assets/graph.webp
to swelljoe below. they stealth edited it after my comment. see other examples.
Anyway, this particular harness isn't doing anything unique, but the combination of an official agent intentionally keeping the data and making it accessible to the user and a model API that provides all the information is unusual and worth calling out. It used to be common, most APIs and models and agents showed the reasoning, or could be configured to do so. Most no longer offer it.
For those who want to know what it achieves: it adds hot-reload and dynamic enable/dispose capabilities to a plugin system, like the one in Pi agents, though they push the boundaries further, to the UI components and so on.
For those who want to know what it does: if you have some PLT knowledge, ask your agent to explain the algebra to you better; for those who aren't familiar, the framework requires each plugin to provide how it initializes and how it destructs (like C++'s RAII, Rust's Drop trait and so on), and the runtime will then properly handle the lifecycle events and the common pitfalls. In addition, it provides a clean way to declare the dependencies between plugins, and the runtime will also properly process the lifecycle changes on a broader plane.
I think it's worth reading if you are not familiar with OSGi, iPOJO, React's useEffect and so on (which the paper itself mentions); for others, a skim is enough: it does point out the gotchas for some common problems, but the algebra may not help you further.
If anybody has tried it, does it let you preview components in any frontend framework with perfect fidelity? That would be a big win.
That actually sounds amazing.
Yep sounds just like the Eclipse IDE plugin system indeed. Nice example of things being rediscovered every generation I suppose.
a plugin's registrations returning individual cleanup handlers is nice. in pi, you clean up all registrations in one go in the session-shutdown handler.
i also like the use of generator to to clean up partial registrations nicely.
the cross-plugin dependency injection and resolution i'm not so sure about. it comes with a lot of footguns and limitations as pointed out in the paper.
works ok within a single compilation unit, i.e. a plugin with many modules. does not help with typing of cross-plugin dependencies.
most plugins do not have dependencies on each other, so this more complex system doesn't win you much, e.g. with load order and conflicting registrations (i.e. two plugins registering the same tool).
being able to reload a single plugin on change while letting the others not in its dependents list jug along is neat. but that also only works if plugins actually declare dependencies (see last paragraph), and also has a lot of limitations. and the simple case, a plugin with no dependencies or dependents, which i'd say is the 90% case, does 't need that complexity either.
definitely cool stuff tho! remains to be seen how well it works in a real plugin ecosystem.
> they push the boundaries further, to the UI components
can you elaborate on this? pi extensions support contributions to the UI. in pi v1, they are limited to in-process UI. v2 splits server and client, and with that UI.
Every product relying on "community plugins" for their features implies it works fine the 6 first months, then it's a nightmare of incompatible, deprecated, incompatible plugins, with no consistency and no governance.
I understand how attractive it can be to companies to think, hey, let's make a very small product and rely on other people to make features, and I hope it works, but I'm personally staying away from that.
AI can write custom plugins for you. So this means the tool is infinitely flexible for you, even without any community.
Compare this to Zed where I can't make a hexviewer for binary files or player for audio files for myself without recompiling Zed's source code.
Truth is that useful dev workflows and tooling probably coalesces in a tight band. There's really no point in re-inventing the wheel over and over again at this level (the raw tooling).
Eclipse has been thriving since 2002 mostly by virtue of being able to coordinate developers via plugin's and a business-friendly license.
They did need to upgrade early plugins into OSGI, and most of the new plugin designs benefit from copying OSGI, et al. The key is SAT solvers for dependencies and namespace separation, not forcing clients into the same dependency version.
But as you suggest, relying on the community is a moral hazard. In Eclipse there were big players willing to fund key use-cases for their own purposes; elsewhere I've seen sufficient monetization of plugins to offer incentives and stability.
I would add that VSCode plugins follow a different development model. While any OSGI/Eclipse plugin can provide an interface, I believe in VSCode you're limited to the API's they give you (and they make a mess of them, so there's more inconsistencies e.g., in LSP support that anyone can enumerate).
Aren't VS Code, Claude Code, Hermes Agent, Goose or Letta harnesses, but with UI, too?
"Everything the model sees is recorded in an append-only session log: system prompts, reasoning, tool calls and results, subagent scheduling, and every context injection. In the Trajectory view, you can inspect these records by source. Resume, fork, search, and replay all operate on the same event stream."
Seems pretty helpful - have sort of wanted something similar (I use Pi).
They also released this research paper that backs their whole plugin composability system that seems pretty cool: https://github.com/cordiverse/paper
But the future is here and thus it's called "Agentic causality's reified temporal traceability."
[0] https://www.dreamcoder.ai -> scroll down to the event graph.
<quote> I'm glad they're doing this also and that more people are adopting it. Event sourcing [0] is the right way to represent informaiton like tool calls, user interactions, etc. --- it makes it easy to fork conversations and maintain a cohesive conversation stream and stable message history that does not break the cache. </quote>
why dont you mention that its your site instead of prenteding like something you discovered?
Good to know I was not the only one confused. Reads like word salad!
Honestly I would not be surprised when it actually IS claude using those resources... It is very clearly vibed
But modern bloat manages perfectly well to make apps that wait for network calls run poorly enough to give you a bad experience.
(I actually have/am writing a harness in Java fwiw, but mostly as a hobby/experimentation)
Having Oracle's tramp-stamp on it may have been the final kiss of death in terms of totally-superficial "coolness" factor.
IMHO, Microsoft made the correct approach on .NET.
For LLMs, I prefer C# and C++ instead of TypeScript, JavaScript or Python as the static + compiled language factor keeps the coding agents on track. Plus, they have a true threading/async implementation.
The actual physical RAM is still entirely available to other applications. It's just made the OS know it might want that many pages. Until there's data in the pages, they will not count towards total RSS.
It's the kind of things some sysadmins used to gripe to me about and I would question whether they should be in charge of a machine at all.
To repeat: just because an application mmaps a large region doesn't mean the OS has actually given it all that physical RAM. It's merely made sure the pagetable knows about it.
If the program is actively using that allocation, that's fine. My problem is with the runtime hoarding RAM when it should have been freed after GC back to the OS.
Then there's also the JVM not handling peaks well because it hit the max heap size, while you still could rely on the OS doing its job to shuffle stuff to swap temporarily. I still see JVM OOMs in my $dayjob's product while the OS has plenty of free physical memory. It is stupid.
I mean, we have malloc() and free(), they are in the stdlib for a reason :)
The JVM seems to follow a philosophy where it assumes it is the only process running besides PID 1, which is valid for some scenarios, but not for others.
In pretty much every other runtime's case you are stuck with whatever their GC uses, and almost every other GC is far less advanced than the JVM's implementationS. And manual memory management is not free of tradeoffs either, e.g. RAII can have pretty long destruct chains in both C++ and Rust.
Because java is still one of the top 3 languages by any ranking worth its salt (not you, tiobe).
Not saying that's a good or a bad thing. I still think the JVM is remarkable.
fast iteration is for POCs. once you have the app built and working, you need performance and stability much more than fast iteration
You do know that node is event driven, right?
I stopped paying attention the third time they redefined matrix arithmetic semantics. That happened to be around the 100th time I was sent a script and it only ran on the author’s machine. Maybe they will fix it some day. When they do, I will not believe it.
In contrast, TS has a much nicer type system and better async support. It runs well on web, mobile, desktop and server. Yes, sometimes you have to ship node.js or a whole web browser, but the tooling for that is slightly less insane than the analogous tooling for python.
Its language interoperability story is slightly nicer too (invoke native code, or use wasm). It’s UI story is much, much better since it reuses all the web stuff.
Pip practically invented the supply chain attack; npm perfected it. That’s probably a draw.
Of course, if you care about performance, then other choices make more sense. If you’re training a model then python probably still wins, but very few customers have a $1M+ machine.
Any reason why it should not be written in nodejs?
For web stuff, sure.
But for CLI, it never made sense to me. Especially when Python and Go exist.
But why? Not saying node is better, just want to know where you are coming from for my own knowledge.
Bc I would have picked typescript + node too. It has types (where python just has type hints) and a lot of developers know it already (where go is more niche).
I was going to say you cannot easily distribute a nodejs based CLI app, but that’s of course not true. devcontainer-cli is a nodejs app and so are many of the coding agent harnesses.
Yeah, thanks for pushing back. I guess my view was irrational.
There’s an interesting counter example for DeepSeek called CodeWhale, though:
1. The first significant agentic harness was made by Anthropic.
2. One of the most senior developers of client-side software at Anthropic is Felix Rieseberg, one of the original creators of Electron. [1]
3. After Claude Code blew up, everyone else copied Anthropic.
---
1: https://daringfireball.net/2026/07/claudes_criminally_bad_ma...
smol has implementations in Go, Python, Clojure, PHP
https://github.com/smol-env/smol
out of the box an agent only needs to be able to do http requests and call tools (which might again be just http requests or shelling out)
there is no inherent reason for why an agent has to be in JavaScript or Typescript
but they are popular languages and come with runtimes and libraries for http requests, steaming, TUI (terminal ui) and so on which can help
Edit: okay I read the code, it's actually four separate implementations
I'm currently working on more 'feature-full' but still minimal variants
e.g. a python variant with automatic compaction + truncation of sh output
https://x.com/__tosh/status/2087606344035479632
i also got quite a lot of requests to provide the code in non-golfed form to make the implementation more approachable and idiomatic in each language (will do!)
I want something that actually has an opinion and gives me productive value without having to spend days reconfiguring it first.
it was widely ridculed at that point but now i am not so sure.
oh i mean 'now i am not sure if it would be ridiculed'
ppl are not doing this right now . right?
if ai can really do this then we dont even need all this glue software . everyone can just use lovable.
But at the core science/tech of AI it's probably the most amount of innovation I've ever witnessed in a field. The pace of new developments is staggering.
Second, if the repo had hooks and instructions for the LLM or user to blindly install/enable the hooks, we'd instead be complaining about security risks and what might happen if the repo is compromised at some point in the future.
Third, sometimes you don't want to mechanically enforce things via git hooks because it impacts your use when what you're really trying to codify and enforce are the LLM's actions. In that case you can enforce mechanically via hooks at the harness level.
And finally, git hooks are a great solution for upstream repositories to enforce quality and protect branches. But it means that the upstream is the one running the checks. It makes the upstream a potential bottleneck - better to have the leaf nodes run the checks locally and fix any issues before pushing it upstream rather than push upstream, wait for results, make changes, push upstream, wait for results, make changes.
Second, the point isn't about a specific repo, it's the general tendency to rely on fuzzy .md files scattered all over the place. And I really don't see how letting the output of a language model run a one time command is more secure than running a script.
Third, "nothing applies in all context"? Yeah, obviously. And harness hooks (at least with Claude code) are still more suggestions than anything else. The only way I've found is literally rejecting a tool use and forcing it to recall in the proper way, which of course makes for more token usage. I wonder who benefits from that.
Finally, no idea what you are arguing against. Use git hooks where they make sense, local or remote.
"this, like all other problems in Computer Science, can be solved by one more level of indirection." Roger Needham, circa ~1981
9 out of 10
Edit: After creating an app it works as expected, no complains, lots to celebrate, being version 0.1 there is room for more surprises but right now it's the perfect tool for those initiating in agentic coding with one of the most affordable and powerful AI. It is really wonderful.
Using memory to track inverses does not scale.
This other day I was looking at that “caveman” skill, and was shocked to see it evolved to become a company, and, in one of its modes, the highest form of compression seems to be “Wenyan” which is Classical Chinese.
Should I get started on learning Chinese?
Also, "less tokens" is not always straight forward. I doubt it's a coincidence that the cavemen skill (or now proxy, I guess) has lots of numbers, but not a single benchmark on model performance or actual per-task token savings
For example one paper I remember found that without CoT, just stating your prompt twice increases model performance. With CoT, the same function is served by the CoT restating the important parts of your question. Something about which tokens can affect which other tokens in attention implementations
I'm finding more and more there seem to be sort of niche prompting skills that are important to be aware of
Why I left that idea is because as a developer I know that was needed but I have limited time so I need to build that is really next path forward.
I am working on whole dev space that can run on my Mac M4 or similar specs. I needed to revamp everything (LLM thinking) from ground up even models. My idea is mixing deterministic nature of existing tooling (non-LLM tooling) with non-deterministic nature of LLMs.
This is the fundamental idea behind every LLM harness.
It was very easy to connect the harness to the local model and it seems to run quite fast, compared to other harnesses that I have tried.
I see that it works with many different providers out of the box and that's a great thing. It also makes it easy for me to build a plugin for the role-model router and have it work properly, so you can route between models automatically. Will be out later today.
Do the first party harnesses really have an advantage when paired with the maker's model?
I also just do a bit of hand-coding to guide the agent still.
I worry the $200 / month plans are loss-leaders encouraging you to maximize token usage to churn out slop, rather than thoughtfully use coding agents in a way that still engages your brain, and produces good software.
Anyway very happy with it, I use it as a plugin to RubyMine and Webstorm.
One of the primary advantages is being able to choose your model - and it often has free deals for newer models that are running promotions. Whenever I switch to Claude Code it seems clunky. Would rather use Claude with Cascade.
I want to use the same consistent working surface across models in the same way I want to use the same text editor across all different languages
Just like Obsidian, there's also hot loading.
Does what it says on the tin. Great work.
I consider my own coding agent bloated at just 1mb (yes 1mb) because it uses postgresql package as db tool, and it works wonders.
* edit 1: Upon further scrutiny, 35 dependencies make up for 1.4gb, what they are for? I don't even see postgres in there so I guess that would be another plugin. 1.5gb of basic functionality?
* edit 2: Most of the time I use the terminal but also developed a web ui for my agent [1] and it is only 20mb with postgres, git, web, file tools, etc I definitely want to know why the bloat
UI is simple, we shouldn't complicate stuff
Everything is a skill, backed by a CLI tool that both I and the agent can use and debug.
To install the harness, first use npm...
And tab is closed. No thanks.
in the era of AI, telling me that the core design is a plugin system that can be reloaded and extended easily is just not exciting.
it is something you feel excited 20 years ago back in the 2000s, in 2026, the expectation is agentic capabilities and self improving.
Did they discover Unix pipes?
oof
this looks like a genuinely new one
The documentation, built from repo, is available here: https://deepseek-harness.github.io/deepseek-harness/en/guide... (I find the development and reference sections easier to read and navigate)
Deepseek Harness supposedly has nice plugin system which others do mostly lack.