HNHacker News
TopNewBestAskShowJobs

schmuhblaster

304 karma · joined November 15, 2025

submissionscomments
schmuhblaster··on How Kepler built verifiable AI for financial services with Claude
Shameless self-plug: https://github.com/deepclause/deepclause-sdk/

The idea is to take markdown instructions and "compile" them into a Prolog-based DSL that orchestrates both deterministic and LLM-based components. The (meta-)interpreter of the DSL automatically tracks the entire execution process, so that the final ouput becomes observable and more explainable. Still at an early stage, but I am having lots of fun with it and would love to explore possible use cases.

schmuhblaster··on Agent Safehouse – macOS-native sandboxing for local agents
I am experimenting [0] with compiling markdown to a DSL first. Then running a static analysis on the DSL code. Still at an early stage though.

[0] https://deepclause.substack.com/p/static-taint-analysis-for-...

schmuhblaster··on Ask HN: What are you working on? (February 2026)
https://www.github.com/deepclause/deepclause-sdk

"Compile" Markdown specs for SDD or Subagents into executable logic, e.g. CodeAct+DSPy+Prolog. Not sure how and if I will continue, but it's been lots of fun.

schmuhblaster··on OpenClaw is changing my life
I'd be curious if a middle layer like this [0] could be helpful? I've been working on it for some time (several iterations now, going back and forth between different ideas) and am hoping to collect some feedback.

[0] https://github.com/deepclause/deepclause-sdk

schmuhblaster··on Show HN: AgentVM – Safe, Sandboxed Linux VM for OpenClaw and AI Agents
Network can be closed down. It also uses a userland network stack, so future iterations might include being able to define rules for ingress and egress.
schmuhblaster··on The Missing Layer
I've been recently experimenting with using a Prolog-based DSL [0] as the missing layer: Start with a markdown document, "compile" it into the DSL, so that you obtain an "executable spec". Execution still involves LLMs, so it's not entirely deterministic, but it's probably more reliable than hoping your markdown instructions get interpreted in the right way.

[0] https://github.com/deepclause/deepclause-sdk

schmuhblaster··on Coding Agent VMs on NixOS with Microvm.nix
Or try this: https://github.com/deepclause/agentvm, it's based on container2wasm, so the VM is fully defined by a Dockerfile.
schmuhblaster··on Sandboxing AI Agents in Linux
My attempt at a portable solution: Linux VM inside WASM for sandboxed execution: http://agentvm.deepclause.ai

Minimal dependencies, but not as fast as containers or bubblewrap.

schmuhblaster··on Show HN: DeepClause CLI – Compile Markdown specs into executable logic programs
Hi all, been building this in public for some time. This is my latest iteration, hoping to get some interesting comments and feedback here, before I go down yet another rabbit hole and unnecessarily burn through even more premium tokens.
schmuhblaster··on Show HN: Amla Sandbox – WASM bash shell sandbox for AI agents
Thank you! When I started working on agentvm my original goal was similar to yours, build a kind of Mingw or Cygwin for WASM. However, I quickly learned that this wouldn't really be feasible with reasonable amounts of time/token spend, mostly due to issues like having to find a way to make fork work, etc. I am no expert for WASM or Linux system programming, but it's been a lot of fun working on this stuff. I hope that the WASI standard and runtimes become more mature, as I feel that WASM sandboxes make a lot of sense in environments where containers are not an option.
schmuhblaster··on Show HN: NPM install a WASM based Linux VM for your agents
linux-wasm is an awesome project, but relies on compiling the kernel itself into WASM. This seems to work in principle, but is still a bit unstable. But I do hope that eventually one can get rid of the emulator in the middle as is done in c2w.
schmuhblaster··on Show HN: NPM install a WASM based Linux VM for your agents
Thought this could be useful. Opus 4.5 helped me build most of it, including a simple network stack so that the VM may access the outside world. Still at a very early stage, but I think it looks promising.
schmuhblaster··on My Gripes with Prolog
Opus 4.5 helped me implement a basic coding agent in a DSL built on top of Prolog: https://deepclause.substack.com/p/implementing-a-vibed-llm-c.... It worked surprisingly well. With a bit of context it was able to (almost) one-shot about 500 lines of code. With older models, I felt that they "never really got it".
schmuhblaster··on My Gripes with Prolog
>> I think much of the frustration with older tech like this comes from the fact that these things were mostly written(and rewritten till perfection) on paper first and only the near-end program was input into a computer with a keyboard.

I very much agree with this, especially since Prolog's execution model doesn't seem to go that well with the "successive approximations" method.

schmuhblaster··on My Gripes with Prolog
I also have a strange obsession with Prolog and Markus Triska's article on meta-interpreters heavily inspired me to write a Prolog-based agent framework with a meta-interpreter at its core [0].

I have to admit that writing Prolog sometimes makes me want to bash my my head against the wall, but sometimes the resulting code has a particular kind of beauty that's hard to explain. Anyways, Opus 4.5 is really good at Prolog, so my head feels much better now :-)

[0] http://github.com/deepclause/deepclause-desktop

schmuhblaster··on Cowork: Claude Code for the rest of your work
Is there any reasonably fast and portable sandboxing approach that does not require a full blown VM or containers? For coding agents containers are probably the right way to go, but for something like Cowork that is targeted at non-technical users who want or have to stay local, what's the right way?

container2wasm seems interesting, but it runs a full blown x86 or ARM emulator in WASM which boots an image derived from a docker container [0].

[0] https://github.com/container2wasm/container2wasm

schmuhblaster··on How to code Claude Code in 200 lines of code
As an experiment over the holidays I had Opus create a coding agent in a Prolog DSL (more than 200 lines though) [0] and I was surprised how well the agent worked out of the box. So I guess that the latest one or two generations of models have reached a stage where the agent harness around the model seems to be less important than before.

[0] https://news.ycombinator.com/item?id=46527722

schmuhblaster··on Implementing a (Vibed) LLM Coding Agent in Prolog
Thank you! I went with a Prolog base, because I was interested in what might be possible when combining its execution model with LLM-defined predicates. For anything related to modelling and querying data, a Datalog dialect might indeed be a better choice. I've also used Logica [0] as an intermediate layer in a text2sql system, but as models get better and better, I believe there is less need for these kinds of abstractions.

[0] https://logica.dev/

schmuhblaster··on Recursive Language Models
Hi, I stumbled on this article in my twitter feed and posted it because I found it to be very practical, despite the somewhat misleading title. (and I also don't like encoding agent logic in .md files). For my side project I am experimenting with describing agents / agentic workflows in a Prolog-based DML [1]

[1] https://www.deepclause.ai

schmuhblaster··on GPT-5.2
Sounds interesting, could you elaborate a bit on this? (I am experimenting in a similar direction)
schmuhblaster··on The "confident idiot" problem: Why AI needs hard rules, not vibe checks
This looks like a very pragmatic solution, in line with what seems to be going on in the real world [1], where reliability seems to be one of the biggest issues with agentic systems right now. I've been experimenting with a different approach to increase the amount of determinism in such systems: https://github.com/deepclause/deepclause-desktop. It's based on encoding the entire agent behavior in a small and concise DSL built on top of Prolog. While it's not as flexible as a fully fledged agent, it does however, lead to much more reproducible behavior and a more graceful handling of edge-cases.

[1] https://arxiv.org/abs/2512.04123

schmuhblaster··on Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022)
> But my bet is that the proposed program-of-thought is too specific

This is my impression as well, having worked with this type of stuff for the past two years. It works great for very well defined uses case and if user queries do not stray to far from what you optimized your framework/system prompt/agent for. However, once you move too far away from that, it quickly breaks down.

Nevertheless, as this problem has been bugging me for a while, I still haven't given up (although I probably should ;-). My latest attempt is a Prolog-based DSL (http://github.com/deepclause/deepclause.ai) that allows for part of the logic to be handled by LLMs again, so that it retains some of the features of pure LLM_based systems. As a side effect, this gives additional features such as graceful failures, auditability and increased (but not full) reproducibility.

schmuhblaster··on Claude Advanced Tool Use
I've been experimenting with giving the LLM a Prolog-based DSL, used in a CodeAct style pattern similar to Huggingface's smolagents. The DSL can be used to orchestrate several tools (MCP or built in) and LLM prompts. It's still very experimental, but a lot of fun to work with. See here: https://github.com/deepclause/deepclause-desktop.
schmuhblaster··on Solving a million-step LLM task with zero errors
Thank you, https://catala-lang.org/ looks very interesting. I've experimented a lot with LLMs producing formal representations of facts and rules. What I've observed is that the resulting systems usually lose a lot of the original generalization capabilities offered by the current generation of LLMs (Finetuning may help in this case, but is often impractical due to missing training data). Together with the usual closed world assumption in e.g. Prolog, this leads to imho overly restrictive applications. So the approach I am taking is to allow the LLM to generate Prolog code that may contain predicates which are interpreted by an LLM.

So one could e.g. have

is_a(dog, animal). is_a(Item, Category) :- @("This predicate should be true if 'Item' is in the category 'Category'").

In this example, evaluation of the is_a predicate would first try to apply the first rule and if that fails fallback on to the second rule branch which goes into the LLM. That way the system as a whole does not always fail, if the formal knowledge representation is incomplete.

I've also been thinking about the Spec->Spec compilation use case. So the original Spec could be turned into something like:

spec :- setup_env, create_scaffold, add_datamodel,...

I am honestly not sure where such an approach might ultimately be most valuable. "Anything-tools" like LLMs make it surprisingly hard to focus on an individual use case.

schmuhblaster··on Solving a million-step LLM task with zero errors
My own attempt at "chain-of-code with a Prolog DSL": https://news.ycombinator.com/item?id=45937480. Similarly to CodeAct the idea there is to turn natural language task descriptions into small programs. Some program steps are directly executed, some are handed over to an LLM. I haven't run any benchmarks yet, but there should be some classes of tasks where such an approach is more reliable than a "traditional" LLM/tool-calling loop.

Prolog seemed like a natural choice for this (at least to me :-), since it's a relatively simple language that makes it easy to build meta-interpreters and allows for a fairly concise task/workflow representations.

schmuhblaster··on Learn Prolog Now (2006)
This is my own recent attempt at this:

https://news.ycombinator.com/item?id=45937480

The core idea of DeepClause is to use a custom Prolog-based DSL together with a metainterpreter implemented in Prolog that can keep track of execution state and implicitly manage conversational memory for an LLM. The DSL itself comes with special predicates that are interpreted by an LLM. "Vague" parts of the reasoning chain can thus be handed off to a (reasonably) advanced LLM.

Would love to collect some feedback and interesting ideas for possible applications.

← PreviousPage 2 of 2