HNHacker News
TopNewBestAskShowJobs

mickdarling

286 karma · joined September 12, 2010

Builder, and Founder

Dollhouse Research — composable AI tools, mostly open source (dollhouseresearch.com):

- DollhouseMCP: Open-source AI customization with AI safety as at it's core. Turn prompts into modular building blocks you can create, activate, combine, version, and share — Personas, Skills, Templates, Agents, Memories, and Ensembles. (dollhousemcp.com)

- Dollhouse Collection: community catalog of those elements (collection.dollhousemcp.com)

- MCP-AQL: a protocol spec that gives MCP tool calls semantic meaning — Create/Read/ Update/Delete/Execute endpoints that cut token overhead and make safe-vs-destructive operations explicit at the protocol layer (mcpaql.com)

- AILIS: an OSI-inspired, layered model that gives AI systems a shared vocabulary, so teams can tell a routing problem from a tool-invocation, memory, or UX one

- Also check out Merview.com, a local-first Markdown and Mermaid editor and viewer that runs entirely in your browser, no login

Elemental Surveys (elementalsurveys.com):

- Fast-turn market research — turn a brief into a QA'd questionnaire, analysis, charts, and a branded report in hours instead of weeks

Some Previous Projects:

- Created Tomorrowish.com a social media DVR used by Hulu, Fox, Turner Broadcasting, and more

- Cofounder and CEO of TVDuffle a Travel and booking app for concert goers and sports fans.

- Cofounder Biostrut 3D Printed organic tissue scaffold

ping me at my firstname@myfullname.com

My personal blog can also be found at http://mickdarling.com

submissionscomments
mickdarling··on Hister – A private, full content search index that you control
Literally the first thing I thought of. Is he trying to reference an Antichrist?
mickdarling··on The New MCP Roadmap
I created MCP AQL, which is an extension to the MCP spec, specifically to reduce the bloat for MCP tools.

It only has five CRUDE endpoint: Create, Read, Update, Delete, and Execute using a GraphQL-like structure for tool calling of the operations within the endpoints. It's very efficient, and robust. there's all kinds of exemplar tools and components to make adapters for any MCP server. You don't even need to rewrite your own MCP server. Just create an adapter for it.

All open source at MCPAQL.com

mickdarling··on Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
If it stores unique information, by definition it can store arbitrary information because it can point to arbitrary information. So they can have it relate to anything they want. Even a full breakdown of the original text if they choose.
mickdarling··on Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
Watermarks are context poisoning.
mickdarling··on Anonymous GitHub account mass-dropping undisclosed 0-days
If only the AI tools didn't shut you down every time you were trying to red team your own tools. I've had to come up with all kinds of workaround scenarios, effectively bypassing the AI security processes in order to stress test my own systems.
mickdarling··on If Claude Fable stops helping you, you'll never know
No, this is their get out of jail free card if people start complaining about the model being dumb or forgetful or lying, they can just say, oh well, you must have been doing something that triggered its distillation prevention technique.

And, they can say that for anybody at any time, and you'll never know why, and there's no way to prove it.

Everyone needs a flight data recorder to prove... "here's what I was actually doing and why it was not distillation." And now you're having to prove your innocence instead of them having to prove you're guilty, and really at the end of the day, it's just the model being stupid that they're protecting themselves from.

mickdarling··on Claude Fable 5
The tags are actually displayed in raw text not rendered.
mickdarling··on Claude Fable 5
Below is the EXACT text in Claude Desktop introducing Fable 5, including the very professional looking break tags, and at least I know where the links begin and end by looking at the anchor tag there.

They obviously put their best model on the job to build that.

----------------------

Fable 5: Our most capable model yet Our newest model tackles your biggest challenges with fewer check-ins needed.

• <b>Included in your plan limits until Jun 22</b><br><br>Fable takes 2× the usage of Opus. • <b>Switch models when a message is flagged</b><br><br>When safety measures flag a message, automatically switch to a different model to keep chatting. When off, your chat will pause instead. <a href="https://support.claude.com/en/articles/15363606" target="_blank" rel="noopener noreferrer">Learn more</a>

mickdarling··on Claude Opus 4.8
I effectively distill the frontier models by building whole sets of skills, personas, and other artifacts that I can then run on smaller models and get 10% even 20% improvements on models like haiku or local models.

There's a lot of room for improving the smaller models at many levels of the stack.

mickdarling··on Thoughts and feelings around Claude Design
I used it today to take a look at my previously built design system with Logos, branding, fonts, and everything else. After a lot of annoying tweaking back and forth, finally, I got something that was satisfactory.

Then I looked at the usage and it said I had used 95% of my Claude design usage for the week!

This isn't a real tool. This is a plaything, if that's what they're providing as examples.

mickdarling··on Simple self-distillation improves code generation
I'm working on a tool to determine which portions of an LLM process can be optimized, and how to measure that optimization and check whether it's optimizable at all. The shaping pattern that they talk about here is directly relevant and makes a whole lot more processes potentially optimizable by looking at the pattern rather than if the metrics just go up or down.
mickdarling··on Claude Code's source code has been leaked via a map file in their NPM registry
Hey, LLM, take a look at these multiple hundred emails and docs in my docs folder from the last few years, before I started using AI, that I wrote personally. create a list of all of the idiosyncrasies that I have in my writing. Create a file to remember that. And then use that to write any new text that'll be published so it sounds like my authentic voice. Thank you.
mickdarling··on The risk of AI isn't making us lazy, but making "lazy" look productive
Maybe for you reading a paper deeply is the most constructive way that you have to absorb information.

For me, it is having a document and interrogating it. Maybe having many sets of documents about a whole category of information. Getting the bullet points. getting the high level and then interrogating and digging down and being able to get bubbled up information as I need it.

That is the learning style that matches how I learn.

I have never been able to skim, so reading a large document WILL teach me that topic, but getting through that doc is tough.

I can dump a very large set of docs in a reader that lets me interrogate the whole data set and I can fly through looking for what is interesting to me, and what I may need, and along the way I will likely dive into other parts too. Asking questions keeps my hyperfocus active.

I think it is just a different style. I have synesthesia and a hard time not working on three to five things at once. I am use to knowing I learn differently than others.

mickdarling··on [dead]
Might be worth something to create an AI summary of the actual documents that are behind the curtain. That way the Client can have some assurance of what the deliverables are. It's a nice ability for the escrow service to be a trusted third party to verify something is there that matches about what is expected. If you do it right, you can even have the LLM prepare the content, encrypt it, and you never even have access to any of that information.
mickdarling··on OpenClaw is a security nightmare dressed up as a daydream
I don't use Claw. It is way too dangerous. I built my own system where I know the ins and outs and how they can break.

When it comes to agents' tasks, I tend to focus on things that I couldn't do before without automated agents, at least at the going price.

The kind of automation I'm doing is more like building a set of agents to generate marketing surveys for me. They take free form input from me and my project. They aren't particularly sexy but they go off and do something valuable that I literally would never pay for at the prices that they are normally.

mickdarling··on Zulip.com Values
You could literally drop this into Claude Code or Codex and point it at a local fork of Zulip and have it build your bimodal version with triage and grazing styles.
mickdarling··on GitHub Agentic Workflows
I use an LLM behavior test to see if the semantic responses from LLMs using my MCP server match what I expect them to. This is beyond the regex tests, but to see if there's a semantic response that's appropriate. Sometimes the LLMs kick back an unusual response that technically is a no, but effectively is a yes. Different models can behave semantically different too.

If I had a nice CI/CD workflow that was built into GitHub rather than rolling my own that I have running locally, that might just make it a little more automatic and a little easier.

mickdarling··on GitHub Agentic Workflows
It looks like it does have an MCP Gateway https://github.com/github/gh-aw-mcpg so I may see how well it works with my MCP server. One of the components mine makes are agent elements with my own permissioning, security, memory, and skills. I put explicit programatic hard stops on my agents if they do something that is dangerous or destructive.

As for the domain, this is the same account that has been hosting Github projects for more than a decade. Pretty sure it is legit. Org ID is 9,919 from 2008.

mickdarling··on LLMs could be, but shouldn't be compilers
This is where the desire to NOT anthropomorphize LLMs actually gets in the way.

We have mechanisms for ensuring output from humans, and those are nothing like ensuring the output from a compiler. We have checks on people, we have whole industries of people whose whole careers are managing people, to manage other people, to manage other people.

with regards to predictability LLMs essentially behave like people in this manner. The same kind of checks that we use for people are needed for them, not the same kind of checks we use for software.

mickdarling··on Clawdbot - open source personal AI assistant
I'm looking at it right now as a tool I can hollow out and stuff in my own MCP server that also has personas, skills, an agentic loop, memory, all those pieces. I may even go simpler than that and simply take a look at it's gateway and channels and drag those over and slap them onto the MCP server I have and turn it into an independent application.

It looks far too risky to use, even if I have it sequestered in its own VM. I'm not comfortable with its present state.

mickdarling··on Cursor's latest “browser experiment” implied success without evidence
Thank you for the note. It's not a site I used all that often.

Whether you had anything to do with it or not, I have no idea. And, since you didn't follow best practices and tell me directly rather than trying to score points here, there's really no way of knowing whether you're the one who caused the problem in the first place.

I built a new site without Wordpress. That took in less than a day.

I don't imagine you will alter your behavior to align with general best security practices anytime soon.

mickdarling··on Cursor's latest “browser experiment” implied success without evidence
Wasn't my project to manage. That was a consulting gig. And I fired the client right after this.
mickdarling··on Cursor's latest “browser experiment” implied success without evidence
Had humans not been doing this already, I would have walked into Samsung with the demo application that was working an hour before my meeting, rather than the android app that could only show me the opening logo.

There are a lot of really bad human developers out there, too.

mickdarling··on First impressions of Claude Cowork
I think Claude Cowork should come with a requirement or a very heavily structured wizard process to ensure the machine has something like a Time Machine backup or other backups that are done regularly, before it is used by folks.

The failure modes are just too rough for most people to think about until it's too late.

mickdarling··on Show HN: Ferrite – Markdown editor in Rust with native Mermaid diagram rendering
This looks cool! And, to add to the list of shameless plugs for OSS markdown editors/renderers With mermaid support, I will add mine to the list:

https://merview.com with full source code at https://github.com/mickdarling/merview

mickdarling··on Claude Code Unable to generate a AGPLv3 license due to content filtering policy
It certainly doesn't seem to have a trouble creating MIT licenses, that's for sure. I've had it insert an MIT license against my express direction instead of the AGPL license.
mickdarling··on Claude Code Unable to generate a AGPLv3 license due to content filtering policy
Claude code refuses to generate AGPL 3.0 licenses or modify them. I've run into this with several projects.
mickdarling··on Tokamak experiments exceed plasma density limit
That's a fascinating path forward and most frustrating of all, if I understand this right, this was something that could have been discovered 20 years ago.
mickdarling··on Show HN: AutoLISP interpreter in Rust/WASM – a CAD workflow invented 33 yrs ago
Ah, AutoLisp, that brings back the memories.

I can tell you that a good number of the design drawings for the higher floors in the Venetian resort in Las Vegas were assembled with AutoLisp scripts. The scripts I created grabbed components from other drawings that were already made to assemble a first pass set of drawings for floors that hadn't been fully designed yet, since the floors all had components of other floors.

They were still in the design process for the upper floors, while the lower floors had already been finished and they were moving up the building.

mickdarling··on I didn't realize my LG TV was spying on me until I turned off Live Plus
I used to work in the industry. I know the guys responsible for real-time data capture from various platforms like Roku and Visio.

I 100% agree, and I own very nice LG TVs. They are not connected to the internet. They each have an Apple TV and that is their only way that they get video, and can't send data out.

Page 1 of 5Next →