HNHacker News
TopNewBestAskShowJobs

charlie-haley

46 karma · joined June 9, 2025

submissionscomments
charlie-haley··on Show HN: Marmot, context layer for agents and humans
Hey, good question. Marmot is designed to be as generic as possible. An "Asset", whether it's a database, glossary term, topic, API or anything else, has the exact same schema, API endpoint and MCP tool. The MCP server exposes 3 tools: discover_data, find_ownership and lookup_term. Scoping happens as filters and arguments to discover_data, and it's summary-first, so a broad query returns counts or provider breakdowns rather than dumping every asset into context, and the agent narrows from there. The search step is really the primary interface here.
charlie-haley··on Postgres: One Database to Rule Them All
I wrote this after spending weeks figuring out how far Postgres could go before reaching for dedicated search indexers or a graph database.

Pretty far, as it turns out. pg_trgm can be used for fuzzy matching, there's built-in full-text search with tsvector and GIN indxes. You can also use recursive CTEs for basic graph traversal.

With the awesome k6, I managed to load test the performance too, easily handling 500k assets and around 100 concurrent users performing various tasks on cheap Hetzner nodes with no external infrastructure. It's really impressive how capable Postgres is for a multitude of tasks.

charlie-haley··on Show HN: Marmot – Single-binary data catalog (no Kafka, no Elasticsearch)
Hey, that's a good question! At the moment, it treats the latest run as the desired state. So any new changes to a schema will simply overwrite the old version. I'd like to version these so people can navigate schema versions in the UI. If using a plugins, they currently are triggered either via the CLI or a schedule on the UI, so updates will only appear in the catalog after a plugin has run.

I'd also love to have some native integrations beyond Airflow. Once I've matured the existing plugin ecosystem a bit more, it's high on my list (along with column-level lineage).

charlie-haley··on Show HN: Marmot – Single-binary data catalog (no Kafka, no Elasticsearch)
Postgres has a lot of features such as trigram-based search which is pretty essential if I don't want to use a dedicated search indexer. It's also much better at handling concurrent writes than SQLite.
charlie-haley··on Show HN: Marmot – Single-binary data catalog (no Kafka, no Elasticsearch)
It can handle discovery within a plugin if the asset types are related. You can also manually add lineage via the UI or use Terraform to create lineage links via IaC. It's pretty complicated to automatically handle discovery of asset lineage, I'm yet to find a nice way of doing it that can work for many use-cases
charlie-haley··on Show HN: Marmot – Single-binary data catalog (no Kafka, no Elasticsearch)
It supports either, I didn't want to restrict people to just one method of getting their catalog populated. The CLI and Plugin system works on needing read credentials to a given Service, it then populates the catalog with those assets. Any lineage links currently need to be done manually (unless they're part of the same plugin). Otherwise, you can integrate with your existing IaC pipelines using Terraform or Pulumi to populate the catalog at deploy time instead of needing to scrape a bunch of services.
charlie-haley··on Show HN: Marmot – Single-binary data catalog (no Kafka, no Elasticsearch)
Hey, there's some documentation around creating plugins here. It's relatively simple and involves adding a new Go package to the repo. Currently they have to be compiled into the Binary but I'd like to support external plugins at some point https://marmotdata.io/docs/Develop/creating-plugins

Also, thanks for pointing out the issue with the docs, I'll get that fixed!

charlie-haley··on Show HN: Marmot – Single-binary data catalog (no Kafka, no Elasticsearch)
That's great to know, I wasn't aware anybody even attempted to used it yet! I'm currently in the process of overhauling the Plugin system, it's been quite hard to test some enterprise closed-source integrations like Tableau and Snowflake to build out plugins.

SSO is sort kind of available, but undocumented, it currently only supports Okta but I'm working on fleshing out a lot of this in the next big release (along with MCP)

charlie-haley··on Show HN: Marmot – Single-binary data catalog (no Kafka, no Elasticsearch)
It depends on your ecosystem. If everything lives under one vendor their native catalog will probably work really well for you. But most of the time (especially for older orgs) there's usually a huge fragmented ecosystem of data assets that aren't easily discoverable and spread across multiple teams and vendors.

I like to think of Marmot as more of "operational" catalog with more of a focus on usability for individual contributors and not just data engineers. The key focus being on simplicity, in terms of both deployments and usability.

charlie-haley··on Show HN: Marmot – Single-binary data catalog (no Kafka, no Elasticsearch)
Hey HN, I wanted to show off my project Marmot! I decided to build Marmot after discovering a lot of data catalogs can be complex and require many external dependencies such as Kafka, Elasticsearch or an external orchestrator like Airflow.

Marmot is a single Go binary backed by Postgres. That's it!

It already supports: Full-text search across tables, topics, queues, buckets, APIs Glossary and asset to term associations

Flexible API so it can support almost any data asset!

Terraform/Pulumi/CLI for managing a catalog-as-code

10+ Plugins (and growing)

Live demo: https://demo.marmotdata.io