46 karma · joined June 9, 2025
Pretty far, as it turns out. pg_trgm can be used for fuzzy matching, there's built-in full-text search with tsvector and GIN indxes. You can also use recursive CTEs for basic graph traversal.
With the awesome k6, I managed to load test the performance too, easily handling 500k assets and around 100 concurrent users performing various tasks on cheap Hetzner nodes with no external infrastructure. It's really impressive how capable Postgres is for a multitude of tasks.
I'd also love to have some native integrations beyond Airflow. Once I've matured the existing plugin ecosystem a bit more, it's high on my list (along with column-level lineage).
Also, thanks for pointing out the issue with the docs, I'll get that fixed!
SSO is sort kind of available, but undocumented, it currently only supports Okta but I'm working on fleshing out a lot of this in the next big release (along with MCP)
I like to think of Marmot as more of "operational" catalog with more of a focus on usability for individual contributors and not just data engineers. The key focus being on simplicity, in terms of both deployments and usability.
Marmot is a single Go binary backed by Postgres. That's it!
It already supports: Full-text search across tables, topics, queues, buckets, APIs Glossary and asset to term associations
Flexible API so it can support almost any data asset!
Terraform/Pulumi/CLI for managing a catalog-as-code
10+ Plugins (and growing)
Live demo: https://demo.marmotdata.io