HNHacker News
TopNewBestAskShowJobs

tfearn

1 karma · joined March 30, 2026

30 years building data infrastructure for financial institutions (Goldman Sachs, Bridgewater, Deutsche Bank, others). Founded several venture-backed startups. Currently building Datris, an open-source agent-native data platform with a remote MCP server. Also run IData Corporation, a data pipeline consulting firm.
submissionscomments
tfearn··on [dead]
Short demo showing Claude desktop connected to a locally running MCP server (32+ tools exposed). I asked it to set up an end-of-day stock price feed and it sequentially called tools to list existing pipelines, generate a schema, create a pipeline, write and schedule a tap (yfinance → MongoDB) on a cron, run a test ingestion, and verify the result.

Crazy part is this works with any endpoint — internal APIs, SaaS tools, whatever you can describe.

The MCP server is open source: https://github.com/datris/datris-platform-oss

Demo: https://datris.ai/videos/supercharge-claude-data-engineer

tfearn··on Data Discovery – plain-English to discovering and acquiring data using AI
I've been building data infrastructure for 25+ years across Goldman, Bridgewater, Freddie Mac. The same problem exists everywhere: getting a new external data source wired up takes days. You have to figure out what datasets the source even exposes, write the ingestion code, build a pipeline, wire up scheduling, and test it — before a single byte of data lands anywhere useful.

The data discovery process that we just added to our platform (open-source) collapses that into one session.

You describe what you want in plain English ("company earnings", "option chain data"," "SEC EDGAR company filings"). The AI identifies the source, enumerates every dataset it exposes — grouped by category with parameters and auth requirements — and you select what you want. From there, Discovery generates Python tap scripts in parallel, runs them immediately as a test, self-heals on failure (up to 3 attempts), creates the pipelines, and optionally schedules everything. The whole thing drops into a Data Catalog that groups related taps and pipelines together.

The artifact isn't a one-time wizard output — it creates real, editable tap scripts and pipeline configs you can modify afterward. Parameters can be sourced from a table you already have ingested, a file upload, or an AI-generated list ("give me S&P 500 tickers"). Date tokens like `{{TODAY}}` are substituted at runtime so daily snapshots just work.

The same flow is also exposed as an MCP tool (`discover_source`), so external AI agents can drive Discovery programmatically — ask "what datasets are in polygon?" and get back the same structured dataset catalog the wizard uses.

Destinations: MongoDB, PostgreSQL, Kafka, MinIO, pgvector, Qdrant, Milvus, Chroma, Weaviate, ActiveMQ, REST endpoints.

Full demo walkthrough: https://datris.ai/videos/data-discovery-ingestion-consumptio... Docs: https://docs.datris.ai/discovery OSS (AGPL): https://github.com/datris/datris-platform-oss

tfearn··on Datris – Open-source data platform built around MCP
Hi HN,

I've spent 25 years building data infrastructure in financial services — Goldman Sachs, Bridgewater, Deutsche Bank, Freddie Mac. Today I'm open-sourcing Datris, a data platform built around MCP (Model Context Protocol) from day one.

The idea: As far as I know, this is the first data platform where MCP is the primary interface, not an afterthought. We built the MCP server first and made everything accessible through it. The platform has 30+ MCP tools — any agent (Claude, Cursor, your own framework) can create pipelines, ingest data, validate with plain English rules, transform, query databases, search vector stores, and monitor jobs. The API and UI use the same pipeline engine, but MCP is the native interface.

What an agent can do: - Create a complete pipeline from sample data in one call (schema auto-detected) - Upload and process CSV, JSON, XML, Excel, PDFs, Word docs - Validate and transform data using natural language instructions - Query PostgreSQL and MongoDB - Semantic search across 5 vector databases (Qdrant, Weaviate, Milvus, Chroma, pgvector) - Ask questions in natural language — SQL generated and executed automatically - Full RAG pipeline — extract, chunk, embed, and search documents - Profile data quality, diagnose errors, explore metadata

There's also a CLI: bash datris ingest sales.csv --dest postgres datris query "SELECT * FROM public.sales" datris query "top 5 stocks by volume" --table trades datris search "return policy" --store pgvector --collection docs

Install: git clone https://github.com/datris/datris-platform-oss cd datris-platform-oss cp .env.example .env docker compose up -d

CLI: brew tap datris/tap && brew install datris MCP server: uvx datris-mcp-server

Stack is all open-source: Spring Boot, Docker, MinIO, MongoDB, ActiveMQ, Vault, PostgreSQL. AI via Anthropic, OpenAI, or Ollama (fully local).

GitHub: https://github.com/datris/datris-platform-oss Docs: https://docs.datris.ai

Happy to answer questions.