Show HN: Capillaries: Distributed data processing with Go and Cassandra
capillaries.io
Capillaries is a built from scratch, open-source Go solution that does just that: ingests data and applies user-defined transforms - Go one-liner expressions, Python formulas, joins, aggregations, denormalization - using Cassandra for intermediate data storage and RabbitMQ for task scheduling. End users just have to provide: - source data in CSV files; - Capillaries script (JSON file) that defines the workflow and the transforms; - Python code that performs complex calculations (only if needed).
The whole data processing pipeline can be split into separate runs that can be started independently and re-run by the user if needed.
The goal is to build a platform that is tolerant to database and processing node failures, and allows users to focus on data transform logic and data quality control.
“Getting started” Docker-based demo calculates ARK funds performance, using EOD holdings and transactions data acquired from public sources. There are also integration tests that use non-financial data. There is a test deploy tool that uses Openstack API for provisioning in the cloud.