HNHacker News
TopNewBestAskShowJobs

kleineshertz

41 karma · joined May 10, 2023

submissionscomments
kleineshertz··on Show HN: Capillaries: Distributed data processing with Go and Cassandra
Right. And this is where our paths go separate ways. As I see it, Capillaries users are not necessarily tech companies and they do not have much appetite for writing and maintaining a lot of code. All they want is to run, say, 50 kinds of workflows on a regular basis and to keep those workflow definitions very formalized and stable. It would be hard to sell the idea of maintaining 50 different codebases to their management.

As for more complex calculations: in Capillaries, Python is not a programming platform, it's just a scripting engine.

kleineshertz··on Show HN: Capillaries: Distributed data processing with Go and Cassandra
Storm positions itself as a stream processing solution, while Capillaries is 100% batch-oriented.
kleineshertz··on Show HN: Capillaries: Distributed data processing with Go and Cassandra
It's always a balance. I have been working with teams on both side of the fence and I think I am well aware of the dangers of both: keeping the custom wheel running for years vs fighting the particularities of a third-party tool (up to the point they start dictating architectural decisions). Most of the operations Capillaries is intended to perform are row-based, and stellar Spark map-reduce capabilities were not a big selling point, while tech lock-in price seemed pretty high.

On a more general note (Spark discussion aside), I like working with third-party solutions that can do only one thing, but they do it perfectly. And I am ok supporting in-house-built frameworks that behave the same way and do not pretend to be a world peace solution.

kleineshertz··on Show HN: Capillaries: Distributed data processing with Go and Cassandra
Regarding using in-memory storage. Early prototype of Capillaries used Redis for storage and the performance was stellar. I decided to drop it for two reasons. First, indexing mechanism required a root-level sorted set, and Redis cannot partition it. Second, most of intermediate data is supposed to be available until the end of the run, which means hours, and I was not sure that typical Capillaries users would agree to carry the cost of providing so much RAM vs disk space. Am I willing to return to the discussion about replacing Cassandra with some in-memory storage? Maybe.
kleineshertz··on Show HN: Capillaries: Distributed data processing with Go and Cassandra
Parquet support is on the radar for sure, and I would like to have it before diving into database connector development.
kleineshertz··on Show HN: Capillaries: Distributed data processing with Go and Cassandra
ScyllaDB is definitely on the radar. The main reason I picked Cassandra on the prototyping stage was because default Cassandra configuration gave me much better performance then ScyllaDB (I know, it is supposed to be vice versa). Another obvious reason was Cassandra's maturity and community support. If gocqlx is indeed a drop-in replacement for gocql, I can't see problems having a separate config/fork using ScyllaDB along with Cassandra.
kleineshertz··on Show HN: Capillaries: Distributed data processing with Go and Cassandra
If it's not invented here, it can't be any good.
kleineshertz··on Show HN: Capillaries: Distributed data processing with Go and Cassandra
Maybe. The scenarios Capillaries is intended for do not need complex/flexible workflow, we just need some basic dependency rules (easy to implement) and really reliable scheduling (RabbitMQ).
kleineshertz··on Show HN: Capillaries: Distributed data processing with Go and Cassandra
Temporal is a different ecosystem (and a much more ambitious solution), but one of the principles is the same: users want a platform that solves scalability issues and lets them focus on biz logic and customer value.
kleineshertz··on Show HN: Capillaries: Distributed data processing with Go and Cassandra
Nothing wrong with this question. I do not have any experience with Spark, but I guess Capillaries belongs to the same or similar ecosystem. My understanding is that Spark is way more generic framework that revolves around DAG-defined workflow and map/reduce-style functionality.

Capillaries is about:

- taking a very structured, stage-by-stage, approach to batch data processing with the possibility to control the results of a specific stage (although some kind of workflow DAG is there as well); - executing a SQL-style aggregation and denormalization on data in Cassandra; - executing workflows without actually writing code (besides one-liner Go expressions and Python math formulas when needed).

Sorry if I am missing the point with Spark, as I said - I never worked with it.

kleineshertz··on Capillaries: Distributed data processing with Go and Cassandra
Capillaries is a distributed data processing platform that: - works with structured row-based data - splits data into batches that can be processed as separate jobs on multiple machines in parallel - allows scenarios that involve human operator supervision and data validation - has ETL/ELT capabilities - has SQL-like join, grouping, and aggregation capabilities allows custom data processing plugins