Proton, a fast and lightweight alternative to Apache Flink
github.com
github.com
Okay, uhm, does this mean the 500 MB is the memory consumption footprint?
No way someone is shipping a giant 500 MB binary hunk, and the selling point is that it's lightweight?
But 500MB of RAM for a DB also seems tiny. That can't be right..
I went through the FAQ and read through the README on the github page, however I'm unable to figure out some really ideal usecases for proton. Ideally from a product perspective.
I like to have some examples so when I am building in the future I can quickly know off the top of my head, "oh yeah, lets just use proton for this."
Does anyone have a few different applications types which they've built or problems they've solved with proton? Bonus points for external examples with some code.
https://docs.timeplus.com/showcases lists a few real world use cases we solve with our customers. Proton is core engine of Timeplus Cloud, which adds extra UI, sources/sinks, multi-tenant etc.
https://docs.timeplus.com/proton-howto lists some of the common tasks for processing real-time data. Basically if you're in the data engineering space, Proton can do many things that Flink/Spark/ksqlDB can do, just faster and more lightweight.
Good point of "oh yeah, lets just use proton for this." Will add more docs and examples for such newly-open-sourced project (we have ran the cloud service for 1.5 years)
A very typical scenario is live query / analysis data in Kafka topics with Proton Streaming SQL and materialize the processed / aggregated results to a downstream system like ClickHouse or a target Kafka topic. What we need to do is 1. curl download the binary 2. create an external stream to point to your Kafka topic 3. run the streaming SQL.
The important part of Apache Flink is the stateful streaming and fault tolerance characteristics of Flink. This is not an alternative to Flink as a distributed runtime for event driven applications.
Proton save the state in memory and local file system.
The cluster and HA features are in Timeplus Platform which build on top of Proton
While Flink DataStream or Table API are powerful and flexible, you can easily setup Proton to learn or use streaming SQL, or query historical data. Thousands of ClickHouse functions are available to use. If you like ClickHouse or Redpanda, probably you will like Proton too.
And boy, Kafka more and more seems like an early warning system for architecture astronautics. (Not necessarily bad by itself, but often part of hectoliters of manure poured in size 6 wellies)
Also it's very popular in China.
TBH, I'm amazed that Proton wasn't already taken.
https://github.com/risingwavelabs/risingwave https://github.com/MaterializeInc/materialize
In short, data streaming is getting popular. Apache Flink, Apache Spark, ksqlDB are traditional players. They work well in some cases, but have challenges on easy-to-use, easy-to-deploy, on even performance sometimes
Proton, RisingWave, Materialize, etc are the alternatives. In the end, it's always 'case by case' to pick up the best tools working for your team in your projects.
If you think this is too general, well - Materialize no longer provide the latest code as an open-source software that you can download and try. It turned from a single binary design to cloud-only micro-service - RisingWave has been open-sourced for 1.5 years and released their cloud mid of last year. It builds its own row-based historical storage, while Proton leverages ClickHouse for faster OLAP-like workload. Also there are more SQL functions supported in Proton, because it's powered by ClickHouse. Proton uses local SSD as the main storage and has an option to send older data to object storage. While RisingWave primarily uses object storage, and use local disk as cache. So in general Proton can achieve lower latency for both streaming or historical queries.
I am a big fan of coffee and drink different kinds of coffee over time, even on the same day. It's all about the use cases and preferences. This could be applied to choose open-source tools to query live data. In the last Current conference, me and Gang had a talk about different tools and different coffee. You may check https://www.timeplus.com/post/query-kafka-with-sql-current23
Materialize CTO here. Just wanted to clarify that Materialize has always been source available, not OSS. Since our initial release in 2020, we've been licensed under the Business Source License (BSL), like MariaDB and CockroachDB. Under the BSL, each release does eventually transition to Apache 2.0, four years after its initial release.
Our core codebase is absolutely still publicly available on GitHub [0], and our developer guide for building and running Materialize on your own machine is still public [1].
It is true that we substantially rearchitected Materialize in 2022 to be more "cloud-native". Our new cloud offering offers horizontal scalability and fault tolerance—our two most requested features in the single-binary days. I wouldn't call the new architecture a microservices design though! There are only 2-3 services, each quite substantial, in the new architecture (loosely: a compute service, an orchestration service, and, soon, a load balancing service).
We do push folks to sign up for a free trial of our hosted cloud offering [2] these days, rather than trying to start off by running things locally, as we generally want folks' first impressions of Materialize to be of the version that we support for production use cases. A all-in-one single machine Docker image does still exist, if you know where to look, but it's very much use-at-your-own-risk, and we don't recommend using it for anything serious, but it's there to support e.g. academic work that wants to evaluate Materialize's capabilities to incrementally maintain recursive SQL queries.
If folks have questions about Materialize, we've got a lively community Slack [3] where you can connect directly with our product and engineering teams.
[0]: https://github.com/MaterializeInc/materialize/tree/main
[1]: https://github.com/MaterializeInc/materialize/blob/main/doc/...
Proton is written in C++, adding streaming processing to ClickHouse and make stream the main concept. It is designed to support streaming analytics where aggregating large number of data over time is the focus.
In geneneral, these three are very similar.
Please check this PR from the proton team: https://github.com/ClickHouse/ClickHouse/pull/54870
Team, what do you think building that with proton?
A:
Sorry, what? We cant do that with email!!
B:
I don't understand, if you want that to run on windows we should use windows, proton is mainly for games.
C:
Why are we talking about reverse osmosis now? Is that software not for a postal-service?
D:
Yeah ok, let us use AWS Proton for deployment, but that's far in the future.
Yes, it lacks context, but can you get the context from Flink or Spark ? both of these are just data processing tools, supporting streaming as well.