Building the world’s fastest website analytics (2021)
usefathom.com
usefathom.com
[1] https://clickhouse.com/blog/we-stand-with-ukraine
Organizations like Facebook, Amazon, Netflix, Cloudflare, and many, many others have been combining specialized data stores for many years. Stonebraker, Madden, and colleagues summarized the argument for this approach in 2007. [0] At this point I can't think of any major SaaS application I've seen recently that did not have multiple data store types. That includes event streams, which are data stores as well.
[0] http://nms.csail.mit.edu/~stavros/pubs/hstore.pdf
Disclaimer: I work for Altinity.
My own company is not especially large, but we run MySQL, ClickHouse, and Prometheus to manage our own SaaS platform for ClickHouse. We are happy with the result. (Though it would be nice to use ClickHouse as a Prometheus backend.)
Well... that's already possible and it works great! As you might know https://qryn.dev turns ClickHouse into a powerful Prometheus *remote_write* backend and the GO/cloud version supports full PromQL queries off ClickHouse transparently (the JS/Node version transpiles to LogQL instead) and from a performance point of view its well on par with Prometheus, Mimir and Victoriametrics in our internal benchmarks (including Clickhouse as part of the resource set) with millions of inserts/s and broad client compatibility. Same for Logs (LogQL) and Traces (Tempo)
Disclaimer: I work on qryn
[1] https://www.crunchbase.com/organization/clickhouse/company_f...
Then why are you opining on the subject?
> We have no operations in Russia, no Russian investors, and no Russian members of our Board of Directors.
It's disappointing after 11+ yrs of failing to find traction for MemSQL, this is the hill you die on? You're running your mouth on some random comments thread about a shitty 2-person analytics platform where the author clearly doesn't know the first thing about databases and just went with the first Twitter ad he came across.
Under Raj's "leadership" you know damn well MemSQL/SingleStore will be dead within two years. The hyper-aggressive sales pressure had some success during the lofty bull market due to clueless IT managers with massive budgets, but that time has come to an end. You guys missed the window to IPO to get your own exit and your runway is quickly diminishing.
Yandex N.V. (A dutch holding company of Yandex Russia) is listed as an investor of Clickhouse Inc. This is all I'm stating.
Looking back, I’m glad we went in this direction. Fathom has grown beyond what we could’ve imagined. Let’s see what happens over the next five years.
I know a company that forked Clickhouse but had to rebuild virtually everything to be a proper database.
There are already plenty of technologies, more mature than Clickhouse for addressing such scenarios.
For example, here is news from today about Splitbee: https://news.ycombinator.com/item?id=33334104 No surprise it is using ClickHouse: https://splitbee.io/blog/new-pricing
> The biggest mistake I made was that I kept the UPDATEs in our code (we update the previous pageview with the duration, remove the bounce, etc.). In a few weeks, I'll be moving to 100% append-only by utilizing negative numbers.
> For example, if you want to set bounce_rate to 0%, you would write a 1 for bounce_rate on the first pageview and then insert a duplicate with -1 for bounce_rate. And for the duplicate row, you'd have nothing set for pageviews, visits and uniques, so it would all group nicely.
This is nearly a copy-paste from ClickHouse documentation.
> We shard on UUID, and then we set SiteId as the sort key. We do this because we want to utilize something called "local joins" in SingleStore. Long story short, events can be joined with event_properties (allowing you to have thousands of dynamic properties per event you track), and it's fast.
This statement is slightly disappointing as well. No magic, no advantage to ClickHouse.
The outcome can be read as follows: "use SingleStore exactly as you'd use ClickHouse and maybe it will work alright".
SingleStore advantages:
- good compatibility with MySQL dialect;
- offers good UPDATE/DELETE for fresh data (invalidated by this article, as they dropped the usage of UPDATE for better performance);
- good support for JOINs (invalidated by this article, as they ended up using local JOINs);
SingleStore disadvantages:
- not open-source, worse brand recognition, the old brand (MemSQL) did not work well too;
- performance claims don't validate;
- older company and many initial developers no longer work there;
Fathom sounds like a toy use case as the data amount is too low. ClickHouse is almost universally selected for similar and larger use cases, with hundreds of billions of events each day.
PS. Nevertheless, the post is hilarious:
> I was suffering from low energy in the two weeks leading up to this migration, and I was feeling awful throughout migration week, especially on migration day. On Sunday 14th March 2021, two days after we had finished the migration, my wife showed me an already-half-used bag of coffee beans I had been drinking through most of this migration, and the bag said "DECAF." Divorce proceedings are underway.
Beyond this, it has a naive query optimizer and limited ability to run distributed joins. This is why you won't find well known analytical benchmark results like TPC-H and TPC-DS for clickhouse vs other SQL data warehouses.
SinglestoreDB has been doing real time analytics much longer then clickhouse and has a much more mature feature set at this point[1][2]. Our go to market is more big enterprise driven so we are definitely much less well known among developers. Revenue share-wise I suspect we are the leader in real time analytics (see disclaimer below - I'm biased but our revenue is reaching thresholds for IPO readiness).
Also, if you follow Jacks later posts you'll see he replaces DynamoDB and Redis with SinglestoreDB as well. Now were getting into places Clickhouse doesn't tread at all...
Disclaimer: I'm one of the cofounders of MemSQL/SingleStoreDB (still working away on making it better all these years later...).
[1] https://www.singlestore.com/blog/the-technical-capabilities-...
Most other cloud DWs have public TPC-H or TPC-DS results that are easily googlable. Clickhouse is missing for a reason...
- https://research.gigaom.com/report/data-warehouse-cloud-benchmark/
- https://www.databricks.com/blog/2021/11/02/databricks-sets-official-data-warehousing-performance-record.html
- https://celerdata.com/blog/starrocks-queries-outperform-clickhouse-apache-druid-and-trino
- https://aws.amazon.com/blogs/big-data/amazon-redshift-continues-its-price-performance-leadership/
Our results are here: https://www.singlestore.com/blog/tpc-benchmarking-results/It's straightforward to build applications that do transaction processing in MySQL and analytics in ClickHouse. You can migrate data out of MySQL that does not belong there. You can also use CDC to mirror transaction data into ClickHouse for analytic processing. [0] The resulting applications are quite robust and have the advantage that different parts can scale independently.
Building applications that use the strengths of each of the component databases is a natural way to scale systems.
Disclaimer: I'm the author of the cited article and I work for Altinity.
[0] https://altinity.com/blog/using-clickhouse-as-an-analytic-ex...
I just started working at FeatureBase a few weeks ago, here in Austin. It's Open Source and we have a 5 minute guide to get it running here: https://docs.featurebase.com/. We also have a cloud/serverless option.
FeatureBase is a B-tree database which uses Roaring Bitmaps. This makes it suitable for doing analytical queries on massive data sets immediately after ingestion. You do have to model the data properly, however.
Clickhouse updates are painful. I looked on postgres with citus extension and it only supports appends.
Their paper on how they organize the columnstore and row store buffer is pretty neat. https://images.go.singlestore.com/Web/MemSQL/%7Bd1e77ba1-0e3...
It’s really cool to see usefathom take off doing millions in ARR with two engineers. SingleStore has to been a high leverage choice so they can focus on product. Power to them.
Meanwhile we’ve been grappling with a mixture of Mongo, Snowflake and Postgres for more than an year with a decent sized team. The curse of VC funded startups is that with more money and more engineers, a likely outcome is more complexity. The infra becomes a complex beast that takes many quarters/years to build because there isn’t that market pressure to be time and money efficient.
Building a reliable DBMS with any kind of decent performance can easily take years. And by then, it's still not sure that you're doing any better than any of the existing solutions.
Though, I’d be sorely tempted to spin up a bfc of clickhouse on bare metal.
Attempting to make your own db while being ignorant of competing products would be an absolute disaster
The query was:
SELECT SUM(pageviews) as total, pathname FROM pageviews GROUP BY pathname ORDER BY total DESC LIMIT 10.
If we don’t hear an answer from you, I’ll be really upset. Otherwise, we may have to add your ideas to the article!