238 karma · joined March 30, 2015
https://github.com/dittofeed/dittofeed/blob/main/packages/is...
If you're interested in how to segment users by processing semi-structured event logs at scale, keep reading!
Article: https://dev.to/dittofeed-max/how-we-stopped-our-clickhouse-d... Dittofeed's primary repo: https://github.com/dittofeed/dittofeed
This article includes code samples, as well as a link to a git repo you can use to run tests and follow along. These code samples demonstrate a simplified user segmentation setup in ClickHouse.
Accompanying repo for the article: https://github.com/dittofeed/clickhouse-segments-tutorial
User segmentation is key to understanding your business or application’s users’ behaviors. For example, you might want to keep track of which users have gotten stuck in an onboarding flow or which users haven’t logged in within the last 30 days.
Here's a quick rundown of what I covered:
Diving into ClickHouse: Why we (and you) should use ClickHouse for your analytical workloads, and user segmentation in particular.
On User Segmentation: What are user segments, and how do we calculate them. We cover some basic pseudocode.
Tackling Idempotency: To avoid duplicate data messing up our calculations, we've made sure our setup is idempotent.
Scaling with Micro-Batching: Simulate stream processing by incrementally aggregating user events in small batches. This is really the core of the article!
"Event Time" vs. "Processing Time": When did an event happen and what are the different ways we can measure time? How do we accommodate different measures of time in our application?
I'd really love it if you could take a moment to check out the article and maybe even drop a star on the Dittofeed GitHub repo. Your support and feedback would mean a lot!
Catch you all soon!
Would love to learn about your use case and pain points.
https://join.slack.com/t/dittofeed-community/shared_invite/z...
Infrastructural concerns are going to have to be embedded deeply into product development and vice versa.
Beyond that, I think our focuses are somewhat different, as our current roadmap is focused on serving growth engineers. Still happy to see them working in the space.
Re-"developer as the customer", the mismatch between who makes the buying decisions for this kind of tech (CMO's etc.), and the devs who do much of the heavy lifting to make it work, is a real challenge.
Also, your point re the economies of scale of self-hosting vs using saas is valid. For small to mid sized orgs, using a cloud offering can be more economical.
However, we've observed that larger orgs often migrate off of saas to use open-source or build software in-house. This occurs for a number of reasons:
- They will have more engineering resources to allocate, in this case to marketing / growth.
- At their scale, the fixed costs of allocating engineers to implementing solutions are often exceeded by the variable costs of saas products, which commonly have volume based pricing.
- They often have more unique requirements that are not served by any particular saas product, and closed source saas is not extensible.
Some recent examples of this:
- Several large orgs are migrating off datadog in favor of open source observability tooling.
- Airbnb recently implemented the equivalent of Dittofeed internally (they responded in this thread).
Still, you raise legitimate concerns, and we're still figuring things out. Would love to get in touch, to better understand your perspective if you have time!