How we saved money by replacing Mixpanel with BigQuery and K8S
blog.doit-intl.com
blog.doit-intl.com
Five weeks of, say, three engineers with $150k salaries ($200k fully loaded) is $60k. People say that the majority of software is after initial development. Let's say 20% is development [0]. That means this project costs $300k total and lives for 5 years before the team decides to rewrite it. That brings us to $60k / year.
This is good savings! My general guidance to folks thinking about this sort of project is that you should implement a vendor and see how you like it. The reason being, MixPanel does a LOT more than this tool, and buying a vendor can help you figure out the subset of features that you actually need. Once you have a clear requirement list, you have a clearly scoped project and can hammer it out relatively easily to get that cost savings.
Beware of thinking that you can do this for $X,000 / year on your own, or get it right on the first shot. There's a reason MixPanel (and any other vendor) has so many engineers. You're paying them to have already made the mistakes that you're about to make.
[0] https://stackoverflow.com/questions/3477706/development-cost...
If things flip around again and you find a new library that solves your needs better than your homegrown lib, you can drop it in forking your lib and turning it into a wrapper on the external lib.
At least this is what I have learned implementing high flexibility cross-platform mobile video playback solutions in our app over the last 5 years, where the landscape of libraries (and our own tools) has kept changing.
Personally I much prefer 'pay me now', because 9 times out of 10, IME, some form of technical debt underlies the reasoning about why a team can't do something. 'It always pays to take the pain up front', basically.
Be limber enough to eschew the tools that aren't working. But "good enough" is, well, good enough for most folks. Perfection the enemy of progress and all that.
Unfortunately, the part "implement a vendor and see how you like it" went only with the first half. Complaints of developers supposed to be using the tool were brushed aside as a simple resistance of change. In the end it was implemented by decree from higher management.
Only after all the paperwork was signed, we started to see that the vendor doesnt support our usecases, that their hardware is underpowered for the load we were expecting, e.t.c ...
In the end, I have heard at least 20 various senior and beyond senior engineers from various teams complaining that they spend 20 to 50 % of time working around the vendors bugs, around once a month bringing down vendors infrastructure, because our use-case mandated uploading more data than the infra could handle.
I definitely agree that clear requirement list and clearly scoped project and implementing a vendor to see how you like it is usually better than in-house
On the other hand in-house is usually better than implementing a random vendor because management feels like it ;)
The real reason Mixpanel (et al) have so many engineers is:
A) They need to scale their service to reliably provide to N corporations which is frequently non-trivial.
B) They need to implement the 80% of features that aren't critical path for Company A because some subset of them are for Company B,C,D,E,F,G,H,I,J,K,etc.
C) I've seen software vendors make some pretty terrible and painfully obvious mistakes.
Idk, maybe I'm crazy but I have very little faith in this idea that more engineers = better decisions were made. I've just never found quantity has an impact on quality of the end product, only the speed at which its produced.
Yeah, and in my experience, that frequently fails to materialize which why I mentioned point C.
Now, that could be arrogance on my part but it doesn't change my experience on the quality of the vast majority of outside software vendors.
1. Right off the bat, there's the cost of building all of this + maintaining it. Something like Mixpanel at scale requires at least two, if not three, engineers: one for client libraries, one for infrastructure (which the OP blogs about), and another for the web app (dashboard, real-time user stream, etc.) To be sure, not everyone needs all features of Mixpanel. To be really sure, nobody needs all features of any SaaS tools. But if you really want to compare apples to apples, then you have to account for these. To hire competent software engineers who can collaborate closely and maintain a complex piece of analytics software requires at least $100k per year, if not twice that based on locale. That's at least $300k and more like $600k right there.
2. The whole point of something like Mixpanel is to make analytics accessible to non-engineers. In their case, it's primarily product people and secondarily marketing/customer success. In any case, building an analytics/data product consumable by non-technical people is hard and takes way more than assembling a couple of cloud infrastructure together. If there's one reason Mixpanel is still in business, that's because of this.
3. Finally, the OP has a valid point which should have been highlighted more, if not to make their own biases more clear: Mixpanel's diminishing differentiation is dev shops/consultancies' opportunity. It is indeed incredible that a dev shop can build even a third of Mixpanel's functionality by leveraging GCP components. Mixpanel had to build a lot of its core backend systems from the ground up, including its original key-value store. Just this year, they fully migrated to Google Cloud Platform themselves, suggesting that there's really little room for differentiation among analytics vendors at the level of infrastructure components (Mixpanel's arch-nemesis, Amplitude, leverages Apache Kafka and various AWS components, most notably Amazon Redshift, heavily)
With all of this being said, one thing remains true: the most expensive cost of any software is people running them and the dependencies created around them. These may not show up as line items, but they sure are deeply embedded in your total cost.
What happens if the system collecting all of your data has vulnerability? suddenly all of your apps and data are owned.
What if it's inefficient? then you have DDoSed yourself.
Non-functional requirements cannot be solved with a functional requirement mindset.
What if you are losing data? what if you aggregate it incorrectly? what if data is ingested in the wrong order? ... you get the idea.
It's not a 3 people problem. When you take on a hard problem like this you have 10 different things to cover and you need to pick all of them.
Also, this post only addresses data flowing in one direction into BQ. Presumably they're then using something like Tableau or the Google/AWS offerings to build dashboards based on this data. There aren't many places for bugs to hide unless the BQ perms are wide open.
I don't see anywhere that the article says they have replaced all of Mixpanel. It only says that they were able to replicate the functionality their client actually needed. I don't see them making claims that the solution they built is anywhere comparable to the full solution from Mixpanel. Only that it provides all the features their client needed.
That last part is the most important. It's why you're really using Mixpanel/Amplitude/other vendors. To say "we replaced Mixpanel" without addressing the front end feels a bit disingenuous to me.
Our goals were to get high volume web data (pageviews, clicks, etc.) alongside application data already saved in our Firebase DB and synced to BigQuery.
We picked Keen because it has an open source web tracking lib https://github.com/keen/keen-tracking.js/ that easily plugs into our React/Redux stack.
They also have built-in streaming to BigQuery: https://keen.io/docs/integrations/google-bigquery/
Keen pricing is about 10% of mixpanel, so for our limited needs it has been working well.
Long term if our volumes really grew the original post looks like a good option, but figured we'd pass along this lower dev approach.
We saved 60% (which was no small amount of money) by realizing that sending more than 1 of an event per user per hour was: A) Not necessary B) Not even something you could query in MixPanel.
more thoughts in: https://blog.ratelim.it/blog/how-to-save-money-on-event-trac...
All of these are issues (that I have faced). From what I've seen, people just keep downloading Mixpanel data and uploading to their DWH, sort of voiding the reason to implement Mixpanel in the first place.
The tool is lovely, their export policy is great, but there's something about actually owning your data.
Using Metabase for visualization.
The other key benefit is you can warehouse the rest of your data in BQ as well -- where it can be easily linked to your non-event data; marketing data from other systems, etc.
https://discourse.snowplowanalytics.com/t/porting-snowplow-t...
Unless you have restrictions about how your data moves and where it is stored, or need to have a trusted computing base with no externally developed software, or have very strict requirements that no available service implements (unlikely), you are better off just using an open source solution or paying for a service.
# Real cost
Even if you have to pay for a commercial service, in comparison, the TCO of developing a similar solution in-house is very large.
When you create a system in-house, you are paying for: design, implementation, testing, maintenance, deployment, infrastructure, security audits, training for users, documentation, and sometimes costs go beyond engineering, e.g: UX and graphic design and such.
After you are done spending all that money, you end up with a custom built service that is far from the actual main activity that supports your business.
# Quality
If you are to authorize a team to do something like this, audit their code constantly and impose a higher quality standard than you do for the rest of your applications. This is because this system will be a dependency for all your applications.
Even in large companies, it is unlikely that you have enough resources to have a large dedicated working on something like this. Because of this, requirements will need to be deprioritized or just neglected.
Because of this, you can lose all hope of selling this solution externally.
# Users
Then, since you are committing significant resources to your internal tool, it is likely that every single team will be forced to use it. That in itself is also a problem. What is better?
a) Learning how to use MixPanel, and put it in your resume (a skill that has market value and can be traded).
b) Learning how to use a proprietary system that is only used internally within a company. A skill that cannot be traded in the job market.
If I am a user, it is against my own self interest to push for an internally developed tool.
May I ask why? Seems like exactly the kind of data where some loss is fine.
Just as a company gets big enough to offer a large revenue stream to Mixpanel is the same time when it probably makes sense to ditch Mixpanel and and just run all this in house for collecting, storing and processing massive data streams. I see a lot of this sort of decision making taking place now.
See, once you have the data pipeline in place, you can bring front-end/ reporting platform to generate any reports you wish, from free (e.g Google Data Studio) to self-hosted open source based, up to SaaS solutions which available few hundred bucks per month. Yet, you have full control/access to your data. This is the main benefit IMHO.
In our region, Tel Aviv, where Jelly Button and DoIT are located, $240K per year == 2x great s/w engineers annual salary.
Perhaps it is time for MixPanel's product/exec. people to reconsider pricing model at large scale. I assume there was a discussion between the parties before they went on let's build our own.
Is 500 events per second considered high throughput over here, or do other people feel they really just needed fault tolerance from Google cloud?
I was thinking GKE = Kubernetes, and GCE = was their "run container on a VM" thing, but I think GKE and GCE are the same thing...
MixPanel costs $X per month. That cost only increases. Self-based solution costs $Y to build and $Z to operate.
delta = $Z x n - ($Y + $Z x n)
Solve for that delta.
* FREE * 999 a year * Contact us
That number 3 must be really spooky!
(The math may be different in different job markets, in which case, yes, it makes sense not to pay a startup that's paying SF wages to its own employees.)
$192K ($240k/1.25) is definitely common; $120K ($240k/2) is low in these markets for someone who knows what they're doing.
One of my goto none coding interview questions is to have a candidate design a mobile analytics solution.
From the article, I've learned Google has a reference architecture: https://cloud.google.com/solutions/mobile/mobile-gaming-anal...
More valuable metric is person hours.
I recently implemented 99% of what mixpanel provides using S3 and lambda. I stored the events on S3 (actually get requests to an S3 bucket with logging turned on and what I wanted to track seeialized as JSON in the query string of the request.
From here it's as easy as writing lambda functions to process the logs and output the results in some location where a static web app can visualize them.