Safely Rewriting Mixpanel’s Highest-Throughput Service
engineering.mixpanel.com
engineering.mixpanel.com
Could someone from Mixpanel elaborate maybe?
As a startup with limited resources, it's important for us to invest all our engineering strength into the things that create direct value for our business. We'd rather pay Google to manage machines and run services like Kubernetes, Spanner, Pub/Sub and others and free up the engineers to work on our core analytics platform.
* analyze every datapoint you receive
* when the dimensional cardinality is high
* you want to analyze behaviors over time (e.g. the output depends on the orders of events followed - like creating a funnel report)
There's no off-the-shelf solution that does this at the scale at which we operate - hence the need to write our own custom solution.Then I used that test/spec to do a TDD type development of the new service. It was the easiest rollout I've ever done. Everything just worked when it went into production. I even ended up giving some internal presentations on the process.
I also tested with logged input from the source program. It's neat to see this technique is common.
Just out of curiosity, what were some of the bugs you found? Were they related to semantics of python not carrying over to go? or was it that you tried using new go features like goroutines and they didnt work as expected?
if val: # ...
in Go.
Both max and avg p99 latency became much more stable. Max appears to have gone down a little too.