Whom the gods would destroy, they first give real-time analytics (2013)
mcfunley.com
mcfunley.com
Emphasizing working on the homepage was also analytically dumb in a more meta way, since item/shop/search were really nearly all of traffic and sales back then. Anyway, I felt motivated to get that person to think first and fire the code missiles second.
At the end of the day, I think back on it fondly even though it was ludicrous. Shipping that much stuff to production that quickly and that safely was a real high water mark in my engineering career and I've been chasing the high ever since.
You can ship A/B tests quickly and many websites do, but decisions are made after a statistically relevant time period.
That said, once you build a system for operational metrics (i.e. what you need to detect anomalies that indicate outages, security concerns, etc.) you're already a huge way there towards having real-time analytics. I still wholeheartedly agree with the author that these real time metrics should only be in the service of operations, not product planning.
Build for the future you want, don't just head for the local minima. Take risks and damn the statistics.
Whom the Gods Would Destroy, They First Give Real-Time Analytics (2013) - https://news.ycombinator.com/item?id=15379660 - Oct 2017 (70 comments)
Whom the Gods Would Destroy, They First Give Real-time Analytics - https://news.ycombinator.com/item?id=6515805 - Oct 2013 (1 comment)
Whom the Gods Would Destroy, They First Give Real-time Analytics - https://news.ycombinator.com/item?id=5032588 - Jan 2013 (55 comments)
As an engineer on many teams shipping features I've found that it's somehow underwhelming to finally launch something after months of work. You launch and the only thing you get to celebrate is some donuts in the office and if something goes wrong a notification from Sentry or Datadog :P
I've spent the past 3 years building a product analytics tool (https://june.so) and I think product analytics can deliver some real-time value to teams.
Some of the ways we've built our product to do this is:
- Live notifications in Slack for important events - to get pinged in a Slack channel when users use the feature you just launched
- Achievements on reports for your features - to celebrate the first 5, 10, 25 and 50 users using your product, see the progress live
I think for team morale, especially in the earlier days of a company it's great to celebrate small wins and as engineers we should be more connected to what happens inside of the products we build - not only when things go wrong.
Just saying that as in the conclusion of the article OP says “what do you need this information for?” and my understanding is that the people asking for “real time metrics” aren’t trying to do anything complex but get a pulse of the product
Seeing people using a hard-to-build feature a couple times a day, then more, until eventually you have to mute notifications to focus on work is a great way A/ to feel the progress, and B/ notice trends you can't pick out in averages.
Example for A: Just yesterday our CTO wrote in a feature-specific channel: > This page is now unreadable due to volume of usage pings! Go team!!
Example for B: Intuitively noticing whether your tool, that has say 6 DAUs on a team, is being used once by all 6 people, or in 3 pairing sessions, or something in between. Yes could run an analysis for this, but at an early stage co it's easier to just notice.
We became June users at our pre-launch co a few months ago, and the feature 0xferruccio mentioned is part of what sold me initially.
Not sure how long it'll remain useful but loving it for now.
Yes you can measure some related things that give you some hints about the thing you care about, but they are fragile. To borrow from Goodhart, if you make the related things a target, they will stop giving you even these hints.
This applies not just in software development, but life in general.
The problem is analyzing why or why not it's hitting your desired value.
This is why I think the idea of technocracy or "evidence based politics" is ultimately a mirage. Sure you can maybe assess some policy but the metrics you're choosing to measure or optimise for are political by their very nature. One's evidence based policy isn't the same as mine.
Health-outcomes-wise it would be better to force everyone to eat salad or whatever but that's only one dimension to optimise on at the expense of freedom and life enjoyment.
Tying it back to tech maybe going down market improves your conversion and lowers your CAC but maybe you've just acquired a bunch of customers with low value, high churn and high costs.
Maybe the sales of Amazon Prime are showing gangbusters returns with the dark patterns but now people loathe your brand and are hoping to see you hit with an FTC banhammer.
There's no silver bullet to this stuff, sure you should probably measure it but ultimately you have to make a decision and be guided by gut instinct and beliefs.
Horrifying but all too common. A wise man once taught me that humans feel considerable discomfort in the presence of uncertainty, and will tend to jump on the first solution that presents itself. His answer to this was to strive to stay in 'exploration mode' for as long as possible - explore the solution space until you've hit the sides and the back and only then make a decision.
[1] "The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data Processing" https://research.google/pubs/pub43864/
The place its truth has been most obvious to me is in analyzing subscription businesses. Your customers pay once a month, or once a year. Nothing, absolutely nothing you do on a minute-by-minute basis is relevant to strategic business decision making. Yet these businesses will invariably want real-time analytics. It serves no purpose! You simply cannot look at your fancy real-time dashboard, and then take action based on it on that same time scale. Meanwhile it costs you 10x what daily data would.
For a longtime there has been a lot of resistance "real time is too hard" used to get thrown around a lot- what is really driving it in last two years or so are Machine Learning and computer vision applications. There has been a huge push to integrate ML models (for example live defect detection) into our operating process which has necessitated low latency access to real time plant data.
The data pipelines we've had to build to enable these ML applications have bough latency way down all over the place and have kind of bought other applications like real time dashboards with them for "free".
It's incredible how much of what he says (especially the stuff about how production workers feel) maps directly to the technology industry.
There is also the question of how long you'd leave multiple treatments out there. Presumably, even if there is no difference in outcomes, there can be benefits to having fewer deployed behaviors.
I'm now also curious if there are non-transitive situations. For example, three treatments together that all act fair if all deployed, but for reasons any two of them deployed alone will show a preference. Ideally, of course, treatments should be done such that this can't happen, but mistakes are often made.
Edit: Fully cede that this is likely chasing edges. The motivation for fewer deployed arms is far more compelling than the edge cases.
1. Product didn't have any idea how to interpret its behavior and therefore never made any decisions based on it
2. Experimentation != product design. It's one thing to look at the results of a test, it's another thing to consider patterns of user behavior observed over months or years, which is what Product Analytics is actually for.
The real benefits here are getting a better understanding of what levers drive your product metrics, as you'll inevitably mess up the first n or so experiments (if I could give you only one piece of advice, it would be to use stratified randomisation, but everyone seems to have to make this mistake for themselves).
Heck you don't even have to do MAB if you don't want to, just don't use NHST. The Bayesian "flavor" of NHST (credible intervals around posterior expected values) has absolutely no problem with optional stopping. Run the experiment until you've got a precise enough estimate, then sit back and make your product decisions.
I guess where I'm going with all this is that it seems like the post's strongest point is "good product decisions require time, and realtime analytics bamboozle us into thinking fast decisions are better". All the stuff about NHST seems kind of tangential. Looking at it again I see that it's like a decade old, so I think this is the best explanation for why they were targeting NHST more aggressively. I would hope in our post-replication crisis world (hopefully "post", anyways) data scientists and A/B testers are more prudent about some of these better-known pitfalls.
This is the crucial misunderstanding: in actuality, you are running a panel.
(There is no such thing as an A/B test outside of marketing. Running a meaningful panel requires some information on the population, your samples, the homogeneity of those, etc, just to pick the right test, to begin with. Also, you need a controlled setup, which notably includes a predetermined, fixed timeframe for your panel to run. Before this is over, you have no data, at all. You are merely tossing coins…)
(In this specific case, as a data analyst, you will probably have an intricate understanding of your population, i.e. your data, the structure of the samples you're running against the algorithms, which you have tailored according to this understanding in the first place. However, while we may assume the best, this may still be what's called pre-scientific knowledge in statistical terms.)
In my entire career I’ve never come across a situation where a company either needed or could even use real time data analytics.
Real time data: occasionally but definitely. Real time data analytics: never.
Me: or you know, we could maybe see what the users biggest issues are first and try to build stuff to solve those problems.
Bart: Isn't that the wrong way?
Homer: Yeah. But faster!
- "Homer to the Max" "
Now that's funny! <g> :-)
What would reasonably costed look like to you?
More on my view on real-time analytics here: https://news.ycombinator.com/item?id=15380607
I don't really know what reasonably costed would've looked like for the time I used Amplitude (2018-2019), but I know that the value my team extracted from it was not commensurate with its cost. Whether that was because of overzealous assumptions on our part, or something else, I don't really know, but I know we signed up, used for a few months and then canceled/downgraded to a cheaper service whose name is escaping me.
I was mostly challenging the "Just use ____" notion of the commenter above, not really that Amplitude is worth the money for correctly-intentioned businesses. Regardless I appreciate the ask.