The pitfalls of A/B testing in social networks
tech.okcupid.com
tech.okcupid.com
They managed to deal with touchy dating sterotypes, race, sex. Reading it, you could tell that they were genuinely smart (as opposed to just sophisticated) in their data/analytical mindset. You also got a feel for how the author thinks and where that mindset translates into making OKC.
Anyway, this reads like a linkedin article. Safe, vanilla company blog post. Big contrast. I wonder if the OKC guys writing that old blog 10 years ago new how unique their position was.
[1] Got taken down after they got bought by the target of the post. Accessible here: http://static.izs.me/why-you-should-never-pay-for-online-dat...
[2]https://theblog.okcupid.com/the-case-for-an-older-woman-99d8...
Still, I imagine there's an underlying shift in company character that being expressed. The old blog felt like it was interesting to the writer.
This is in no way a condemnation of the post. I hope I haven't insulted the author.
If anyone has found a good way to find that balance and create really solid in-depth data-driven content for there org, I'd love to learn more.
Censoring past published works is quite another.
When it comes to actual messaging activity, both graphs show a 'belly' of preferring to message younger partners (except near the origin); it's just that the men's one is considerably more pronounced. Both genders are sending a lot of messages to people younger than their own minimum settings, too.
In general I liked the OKC data articles, but this particular one always seemed to me like they were forcing a story onto the data.
- Facebook data scientists have an extensive list of publications on the topic: [1][2]. (Dean Eckles is now at MIT).
- LinkedIn's experimentation team has also published papers on the topic: [3][4]
- Data scientists at Google have also worked on this topic. One external publication/collaboration I'm aware of is: [5]
[1] http://www.pnas.org/content/113/27/7316
[2] https://arxiv.org/abs/1404.7530
[3] http://www.kdd.org/kdd2017/papers/view/detecting-network-eff...
[4] https://dl.acm.org/citation.cfm?doid=2783258.2788602
[5] http://proceedings.mlr.press/v51/basse16b.pdf
Disclaimer: this list is heavily biased by experiences I've had through collaborations/internships. I am an author on [3]. Edit: formatting.
A/B testing is really hard, even when there is no social interaction internal to your application (like a game). The easiest thing to test is marketing for new users. But even that can be tricky, where one group of people might click fewer ads but are higher quality. This is especially troublesome when you can't tell how many friends they invite through word of mouth. And then if you want to implement other features or fix bugs over the duration of the test, that also poisons the results.
And then the number of incoming users you need is massive.
To do a good A/B test that doesn't mislead you is extremely difficult. Almost to the point that for most companies, I don't think it is worth doing.
> So... I guess my final recommendation is that you should hire some data scientists that like doing experiments.
Like this guy says, you basically need a data science team to do it effectively. If you've got a small team (like I work with), don't waste your time with it. You're better off just adding new features and fixing bugs.
- it's easy to push the wrong people down the funnel
- use quantitive observation to train your qualified decisions
This is why I always pull performance metrics in addition to impressions/clicks when running A/B tests. Not being in the media/creative department means I often don't see the creatives i'm asked to analyze the performance of so there are occasionally different creatives that have a stronger resonance with people with a higher propensity to convert. This is itself a learning and could lead to using that sort of creative more for acquisition focused campaigns, rotating it out of the upper funnel ad rotation and coming up with a new upper funnel creative to test for the purposes of building large audience pools.
> This is especially troublesome when you can't tell how many friends they invite through word of mouth.
Could you provide an incentive (gold/credit/whatever) for each new active user referred which could feed into the evaluation criteria of the test as far as the "value" that a particular conversion brought with it? Then I suppose it turns more into a LTV study.
Many threads like this one:
https://www.reddit.com/r/OkCupid/comments/75r2ay/seriously_o...
I'm betting that they run tests for everything and change into whatever direction the results tells them. If people like "slide to the left/right" and a simple UI, no walls of text, quizzes, etc, so be it.