Musk: Twitter legal said I violated NDA by revealing bot check sample size = 100
twitter.com
twitter.com
(The key thing is that the sample has to be truly random, which is usually very hard to do, but Twitter can easily pick 100 random accounts in their own database.)
I'm not sure why the standard deviation worse case is 5%. Presumably the accuracy of the bot decision plays into it. But I would imagine a couple thousand would get a far more accurate answer.
Now, we want to know the standard error for our estimate of the bot proportion. That is sqrt(p(1-p)/n). Suppose 50% of accounts are bots (I assume that would be very high), then our estimate of p would be 0.5 and our standard error with a sample of 100 would be 0.05. Hence, our 95% confidence interval is roughly 0.4–0.6 in the worst case (with a sample of 100).
If the proportion is under 0.1 (let's assume 0.05), then the standard error would be sqrt(0.05(1-0.05)/100) = 0.022. Our 95% confidence interval in this case would be roughly 0.01–0.09.
These seem like large ranges to me. Hence, I would expect them to use a larger sample too.
For presidential elections it's common to see samples sizes like 1,000 or so, which have standard deviations of around 1.5%. That's better than 5%, and it makes sense since elections are often won by just a few %. But here, IIUC, the goal was to see if bots are a big part of Twitter or not, and so the answer "there are fewer than 5% bots" is enough. That is, we don't care if the % of bots is 3.5% or 4.7%.
What is your 5% value referencing?
https://www.calculator.net/standard-deviation-calculator.htm...
> That has a standard deviation of 0.5 in the very worst case
followed by
> the standard deviation of a random sample scales like 1 over the square root of the sample size, so 0.5 divided by 10 => 0.05 (5%).
I vaguely do remember the standard way of coming up with a reasonable sample size.
Assuming:
1) We want the commonly used confidence level of 95% (corresponding to a a Z-score of ~1.96).
2) We want a margin of error of 1 percentage point (kinda reasonable, since their estimate is stated as a percentage without explicitly stating the margin of error).
3) We don't know anything about the expected percentage of bots a priori.
4) We ignore the error in determining an account is a bot or not.
Then, we'll need a sample size of 1.96^2/(2*0.01)^2 = 9604.
> We want a margin of error of 1 percentage point (kinda reasonable, since their estimate is stated as a percentage without explicitly stating the margin of error)
If the question they want to answer is "give me the exact % of bots on Twitter", then yes, you probably want a standard deviation of less than 1%. That would take many thousands of samples, as you say - much larger than sample sizes used in presidential elections, even.
OTOH, my guess is that the question for them is "are bots a big part of the Twitter population?" In that case, 97% of accounts being real versus 96% doesn't matter much, since both show bots are a very small group. (But maybe bots write a disproportionately large # of tweets?)
I'm interested to hear how Musk violates NDA, not whether his bot checking approach is good or bad.
People focused on the bot checking approach because the NDA question is pretty dull and whether he technically did or not depends primarily on how paperwork was filed.
This is the trade secret of twitter bot checking approach?
Wow, this is so mundane.
It's not like the NDA says "keep our secrets unless you think they're too boring to keep". It says "anything marked in the following way is a secret you cannot reveal." If Musk thought that was public knowledge, he could fight them in court. Of course, that would torpedo his "this is new information to me, I need to get out of the TWTR acquisition" claims.
If you give anyone 5 seconds to come up with a methodology to check for bot, they would probably come up with "yeah, let's sample 100 of our followers and check manually".
It is the most obvious simplest method that it surprises me twitter can even call this trade secret under NDA.
Similarly, the fact that Twitter used 100, instead of 1000 or 2700 gives you error bars. It means something.
Semi-unrelated: a new stat grad would almost certainly not choose 100. They would feel compelled to run some math and come up with a different sample size that matched the results of some equations for power analysis, etc.
So, I guess it is not that mundane.
Downvote it if you like.
I'm interested in the discussion, so ....
Therefore, I'd prefer discussing it here.
Your opinion is certainly valid, so is mine.
If there was only some sort of democratised mechanism to settle this as a community and a way to easily hide the threads one isn't interested in.... I could only wish.