104 karma · joined April 23, 2017
For some DS, this is true. For other DS, they are there to create or maintain a rigorous process so that we can reliable causal inference.
Example: how effective / safe is the covid-19 vaccine on trial. Thinking most of the statisticians in FDA and pharmaceutical companies.
Say you buy a stock for $1000 in year 1. Suppose monthly cost for living in that year is $1000.
In year 10, you sell the stock for $2500. Suppose monthly cost for living in year 10 is $2000. Your investment essentially just beat the inflation by $500 at year 10. In other words, real capital gain is not $1500.
Comparing Tesla cars crash rate with that of the overall population is dishonest:
1. drivers are biased population 2. the age of the car is biased
1. The accident rate does not take into account of drivers age, credit score and prior safety record, or the age / value of the car.
2. Most people only turn on autopilot when driving is easy (e.g. on a highway).
Is this a meaningful comparison? These benchmarks repeatedly call the function (with same input file) many times and then take the average of the time spent. However, in real world, we only read a specific file ONCE -- if we call the csv reader several times, it is almost always for different files.
If you read the dict union operator "|" as something similar to "or", (like me), you are in for a surprise:
>>> x = {"key1": "value1 from x", "key2": "value2 from x"}
>>> y = {"key2": "value2 from y", "key3": "value3 from y"}
>>> x | y
{'key1': 'value1 from x', 'key2': 'value2 from y', 'key3': 'value3 from y'}
What I expect is
>>> x | y
{"key1": "value1 from x", "key2": "value2 from x", "key3": "value3 from y"}
Correlation does not imply causation[1]. Looks to me this is one of the studies that I would label as "shallow" or "problematic".
https://en.wikipedia.org/wiki/Correlation_does_not_imply_cau...
Depends on how "accurate" you want to measure a ratio. Here the ratio is the percentage of Americans that support net neutrality.
In this case, sample size is 1000, estimated ratio is 76%, the 95% confidence interval is 76% +/- 2.65%, which means if you repeat this survey again, you have a 95% chance that new estimated ratio is within 76% +/- 2.65%.
Edit: 99% confidence interval is 76% +/- 3.48%
Money laundering is the process of transforming the profits of crime and corruption into ostensibly 'legitimate' assets.