Introduction to Statistics using NumPy
mubaris.com
mubaris.com
However, by default numpy gives you population variance, with N as the divisor. This assumes X is the entire population, which is different from R and Matlab.
As a result, the answer in your code is different from the formula just above it.
np.var(X)
gives 207, while np.sum((np.array(X) - np.mean(X))**2) / (len(X) - 1)
gives 236. If you want the N-1 corrected sample variance from numpy you'd use np.var(X, ddof=1)Then, readers learn that statistics is about applying a few summing and averaging procedures, and trust the numbers they get as a result.
Or E.T. Jaynes - Probability Theory, The Logic of Science (but this one is much more detailed and has lots of philosophy)
Or Ian Hacking - An Introduction to Probability and Inductive Logic (this is the light approach written for philosophy students)
http://nbviewer.jupyter.org/url/norvig.com/ipython/Probabili...
from statistics import mean, median, stdev, varianceSecond is that everyone uses NumPy for this. It is proven, it is fast, and using it for simple things is a great way to get started on a path toward using it for more complex things. For example I use NumPy for tabular text processing sometimes, as it is much faster than the built-in stuff if you have a lot of data.
Missing a close paren to open(); and splitlines() is misspelled. Code should be:
with open('salary.txt') as f: X = f.read().splitlines()
https://news.ycombinator.com/item?id=15088840
Original Responses - https://docs.google.com/spreadsheets/d/1cjwB_s4ya57auTjIw-OV...
Then I went on to refine the data for my purposes - https://docs.google.com/spreadsheets/d/1AAJmOWOE-zYydNR_HzLU...
I started to look at the dataset a little, if someone is interested: https://github.com/davidgengenbach/developer-salary-statisti...