77 karma · joined January 2, 2014
A quick Google search suggests the prevalence of pancreatic cancer in the population is 13 per 100,000.
So if you gave this test with a 0.005 false positive rate and 0.5 true positive rate to 100,000 people it would miss diagnose 500 people and only correctly detect 7 cancers.
So given you had a positive test result there would be a 1-(7/500)=98.6% chance you did _not_ have pancreatic cancer.
Doesn't seem very useful in that light...
The language statistics show a good deal of FORTRAN (%24.5) [0], however that is largely skewed by the included LAPACK code [1], which accounts for 221,921 / 259,773 lines of FORTRAN in R.
[0]: https://github.com/wch/r-source [1]: https://github.com/wch/r-source/tree/trunk/src/modules/lapac...
This makes Super Hexagon more a game of quick pattern recognition than reaction time.
When I play it I am generally focused at the edges of the screen to quickly identify the next pattern and only using my peripheral vision to maneuver around the walls in the center.
0: https://github.com/jimhester/dotfiles/blob/master/colemak/vi...
I have used QWERTY, Dvorak and Colemak for multiple years and Colemak is the clearly the best of the there for me.
[1]: https://colemak.com/
[2]:https://colemak.com/FAQ#What.27s_wrong_with_the_Dvorak_layou...
They actually display fairly well in vim
2. The author posted the exact search terms used for all languages in an earlier post [0]
[0]: http://r4stats.com/articles/how-to-search-for-data-science-a...
awk -e '!a[$0]++'
This also preserves the original input order, which is a nice property.I used Dvorak for ~2 years and then switched to using colemak for the last 3+. OS support for both is widespread.
You can get back up close to your QWERTY speed in about a month or so (maybe less if going from QWERTY straight to colemak).
I switched to the alternatives to reduce RSI rather than speed and found it helped me with both.
[1]: http://colemak.com/
An alternate (simpler) implementation of the rvest web scraping example is at https://gist.github.com/jimhester/01087e190618cc91a213
It would be even simpler but basketball-reference designs it's tables for humans rather than for easy scraping.
On the topic of the article I have per directory history going back to ~2013 when I wrote the script.
Also note that RealWaitForChar has been completely removed from neovim since April of 2014 (https://github.com/neovim/neovim/pull/474/files)
> The x axis (log scaled) gives the number of speakers (plus one, so as not to make dead languages fall off the scale). The y axis, also log scaled, shows the adjusted wikipedia size.
As is the real ratio
> real ratio, defined as the number of ‘real’ pages divided by the total page count
Not sure why they couldn't put the above in the figure legend (or better yet label the plot directly). I agree it is poorly done.
As far as Julia goes it aims to bring strong typing and very fast native performance for numerical operations. Neither of which R, Matlab, or Python provide. Seems perfectly reasonable to me.
[1]: http://en.m.wikipedia.org/wiki/S_(programming_language)
[2]: http://en.m.wikipedia.org/wiki/MATLAB
[3]: http://en.m.wikipedia.org/wiki/Python_(programming_language)
awk 'END{for(i=0;i<1000000;i++){ print int(rand() * 1000000) } }' </dev/null > data.txt
time datamash sstdev 1 < data.txt
288619.72189328
0.72s user 0.01s system 99% cpu 0.736 total
time R --vanilla --slave -e 'sd(scan("data.txt"))'
Read 1000000 items
[1] 288619.7
1.09s user 0.04s system 99% cpu 1.134 total
R read.table read performance is fairly slow by default because it has to infer the types of columns and check for inline comments, quotes ect.This seems like a better replacement for awk and bash one-liners to me than tasks I would use R for.
For instance counting unique elements.
#naive approach
time (sort data.txt | uniq | wc -l)
632209
13.09s user 0.04s system 101% cpu 12.984 total
#using hashing
time (awk '!a[$0]++' data.txt | wc -l)
632209
1.34s user 0.03s system 100% cpu 1.360 total
#R
time R --vanilla --slave -e 'length(unique(scan("data.txt")))'
Read 1000000 items
[1] 632209
1.20s user 0.04s system 99% cpu 1.244 total
#datamash
time datamash countunique 1 <data.txt
632209
0.83s user 0.01s system 99% cpu 0.840 total
Quite good performance in that case, although R surprised me here as well.