Most commonly used statistical tests and implementation in R
r-statistics.co
r-statistics.co
Erm, no. P=0.05 is borderline meaningless, there could as much as 30% chance you are wrong about the actual difference being there depending on the true probability of the initial hypothesis.
P-values should be used with strong caution.
FiveThirtyEight (and Scientific American, and others) did some pretty interesting articles about this recently if you haven't seen it:
http://fivethirtyeight.com/features/science-isnt-broken/
Just from personal experience, the use of p-values is really broken in biology/chemistry. The things I've heard principal investigators say...
I'm having trouble parsing this, are you talking about the power of the test?
This is a common problem in many fields, and you can use false discovery rate control methods to account for it: http://www.statisticsdonewrong.com/p-value.html
Everything is just so much more sensible if you allow yourself to assign probabilities to hypotheses, rather than assuming a hypothesis from the outset and computing opaque statistics relating to your data.
For example, the chisq.test has optional built-in Monte Carlo testing, and none of the other functions do, oddly.
R has some really good GUI layers now. I struggled and struggled for years trying to learn the command line methods, but it was too much for me. The following do a great job (these are alternatives)
- Deducer
- R Commander
- RKWard
The company I work for, Domino Data Lab[5], let's you fire up a lot of these notebooks in a nice hosted environment on big cloud servers with minimal cost and effort. It's a fun way to learn how all these new environments can work together. From RStudio for exploratory analysis, to Jupyter notebooks for presenting a topic. The other two I haven't really found the superior use-case. The tools in this space are just getting better and better.
1. https://www.rstudio.com/ 2. http://jupyter.org/ 3. http://blog.yhat.com/posts/introducing-rodeo.html 4. http://beakernotebook.com/ 5. https://www.dominodatalab.com/
Jupyter and R is a bit iffy since the R kernel is not native. Although the kernel works fine, setting it up has a ton of manually-installed dependencies, and in-line plots flat-out give unexpected output. (I've had to cheat by embeding charts via Markdown. Although that has the benefit of having the charts be responsive)
The important perk is that Jupyter notebooks are now rendered natively on GitHub, which I've made considerable use of: https://github.com/minimaxir/sf-arrests-when-where/blob/mast...
You know, to be completely honest, I've never used it directly. I've always used it on our platform. It's very possible that our engineers already did all that setup so it "just works." I took the original post: http://r-statistics.co/Statistical-Tests-in-R.html and reimplemented it in an R notebook with some simple plots at the end, but yeah, the plotting just sort of works for me. I didn't realize I had an incomplete view of the complexity of getting that working :(
https://app.dominodatalab.com/earino/statistical_tests/view/...
We also render the notebooks. The difference is that we also let you run them :)
Edit: the closest thing I can think of in RStudio to that is installing the manipulate package which allows adding sliders and such to plots for some custom plotting controls.
1. If you're interested in running a shiny server, use http://rstudio.github.io/shinydashboard/! I have used it to build professional high quality dashboards VERY quickly.
2. You can use an API server like Domino's API end points or OpenCPU to expose R APIs and build the interface using JavaScript at plot.ly! This really can be incredibly elegant and you can do really neat dynamic dashboards.
Is opencpu something you recommend for production? I'm just starting to work with analysts who work in R, and I have struggled with the question whether we should wrap existing R code as an api...or port to python.
Low volume right now, so not really concerned with performance... But rather that can R deployment play nicely with things like supervisord,etc in production
As for OpenCPU, I know that the guy who wrote it, Jeroen Ooms is genuinely quite brilliant. I know it was his project during his PhD, and I don't know what his plans are for continuing to support it. It's up to you to determine what that means for your "production" needs.
There are lots of people who have rolled their own solutions for production deployment. Including nodejs !
I keep banging my head against issues around persistent data storage and app customisation at a user level. Unless one pays for Shiny Server Pro, the free Shiny Server doesn't support user authentication. Hosting on shinyapps.io doesn't really support persistent user data, unless it's offloaded elsewhere such as Dropbox or a remote SQL database, which brings into play a bunch of security questions.
Shiny is good, but not quite yet outstanding.
I was really interested in an article I read recently about using jQuery.ui widgets and R to build interactive web app.[1] I'm keen to explore this as a potential way forward.
[1] http://www.r-bloggers.com/creating-multi-tab-reports-with-r-...
I am willing to bet that he would rather use R api and excel to build a dashboard, rather than anything else.
I just wanted to point out that there are limitations with this route, when the apps start to become more complex, with multiple users, various access permissions and personalisation requirements.
So I made a small demo for you. If you go to https://app.dominodatalab.com/earino/d3_dashboard_demo/raw/6... you will see a very simplistic, d3 powered, R backed dashboard. It just draws a pie chart, a line chart, and a bar chart. Every 10 seconds, it polls an R API endpoint to get new data.
The code for the R endpoint is https://app.dominodatalab.com/earino/d3_dashboard_demo/view/.... It's a simple R function which generates synthetic data. It could be used to pull much more complex data, generate predictions from an ML model, etc...
Please don't hesitate to reach out if you have any questions!