Data Analysis and Visualization Using R (2014)
varianceexplained.org
varianceexplained.org
If you are interested in learning R, you may want to read the R for Data Science book (http://r4ds.had.co.nz/) book by dplyr (and ggplot2) author Hadley Wickham.
Relatedly, I have my own (slightly more complicated) notebooks using R/dplyr/ggplot2, open-sourced on GitHub, if you want further examples of real-world analysis with publically-available data along the lines of the Trump Tweet analysis:
Processing Stack Overflow Developer data: https://github.com/minimaxir/stack-overflow-survey/blob/mast...
Identifying related Reddit Subreddits: https://github.com/minimaxir/subreddit-related/blob/master/f...
Determining correlation between genders of lead actors of movies on box office revenue: https://github.com/minimaxir/movie-gender/blob/master/movie_...
http://varianceexplained.org/r/trump-tweets
By the author of the tutorial, not the poster of the link.
I'm working with DataCamp to develop an R course that covers dplyr, tidyr, and other newer additions to the R language.
1. Debugging seems way more primitive than in other languages; I get cryptic messages and really struggle to pinpoint what is happening. Debugging in (free) shiny is even harder, the page says connection closed and I have to guess what has happened.
2) Code structure. R is simply fantastic in REPL and/or RStudio mode for digging around in data, but longer programs remind me of COBOL (yes, I have programmed in COBOL) longer programs written by other people remind me of the need to drink alcohol. Creating good code with R is vastly harder than Julia, in Julia the challenge is not to create working clean code - that's natural, the challenge is to create the best code that it's ever possible to have. In R the challenge (for me) is to make it work and not make a plate of spaghetti.
And you probably need to be more assertive.
https://cran.r-project.org/web/packages/assertive/index.html
https://www.youtube.com/watch?v=JWjiMvlfCwk
(and you do use testhat for unit testing, right?)
Additionally you want to write more modular code. There is lots of infrastructure around that in R, but people just don't use it often enough because a lot of them aren't programmers.
mlr provides very convenient infrastructure for building data mining pipelines where you can fuse steps with each other.
http://mlr-org.github.io/mlr-tutorial/release/html/
For non-model building activities, i.e. inference or exploratory analysis, mason is a great way to do it.
https://cran.r-project.org/web/packages/mason/vignettes/spec...
Ad 2. There are reference classes which also provide limited type checking for fields. You can also encapsule your code in environments which is more R-style but doesn't work well with roxygen.
https://github.com/klmr/modules
https://github.com/wahani/modules
Neither is perfect, but I've found them helpful in my projects. They involve much less overhead than writing packages, especially when the modularization I'm trying to achieve is purely internal to my project and I don't intend to publish the code. At the same time, they provide much better encapsulation compared to `base::source`.