Rvest: Easy web scraping with R
blog.rstudio.org
blog.rstudio.org
The core language is still a confusing mess (I'm still never sure when to use a matrix, a dataframe, a list..), but if you use their tools you can ignore it for the most part.
In under 10 lines you can massage data and generate fantastic graphics.
A little off topic: but does anyone know what their business model is? Are they going to run out of money and burnout in a year or two?
And no, we're not planning on burning out. We currently sell three things:
* RStudio Server Pro. An commercial version of the open-source server version that provides stuff that corporate IT wants (e.g. monitoring, more auth options, ...)
* Shiny Server Pro. A more flexible version of the open-source shiny server that offers more configurability (e.g. number of R processes per app), and again other stuff that corporate IT wants.
* Right to use the RStudio desktop IDE to companies who don't want to use AGPL software
Deleted comment
matrix - If you have data that would make sense to be in a spreadsheet-type format and all your data are numbers.
dataframe - If you have data that would make sense to be in a spreadsheet-type format and some columns are numbers but other columns are something else (character strings, dates, TRUE/FALSE); but each column is only one thing. That is, you have one column that's all dates, another column that's all numbers, yet another column that's all character strings, etc.
list - if you need to mix data types within a certain entity (vector or column of data).
One useful heuristic worth asking is "Does it make sense to sort this data by something". In that case, you have a data frame. Whereas if you want to perform matrix math on something (inverting it, multiplying it by another matrix, reducing it, etc.), you have a matrix. Things that I use a matrix for can generally also be expressed as a data frame with columns rowId, colId, and value. If it doesn't make sense in that format, a matrix is generally not the appropriate structure.
Performance is also historically an R bugaboo, but with changes to R's copy on write semantics and other optimizations in the base language, current benchmarks show it behaving on par with Python and other dynamic languages (if not even slightly better with tools such as dplyr and data.table.)
The maggritr package's implementation of a "pipe semantic" (often considered the only truly successful implementation of a 'component architecture') and the adoption of the model for tools such as Rvset are really allowing for the functional, vectorized nature of R to shine through. These are really darned exciting times to be a part of this community!
Performance isn't currently a huge focus for dplyr. In my opinion dplyr is fast enough that the bottleneck becomes mostly cognitive - you spend more time thinking about what you want to do than actually doing it.
(Also you should use UTC and not GMT)
> as.chron("1970-01-01")+unclass(as.chron("2001-04-01")) [1] 04/01/01
> as.POSIXct("1970-01-01","EST")+unclass(as.POSIXct("2014-06-01","EST")) [1] "2014-06-01 05:00:00 EST"
If there is any conversion necessary it is difficult to get back the original intended time.
It looks like rvest intends to be the equivalent of Mechanize, with stateful navigation in the works. Is there an R equivalent to just Beautiful Soup or Nokogiri?
One of the reasons I learned Python for data scraping was that R in general does not play nice with https (RCurl requires a certificate and even then it's pretty fussy)
A true 10xer.