Allstate compares SAS, Hadoop and R for Big-Data Insurance Models
r-bloggers.com
r-bloggers.com
2) It is dishonest ad copy
This model could have been estimated with the biglm library. Revolution's claim that they are the only game in town for big data computing with R is bullshit.
The SAS performance reported is surprisingly bad and likely is a result of the data processing steps involved. A SAS procedure is much more similar to a program than a function. While SAS wasn't the fastest solution, most of its slowness was due to building the design matrix rather than running the actual regression. There are a number of options for both tuning how this is done and caching this preprocessing. The biggest different would be using SSD's on the machine which seems unlikely given the results.
They also won't even be seriously considered at most large companies because the typical person in that role doesn't have the skills and a single-person using a different solution makes it difficult to transition work.
If you're fitting a model with only 70 degrees of freedom then analyzing 150M records is a complete waste of time.
That said, 50000 is too few. For a dataset of this size, 20 million records is likely more reasonable. The actual answer depends on the variance of the individual predictors and their correlation with each other.