R, the master troll of statistical languages (2012)
talyarkoni.org
talyarkoni.org
Instead you must use a huge number of special tools that do only a few things. The code is hard to read, hard to write, slow, the compiled version is big. It is also error prone since you must use a large number of different paradigms. Some might have their arguments slightly differently. It's produced by googling for everything (or if there's a decent builtin help system, using that).
It seems strange that such concepts, like generic data types, operators and functions are not more widely spread. For example if you can calculate "max(a-b)" even if "a" is a matrix and "b" is a scalar, that's already quite nice. This can increase productivity a lot.
Since such design models are so fundamental, they are rarely found in "matlab-like" extensions for other programming languages: instead you must constantly do type conversions and transformations by hand and use awkward middle man functions to access data types. Most of your code is housekeeping and little actual progress.
See for example the problems with Julia where column and row vectors are totally different types: http://2pif.info/op/julia.html
That was one year ago. They are still different types, but they behave as you would expect wrt arithmetics.
julia> x = [1,2,3]
3-element Int64 Array:
1
2
3
julia> y = [1 2 3]
1x3 Int64 Array:
1 2 3
julia> y'
3x1 Int64 Array:
1
2
3
julia> x + y'
3x1 Int64 Array:
2
4
6
Their size and dimensions are still different, though, but it is mostly a concern for library writers. julia> ndims(x)
1
julia> ndims(y)
2
julia> ndims(y')
2
julia> size(x)
(3,)
julia> size(y)
(1,3)
See https://groups.google.com/forum/?fromgroups=#!msg/julia-dev/... for the discussion that resulted from that post.Edit: I should add that there's now a mailing list dedicated to statistics with Julia: https://groups.google.com/forum/#!forum/julia-stats
>> z=size(a)
(3,)
>> z[2]
BoundsError()I saw an example of R code using underscore for assignment posted on twitter a couple of months ago. See if you can make heads or tails of this: http://www.stat.washington.edu/hoff/Code/hoff_raftery_handco...
... and since the dot is legal in identifiers, they use $ as lookup operator.
> ice.cream$flavor
[1] "chocolate" "vanilla" "butterscotch advocado surprise"Huh, early Smalltalk mapped the left arrow to the underscore key; I'm not sure if Squeak and Pharo still do it, though.
>> I won’t bother to explain all of these; the point is that, as you can see, they all return the same result (namely, the first column of the ice.cream data frame, named ‘flavor’).
Having many ways to do one thing is good - whatever float your boat you know. Just for the given examples, I see how one may be better in a loop (the one where you use 1) while another one may be better when you type it on the command line to check stuff (the one with flavour)
>> The answer is that when you’re trying to learn a new programming language, you typically do it in large part by reading other people’s code–
No. You learn by doing. You write some stuff, test it, and if you don't get the result you expect go back to try and figure out what's wrong.
And even before doing that, you set aside some time to LEARN about the language.
But maybe, just maybe, the problem is not with the tool but with the person using it?
>> I have to confess that I’ve never set aside much time to really learn it very well; what basic competence I’ve developed has been acquired almost entirely by reading the inline help and consulting the Oracle of Bacon Google when I run into problems. I’m not very good at setting aside time for reading articles or books or working my way through other people’s code (probably the best way to learn), so the net result is that I don’t know R nearly as well as I should.
At least that's honest. Maybe you don't do well with a language (any language!) because huh, you didn't take time to study it and expect it to magically work??
I know I'm no good in java, but I also know why - I never took some time to actually learn it. I can do things with java, but not very complex things, and when I fight my way out of a problem I caused, I won't blame java but myself and my lack of knowledge of java.
It's funny how similar the arguments are.
They can both be described as a language with some serious gotchas, but that's simple enough and has such a great package collection that it's easy to get a lot of things done very quickly.
That description could also fit Javascript.
R is a bit of a train wreck. It makes hard things easy (if you know the right library or idiom), but often makes easy things hard.
Perl is the duct tape of languages, R is more like the A-10 Warthog, ugly but powerful for a specific job.
No. This is how you learn to do whatever simple things you want in a language. The way you learn how to do things correctly is mainly by reading code. If you have an expert to review your code that's better, but almost no one has that opportunity.
Edit:
I guess I should come right out and say what I'm thinking: not everyone's opinions about a programming language are equal. I suppose this is why people are constantly having discussions about the pluses and minuses of "dynamically typed" languages, while type theorists don't even recognize "dynamic typing" as a form of typing. Expert analysis has consistently shown problems with the R language design and implementation. It's not a good language. Its features are often misused or poorly used and it doesn't have a strong sense of what support it wants to give to its users.
http://channel9.msdn.com/Events/Lang-NEXT/Lang-NEXT-2012/Why...
It may be a terrible language by some objective external computer science criteria. But as a domain specific language, that's not really important. What is important is that the domain specific aspect works well for people with domain knowledge.
It's not the right tool for every job or every person, but it's incredibly expressive for people working with a specific set of problems.
I was quite clearly criticizing the language, which is bad. That doesn't mean you shouldn't use it if it's the best option for your use case, just that it's bad.
Dynamically typed languages have a typed runtime, do not a type in the static program text. Hence the commonly used correct qualifiere: "dynamic" and "static".
http://research.microsoft.com/apps/pubs/default.aspx?id=6752...
A type system is a tractable syntactic method for proving the absence of certain program behaviors by classifying phrases according to the kinds of values they compute.
The word “static” is sometimes added explicitly--we speak of a “statically typed programming language,” for example--to distinguish the sorts of compile-time analyses we are considering here from the dynamic or latent typing found in languages such as Scheme (Sussman and Steele, 1975; Kelsey, Clinger, and Rees, 1998; Dybvig, 1996), where run-time type tags are used to distinguish different kinds of structures in the heap. Terms like “dynamically typed” are arguably misnomers and should probably be replaced by “dynamically checked,” but the usage is standard.
Perhaps I should have been more clear that we use the words "dynamically typed" because they have entered the lexicon, but they are misnomers -- they do not properly capture what type theorists mean when they say "type".
Could people here list any resources on type theory that comes to there mind? Books, blogs, people, etc. A book that explains the fundamentals would be great.
You can check it online while waiting for the Pierson ones: http://research.microsoft.com/en-us/um/people/simonpj/papers...
Also, thanks for the link to the article.
But maybe you just don't like "dynamically typed" languages?
As for sharing, the semantics cleary demonstrates that R prevents sharing by performing copies at assignments. The R implementation uses copy-on-write to reduce the number of copies. With superassignment, environments can be used as shared mutable data structures. The way assignment into vectors preserves the pass-by-value semantics is rather unusual and, from personal experience, it is unclear if programmers understand the feature
And no, I did my master's thesis in Racket. I'm fine with unityped languages, I just describe them properly.
On a related note, have you read the riposte paper? http://www.justintalbot.com/wp-content/uploads/2012/10/pact0...
;)
But seriously, every under-used feature harbors a ton of bugs in the compiler and drains resources for implementation. Promises, for example, seem like a "feature" in R that only really serves to destroy performance and no one uses it.
> Having many ways to do one thing is good - whatever float your boat you know.
I can't get the article itself to load, but it cracks me up that foo[range] and foo[range, ] do completely different things.
--someone who's writing numeric/ data analysis tools in Haskell.
I am quite literally building a full data analysis stack (as a product) in haskell, some parts of which will be available as a sort of proprietary augmented version of the haskell platform, and some parts are / will be open source.
I do think that there are compelling reasons to consider Haskell / GHC for analytical workloads, but depending on the details it really depends.
The principal cliff is just the HUGE number of (mostly poorly designed) libraries for many standard analyses written in R. Theres some nice engineering approaches to circumvent this, and theres some really exciting libs that a uniquely awesome and handy in haskell land.
A notable example is AD, a really easy to use auto differentiation lib by Edward Kmett, which has a really exciting refactor thats nearly done that will make it useable by mortal Haskellers :) http://hackage.haskell.org/package/ad and https://github.com/ekmett/ad (I've some neat bits i'll be hopefully adding to AD myself in the next month)
It's like every line I wrote there was some catch or trick I had to know just to get it to work. Pretty frustrating.
I've seriously got something like a 100 textbooks on R on my system, intro to stats with R, machine learning with R, questionnaire analysis with R, bayesian stats with R... There are Coursera classes, the amazing r-bloggers.com, etc. For someone who is simultaneously learning statistics and a tool, this is invaluable.
For Python, there's still very little. There's Wes' book, which is mostly about pandas and a lot of finance/time series stuff... I haven't seen a single "intro to stats in social sciences with Python"... There must be one book out there that is open source, which could just be rewritten with Python examples? (The most useful thing I've seen is people reworking examples from Machine Learning for Hackers, or one of the Coursera R courses, in Python).
I agree it's frustrating. I've been working with R on and off for more than five years and almost exclusively for 1-2 years, and I still need google every ten minutes to remember some esoteric command that is perfect for the problem at hand. In my experience this is still quicker and less error-prone than python, which often requires you to effectively roll out your own solutions for small data processing needs.
I still struggle sometimes to write nice looking code in R, but I've known people who are pretty good at managing it. If you have to inherit an R program from someone else though, it's almost guaranteed to be a nightmare to comprehend.
* The lack of built-in dataframes and libraries to work on them (like plyr). Pandas seems to be getting pretty good, but it's still not as mature as R's solutions.
* Visualization. ggplot for R is great, matplotlib for python not so good IMO. I've heard Bokeh and rplot are attempting to bring ggplot functionality to python. Again, not nearly as mature as R's solution.
I'd love to move to Python because R is not a fun language to develop software in. But at this point, R is by far the better tool for working with data (for my needs at least).
There are many little idiosyncrasies in R's syntax, I feel like I never grokked the language. For example, pretty much anytime I see '~' I have to relearn what is going on. From a mathematical perspective I appreciate that vectors are indexed from 1 instead of 0, but from a programming perspective it can be annoying.
BTW, thanks for contributing so many great things to R, I owe a lot of what I do to you.
Part of the problem is the base packages: as soon as you open R you have ~1600 functions you can use, and you obviously can't memorise a significant proportion of them. Learning R is as much about what you don't learn as what you do.
Actually one thing that has helped in the past year was reading your split/apply/combine paper and using the plyr package more.
I have a function in R that given a list of keywords, for each word (~100 keywords), it checks a data frame that stores a list of keywords (~5000 lists of keywords, ~10 keywords/list) to see which keywords matches the keywords in each row of the data frame.
Efficiency using straight for loops would be O(mno) and could take awhile. But in R, all those operations are vectors, and it takes less than a second.
Exactly that. I've always found R a horribly confusing mess, compared to general programming languages (I'm proficient in Ruby, JS, Obj-C).
On the other hand, compared to some of the other commercial stats packages, it's beautiful and logical and reasonable. I regularly use Stata, where you're only allowed one data table in memory at once, and almost everything relies on side effects and Byzantine macros. E.g want to calculate a mean? First, 'summarise' the variable, then assign 'r(mean)' to a var name, then quote that the right way to be substituted into an expression where it's needed.