The plots and chi-square tests are a single function each, it's the data preparation that takes up so much space. Alas, most statistical analyses involve a lot of preparatory steps, which are rarely shown in the final write-up.
d <- read.csv(file='~/stuff/earlh/Iran_2009.csv', header=T, sep=',')
lastDigit <- function(v){
v - 10*floor(v/10)
}
digits <- lastDigit( c(d$Ahmadinejad, d$Karroubi, d$Mousavi, d$Rezaee))
hist(digits, breaks=10)
#chi2 gof
tab <- table(digits)
n <- length(digits)
model <- chisq.test(x=tab, p=rep(0.1, 10))
model
# hand generated -- check our work above
ts <- 0
for(i in 1:length(tab)){
ts <- ts + ( tab[[i]] - 0.1*n)^2 / (0.1*n)
}
qchisq(p=1-0.076, df=9)
and to be more specific: (def regions (sel votes :cols "Region"))
(def ahmadinejad-votes (sel votes :cols "Ahmadinejad"))
(def mousavi-votes (sel votes :cols "Mousavi"))
(def rezai-votes (sel votes :cols "Rezai"))
(def karrubi-votes (sel votes :cols "Karrubi"))
# -or-
attach(d)
(def ahmadinejad (map first-digit ahmadinejad-votes))
(def mousavi (map first-digit mousavi-votes))
(def rezai (map first-digit rezai-votes))
(def karrubi (map first-digit karrubi-votes))
# -or-
digits <- lastDigit( c(d$Ahmadinejad, d$Karroubi, d$Mousavi, d$Rezaee)) # not even attached -- could drop d$
etc. While it's obviously a matter of taste, it looks horridly verbose.