For this example (1 column data), you get much closer results using R's scan function rather than read.table
awk 'END{for(i=0;i<1000000;i++){ print int(rand() * 1000000) } }' </dev/null > data.txt
time datamash sstdev 1 < data.txt
288619.72189328
0.72s user 0.01s system 99% cpu 0.736 total
time R --vanilla --slave -e 'sd(scan("data.txt"))'
Read 1000000 items
[1] 288619.7
1.09s user 0.04s system 99% cpu 1.134 total
R read.table read performance is fairly slow by default because it has to infer the types of columns and check for inline comments, quotes ect.This seems like a better replacement for awk and bash one-liners to me than tasks I would use R for.
For instance counting unique elements.
#naive approach
time (sort data.txt | uniq | wc -l)
632209
13.09s user 0.04s system 101% cpu 12.984 total
#using hashing
time (awk '!a[$0]++' data.txt | wc -l)
632209
1.34s user 0.03s system 100% cpu 1.360 total
#R
time R --vanilla --slave -e 'length(unique(scan("data.txt")))'
Read 1000000 items
[1] 632209
1.20s user 0.04s system 99% cpu 1.244 total
#datamash
time datamash countunique 1 <data.txt
632209
0.83s user 0.01s system 99% cpu 0.840 total
Quite good performance in that case, although R surprised me here as well.