julia> import CSV (takes ~1-2 seconds)
julia> data = CSV.File("/User/orbifold/Downloads/FL_insurance_sample.csv) (takes seconds again...)
I want to explore the data. I vaguely remember there is Gadfly to do that. For some reason it depends on FFTW and a whole bunch of other packages, but the dependency story in Python is equally insane, so whatever. As I'm typing this, I'm waiting for Gadfly to precompile.
julia> @time import Gadfly (104.460307 seconds)
python> import matplotlib.pyplot as plt (~10s if you are unlucky and it has to generate the font cache)
Maybe it's a good idea to do a scatter plot of two of the fields? Let's find out
julia> @time Gadfly.plot(x=data.eq_site_limit, y=data.hu_site_limit) 0.320751 seconds
but that number is not accurate, in fact it took so long for the browser to open the result, that I had time to check the documentation on backends and see whether I'm missing something.
python> %time plt.scatter(data['eq_site_limit'], data['hu_site_limit']) 41.9 ms, a bit more for the window to open
Turns out that is not a useful plot... maybe a histogram?
julia> @time plot(x=data.eq_site_limit, Geom.histogram(bincount=10)) 0.042802 seconds
It in fact again takes several seconds until the browser shows a plot.
python> %time plt.hist(data["eq_site_limit"]) ~100ms
and the window with the plot opens instantly.