278 karma · joined July 27, 2012
It would be amazing if the world of open-source column stores matured a little bit relative to where we are today...
In terms of radically different takes on workflow engines, I'm very interested in reflow. I haven't used it enough to know if the rough edges are a deal breaker.
So I'll give my opinion as a datascientist: null hypothesis significance testing is broken, the whole thing needs to get chucked. It's not fixable, and it's not only p-hacking that's the problem. Go read Frank Harrell, I like his writings on the topic more than Gelman's.
One point of comparison is Cam Davidson Pilon's Bayesian Methods for Hackers, they have a similar vibe: practical applied advice from a field that tilts towards the academic...
Naive bayes can be parallelized in ways that SGD can't, that's a whole other conversation.
The closest IME has been ipywidgets+jupyter and the jupyter dashboards server. Jupyter is really nice for developing, but deploying is another question. The dashboards server works but development on it has been stalled, and deploying it is a pain. Also, if dash can leverage the react ecosystem, that could make it pretty compelling compared to ipywidgets.
FWIW I also found YAML to be very confusing syntactically at first; things got easier once I realized that it's basically 1-to-1 with JSON, and could convert to JSON to get intuition for the file structure (simple yaml2json script here: http://pastebin.com/TpjZLnLa). Thumbing through the book Ansible Up and Running also helped.
In the long run, putting a declarative idempotent layer atop the same old mutable infra is tough but a necessary compromise right now. It'll be great if the immutable-first tools of today (nixos et al) mature and we can leave this behind.
This makes our CI builds painfully slow. It's not entirely CircleCI or Travis's fault, it's the interaction between them and docker.
I also find that multiple dispatch is something that's occasionally handy, but single dispatch is truly everywhere. I think if you broke down multipledispatch uses in languages which embrace it (e.g. CLOS), the mix is like 95% single, 5% >single. For single dispatch in python, check out https://docs.python.org/3/library/functools.html#functools.s... which was added in to the Python 3.4 stdlib. The alternative is the OOP subclassing style, which works but IMHO isn't very readable.
Personal accounting is about data capture; if you don't gather good data, you end up with garbage in garbage out analysis. Having a simple, non-propretary file format enables so many things. The data capture side (OFX/QIF/CSV files, bank websites, etc. etc.) is still clunky and painful but that's a hard problem to solve.
It works great, but it really bugs me that we had to do that. The default download speeds from our buckets on S3 are atrocious. We store big datafiles in S3, and our development flow involves downloading them lot. If Amazon had an upgrade to S3 so downloads by chosen users weren't throttled or slow, we'd pay for it in a heartbeat.
That's my preferred design too. The first job of any external data capture process is to capture the full fidelity source data. Everything else belongs in a followon job.
The schema-based serialization libraries (Thrift, Protobufs) or MsgPack are a good way to avoid that too. They come with a lot less baggage than say HDF5. Also, if efficiency isn't paramount -- just use SQLite! Amazing tool when it's in its sweetspot.
Lots of tradeoffs when dealing w/ serialization and file formats, no easy answers.