Data Science at the Command Line
datascienceatthecommandline.com
datascienceatthecommandline.com
All this attention (read: likes, shares, and page views) is making me wonder whether it's worthwhile to write an update (or even a second edition). What do you think? What would you like to see changed or added?
https://stackoverflow.com/questions/28888719/multi-threaded-...
At best you’re constantly restarting your kernel and clearing output. More likely, output from cell #7 has modified output [138] but you haven’t updated the chart produced in cell #17 (or some similar craziness). Not much better than programming with GOTOs.
“But they’re great for reporting and visualization!” you might say. If you’re building any report of value, though, it will influence important decisions. That’s the reason your code shouldn't live in a notebook. It should be in a library, covered by unit tests, so that those decisions aren’t based on faulty logic.
Dynamic data/reporting is a different thing entirely, at which point things like business intelligence software and dashboards come into play, and outside the scope of a command line anyways.
If reproducibility is important, like you say, then a notebook is the last thing you want. Your code needs to be tested and designed like the software it is, instead of tossed into some notebook that does not fit into classic software practices.
It's possible (actually very easy) to have code which works as you're making it but not if you run it from scratch.
In all my years of work with Notebooks, I've never had an issue with downstream dependencies.
This is no different than engineers who are, by convention, supposed to write bug-free code. Even with this convention, however, devs still rely on testing to decrease the likelihood of bugs.
No similar tooling exists for notebooks, which is why I recommend moving re-used logic into a tested codebase.
I was with you until the end. The important part is that a script is clear and commented, and can be written to fail informatively. These are benefits of using a scripting language.
Who cares if an IDE was used rather than a traditional editor?
Dynamic notebooks and the JSON mess they generate are a personal peeve of mine, and IME, the enemy of reproducibility.
If it’s for a one-off analysis, but processing is complicated enough that scripts don’t make dependencies between processing steps clear, use GNU Make.
If it’s a data product that’s running in the background, consider something like Airflow.
Still, I find coreutils and friends incredibly useful for interactively sizing up text data.
EDIT: Also I should note that notebooks are not the only (or in my opinion best) way to present a reproducible analysis.
So my scripts focus on the data, but around the code block are comments and notes about what is being done, etc.
I've recently been looking at how to integrate makefile into this since watching the following video: https://www.youtube.com/watch?v=sd0HhW8vkSQ
All that said, kudos to the author for this, reading through it I see a few things I hadn't even though of yet, and I live on the command line in general. If it werent for browsing the internet (eww in emacs is nice though) and gaming, I don't think I'd even need a desktop environment.
I switched to EXWM[1]; Firefox looks and acts like an Emacs buffer.
[0]: https://blog.dominodatalab.com/lesser-known-ways-of-using-no...
The thing is that people who are not professional programmers often don't realize that this is a danger, and their thing doesn't work and they don't know that it's because they're in a weird state emanating from the random execution of code blocks. So, I mean, notebooks are useful and cool, but they're definitely dangerous, especially for people who aren't software engineers, which is likely a huge fraction of their users
Cookiecutter: https://drivendata.github.io/cookiecutter-data-science
DataVersionControl: https://dataversioncontrol.com/
Drake: https://github.com/Factual/drake
Luigi: https://github.com/spotify/luigi
Pachyderm: http://www.pachyderm.io/
Sacred: https://github.com/IDSIA/sacred
They all focus in slightly different ways on the issue of managing data science/machine learning workflows, so I wonder if someone has a clear preference for one of those over any another.
EDIT: added Luigi
It is at version 1.0.3 though so it could be that it's considered finished. Seems strange to leave the issues open if it was though.
Indeed, for some of the tools I listed, they barely have any more functionality than I'd get out of Make & Git alone. And for Make, I'm pretty sure development & support will stick around for a few more years...
To me, only Luigi (Hadoop integration), Pachyderm (containers, production deployments in the Enterprise version) and Sacred (Python & TensorFlow integration) really stick out as differentiating themselves. But maybe I'm overlooking something?
Give it a try: http://bigbash.it