How to separate your data from your code
negfeedback.blogspot.com
negfeedback.blogspot.com
Sorry, not buying.
Also, the whole post could have been written as "Don't hardcode filepaths, use a configuration file." but again, that's not what configuration files are for.. If you want to run a program on a specific set of files, you're better off wrapping the invocation in a script.
But I'm also a stupid data scientist.
What do you propose as the long-solved solution?
You generally don't open a notebook from the commandline, you open it from the notebook file picker inside your browser.
You generally don't "run" a notebook, like you would a script. It is an interactive programming environment.
It's more like someone who started by recommending my students use parametrised scripts, noticing that they will not and working through alternatives.
My students are relatively familiar with the notebook interface, but not at all comfortable with writing a script which will accept command line arguments or read from standard in.
PS: since I mention .gitignore explicitly in the post, why would you thing I don't know about .gitignore?
I work with producing some analysis, and the writeup is absolutely tied to the data file used.
I want to ensure these are available and tied to the code & output. However, they're too big for git.
There are better solutions though. I've been using dvc.org recently.
For the cases when you start with huge dataset, and want to generate dozens graphs and tables on the screen, each changeable interactively without data reload. And leave python interpreter running, so you can run more code if you have more questions. Oh, and do all the hard work on the cloud machine with hundred of gigs of RAM, while displaying the data on your laptop over slow internet link.
I have tried to implement those things by hand (intermediate state, HTML generation, shotgun plotting, pdb interaction), and it is a major pain.
Notebooks work great in this situation. There are some downsides -- like inability to ask filenames from user -- but those could be easily bypassed, as this article shows (my solution will be a user-created symlink btw, not an ini file)
For the situation I show as an example, you wouldn't want to bring up a dialog box every time you run the cell as in most cases you're running it many times to get things right. Even when the code is working perfectly it isn't useful to pop up a dialog during a long-running analysis.
The target audience of this post are not comfortable with running scripts from the commandline. I have tried teaching that approach and failed. This is a half-way point.
#!/usr/bin/env python
# process-csv.py
import pandas
def expensive_operation(data):
# TODO
return data
input_data = pandas.read_csv(sys.stdin)
results_data = expensive_operation(input_data)
results_data.to_csv(sys.stdout)
Or maybe even shorter (I think this is valid Python?): #!/usr/bin/env python
import pandas
def expensive_operation(data):
# TODO
return data
expensive_operation(pandas.read_csv(sys.stdin)).to_csv(sys.stdout)
And then use it like so: $ python process-csv.py <input.csv >results.csv
Or like so: $ csv-generating-command | python process-csv.py | csv-consuming-command
You could even do a quick $ chmod +x process-csv.py
And then call it directly $ ./process-csv.py <in.csv >out.csv
Or copy it to somewhere in your $PATH $ cp process-csv.py ~/bin/process-csv
At which point you can run from anywhere $ cd /literally/anywhere/else/
$ sudo pip install csvcat # or something
$ csvcat *.csv | process-csv >~/Documents/results-201904021616.csv
Of course, we could get even crazier with a custom Pip package or whatever the actual terminology is (egg?) and do all the setup.py and requirements.txt and whatnot that Python packaging entails and yadda yadda yadda, but that's probably overkill for a one-off script. Point is: your fancy Macbook has a full-blown actual UNIX™ on it, so might as well put it to good use :)And of course, an stdio-based approach works perfectly fine for Dropbox:
$ cd ~/Dropbox/csv/data/folder && process-csv <in.csv >out.csv
Or for S3: $ s3fs my-bucket ~/mnt/my-bucket # macOS has FUSE support, right?
$ cd ~/mnt/my-bucket/csv/data/folder
$ process-csv <in.csv >out.csv
Or for pretty much any other filesharing service that provides a folder on the filesystem, really.I am a huge fan of this approach and use it extensively in my own code. This natural commandline interaction is completely alien to my students. They are barely holding on with trying to learn Python in the notebook environment, telling them that the solution to their problem is to learn another language has not woked well for me.
In that case, I'd probably more readily recommend just adding a *.csv line to .gitignore and letting students add/remove files accordingly (unless the lesson is specifically about how configuration files work, of course, in which case by all means your approach is reasonable).
I am aiming to minimise the manual steps involved in running the latest version of the code on the latest version of the input data. With my system you just pull the latest code and run it.
Yet another way to skin this particular cat might be to write a script as a drag-and-drop target. I don't think Jupyter supports this, so your students would have to venture beyond it a bit, but your Windows-using students for example can drag files onto actual .py files and those dropped files will show up in sys.argv (I think macOS requires the extra step of using py2app, but that might be a good opportunity for a py2app/py2exe lesson). I'd post sample code like in my other comments, but I'm on my phone at the moment.
Can also be said, Data and logic are strictly separated, Element level separation of data and logic, data stream processing.
The sea sails by the helmsman and the programming moves toward the data. Initial state, final state, the shortest linear distance between two points. Simplicity is the root of fast, stable and reliable.
https://github.com/linpengcheng/PurefunctionPipelineDataflow
So, I can only comment by the title.