So we're trying to automate recording this metadata, but then give you that metadata in various ways for you to inspect it. One of those ways is actually spreadsheets: https://github.com/replicate/replicate/issues/289
Differences in tools used: (spreadsheets, flat files, logs, pen and paper, human memory). Forgetting to do it. Snippets to do it flying around. Different locations (laptop, group workstation, git repository, cloud sheet). Dissociated from the notebook that produced the model.
Tighter tracking should answer questions like: what notebook ran on which data and produced which model with which parameters and which scores? Then questions like: give me all notebooks that ran on this dataset which produced a model with scores that are [condition].
Once you do that, the "spreadsheet" can just be a "view" of the underlying data. Something you can export as, but not the thing itself.
I think it's good there are tools with this granularity that can be composed.
- [0]: https://iko.ai
Yes, and sorry because rereading my comment I wasn't clear, but this is what I meant. I view Spreadsheets as the view (but also an editor of the view), but I don't mean people should be working with XLS files (TSVs seem to work the best— I personally hate tabs but may have lost that battle). I do think though of Spreadsheets as the primary view, and so always design my data structures with the understanding of "how will this interoperate with spreadsheets". JSON is the pits.
IMO 2-D DSLs with Spreadsheets as the primary view/editor paradigm are the future.
I think Spreadsheets are 1, if not 2 OOM better than notebooks for doing actual work (notebooks can be good for presenting results in a narrative):
- non-linear for both humans and machines (allows for really creative and fast out-of-order parsing techniques on the machine side)
- concise signal with high information density
- unlimited cursors
- fantastically easier for version control and multi-player experiences
I maintain a list of all data science tools to try and stay on top of best practices, and when I see a tool is notebook first, I think "good, less work for me to track this one because they aren't getting the core things right yet." (Though often times notebooks will have really innovative orthogonal features).breck wrote:
>I would always save my experiment results so they were ready to analyze in spreadsheets (and other vis tools).
I should have elicited further before making assumptions. My assumption was that you were referring to results of training machine learning models. Is this correct?
If my assumption is correct, how do you go about training machine learning models?
Also, one question concerning your "scare" of "Throw away your spreadsheets"... Do you mean that you'd like a tool that exports results to spreadsheets or something in that direction?
I think we are addressing different problems, and a large part of that is due to assumptions I have made earlier.
I would set my hyper params in a spreadsheet, which would kick off training runs on a cluster, and report the results back in the spreadsheet hours/days later(really TSVs, but my UI was a spreadsheet). And repeat.
This was primitive stuff, and haven't done DL in a while, so not even sure if hyperparam tuning is still a thing or if that is all automated now.
Also, how did you deal with changing which hyperparameters you used or which algorithms you used? Did you make a spreadsheet per project per model?
Something like:
https://jtree.treenotation.org/designer/#grammar%0A%20inferr......