Sharing data? They had enough problems sharing their notebooks.
As someone who spends a substantial amount of time working with both modes (writing research code in Jupyter notebooks, and writing production code as python modules), notebooks scratch certain itches that IDEs typically don't even come close to. (Some recent progress on add-ons in Javascript-based editors is potentially interesting, because that might help marry the strengths of the two)
In my experience, in the evolution of code from Jupyter notebooks to repositories of production code as part of any project, there comes a "right time" to switch from the former to the latter. And this can typically only be learned with experience.
That is basically what I meant by knowing when to transition from one mode to the other.
Here's a concrete example (maybe somebody considers this an inspiring challenge?), to illustrate how notebooks are infuriating in their primitiveness, but still better compared to using an editor on source files: Imagine a beginner trying to write/learn a sorting algorithm, and who would like to keep experimenting with their code and observing what happens on examples, possibly profiling space/time complexity along the way.
To expand on my point above, there are actually three distinct computational use cases, not just two: Interactive learning -> Sharing insights with others -> Productionizing code.
Unknowable ad-hoc, unversioned spreadsheets running much of the capital of the company.
Notebooks tend to be the same way. It’s a simple GUI-ish was to do many complex analyses in a quick and dirty way.
And many of the arguments for not using Excel are the same as not using notebooks. Each is good at the initial data exploration stage, but are often abused and used in production when everyone knows it is a bad idea. But it still “works” so it is unlikely to be replaced.
(Especially when those that are working with the data don’t always have the skill set to build out a full production workflow.)
(For what it's worth, I feel the same way about people who try to send me RDS files with dataframes stored as R objects).
However, I think that whoever decided to name genes "OCT4" and "SEPT7" have to share some of the blame here too...
We just store the data tables in the project's database on a Postgres server. Then it's just a matter of pd.read_sql_query()
My background is programming (instead of data analysis & modeling) so I'm sympathetic to your idealistic "software engineering" view... but I'm also sympathetic to the academics' side as explained by Yihui Xie's blog post:
https://yihui.org/en/2018/09/notebook-war/
He's convinced me that criticizing non-programmers for using (or over-using) computational notebooks when it should be a "proper" programming language and deployment is like criticizing financial analysts over-using Excel to learn how to program VB or Python and re-write their spreadsheets into a "proper database" like Oracle or MySQL. That's just not reality. This divide between "end user tools" and "proper programmer tools" will always exist because there is no perfect tool in existence that serves the needs of both skill sets. Therefore, the programmers will always be able to say the data scientists or financial analysts are "doing it wrong".
I think this is very much off the mark. For sure plenty of scientists are poor programmers, but that isn't the reason they use notebooks. It is because:
They are not attempting to write something that will run everywhere, and often. They are either analyzing some data or doing rapid prototyping. For the latter, it's like criticizing someone who uses a REPL. It's just that the Notebook is much more powerful than a simple REPL that one can safely stick to it. Imagine you will do 40-50 prototypes and only one of those may end up worthy enough to make a product out of, and you don't know which one that will be. If you used a non-notebook environment, you'd give up in frustration by the time you hit the 15th one.
As you said: At the moment, there simply isn't an alternative that allows for rapid prototyping and is production ready. It's a hard problem to solve - there's a reason no one had solved it for decades (well before notebooks were a thing).
Had notebooks not been invented, you would have the same people handing you MATLAB code asking you to productize it.
Claiming they are beginners/novice programmers is off the mark. Peter Norvig started using notebooks for a reason, and no one would call him a novice. I do SW for a living, but when I need to analyze data and visualize it, I'll pick a notebook over "proper" SW tools any day.
I am very happy to be out of that business.
Have you looked into the domain of "research data management"? Concerns such as "archival", "security" or "share & collaborate" are core to this research domain:
In academics, there's a trend to prepare a "data management plan" up front that creates awareness about these concerns. They are even a requirement in order to get funding:
i.e. https://dmponline.dcc.ac.uk/
So, it's a bit odd to see a study that's focussed on a single technical tool yield the same concerns... but not making that jump to a larger, existing framework on information management.
Looking at the authors, it seems you are located at Oregon State University. A quick DuckDuckGo search yields this service from your colleagues at the University Library:
https://guides.library.oregonstate.edu/dmp
With the context of notebooks themselves, I think the study reflects on "if you have hammer, every problem looks like a nail." Notebooks aren't the only powerful tool to work with data. I think many of the same concerns could be raised with Google Sheets or Excel with heavy VBA scripting. Like others said, this is not a new problem.
Notebooks do have a place in the bigger process of doing iterative research based on data mining techniques. They can help to formulate more accurate questions and perform quick tests without the friction of having to set up complex environments. Moving on from initial data exploration, it's up to the researcher to use a formal method and tools that do mitigate those concerns. RDM is all about providing tools and mitigating (legal) liabilities as far as "what do you do with your data?" is concerned.
I'd also love code completion in notebooks.
I think the cleaning and code reuse problems can easily be mitigated by putting functions into libraries and using auto reload.
My normal workflow is hack something in a notebook until it runs, then refactor and put in a library I import with auto reload. I work on production ML and I use this for both software development and research.
It's not clear who the audience is. It sounds like most people who complain about them are software people and not researchers/scientists.
For someone like me, who once did computational research using MATLAB, and later analyzed data for my job, Jupyter is not worse, and is in most ways superior. Let's take your points one by one:
> Participants stated they often downloaded data outside of the notebook from various data sources since interfacing with them programmatically was too much hassle.
This was the norm with MATLAB, Excel and JMP as well, unless someone wrote code to autodownload (extremely rare - less than 1% of people did that). And if you are going to write code to get the data from somewhere, it's much nicer in Jupyter than in these other tools.
> Not only that, but notebooks often crash with large data sets (possibly due to the notebooks running in a web browser).
I honestly have not seen this, and the reason makes no sense. Your browser is not handling the data. The kernel is. I mean yes, if you try to load several GB of data in pandas, it's possible you will have problems if you run out of RAM, but this has nothing to do with notebooks.
> Once the data is loaded, it then has to be cleaned, which participants complained is a repetitive and time consuming task
This was as much a problem prior to notebooks as it is now. Notebooks did not make this any worse.
> Explore and analyze. Modeling and visualizing data are common tasks but can become frustrating. For example, we observed one participant tweak the parameters of a plot more than 20 times in less than 5 minutes.
It was even worse with MATLAB. Ditto for Excel. JMP is a bit nicer for visualization, though.
> Notebooks do not have all of the features of an IDE, like integrated documentation or sophisticated autocomplete, so participants often switch back and forth between an IDE (e.g., VS Code) and their notebook.
It may be better now, but this was a problem in MATLAB as well.
> While it is easy to share the notebook file, it is often not easy to share the data.
This is as true with MATLAB, JMP, etc. A lot of the complaints about it being hard to reuse notebooks is because notebooks at least attempt to be reproducible, and thus many more people attempt it. Prior to notebooks, I know almost no one who tried to share MATLAB analyses, because it was such a pain to do so.
> Notebooks as products. If a large data set is used, as one might expect in production, then the notebook will lose the interactivity while it is executing. Also, notebooks encourage "quick and dirty" code that may require rewriting before it is production quality.
I suppose some people are trying to make products out of notebooks, and this is where all the recent grief I see is coming from. I do not think it was the primary goal of notebooks, though. They were meant for data analyses and prototyping, not for production use.
Kudos for looking at real people's work and surveying it.
I wonder how much workflow could be improved if researchers would be temporarily paired with developers - who are generally better at modularising and removing friction in their work.
Personally I believe that a bit of clean-code discipline and following known best practices could solve couple of those pain points.
It's also true some could be improved by rethinking how notebooks work; ie. being able to specify input/output of notebook so it can be used as a library; detaching runtime data from the code so it plays better with version control/publishing; maybe even more radical ideas like adding visual/flow view that helps with linking elements; adding built-in excel-like sheets that can be queried/manipulated could also be interesting; built-in, first class support for relational database (sqlite) could also be a big win.
There are many interesting developments happening in this space and there seem to be some unexplored ideas waiting to be tested out.
I found your defense, which basically just says "It wasn't better before - so, no critique allowed?", considerably less valuable than the article.
This is true nowadays with Jupyter, because it is smart about truncating output. But it used to be possible to OOM the browser by e.g. printing in an long-running loop or displaying too long of a list/table.
But I guess if you're manually printing in a for-loop, then I suppose you could make the browser crash - whereas tools like MATLAB wouldn't.
People who design and implement features in notebooks. The conclusions in the blog post and research paper are clear that improving these identified problems could improve user experience.
We use it exclusively for our Data Science and find that it ameliorates all of the pain points you highlight in your article.
See Yestercode [1] and CodeDeviant [2], two tools that I specifically designed for LabVIEW programmers to refactor and test their code without expecting them to behave like traditional software engineers.
[1] http://web.eecs.utk.edu/~azh/pubs/Henley2016VLHCC_Yestercode...
[2] http://web.eecs.utk.edu/~azh/pubs/Henley2018VLHCC_CodeDevian...
What are your thoughts on the best way to address these things?
I don't get to use it in my current role, miss it a lot.
This is a very minor question (and I am not concerned about risk to participants)--when you say they signed consent "in accordance with our institutional ethics board", are you talking about Microsoft, one of the two universities, or all?