Is any real science done with this or is it the Powerpoint for PyCon talks?
Is any real science done with this or is it the Powerpoint for PyCon talks?
Say you have a large file you want to read into memory. That’s step 1, it takes a long time to parse that big json file they sent you. Then you want to check something about that data, maybe sum one of the columns. That’s step 2. Then you realize you want to average another column. Step 3.
If you write a basic python script, you have to run step 1 and 2 sequentially, then once you realize you want step 3 as well, you need to run 1, 2 and 3 sequentially. It quickly becomes much more convenient to have the file in memory.
Reactive notebooks are nice but I’ve accidentally reran slow SQL queries because I updated an upstream cell and that’s painful (using Pluto, ObservableHQ, and HexComputing).
In practice I never see or create notebooks that don’t run when you push the run-all button, it’s a well understood and easily avoidable issue. It’s probably a local optimum but I’m happy with them.
This led a lot of programming environments to where batch loading of the code is basically required. But "image based" workflows are a very old concept and work great with practice. Some older languages were based on this idea. (Smalltalk being the main one that pushed this way. Common Lisp also has good support for interacting with the running system.)
It is a shame, as many folks assume everything has to be "repl" driven, when that is only a part of what made image based workflows work.
Interactive environment without compile nonsense is just too new for folks.
#+BEGIN_SRC python :results file
import matplotlib
matplotlib.use('Agg')
import matplotlib.pyplot as plt
fn = 'my_fig.png'
plt.plot([1, 2, 3, 2.5, 2.8])
plt.savefig('my_fig.png', dpi=50)
return fn
#+END_SRC
#+RESULTS:
[[file:my_fig.png]] import matplotlib
matplotlib.use('Agg')
import matplotlib.pyplot as plt
plt.plot([1, 2, 3, 2.5, 2.8])
Alright, saving the figure at 50 dpi first
plt.savefig('my_fig.png', dpi=50)
Trying a bit more DPI to see if that makes a difference
plt.savefig('my_fig2.png', dpi=150)
Oh, wrong numbers, forgot that the fourth datapoint was going to signify 100, going back to 50 dpi as well
plt.plot([1, 2, 3, 100, 2.3])
plt.savefig('my_fig4.png', dpi=50)
It seems like your example misses the interactivity.R.
No thank you.
Besides the ecosystem, what makes R better than Python or Julia for biostats?
The statistics built in are great. They're just there, less need to find a package (general stats, ttest, chi_squared test...). We tend to use the "tidyverse" packages [1] https://r4ds.hadley.nz/. Bio-python is amazing for manipulating biodata, but once the data is extracted and you need statistics, our scientist seem to use R. I really don't love R's syntax, but I get why they use it. I use python all the time for data wrangling (right now I'm pulling sequences from a fasta file to inject into a table).
Rstudio is like an IDE for your data. You can view the data tables, graph different things etc. If you try the first chapter of the R4data Science book, you can see how get up and graphing and analyzing quite quickly. https://r4ds.hadley.nz/data-visualize.html
Though at this point Python and R are necessary depending on what package/ algorithm you want to use.
There are some good packages for single cell analysis: We use "Seurat".
https://satijalab.org/seurat/articles/get_started_v5.html
Jupyter supports R now with an add in, so its less of an issue.
Nowadays they can even be used for running models/analytics in prod with tools like Sagemaker (though I’m not advocating that they should).
Maybe you’re mistaking Jupyter for a different tool like quarto or nbconvert but your dismissive comment misses the mark by miles.
Advantages...no setup on the students side (+they get reliable compute remotely) and we can prepare notebooks highlighting certain concepts. Text cells are usefull for explaining stuff so they can work through some notebooks by themselves. Students can also easily share notebooks with us if they have any questions/issues.
I also use notebooks for data exploration, training initial test models etc. etc. Very useful. I'd say >50% of my ML related work gets done in notebooks.
I personally prefer when people share code as notebooks because you have code alongside the results. It’s really a good practice to use Jupyter.
The models run on a small cluster and/or a supercomputer, and the data reductions of these model runs are done in python code that dumps files of metrics (kind of a GBs -> MBs reduction process). The notebook is at the very tail end of the pipeline, allowing me to make ad hoc graphics to interpret the results.
https://income-inequality.info/
All the processing is documented with Jupyter notebooks, allowing anyone to spot mistakes, or replicate the visualizations with newer data in the future:
That is why jupyter lab is not the wrong name, it is a bit like a lab. Not meant for production use, but very good for exploring solutions.
Then I can easily put them into a Python file and import them from the notebook.
Easy peasy and very nice for iterative development.
People use it for two reasons: a) because they need to get those graphs on the screen and this is the only way b) running ML code on a remote, beefier server.
The real reason is because it’s a much better workflow for data exploration and manipulation because you don’t always know exactly what code to write before you do it. So having the data in memory is really useful.
Or you just use jupyterlab and the problem is fixed.
As it is now, you typically wind up “programizing” your notebook once it does what it should so you can run in batch and so on.
Do you have a source of this, or is it something you dreamed up? Weird claim as none of those are my use case.
Jupyter notebook is neither the first nor the only implementation of such literate approach.
If some code is stable enough for reuse, you can make it composable as any other code: put it into the module/create CLI/web API/etc -- whatever is more appropriate in your case.