JupyterHub 1.0
blog.jupyter.org
blog.jupyter.org
[1] https://github.com/jupyter/nbgrader/issues/530#issuecomment-...
[1] https://jetstream-cloud.org/
[2] https://zonca.github.io/2018/09/kubernetes-jetstream-kubespr...
[3] https://zonca.github.io/2018/09/kubernetes-jetstream-kubespr...
[4] https://zonca.github.io/2018/09/kubernetes-jetstream-kubespr...
[5] https://github.com/Unidata/xsede-jetstream/blob/master/vms/j...
I'd wager that almost no data scientists write object oriented code.. it's probably mostly done one calculation at a time. executed in the notebooks repl. So the value you get from ide debuggers is tiny, as you're already doing everything one step at a time.
I ended up getting really frustrated with setting it up. Followed several different tutorials, had it blow up in a different way each time.
Going to have to revisit this to see if the documentation has gotten better.
User goes with the web browser to our jupyterhub URL, logs in with our usual credentials, selects a job type (amount of memory and max duration), and jupyterhub takes care of launching a jupyter kernel as a slurm batch job on a compute node in the cluster, and proxies http I/O via the jupyterhub node to the user web browser. In the jupyter notebook, users have access to the same cluster filesystems as if she would login traditionally via ssh.
Then there's the PITA of integration with the site auth system, but that tends to be site specific..
I was trying to read about whether jupyterhub is included in RStudio Connect, or if they are competing products.
1. Anaconda sells support that many companies will value, https://www.anaconda.com/support/
2. Anaconda checks the installed versions of packages in the distribution are compatible.
3. Conda has an “environments” feature so a developer/scientist can switch between many, project-specific development environments, https://docs.anaconda.com/ae-notebooks/user-guide/adv-tasks/...
Edit: Also, there is a distribution for Windows that’s handy when your employer has you using Windows. And, depending on your software approval process, it can be convenient to get one package approved (Anaconda) instead of every package Anaconda includes.
There are certain packages which aren't available on pip that conda-forge provides e.g. for a while OpenCV wasn't available through pypi.
Remember you can use pip within conda, and you can install directly (e.g. setup.py) within a conda env.
Docker is mostly useful if you want to mix and match different versions of CUDA/cudnn etc. If you just want to run an isolated Python environment, then Conda will do the job.
For Python :
$ ipykernel install --user --name myenv --display-name "My environment"
For R : > IRkernel::installspec()
As mentioned previously, nb_conda_kernels allows to automate this step.EDIT: I meant `nb_conda_kernels` https://github.com/Anaconda-Platform/nb_conda_kernels
JupyterHub is more flexible - for example, you could deploy JupyterHub to one beefy server and have Jupyter deployed for many users, which could all read data from a shared filesystem. that kind of thing is not easy to do with SageMaker since everything runs on a separate ec2 instance.
I can't comment on EMR notebooks.