Shareable Jupyter Notebooks That Run on Free Cloud GPUs
blog.paperspace.com
blog.paperspace.com
Giving anyone access to free GPUs and powerful tooling seems like an incredible opportunity. I'd love to hear what you all think!
Am wondering what to use for the data tier? Is there a dedicated backing cloud store supporting BigQuery style syntax? Import form S3 datasets?
Checkout the docs here -> https://docs.paperspace.com/gradient/data/storage
And happy to answer any other questions
Edit: we also have another tool called GradientCI (https://docs.paperspace.com/gradient/projects/gradientci) that might also be of interest. Basically it lets you connect a GitHub repo directly to a project and you can use it to build your container automatically.
Side question, if you're willing to entertain it. Tired me tells a notebook to checkpoint, wanders off to bed, comes back next day and wonders why its taking an aeon to open the notebook..... oh yeah, damn, I've got gigs of images and panda crap in it . -- do you wrangle this problem, i.e. please don't save your notebooks as gihugic files representing a mental state you can't possibly remember ?
As for checkpointing data, that is still a relatively difficult problem to solve and our current recommendation is to use a combination of the persitent /storage directory and the notebook home directory. There are definitely issues with doing 100K+ of small files and committing those to the primary docker layer.
When you get to testing it out don't hesitate to reach out to use and we can try to see what the best solution is for your particular project. To date there isn't a "one size fits all" solution but we are working hard on making more intelligent choices behind the scenes to unblock some of these IO constraints.
I used to also keep a beefy EC2 or GCP instance all set up with Jupyter, custom kernels (Common Lisp, etc.), etc., etc. Being able to start and stop them quickly made it really cost effective.
I ended up, though, getting a System76 GPU laptop a year ago and it is so nice to be able to experiment more quickly, especially noticable improvement for quick experiments.
That said, a custom EC2 or GCP instance all set up, and easily start/stop-able is probably the cost effective thing to do unless just vanilla colab environments do it for you.
Do you guys have a trademark for the name Gradient? I am developing a C#/.NET binding to TensorFlow with the same name: https://losttech.software/gradient.html
If you do, do you mind that clash of names?
Secondly, would you be interested to invest in providing .NET-powered notebooks? I just got it working on Azure Notebooks for F# ( http://ml.blogs.losttech.software/What-New-In-Preview-6.4/ ), but I feel like there would be more interest from C# developers. There is a good C# Jupyter kernel out there.
Other similar services include Kaggle (way more locked down) and Baidu's free GPU powered notebook platform (only open to Chinese citizens)
But as a cloud provider, I would be worried about abuse by cryptojackers. I hope that's not a problem and this is sustainable.
Google Colab is decent and free.
But you can do computing on cheaper hardware too. The CPU is good enough for learning and a GPU is not outside the budget of many people that presumably already own laptops and such.
I know a local group that shares an i9 / dual GTX "server" and are learning on this shared hardware. I think it's great!
I had a small budget for this learning curiosity and bought a Ryzen and a GTX with only 4GB of RAM. Got a job offer after a while which I had to turn down as it seemed to actually kill my interest in the field. Doing some small personal project now without much fuss to rekindle the fire. And using the CPU for it since it's so small.
Do you think there's a format which is better suited to sharing data stories?
You are right that versioning is still an issue and we largely punt on it by using the docker container (with layer commits on each notebook teardown) as the versioning mechanism. Maybe not the best solution but it does have it's advantages.
I agree with that problem, yet somehow like the idea of notebooks. Maybe the real tragedy here is that jupyter notebooks are saved as json and not as a valid program with comments, that can be run "as is" from the command line.
You like the tight feedback loop of REPLs, as do I. You don't need the clunky machinery of jupyter to effect REPLs with emacs and ipython.
> real tragedy here is that jupyter notebooks are not [saved] as a > valid program with comments
Here's a simple way of doing that: write your code as valid programs with comments.
That's what I do! Then I have a script that converts my python program to shitty json that my colleague--who can only conceive to work inside a notebook--can run it. Finally, another script translates back the json to a readable code; and more importantly, to something that can meaningfully be put into git.
I would love if the jupyter interface allowed to save the notebook directly into a program with comments. Then all this silly sorcery would not be necessary.
But that's besides the point. The jupyter ecosystem, like other widespread and unwieldy formats such as Microsoft Word and PDF, poses a brutal obstruction to Unix workflows.
Yes, my script simply calls the nbconvert library, with some trickery to ensure that the result is idempotent. But I would like not to need this script, that instead the jupyter interface worked with valid python files directly (maybe after enabling some option).
> poses a brutal obstruction to Unix workflows.
It's not as much the notebook itself, but the file format chosen by default by the notebook. If it was a human-editable textfile there would be no problem, and there is no practical obstruction for that (other that young programmers today cannot conceive a different "serialization" format than json).
Data is the second problem - you usually can't run a notebook without it. If the notebook transforms data over time, it can be a real issue.
Here's a pretty complete run down of the first problem, the .ipynb itself: https://nextjournal.com/schmudde/how-to-version-control-jupy...
An integrated solution that versions data, the notebook, and the computational environment (collaborators will make changes over time) is the ideal collaborative platform.
Disclaimer: I use Colab regularly as part of my job in Google.