Grid: AI platform from the makers of PyTorch Lightning
grid.ai
grid.ai
Luckily, that form is changing. There are interesting plans. But they are still just plans.
It’s better to go the other direction, I think. I ported pytorch to tensorflow: https://twitter.com/theshawwn/status/1311925180126511104?s=2...
Pytorch is mostly just an api. And that api is mostly python. When people say they “like pytorch”, they’re expressing a preference for how to organize ML code, not for the set of operations available to you when you use pytorch.
TPU support is VERY real... but yes, sometimes it breaks but PyTorch and Google are working very hard to bridge that gap.
But we have dedicated partners at Google on the TPU team working to get Lightning working seamlessly on pods.
Check out the discussions here: https://github.com/PyTorchLightning/pytorch-lightning/issues...
TPU support is real. Pytorch does in fact run on TPUs. But you don’t support TPU CPU memory, the staging area that you’re supposed to fill with training data. That staging area is why a TPU v3-512 pod can train an imagenet resnet classifier in 3 minutes at around 1M examples per second.
You will not get anywhere near that performance with pytorch on TPUs. In fact, you’re expected to create a separate VM for every 8 TPU cores. The VMs are in charge of feeding the cores. That’s insane; I’ve driven TPU pods from a single n1-standard-2 using tensorflow.
Repeat after me: if you are required to create more than one VM, you do not (yet!) support TPU pods. I wish I could triple underline this and put it in bold. People need to understand the limitations of this technique. Creating 256 VMs to feed a v3-2048 is not sustainable.
We’re not talking about a small 10% reduction in performance here. We’re talking like 40x differences.
If it seems unbelievable, and like it can’t possibly be true, well: now you understand my frustration here, and why I’m trying to break the myth.
Notice not a single benchmark has ever gone head to head in MLPerf using pytorch on TPUs. And that’s because using pytorch on TPUs requires you to feed each image manually to the TPU on demand, from your VM. Meaning the TPU is always infeed bound.
Engineers should be wincing at the sound of that. Especially anyone with graphics experience. Being infeed bound means you have lots of horsepower sitting around doing nothing. And that’s exactly the situation you’ll end up in with this technique.
There’s a way to settle this decisively: train a resnet classifier on imagenet, as quickly as possible. If you get anywhere near the MLPerf v0.6 benchmarks for tensorflow on TPUs, I will instantly pivot the other direction and sing the praises of pytorch on TPUs far and wide.
Picture a person with one arm and without legs. Would you say they aren’t “1:1 in terms of features”? They certainly won’t be winning any races.
And unlike real people, you can’t graft on a prosthetic limb to help this situation. The issue I’m describing here is a fundamental one that everyone keeps trying to sweep under the rug and pretend isn’t an issue. And then everyone wonders what’s going on.
We just need to be a part of the effort to help bridge the big gap and barriers keeping users from TPU adoption.
https://pytorch-lightning.readthedocs.io/en/latest/tpu.html#...
This was mentioned above, but nowhere on that page does it talk about any limitations whatsoever.
I have great respect for the PyTorch-TPU team, but I would recommend not heavily advertising PyTorch-TPU support until this major feature disparity is made up.
https://pytorch-lightning.readthedocs.io/en/latest/tpu.html#...
I just read the linked page and found no references to data loading limitations or performance limitations. Is it only in the video which isn't search indexed and few people would bother watching?
edit: The page literally advertises the speed of TPUs with "In general, a single TPU is about as fast as 5 V100 GPUs!" which is the exact opposite of warning people.
Maybe that's the way you feel, but for me that's very different. Pytorch is much more than just an API (which is also nothing to scoff at).
It's also much cleaner documentation, a very different ecosystem of libraries (mostly better, but sometimes lacking depending on the niche), less magic (which makes it easier to debug). It also has the benefit of less ecosystem churn, while the transitions of TF1->2 as well as the external Keras->internal Keras are a shitshow that's almost as bad as Python2->3.
So you can't blame the PyTorch team. If there's anyone to blame, it's Google Cloud. In the meantime, I don't think there's any harm advertising PyTorch with TPU support if running on TPUs with PyTorch is often much faster than running on GPUs with PyTorch.
The other thing is that stitching together other open source tools like this is simply not enough value. Who will be incentivised to buy?
Saying this as FAANG ML org person where I see the push to open source ops tooling like this.
Google certainly has made a push toward scalable AI training and deployment. However, it is not fun to use in practice, speaking from experience.
Startups beating an incumbent with substantially better UX is always a good story. Improving productivity is an easy winner for potential customers.
But calling Grid an ML Ops startup is like calling Lightning Keras... maybe at a quick blink it looks like that, but that's where the similarities end.
For what it's worth, a lot of what we're building comes from my experience at FAIR.
For us is basically integrating clouds directly into your code so the barrier disappears and the cloud providers become an extension of your laptop.
One of the professors at my lab at NYU CILVR (Kyle Cranmer) i believed was super involved with this. Will definitely sync up with him!
Thanks for the heads up!
It seems like there is an emerging consensus that (a) DL development requires access to massive compute, but (b) if you’re only using off-the-shelf PyTorch or TensorFlow, moving your model from your personal development environment to a cluster or cloud setting is too difficult — it is easy to spend most of your time managing infrastructure rather than developing models. At Determined AI, we’ve spent the last few years building an open source DL training platform that tries to make that process a lot simpler (https://github.com/determined-ai/determined), but I think it's fair to say that this is still very much an open space and an important problem. Curious to take a look at Grid AI and see how it compares to other tools in the space -- some other alternatives include Kubeflow, Polyaxon, and Spell AI.
Not a nice move for the opensource spirit. Also, pretty sure it's a violation of our patent and 100% copyright infringement.
In fact, our PyTorch API makes some significantly different design choices than Lightning does -- e.g., we require users to step optimizers and run the backward pass explicitly, which is a bit lower-level but allows for more flexibility when using the API.
For instance, here is an example of a GAN using our PyTorch API: https://github.com/determined-ai/determined/blob/master/exam...
This is a port of this PyTorch Lightning example: https://github.com/PyTorchLightning/pytorch-lightning/blob/m...
Despite the former being a port of the latter, there are significant differences between the two APIs.
More broadly, we welcome competition in this space and think there's a lot that we can all learn from one another.
If this is the philosophical stance that Grid and Lightning are taking then it's definitely a project I'm going to advice people to stay well clear off. It's the worst flavor of commercialized open source software and potentially a legal liability to touch in any way as you seem way too lawsuit trigger happy.
We want people to build on Lightning and we want companies to deliver value and products for their users.
We place no limitations on how Lightning is use or what products people will build.
Our in-process patent portfolio is not to limit the use of lightning in any way but for defensibility purposes only.
If you have any other inquiries please email legal@pytorchlightning.ai for more details.
As I said in my previous comment, using patents to try and get around an open source license is skeevy as hell.
We want people to build on Lightning and we want companies to deliver value and products for their users.
We place no limitations on how Lightning is use or what products people will build.
Our in-process patent portfolio is not to limit the use of lightning in any way but for defensibility purposes only.
If you have any other inquiries please email legal@pytorchlightning.ai for more details.
IMO, open source is at its best when it's supported by a SaaS as it provides a strong incentive to keep the project up-to-date, and the devs of PL have been very proactive.
Not just a marginal improvement on that experience but a 10x completely different approach.
I wish you luck though. We’ll see in ten years whether programmers are as effective as researchers, or whether researchers are as effective as programmers. In 60 some years of computing, no one has achieved the latter, despite many attempts.
Also, I was surprised that this webpage is basically a waitlist and nothing else. No discussion of technique, no docs, no substance. Just a “you like pytorch? Pytorch rules!” type hype.
I do like pytorch, but I also like knowing one or two substantive points about what the proposal here is. If you want to train a model from your laptop, it’s a matter of applying to TFRC and kicking off a TPU.
The whole ecosystem is in need of massive overhaul. I like the ambition. But I dislike trying to pretend we aren’t programmers. ML is programming, and pretending otherwise will always cause massive, avoidable delays.
I forgot who said it: “In engineering, if you don’t know what you’re doing you shouldn’t be doing it. In science, if you know what you’re doing you shouldn’t be doing it.“
I meant this. Academia-related BS is just a one way to do that.
There's also a ton of bad research being published in second rate conferences, but I don't consider that "research". You won't get "simple extensions of what has been done before" published in a top level conference, the acceptance rates have been extremely low recently.
Easy onboarding for ML tooling is very valuable for the industry as a whole.
100% agree with you that going the other way is likely not the best approach.
Lightning + Grid elevates and turns non experts closer to researchers... ie: focus on building the products and doing science and not the engineering.
That's what lightning excels at today. That's the experience Grid will 10x.
I know the same could be said about Azure and AWS, but the big name cloud providers stake their prestige on having tight security, while a startup has much less to lose.
Another major concern is GDPR - can you reliably track and delete data related to a specific person? You got to think about this right from the start to keep the necessary metadata in the tagging pipeline.
Second, electricity was a great new technology (ie: AI), but you needed the power grid to make it usable - that's grid AI.
What do you think folks?
Also, is there support for distributed training for large datasets that don't fit into single instance memory? or just distributed grid-search/hyper-parameter optimization?
Just that if you use lightning you'll have zero friction. Well as with the others... you might run into issues inherent in the other framework's hard to work with designs.
also, this was probably the first (and maybe still is?) high-level pytorch library that let you train on tpus without a lot of refactoring and bugs which was a really nice thing to be able to do given how the pytorch-xla api was still unstable at that point. <3