It looks like this is for running already-trained networks, and I think that's really the only practical way to do things right now if you want to make sense of something like video in real time because of bandwidth constraints. It looks like training would happen in 'the cloud' or similar.
Local is also useful in situations with limited or no internet access. Say you're trying to do deep recognition on a live video feed: many places this would be useful simply do not have the bandwidth available to stream video.
Also useful in latency-sensitive applications, such as the drone flight demo they use in their video.
Pro: Reliability/longevity (your product still works even six months later when the cloud API provider has been acqui-hired.)