Wide and Deep Learning: Better Together with TensorFlow
research.googleblog.com
research.googleblog.com
I don't know of any specific cases of this being down with a raspberry pi, but many phone apps, for example, have this sort of architecture. Train a model on a powerful server/cluster, send model to phones for use, phones collect more training data that is sent to server/cluster, repeat.
You can certainly send data from your raspberry pi to a central server, and then periodically retrain and update the model on your raspberry pi through whatever update mechanism you create, but that will require your own infrastructure.
Some of the models can be quite large for say, 3G (225mb), but I'm working on various compression techniques now too.
I wouldn't be surprised if a simple ensemble performs better!
Another advantage of joint learning (which the authors mentioned) is that the individual models need not be as big when trained independently since they complement each other. Though the joint model will surely be bigger than each of the individual models.
It strikes me that the example in the blog post is just a general search problem, eg google search could use this: if you type "Brexit" and you want a general overview of a lot of different things, vs you type a specific query and are looking for a specific page.