TensorFlow is designed to be trained on distributed systems, but deployed on embedded systems; in fact, to me, this is the single greatest advantage TensorFlow has currently.
Can you point me to "current memory utilization" numbers you're referring to?