Cool idea, but is this of any benefit?
Isn't this essentially what TensorFlow does internally, except it inserts CUDA primitives at the right positions...
Isn't this essentially what TensorFlow does internally, except it inserts CUDA primitives at the right positions...
"We also have a number of concrete directions to improve the performance of TensorFlow. One such direction is our initial work on a just-in-time compiler that can take a subgraph of a TensorFlow execution, perhaps with some runtime profiling information about the typical sizes and shapes of tensors, and can generate an optimized routine for this subgraph. This compiler will understand the semantics of perform a number of optimizations such as loop fusion, blocking and tiling for locality, specialization for particular shapes and sizes, etc."