What is so appealing about a delayed execution model? Why can't we just perform tensor math as in numpy, and let the library figure out the fastest way to do it behind the scenes? I think the whole "graph" approach is making things needlessly complicated.