Well, what the CNN is doing at each stage is very simple to understand. There is a forward pass which is a matrix multiply (and addition of there is a bias) and then the matrix weights are learned in the backward pass, which is just basic differentiation and chain rule application. Now I am not trivializing differentiation ( when you try take the differntial of a vector, you are tearing your hair apart) but it's fundamentally a simple concept to understand.
Even with this understanding, designing deep neural nets and tuning hyper-parameters is mostly guesswork. Yes, the frameworks have little or nothing to do with this.
What I've found is that TensorFlow is difficult to wrap your head around for programmers because it's more like a DSL. You declare a computational graph and then run it multiple times. So when you are declaring a computational graph, you have no way of debugging that graph, unless you run it. Also the conversion from numpy arrays to Tensors and back is an expensive operation. PyTorch simplifies this to a great extent. You just create graphs and run like how you run a loop and declare variables in the loop. This is great for imperative programming. However, think about it - every time your graph is recreated. Now if it's just a variable re-initialization it's not a big deal, but we are dealing with Tensors so you give up efficiency for flexibility.
Again, all of this is immaterial for learning how to build deep neural nets. I would say, just stick to whatever framework you can wrap your head around. I am learning that my ability to tweak numpy arrays, visualize them in pyplot, load data from csvs using Pandas and the like will take me a lot further in learning deep learning.