1. The pace of published research is pretty fast right now. This makes it difficult to know where the research fits in when solving problems. It'll probably take a few years before we know where to use many of the approaches published last year.
2. Iteration performance (trying many new things quickly) is improving with high-level frameworks but still lots of work to do here. Since we don't know where a lot of research fits it's not always apparent which methods work best (and "best" changes every quarter).
3. We're still missing theoretical foundations for much of deep learning. This is useful, not just for research, but to know what can and cannot work with current approaches.
4. Model architectures are still based largely on trial-and-error, intuition and search.
Some hard long term challenges revolve around cases where you don't have a lot of unlabeled data or examples of classes you care about. There's also the technical challenges of training large scale models without obscene computational resources.
I also prefer to reserve the term AI for generalized AI, which we're still a ways off of, as opposed to modern classification problems etc. that I would call machine learning (though I know that nomenclature is uncommon).
EDIT: jimfleming makes a great point about theory as well - we could likely be much more efficient with better theory for deep neural nets.
That said there is a bigger problem with those bounds because it doesn't incorporate model complexity. The VC dimension is much more insightful because the complexity of your model and the hypothesis space it represents is important for proper training. As an example, add a regularization term to your model and you're no longer doing anything like N/e^D. Convolutions, dropout, etc all prevent NN models from becoming too complex to train.
People often try to intuit the intrinsic dimensionality of a dataset by using techniques like looking at singular values above some threshold or reconstruction error versus changing the output dimensionality of a dimensionality reduction/unsupervised technique like PCA, matrix factorization, or an autoencoder.
An info theory person might argue entropy and compression ratios are also insightful.