1. Meh. PyTorch is close enough to not worry about it, and is better in some places.
2. Meh. All the methods people use in practice for deep learning in particular do not use higher order gradients. Most higher order methods are prohibitively memory expensive, and memory is at a premium in acceleration hardware (and so is the bus bandwidth - so you can't "swap to RAM"). I do agree that higher order gradients are the next frontier in optimization though - current optimizer research seems to have stalled, so people focus on training with huge batches and stuff like that. Most SOTA models in my field are trained with SGD+momentum - super primitive stuff. I don't see how Jax would solve the memory problem though. You still have to store those Hessians somewhere, at least partially.
3. Do agree, that's cool if it actually parallelizes nontrivial stuff which e.g. tf.vectorized_map barfs on. Although in a lot of cases you can "vectorize" by concatenating input tensors into a higher dimensional tensor.
4. Meh. Not sure why I'd want that if I already have tracing and JIT.
5. This is #2
With PyTorch though, you get close enough to Numpy to feel at home in both, and there's so much code written for it already that you can usually find a good starting point for your research pretty easily on github and then build on top of that.
If you need to deploy, there's also tracing and jit, which lets you load and serve models with libtorch.
I see what you're saying regarding "advantages", I'm just pointing out that PyTorch might be "good enough" for most people. If I were on that team, I'd focus on providing comfortable transition from TF 2.x which is a dumpster fire (with the exception of TensorFlow Lite which is excellent). That, IMO, would be the only way for this project to achieve mainstream success unless PyTorch disintegrates over time.