1. Much of the scientific computing/stats community is stuck in the past. Many are still using Matlab! As opposed to the CS community, who are used to learning new frameworks, offering the ability to "import jax.numpy as np" and having their scripts just run is valuable to that community. As is having an API that they've only just started to become familiar with (and has way more documentation about).
2. Once again, this is true for practitioners, but not research. Hessian vector products show up in a decent amount of places. For example, if you have an inner optimization loop (a la. most meta learning approaches or Deep Set Prediction Networks) you have a Hessian Vector Product! Perhaps not prevalent in models that practitioners run but definitely something to keep an eye on in research.
3. My understanding is that it actually does a pretty decent job. Enough that it's useful in the prototyping phase.
4. PyTorch JIT is neat, and is what I meant by Pytorch team is "working" on it. However, the JIT doesn't do full code gen (thus, significant operator overhead for say, scalar networks) and has significantly less man hours poured into compared to XLA.
5. I was specifically talking about how you call grad on a function to get a function that returns its gradient. It's a cleaner API than PyTorch's autograd.
Jax is definitely not meant for deployment or industry usage, and I believe their developers hope they'll never be pushed along that direction :^)
I definitely agree that PyTorch is "good enough" for most people. However, among researchers, there's a decent amount of subgroups it could gain favor in.
You'd be surprised how many papers get submitted to ICML/Neurips that don't use PyTorch or TensorFlow at all, in favor of raw numpy, C++, or even MatLab! I think the numbers I had were something about 30% of papers don't use any ML framework. Jax could easily gain favor in this crowd.
There's also the crowd that cares a lot about higher order gradients. Also, admittedly a specific subgroup, but growing. Meta learning people care a lot. So do Neural ODE people. All it takes is for one of these subfields to blow up for higher order gradients to all of a sudden become a lot more appealing.
And finally, you have Google. Google researchers are never going to use PyTorch en masse (probably). If researchers at Google want to switch from TF, their only option is Jax. This is a pretty big subgroup of researchers :)
I definitely agree that Jax has a difficult hill to climb. But, they have a solid foothold within Google, and several subfields very amenable to their advantages.
PyTorch seems like the predominant research framework currently, but if any framework is going to erode their lead, I'd place my bets on Jax.