If you want to specify the type (for example for aot or just because you want to make it clear) then the call signature is less flexible.
In short, pick any random Python library, you’d find there are very few places you can jit accelerate something effectively. It is for numeric.
Even for numerical code, it is more like writing C functions than say C++ (with classes etc).
But it does makes accelerating vectorized code very easy. Even if you have a function that uses Numpy, it is likely you can speed it up using Numba with a decorator only.
But when it doesn’t work, it might often be not very clear why you can’t until you get some experience.
"Fits very few use cases" LOL okay without numba there's no UMAP and HDBScan and those are pretty popular and important libraries that come to mind just off the top of my head...
Also, claiming Cython is well documented also gets a huge LOL from me as someone whose actually written a bit of Cython.
I've succesfully deployed numba code in an AWS lambda for instance -- llvmlite takes a lot of your 250mb package budget, but once the lambda is "warm" the jit lag isn't an issue.
That said, if you absolutely want AOT you'll have to use Cython or some horrible hack dumping the compiled function binary.
I like numba, but cython is clearly used more in the popular packages
In the past, that was also possible for AOT compilation [1], but that technique broke during some update and it seems like there is no one left who knows how to fix this.
And I agree that it's not actually usable everywhere, since the support for numpy's feature set is actually quite limited, especially around multidimensional arrays. I had to effectively rewrite my logic to make use of numba. Still it is pretty worth it imo, given how it can add parallelism for free. And conforming to numbas allowed subset of numpy usually results in simpler and more efficient code. In my case I ended up having to work around the lack of support for multidimensional arrays but ended up with a more efficient solution relying on low dimensional arrays being broadcasted, reducing a lot of duplicate computations