To reiterate: None of this contradicts my main point that numpy.dot releases the GIL, which it does even if it calls into a BLAS library that uses multiple cores (although of course that would then have its own global lock - potentially even an interprocess lock - to manage concurrent uses). numpy.dot was only meant to be an example in the first place! I could just as easily have talked about some optimisation routine in scipy or something in the standard library (actually I did mention ZipFile.extract).
> I've worked in this field awhile.
This statement undermines your point more than it supports it. There is no "this field" for numpy. Just working in some field that uses it is no qualification for knowing whether users as a whole would like it multithreaded automatically. "Data analysts" and "supercomputer folks" no doubt include many users but certainly not all and possibly not even most (for all we know).
> 99% of people want this: ...
I can believe this might be true. I don't think I ever claimed otherwise. In fact I said "maybe"! I only claimed that I don't want this.
I will make one more claim though: I don't think there's any way that you could know this, even if it's really true.
But yes, I agree and already knew that it's how BLAS libraries work by default, including OpenBLAS which is what the prebuilt numpy wheels use. Even then, as you're probably already aware by the sounds of it, that's only for sufficiently large matrices. If you're just using small matrices and vectors (e.g. in 3D to represent real-world coordinates) then processing will be done directly in the calling thread (because the overhead of using multiple threads would be greater than the benefit).