NumPy Exercises for Data Analysis in Python
machinelearningplus.com
https://www.machinelearningplus.com/101-numpy-exercises-python/
machinelearningplus.com
https://www.machinelearningplus.com/101-numpy-exercises-python/
I've found the following to be quite helpful but would love to know if anyone knows of other resources in a similar vein: https://pandas.pydata.org/pandas-docs/stable/cookbook.html
There's also pandas_exercises by Guilherme Samora (https://github.com/guipsamora/pandas_exercises) which is very good - it's split across multiple notebooks and is more extensive than my repo.
For question 48 it might be simpler to just write
np.sort(a)[-5:]
instead of using argsort() and then using fancy indexing. Better yet, use np.partition(a, kth=-5)[-5:]
which scales linearly with the size of the array.Also, the one-hot encoding puzzle (51) would be more efficiently solved using
(arr[:, None] == np.unique(arr)).view(np.int8)
In general, `for` loops over NumPy arrays should be avoided where at all possible.For example from std::vector::insert [1]:
Complexity
1-2) Constant plus linear in the distance between pos and end of the container.
3) Linear in count plus linear in the distance between pos and end of the container.
4) Linear in std::distance(first, last) plus linear in the distance between pos and end of the container.
5) Linear in ilist.size() plus linear in the distance between pos and end of the container.
[1][http://en.cppreference.com/w/cpp/container/vector/insert]edit: formatting
For example, np.einsum for all its greatness in the past wasn't faster than np.tensordot, but it was more flexible. One can tell einsum to try and use the same underlying BLAS functions that tensordot uses (which can parallelise the computation) if applicable, and it will likely be default for einsum to perform this optimisation automatically once the devs iron out some bugs. But for now, it pays to know how the two methods are different.
One downside is that unless you're doing BLAS-style operations, writing non-trivial transformations of StaticArrays always seems to require generated functions.
Anyway, I think this is a feature that numpy doesn't provide.
np.unique??
the suggestion that jumps out is to just learn about algorithms.
For #3, you can make a boolean array with np.ones/np.zeros with the same dtype arg, saves a little bit of space.
ie np.ones((3,3), dtype=bool)
For #14, you can make use of the same compound boolean statements as you can in pandas to make it a bit simpler.
ie a[(a > 5) & (a < 10)]
For #15, this is a built in numpy function.
np.maximum(a,b).
That's as far as I've made it, but I'm really enjoying them.
However, for No. 15, that is not the point of the exercise.
> Desired Output:
> #> (array([1, 3, 5, 7]),)
Why is (array([1, 3, 5, 7]),) the desired output, and not array([1, 3, 5, 7]) ?(Also numpy has some really nice features over Matlab, like [None,:] broadcasting and being able to index a parenthesized expression or function output without naming it. Ok, the latter is not really a feature, more of an example of how Matlab is broken as a language)