For a more representative comparison, let's make everything 1000x larger, e.g., closing = np.concatenate(1000 * [closing])
Here's how a few representative benchmark change:
- describe: PandasPy was 5x faster, now 5x slower
- add: PandasPy was 2-3x faster than pandas, now ~15x slower
- concat: PandasPy was 25-70x faster, now 1-2x slower
- drop/rename: PandasPy is now ~1000x faster (NumPy can clearly do these operations without any data copies)
I couldn't test merge because it needs a sorted dataset, but hopefully you get the idea -- these benchmarks are meaningless, unless for some reason you only care about manipulating small datasets very quickly.
At large scale, pandas has two major advantages over NumPy/PandasPy:
- Pandas (often) uses a columnar data format, which makes it much faster to manipulate large datasets.
- Pandas has hash tables which it can rely upon for fast look-ups instead sorting.