- To get the most out of Python, you really need efficient libraries like Numpy. I started off my project in pure Python and some Pandas. But soon I realized it was painfully slow and looked for better solutions.
- Polars looked like a good Pandas alternative, but some googling revealed that it could be slower than Numpy. Although it's easy to work with dataframes (e.g., you can select a column by calling its name), I converted all my code to use Numpy and suddenly saw huge speed boost. But that wasn't the end of it. I was now in a rabbit hole...
- Vectorization helped a lot with speed as well. I learnt that using the right tool (in this case, Numpy) guides you to do better and vectorize operations, whereas Pandas didn't enforce that technique on me. Despite the initial hurdle, I realized that vectorization made my code much more readable and easier to follow.
- I then parallelized my code to utilized all 10 cores of the M1 Pro chip. Seeing that CPU graph get full was just pleasing. For this, I used Python's ProcessPoolExecutor, which is better than ThreadPoolExecutor.
- At this point, my simulations had gone from taking several hours to a couple minutes! Pretty impressive, but then I noticed that increasing the resolution of some variables resulted in RAM issues. I have 32GB of RAM and working with huge matrices was just not feasible. I remembered that the OS community has been using quantization techniques to reduce the memory footprint of LLMs. The way they work often involves going from np.float32 to np.float16. To my surprise, doing this solved my RAM issues without much effect on the accuracy of the results.
- Looking for further improvements, I discovered Dask and used it instead of ProcessPoolExecutor (although, I still used `schedulor="processes"` in Dask). Unlike ProcessPoolExecutor, Dask didn't utilize all my CPU cores all the time, but somehow I found it more capable of handling parallel computation.
- But increasing the resolution of parameters even further would still cause problems. Specifically, I noticed that my Python code would halt. Doing CTRL-Z would stop the program, but the RAM wouldn't get released. In this case, I had to manually `pkill -f python` to throw away all the Python objects that were taking up RAM space.
- This is where I discovered garbage collection. Using `import gc`, I was able to `del <obj>` at key parts of my code and immediately do `gc.collect()` to collect garbage. This was the first time I ever did this and I was pleased to see my program run again w/o RAM issues.
- I then optimized some search algorithms in my code which further sped up the program while also reducing the RAM usage.
- I converted all Numpy arrays to ds.array for further improvements, but some functionalities were not immediately available and I went back to Numpy.
- Looking for further improvements, I thought about running my Python code on GPU. I tried `import cupy as np` but got to problems that I couldn't solve.
- So I tried converting all numpy arrays to Pytorch tensors. Pytorch is available on M1 Mac so I thought why not? Unfor, I got memory limits errors. Reducing the resolution of parameters fixed that, but then I got a mysterious error for which I couldn't find any help:
-[_MTLCommandBuffer addCompletedHandler:]:871: failed assertion `Completed handler provided after commit call'
If you know what this means, please let me know!This was an educational journey and I learnt much. I think projects like this are the best way to learn new stuff.