This reminds me of the first internship I did. It was at an AI based anamoly detection startup. Their pipeline for the biggest client used to run in just under 3 hours, most of which was apparently spent extracting meaningful information from sensor data from a proprietary binary format using a shared library provided by the client. Ran it under perf and found that most of the time was spent calculating a bunch of sines and cosines. Used the LD_PRELOAD trick to log the sine and cosine calls and turns out that we were just calculating the sines and cosines of the same set of angles over and over again literally thousands of times. Wrote a wrapper to cache the values of sines and cosines in a table and return the cached value if it is already calculated and brought the running time of the pipeline to under 15 minutes. Which meant that the data science team could run their experiments and iterate on their model much faster and cheaper. Fun times.