HNHacker News
TopNewBestAskShowJobs

lukasz_km

76 karma · joined April 29, 2017

submissionscomments
lukasz_km··on Data Science: Performance of Python vs. Pandas vs. Numpy
There is a lot of great Python and Pandas code snippets, but I am not sure anynone posted a Numpy based solution. Below is mine. It gives 100x speed up over the base (quadratic Python) solution. Still, suggested Pure Python solutions and idiomatic Numpy solutions are much faster. I suspect Numpy has more power than that.

def gen_stats_numpy_l(dataset_numpy): start = time.time() unique_products,unique_indices = np.unique(dataset_numpy[:,0],return_index = True) product_stats = [] split = np.split(dataset_numpy,unique_indices)[1:] for item in split: length = len(item) product_stats.append([int(item[0,0]),int(length),int(np.sum(item[:,2])),float(np.round(np.sum(item[:,3])/length,2))]) end = time.time() working_time = end-start return product_stats,working_time

lukasz_km··on Data Science: Performance of Python vs. Pandas vs. Numpy
I have tried PyPy once and it gives huge speed up. I am not sure how it works with Pandas/Numpy. I have also tried Numba JIT. Eeasy to use, but in my experince it saves just up to 10% of time. There is also Cython (https://pandas.pydata.org/pandas-docs/stable/enhancingperf.h...) and it definitely works with Pandas. Have not tried yet.
lukasz_km··on Data Science: Performance of Python vs. Pandas vs. Numpy
I have had much fun reading this thread of comments today :)

This comparison was never about concrete (optimal) implementations but about comparing exactly the same implementations, doing exactly the same operations. Such approach would be the same for comparing performance of any languages, like C and Java. This is not "the best algo" race.

So, after a day of discussion, the result is the following: regardless if the implementation is quadratic or linear, the relative speed of all 3 technologies is (roughly) the same. I could not agree more as this is primary school math: dividing by the same denominator still means that proportions are the same :)

I plan to update my blog post, based on many interesting comments you have all provided. I definetely think this analysis should be more complex. Thank you for your insights.

I have many great pieces of code for inspiration, but I am struggling with providing exactly the same implementations. So it may cost me some time to provide all 3 optimized implementations that are exactly relevant to each other.

lukasz_km··on Machine Learning: Regression of 911 Calls
Minimaxir thanks for comment. Regarding "accuracy" I agree that this word usage is unfortunate in case of R2. A "goodness" would be better description of R2. Regarding the "one-hot" encoiding of inputs, I would say - it depends. I know it is a common practice to do "one hot" encoding in Machine Learning. There are cases when it really helps the model to converge (for example, sometimes Neural Networks would require such curation of input). However, model that I have used in my post is Decision Tree based. And when you think for a moment about it, Decision Tree based model should handle numerical input just fine. Usually, Decision Tree based models do not even require standardization/normlization of input and work correctly without it.