K-means clustering using sklearn and Python
heartbeat.fritz.ai
heartbeat.fritz.ai
And K-means??? Why not HDBSCAN?
https://hdbscan.readthedocs.io/en/latest/performance_and_sca...
Recently I was asked to participate in a competition to identify brain hemorrhages (https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detect...). It turns out jhoward has published a lot of Kaggle notebooks walking through the entire process of gathering the data, cleaning it, implementing a learning algorithm, and submitting an entry.
If you're trying to do practical end to end machine learning, these are definitely worth studying.
Notebooks:
https://www.kaggle.com/jhoward/from-prototyping-to-submissio...
https://www.kaggle.com/jhoward/cleaning-the-data-for-rapid-p...
https://www.kaggle.com/jhoward/some-dicom-gotchas-to-be-awar...
https://www.kaggle.com/jhoward/don-t-see-like-a-radiologist-...
The last one is particularly interesting, and proposes a new color map for seeing 65536 different greyscale values. Turbo is also an alternative: https://ai.googleblog.com/2019/08/turbo-improved-rainbow-col...
1) there is some very useful stuff there, particularly for someone new to consuming medical imaging data. The writeups are aiming to be fairly complete
2) there are some things jhoward is being naive, e.g. CT image scaling, where the advice could give you trouble.
This is a good exercise for anyone starting out learning about machine learning, but I'd always stick to a well known library if I was actually using it for something else.
> Care is needed to pick the optimal starting centroids and k.
Definitely, and I think that speaks to the laziness of the linked article that they just say "Use the elbow method" for choosing k. In my ~4 years of being a data scientist I've never seen or heard of this working for any "real world" problem. Metrics like silhouette scores are much more useful and quantifiable.
Nice write for someone starting, more details about details of algorithm steps would greatly attract more readers.
Yes, you can. I have studied statistics and I cringe at the watering down of what "machine learning" and "AI" has become; simple statistics.
>Nice write for someone starting, more details about details of algorithm steps would greatly attract more readers.
I disagree, making it even simpler would attract more readers. You see the same with Youtube tutorials that have 22 parts. The first part has 200.000 views, the second 150.000 and the 20th part only has 400 views or so.