As one journeys through his/her career in analytics, some truths start becoming evident over time. And while none of them are ground-shattering, they often surprise novices in the field. So, it’s worthwhile to know 11 absolute facts of data science.
Clustering is a canonical example of un-supervised machine learning methods. Un-supervised, as in, true clusters (segments) don’t exist or aren’t known in advance. Hence method tries to separate observations in different groups without any way to verify if model has done good job or not. There are various ways we can try to measure performance of un-supervised clustering: Within-Cluster-Sum-of-Squares is one, Silhouette Coefficient is another
Coursera’s website is an educational platform that has partnered with top organizations and universities across the globe to offer online courses for anyone for free. You can select any course that suits your profile and sign up for it and learn with your own convenience. Best Analytics Courses online
Curse of Dimensionality refers to non-intuitive properties of data observed when working in high-dimensional space, specifically related to usability and interpretation of distances and volumes.
Gradient Boosting models are another variant of ensemble models, different from Random Forest. In Random Forest (RF) models, goal is to build many-many overfitted models each on subset of training data, and combine their individual prediction to make final prediction. In Gradient Boosting models, goal is to build series of many-many underfitted models, each bettering errors of previous model, and cumulative prediction is used to make final prediction.
Text mining is an analytical field which derives high quality information from text. Text mining is widely used in the industry when data is unstructured. Derived information can be provided in the form of numbers (indices), categories or clusters, summary of text. This article focusses on applications of text mining, workflow and example.