{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <div style=\"color:white;display:fill;border-radius:5px;background-color:#0f438c;letter-spacing:0.5px;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;margin:0;font-size:110%\"><b></b> Tabular Playground Series - Jul 2022  </p></div>","metadata":{}},{"cell_type":"markdown","source":"### Content: \n> 1. Checking correlation of features  \n> 2. Measuring the cluster tendency - **Hopkins Statistic**\n> 3. Selecting best number of clusters - **Elbow Method** \n> 4. Vizualizing data with **PCA**  \n> 5. Evaluating your clustering model without true lables  - **1) Silhouette Coefficient, 2) Davies-Bouldin score, 3) Calinski-Harabasz Index**\n> 6.  Intercluster Distance Maps","metadata":{}},{"cell_type":"markdown","source":"#### Importing libraries ","metadata":{}},{"cell_type":"code","source":"# Basic libraries \nimport numpy as np\nimport pandas as pd \nimport seaborn as sns \nimport matplotlib.pyplot as plt \nfrom sklearn.decomposition import PCA \nfrom sklearn.cluster import KMeans  # basic clustering model \n\n\n# Yellowbrick for Elbow method and inter cluster distance map \nfrom yellowbrick.cluster import KElbowVisualizer  \nfrom yellowbrick.cluster import InterclusterDistance \n\n# Metrics for evaluating clustering models \nfrom sklearn.metrics import silhouette_score\nfrom sklearn.metrics import davies_bouldin_score  \nfrom sklearn.metrics import calinski_harabasz_score   \n\n# For Elbow method (if you want to implement elbow method from scratch)\nfrom scipy.spatial.distance import cdist ","metadata":{"execution":{"iopub.status.busy":"2022-07-09T19:34:28.361061Z","iopub.execute_input":"2022-07-09T19:34:28.361609Z","iopub.status.idle":"2022-07-09T19:34:29.509545Z","shell.execute_reply.started":"2022-07-09T19:34:28.361501Z","shell.execute_reply":"2022-07-09T19:34:29.508524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Importing dataset ","metadata":{}},{"cell_type":"code","source":"dataset = pd.read_csv('../input/tabular-playground-series-jul-2022/data.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-09T19:34:37.340549Z","iopub.execute_input":"2022-07-09T19:34:37.340979Z","iopub.status.idle":"2022-07-09T19:34:38.579663Z","shell.execute_reply.started":"2022-07-09T19:34:37.340931Z","shell.execute_reply":"2022-07-09T19:34:38.578426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Number of rows: {}\".format(dataset.shape[0]))\nprint(\"Number of columns: {}\".format(dataset.shape[1]))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T19:34:38.581621Z","iopub.execute_input":"2022-07-09T19:34:38.581949Z","iopub.status.idle":"2022-07-09T19:34:38.588320Z","shell.execute_reply.started":"2022-07-09T19:34:38.581919Z","shell.execute_reply":"2022-07-09T19:34:38.587152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <div style=\"color:white;display:fill;border-radius:5px;background-color:#0f438c;letter-spacing:0.5px;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;margin:0;font-size:110%\"><b></b> 1. Checking correlation of features  </p></div> ","metadata":{}},{"cell_type":"code","source":"fig,ax = plt.subplots(figsize=(25,10)) \nfig.suptitle(\"Correlation plot\", fontsize=24)\ncorrcoef = dataset.drop(['id'], axis=1).corr()\nmask = np.array(corrcoef)\nmask[np.tril_indices_from(mask)] = False\nsns.heatmap(corrcoef, mask=mask, vmax=.8, annot=True, ax=ax)\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2022-07-09T19:34:41.192515Z","iopub.execute_input":"2022-07-09T19:34:41.192896Z","iopub.status.idle":"2022-07-09T19:34:43.508719Z","shell.execute_reply.started":"2022-07-09T19:34:41.192866Z","shell.execute_reply":"2022-07-09T19:34:43.507955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <div style=\"color:white;display:fill;border-radius:5px;background-color:#0f438c;letter-spacing:0.5px;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;margin:0;font-size:110%\"><b></b> 2. Measuring the cluster tendency - **Hopkins Statistic** </p></div> ","metadata":{}},{"cell_type":"markdown","source":"> The Hopkins statistic (introduced by Brian Hopkins and John Gordon Skellam) is a way of measuring the cluster tendency of a data set. It belongs to the family of sparse sampling tests. It acts as a statistical hypothesis test where the null hypothesis is that the data is generated by a Poisson point process and are thus uniformly randomly distributed. \n\n* H=0.5 -> Random data will tend to result in values around 0.5 \n* H=0.0 -> actual data are highly clustered\n* H=1.0 -> actual data are regularly distributed in the data space (e.g grid) ","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://raw.githubusercontent.com/fukashi-hatake/kaggle_notebooks/main/images/hopkins_statistics1.png\" \n     width=\"50%\" \n     height=\"50%\" \n     /> ","metadata":{}},{"cell_type":"code","source":"def hopkins_statistic(df):\n    from sklearn.neighbors import NearestNeighbors\n    from sklearn.preprocessing import StandardScaler\n    n_samples = df.shape[0]\n    num_samples = [int(f*n_samples) for f in [0.25,0.5,0.75]]\n    states = [123,42,67,248,654]\n    for n in num_samples:\n        print('-'*12+str(n)+'-'*12)\n        hopkins_statistic = []\n        for random_state in states:\n            data = df.sample(n=n, random_state=random_state)\n            nbrs = NearestNeighbors(n_neighbors=2)\n            scaler = StandardScaler()\n            X = scaler.fit_transform(data.values)\n            nbrs.fit(X)\n            sample_dist = nbrs.kneighbors(X)[0][:,1]\n            sample_dist = np.sum(sample_dist)\n            random_data = np.random.rand(X.shape[0], X.shape[1])\n            nbrs.fit(random_data)\n            random_dist = nbrs.kneighbors(random_data)[0][:,1]\n            random_dist = np.sum(random_dist)\n            hs = sample_dist/(sample_dist+random_dist)\n            hopkins_statistic.append(hs)\n            print('*'*25)\n            print('hopkins statistic :'+str(hs))\n        print('mean hopkins statistic :'+str(np.mean(np.array(hopkins_statistic))))\n        print('hopkins statistic standard deviation :'+str(np.std(np.array(hopkins_statistic)))) ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"hopkins_statistic(dataset.drop(['id'], axis=1)) ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <div style=\"color:white;display:fill;border-radius:5px;background-color:#0f438c;letter-spacing:0.5px;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;margin:0;font-size:110%\"><b></b> 3. Selecting best number of clusters - **Elbow Method**  </p></div>  ","metadata":{}},{"cell_type":"markdown","source":"#### We can easly use Yellowbrick library for Elbow method ","metadata":{}},{"cell_type":"code","source":"print('Elbow Method to determine the number of clusters to be formed:')\nElbow_M = KElbowVisualizer(KMeans(), k=10)\nElbow_M.fit(dataset.drop(['id'], axis=1))\nElbow_M.show()  ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### If you want to plot Elbow from scratch, use following code ","metadata":{}},{"cell_type":"code","source":"X = dataset.drop(['id'], axis=1)\n\n\ndistortions = []\ninertias = []\nmapping1 = {}\nmapping2 = {}\nK = range(1, 10)\n \nfor k in K:\n    # Building and fitting the model\n    kmeanModel = KMeans(n_clusters=k).fit(X)\n    kmeanModel.fit(X)\n \n    distortions.append(sum(np.min(cdist(X, kmeanModel.cluster_centers_,\n                                        'euclidean'), axis=1)) / X.shape[0])\n    inertias.append(kmeanModel.inertia_)\n \n    mapping1[k] = sum(np.min(cdist(X, kmeanModel.cluster_centers_,\n                                   'euclidean'), axis=1)) / X.shape[0]\n    mapping2[k] = kmeanModel.inertia_\n    \n### with distortions \nplt.plot(K, distortions, 'bx-')\nplt.xlabel('Values of K')\nplt.ylabel('Distortion')\nplt.title('The Elbow Method using Distortion')\nplt.show()\n\n### Using the different values of Inertia\nplt.plot(K, inertias, 'bx-')\nplt.xlabel('Values of K')\nplt.ylabel('Inertia')\nplt.title('The Elbow Method using Inertia')\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <div style=\"color:white;display:fill;border-radius:5px;background-color:#0f438c;letter-spacing:0.5px;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;margin:0;font-size:110%\"><b></b> 4. Vizualizing data with **PCA**  </p></div> ","metadata":{}},{"cell_type":"markdown","source":"#### Applying PCA with 3 component \n> We have total 29 features and we want to squeeze these 29 features into 3 featues. ","metadata":{}},{"cell_type":"code","source":"X = dataset.drop(['id'], axis=1)\n\npca = PCA(n_components=3) \npca.fit(X) \nPCA_ds = pd.DataFrame(pca.transform(X), columns=([\"col1\",\"col2\", \"col3\"])) \n\n# A 3D Projection Of Data In The Reduced Dimension\nx =PCA_ds[\"col1\"]\ny =PCA_ds[\"col2\"]\nz =PCA_ds[\"col3\"]\n#To plot\nfig = plt.figure(figsize=(10,8))\nax = fig.add_subplot(111, projection=\"3d\")\nax.scatter(x, y, z, c=\"maroon\", marker=\"o\" )\nax.set_title(\"A 3D Projection Of Data In The Reduced Dimension\")\nplt.show() ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 😅 Not informative, let's cluster our data and then plot it again ","metadata":{}},{"cell_type":"markdown","source":"* Appling Elbow method again after PCA for making sure ","metadata":{}},{"cell_type":"code","source":"print('Elbow Method to determine the number of clusters to be formed:')\nElbow_M = KElbowVisualizer(KMeans(), k=10)\nElbow_M.fit(PCA_ds)\nElbow_M.show()   ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### OK! Let's devide our data into 5 clusters as Elbow method suggests 😇","metadata":{}},{"cell_type":"code","source":"NUMBER_OF_CLUSTERS = 5 \nSEED = 42 \nINIT = \"k-means++\"\n\nkmeans = KMeans(n_clusters=NUMBER_OF_CLUSTERS, init=INIT, random_state=SEED)\n\n# Fitting model \nkmeans.fit(PCA_ds) \n\nlabels = kmeans.predict(PCA_ds)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Now let's plot again with label ","metadata":{}},{"cell_type":"code","source":"# A 3D Projection Of Data In The Reduced Dimension\nx =PCA_ds[\"col1\"]\ny =PCA_ds[\"col2\"]\nz =PCA_ds[\"col3\"] \n\n#To plot\nfig = plt.figure(figsize=(10,8))\nax = fig.add_subplot(111, projection=\"3d\")\nax.scatter(x, y, z, c=labels, marker=\"o\", cmap=\"BuGn\")\nax.set_title(\"A 3D Projection Of Data In The Reduced Dimension\")\nplt.show() ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <div style=\"color:white;display:fill;border-radius:5px;background-color:#0f438c;letter-spacing:0.5px;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;margin:0;font-size:110%\"><b></b> 5. Evaluating your clustering model without true lables    </p></div>  ","metadata":{}},{"cell_type":"markdown","source":"#### Silhouette Coefficient ","metadata":{}},{"cell_type":"markdown","source":"> Silhouette Coefficient or silhouette score is a metric used to calculate the goodness of a clustering technique. Its value ranges from -1 to 1.\n\n* 1: Means clusters are well apart from each other and clearly distinguished.\n* 0: Means clusters are indifferent, or we can say that the distance between clusters is not significant.\n* -1: Means clusters are assigned in the wrong way. ","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://raw.githubusercontent.com/fukashi-hatake/kaggle_notebooks/main/images/Silhouette+Coefficient.jpg\" \n     width=\"50%\" \n     height=\"50%\" \n     /> ","metadata":{}},{"cell_type":"markdown","source":"* Read more: https://towardsdatascience.com/silhouette-coefficient-validating-clustering-techniques-e976bb81d10c ","metadata":{}},{"cell_type":"markdown","source":"#### Let's evaluate our previous clustering model  ","metadata":{}},{"cell_type":"code","source":"# Just call this function \nsilhouette_score(PCA_ds, labels)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> **Comment**: the score is not good. For improving results, 1) you apply PCA with different number of components (other than 3). 2) you can apply other clustering techniques like Fuzzy C-means, Gaussian Mixture, AgglomerativeClustering etc.  ","metadata":{}},{"cell_type":"markdown","source":"#### Davies-Bouldin score","metadata":{}},{"cell_type":"markdown","source":"> The score is defined as the average similarity measure of each cluster with its most similar cluster, where similarity is the ratio of within-cluster distances to between-cluster distances. Thus, clusters which are farther apart and less dispersed will result in a better score.\n\n* The minimum score is zero, with lower values indicating better clustering.\n\n* Read more: https://www.geeksforgeeks.org/dunn-index-and-db-index-cluster-validity-indices-set-1/?ref=lbp ","metadata":{}},{"cell_type":"code","source":"# Just call this function \ndavies_bouldin_score(PCA_ds, labels) ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Calinski-Harabasz Index","metadata":{}},{"cell_type":"markdown","source":"> Calinski-Harabasz (CH) Index (introduced by Calinski and Harabasz in 1974) can be used to evaluate the model when ground truth labels are not known where the validation of how well the clustering has been done is made using quantities and features inherent to the dataset. The CH Index (also known as Variance ratio criterion) is a measure of how similar an object is to its own cluster (cohesion) compared to other clusters (separation). Here cohesion is estimated based on the distances from the data points in a cluster to its cluster centroid and separation is based on the distance of the cluster centroids from the global centroid. CH index has a form of (a . Separation)/(b . Cohesion) , where a and b are weights.\n\n* Higher value of CH index means the clusters are dense and well separated, although there is no “acceptable” cut-off value. We need to choose that solution which gives a peak or at least an abrupt elbow on the line plot of CH indices.\n\n* Read more: https://www.geeksforgeeks.org/calinski-harabasz-index-cluster-validity-indices-set-3/ ","metadata":{}},{"cell_type":"code","source":"# Just call this function  \ncalinski_harabasz_score(PCA_ds, labels)  ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> **Comment:** this number does not make sense by looking but when you apply other techniques and other number of componenets for PCA, at that time you can use it for comparison. ","metadata":{}},{"cell_type":"markdown","source":"### 🧐 It is all the time better if you evaluate your performance with different matrices because you can see how you model is doing from different angles ","metadata":{}},{"cell_type":"markdown","source":"### <div style=\"color:white;display:fill;border-radius:5px;background-color:#0f438c;letter-spacing:0.5px;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;margin:0;font-size:110%\"><b></b> 6.  Intercluster Distance Maps </p></div>  ","metadata":{}},{"cell_type":"markdown","source":"> Intercluster distance maps display an embedding of the cluster centers in 2 dimensions with the distance to other centers preserved. E.g. the closer to centers are in the visualization, the closer they are in the original feature space. The clusters are sized according to a scoring metric. By default, they are sized by membership, e.g. the number of instances that belong to each center. This gives a sense of the relative importance of clusters. Note however, that because two clusters overlap in the 2D space, it does not imply that they overlap in the original feature space.\n\n","metadata":{}},{"cell_type":"code","source":"model = KMeans(NUMBER_OF_CLUSTERS)\nvisualizer = InterclusterDistance(model) \n\nvisualizer.fit(PCA_ds)       # Fit the data to the visualizer\nvisualizer.show()        # Finalize and render the figure ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <div style=\"color:white;display:center;border-radius:5px;background-color:#0f438c;letter-spacing:0.5px;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;margin:0;font-size:110%\"><b></b> Thank you for reading!</p></div>   ","metadata":{}}]}