{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 id=\"adda\" style=\"color:white;background:#0087B6;padding:8px;border-radius:8px\"> TPS - JULY 2022 - Unsupervised clustering challenge </h1>\n\n### Problem definition :\n\n* In this challenge, we are given a dataset where each row belongs to a particular cluster. our job is to predict the cluster each row belongs to .\n\n### Data :\n* data.csv - the file includes **continuous** and **categorical** data; your task is to predict which rows should be clustered together in a control state\n\n* sample_submission.csv - a sample submission file in the correct format, where Predicted is the predicted control state\n","metadata":{}},{"cell_type":"markdown","source":"### What you will read here :\n\n* Load data\n* Correlation between Continuous Features\n* Distribution of features\n* Boxplot features to see if there are outliers\n* Feature selection\n* Preprocessing\n    * Scaling\n    * Handling outliers\n* Finding optimal value of K\n    * Without using external library\n    * Using KElbowVisualizer to get the exact value of elbow point\n* Clustering with Bayesian Gaussian Mixture \n* PCA for visualizing clustered data\n","metadata":{}},{"cell_type":"markdown","source":"<h1 id=\"addfdsda\" style=\"color:white;background:#0087B6;padding:8px;border-radius:8px\"> Libraries </h1>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom yellowbrick.cluster import KElbowVisualizer\nfrom sklearn.cluster import KMeans\nfrom sklearn.decomposition import PCA \nfrom sklearn.feature_selection import VarianceThreshold\nfrom sklearn.mixture import BayesianGaussianMixture\nfrom sklearn.preprocessing import MinMaxScaler,RobustScaler,QuantileTransformer,MaxAbsScaler,PowerTransformer,Normalizer\nimport plotly.express as px\n\nsns.set()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-12T09:23:59.380780Z","iopub.execute_input":"2022-07-12T09:23:59.381382Z","iopub.status.idle":"2022-07-12T09:24:02.551691Z","shell.execute_reply.started":"2022-07-12T09:23:59.381329Z","shell.execute_reply":"2022-07-12T09:24:02.550718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 id=\"adhdeyda\" style=\"color:white;background:#0087B6;padding:8px;border-radius:8px\">Loading the data </h1>","metadata":{}},{"cell_type":"code","source":"data = pd.read_csv('../input/tabular-playground-series-jul-2022/data.csv')\nsub = pd.read_csv('../input/tabular-playground-series-jul-2022/sample_submission.csv')\n\ndata","metadata":{"execution":{"iopub.status.busy":"2022-07-12T09:24:02.553642Z","iopub.execute_input":"2022-07-12T09:24:02.554000Z","iopub.status.idle":"2022-07-12T09:24:03.591944Z","shell.execute_reply.started":"2022-07-12T09:24:02.553955Z","shell.execute_reply":"2022-07-12T09:24:03.590892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 id=\"awdjjjja\" style=\"color:white;background:#0087B6;padding:8px;border-radius:8px\"> Info </h1>","metadata":{}},{"cell_type":"code","source":"data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T09:23:49.685220Z","iopub.status.idle":"2022-07-12T09:23:49.686243Z","shell.execute_reply.started":"2022-07-12T09:23:49.685925Z","shell.execute_reply":"2022-07-12T09:23:49.685953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 class=\"alert alert-info\"> Data has no missing values </h3>\n","metadata":{}},{"cell_type":"markdown","source":"# Drop id","metadata":{}},{"cell_type":"code","source":"data=data.drop('id',axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T09:24:03.614342Z","iopub.execute_input":"2022-07-12T09:24:03.614641Z","iopub.status.idle":"2022-07-12T09:24:03.627010Z","shell.execute_reply.started":"2022-07-12T09:24:03.614615Z","shell.execute_reply":"2022-07-12T09:24:03.625423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 id=\"addawwdda\" style=\"color:white;background:#0087B6;padding:8px;border-radius:8px\"> Correlation between continuous features </h1>","metadata":{}},{"cell_type":"code","source":"Continuous_Features = data.select_dtypes(include='float')\nCat_features = data.select_dtypes(include='int')","metadata":{"execution":{"iopub.status.busy":"2022-07-12T09:24:05.647236Z","iopub.execute_input":"2022-07-12T09:24:05.647862Z","iopub.status.idle":"2022-07-12T09:24:05.664127Z","shell.execute_reply.started":"2022-07-12T09:24:05.647824Z","shell.execute_reply":"2022-07-12T09:24:05.663152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = ( 15 , 11 ))\nsns.heatmap(Continuous_Features.corr(),annot=True,fmt=\".2f\", cmap='Blues');","metadata":{"execution":{"iopub.status.busy":"2022-07-12T09:24:06.954081Z","iopub.execute_input":"2022-07-12T09:24:06.954757Z","iopub.status.idle":"2022-07-12T09:24:09.131976Z","shell.execute_reply.started":"2022-07-12T09:24:06.954717Z","shell.execute_reply":"2022-07-12T09:24:09.131067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 class=\"alert alert-info\"> Features are not much correlated with each other </h3>","metadata":{}},{"cell_type":"markdown","source":"<h1 id=\"addawrjda\" style=\"color:white;background:#0087B6;padding:8px;border-radius:8px\"> Continuous features distribution </h1>","metadata":{}},{"cell_type":"code","source":"Continuous_Features.hist(figsize=(20,15),bins=100,color='#000');","metadata":{"execution":{"iopub.status.busy":"2022-07-12T09:24:13.323036Z","iopub.execute_input":"2022-07-12T09:24:13.323412Z","iopub.status.idle":"2022-07-12T09:24:19.601247Z","shell.execute_reply.started":"2022-07-12T09:24:13.323373Z","shell.execute_reply":"2022-07-12T09:24:19.600344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 class=\"alert alert-info\"> Numeric features have approximately normal distribution </h3>","metadata":{}},{"cell_type":"markdown","source":"<h1 id=\"addaggdd\" style=\"color:white;background:#0087B6;padding:8px;border-radius:8px\"> Boxplot features </h1>","metadata":{}},{"cell_type":"code","source":"def my_boxplot(data):\n    tmp_df = pd.DataFrame(data = data, columns = data.columns.to_list())\n    plt.figure(figsize=(20,10)) \n    sns.boxplot(x=\"variable\", y=\"value\", data=pd.melt(tmp_df)).set_title('Boxplot of each feature',size=15)\n    plt.show()\n    \nmy_boxplot(data)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T09:24:19.603085Z","iopub.execute_input":"2022-07-12T09:24:19.603547Z","iopub.status.idle":"2022-07-12T09:24:22.765702Z","shell.execute_reply.started":"2022-07-12T09:24:19.603507Z","shell.execute_reply":"2022-07-12T09:24:22.764757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 class=\"alert alert-info\"> We can see the difference between categorical and numeric features  </h3>","metadata":{}},{"cell_type":"markdown","source":"<h1 id=\"agrydaawdda\" style=\"color:white;background:#0087B6;padding:8px;border-radius:8px\"> Feature selection -  VarianceThreshold</h1>\n\nThis feature selection algorithm looks only at the features (X), not the desired outputs (y), and can thus be used for unsupervised learning.\n\nIt removes all low-variance features.","metadata":{}},{"cell_type":"code","source":"selector = VarianceThreshold(threshold=1.5)\nselector.fit_transform(data)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T10:27:42.340071Z","iopub.execute_input":"2022-07-12T10:27:42.340774Z","iopub.status.idle":"2022-07-12T10:27:42.391067Z","shell.execute_reply.started":"2022-07-12T10:27:42.340733Z","shell.execute_reply":"2022-07-12T10:27:42.389622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,10))\nsns.barplot(x=selector.variances_, y=data.columns,orient='h' ).set_title('Feature selection with VarianceThreshold',size=15);\nplt.xlabel('Variance');\nplt.axvline(x=1.5, color='r', linestyle='--', label='Threshold')\nplt.legend();","metadata":{"execution":{"iopub.status.busy":"2022-07-12T10:35:13.512542Z","iopub.execute_input":"2022-07-12T10:35:13.513226Z","iopub.status.idle":"2022-07-12T10:35:14.084780Z","shell.execute_reply.started":"2022-07-12T10:35:13.513188Z","shell.execute_reply":"2022-07-12T10:35:14.083874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 class=\"alert alert-info\">  Low-variance features : f_00 to f_06  and  f_14 to f_21 </h3>","metadata":{}},{"cell_type":"markdown","source":"## Get selected features","metadata":{}},{"cell_type":"code","source":"Selected_features = list(selector.get_feature_names_out())\ndata[Selected_features]","metadata":{"execution":{"iopub.status.busy":"2022-07-12T10:27:48.516476Z","iopub.execute_input":"2022-07-12T10:27:48.517132Z","iopub.status.idle":"2022-07-12T10:27:48.539189Z","shell.execute_reply.started":"2022-07-12T10:27:48.517098Z","shell.execute_reply":"2022-07-12T10:27:48.538073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 id=\"agrywdda\" style=\"color:white;background:#0087B6;padding:8px;border-radius:8px\"> Preprocessing </h1>\n","metadata":{}},{"cell_type":"markdown","source":"## Scaling","metadata":{}},{"cell_type":"code","source":"#Scaler\n#MaxAbs =  MaxAbsScaler()\n\n#Power_transformer - makes columns more gaussian like \nPower_transformer = PowerTransformer()\n\ndef Scale_transform(X):\n    #X = MaxAbs.fit_transform(X)\n    X = Power_transformer.fit_transform(X)\n    X= pd.DataFrame(X, columns = Selected_features)\n    return X","metadata":{"execution":{"iopub.status.busy":"2022-07-12T11:27:37.130594Z","iopub.execute_input":"2022-07-12T11:27:37.131669Z","iopub.status.idle":"2022-07-12T11:27:37.137503Z","shell.execute_reply.started":"2022-07-12T11:27:37.131621Z","shell.execute_reply":"2022-07-12T11:27:37.136412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_scaled = Scale_transform(data[Selected_features])","metadata":{"execution":{"iopub.status.busy":"2022-07-12T11:27:38.398907Z","iopub.execute_input":"2022-07-12T11:27:38.399518Z","iopub.status.idle":"2022-07-12T11:27:39.897109Z","shell.execute_reply.started":"2022-07-12T11:27:38.399483Z","shell.execute_reply":"2022-07-12T11:27:39.895871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Boxplot after scaling","metadata":{}},{"cell_type":"code","source":"my_boxplot(X_scaled)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T11:02:22.307919Z","iopub.execute_input":"2022-07-12T11:02:22.308792Z","iopub.status.idle":"2022-07-12T11:02:23.863478Z","shell.execute_reply.started":"2022-07-12T11:02:22.308743Z","shell.execute_reply":"2022-07-12T11:02:23.862555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 id=\"adggddweda\" style=\"color:white;background:#0087B6;padding:8px;border-radius:8px\"> Finding optimal number of clusters using K-means </h1>\n\nThe number of clusters that we choose for a given dataset cannot be random. Each cluster is formed by calculating and comparing the distances of data points within a cluster to its centroid. An ideal way to figure out the right number of clusters would be to calculate the Within-Cluster-Sum-of-Squares (WCSS)","metadata":{}},{"cell_type":"markdown","source":"## Kmeans Clustering and storing WCSS for each K","metadata":{}},{"cell_type":"code","source":"#Clustering function\ndef cluster(n_clusters):\n    kmeans = KMeans(n_clusters=n_clusters)\n    kmeans.fit(X_scaled)\n    return kmeans\n\n\n# Clustering and storing WCSS for each K\nWCSS = []\n\nfor k in range(1, 15):\n    kmeans = cluster(k)\n    WCSS.append(kmeans.inertia_)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T18:24:04.739451Z","iopub.execute_input":"2022-07-08T18:24:04.740111Z","iopub.status.idle":"2022-07-08T18:26:19.745866Z","shell.execute_reply.started":"2022-07-08T18:24:04.740076Z","shell.execute_reply":"2022-07-08T18:26:19.744746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Plot WCSS against number of clusters","metadata":{}},{"cell_type":"code","source":"px.line(x=range(1,15), y=WCSS ,\n       labels={'y':'WCSS', 'x':'Number of clusters'},\n       title=\"Investigate K-means clustering\")","metadata":{"execution":{"iopub.status.busy":"2022-07-08T18:24:00.057137Z","iopub.execute_input":"2022-07-08T18:24:00.058149Z","iopub.status.idle":"2022-07-08T18:24:00.078767Z","shell.execute_reply.started":"2022-07-08T18:24:00.058103Z","shell.execute_reply":"2022-07-08T18:24:00.077328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## An easier way: Use KElbowVisualizer to get the exact value of elbow point\nFor more information read the documentation [here](https://www.scikit-yb.org/en/latest/api/cluster/elbow.html)","metadata":{}},{"cell_type":"code","source":"model = KMeans()\nvisualizer = KElbowVisualizer(model, k=(4,13))\n\nvisualizer.fit(X_scaled)        # Fit the data to the visualizer\nvisualizer.show()        # Finalize and render the figure","metadata":{"execution":{"iopub.status.busy":"2022-07-12T09:48:53.402786Z","iopub.execute_input":"2022-07-12T09:48:53.403159Z","iopub.status.idle":"2022-07-12T09:49:57.643325Z","shell.execute_reply.started":"2022-07-12T09:48:53.403127Z","shell.execute_reply":"2022-07-12T09:49:57.642424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 class=\"alert alert-info\"> The Elbow method suggests 7 as optimal number of clusters</h3>","metadata":{}},{"cell_type":"markdown","source":"<h1 id=\"addffswtjuja\" style=\"color:white;background:#0087B6;padding:8px;border-radius:8px\"> Clustering </h1>","metadata":{}},{"cell_type":"markdown","source":"## Bayesian Gaussian Mixture ","metadata":{}},{"cell_type":"code","source":"preds = BayesianGaussianMixture(n_components=7,\n                                covariance_type='full',\n                                max_iter=1000,\n                                n_init=5,\n                                random_state=1).fit_predict(X_scaled)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T11:02:35.083787Z","iopub.execute_input":"2022-07-12T11:02:35.084517Z","iopub.status.idle":"2022-07-12T11:05:02.159677Z","shell.execute_reply.started":"2022-07-12T11:02:35.084471Z","shell.execute_reply":"2022-07-12T11:05:02.158292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create submission file\nsub['Predicted']=preds\nsub.to_csv('submission.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T11:05:02.161668Z","iopub.execute_input":"2022-07-12T11:05:02.162290Z","iopub.status.idle":"2022-07-12T11:05:02.338686Z","shell.execute_reply.started":"2022-07-12T11:05:02.162255Z","shell.execute_reply":"2022-07-12T11:05:02.337671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 id=\"addffswtjuja\" style=\"color:white;background:#0087B6;padding:8px;border-radius:8px\"> PCA & visualizing clusters","metadata":{}},{"cell_type":"code","source":"#PCA\npca = PCA(n_components=2)\n\nPCA_DF = pd.DataFrame(pca.fit_transform(X_scaled), columns=([\"col1\",\"col2\"]))\nPCA_DF[\"Cluster\"]=preds\n\n# Plot \nplt.figure(figsize=(20,10)) \nsns.scatterplot(x=PCA_DF['col1'], y=PCA_DF['col2'], hue=PCA_DF['Cluster'])","metadata":{"execution":{"iopub.status.busy":"2022-07-12T11:06:39.328951Z","iopub.execute_input":"2022-07-12T11:06:39.329299Z","iopub.status.idle":"2022-07-12T11:06:42.264218Z","shell.execute_reply.started":"2022-07-12T11:06:39.329269Z","shell.execute_reply":"2022-07-12T11:06:42.263406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Still working on it . . .","metadata":{}},{"cell_type":"markdown","source":"### If you have time take a look at my other notebooks [here](https://www.kaggle.com/mehrdadsadeghi/code)","metadata":{}},{"cell_type":"markdown","source":"## Credits \n\n* https://www.kaggle.com/code/gemartin/load-data-reduce-memory-usage/notebook\n* https://www.kaggle.com/code/sfktrkl/tps-july-2022/notebook?scriptVersionId=99952091","metadata":{}}]}