{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px\"> Objective </h2>\n\nWe are given simulated manufacturing control data that can be clustered into different control states. Our task is to cluster the data into these control states. We are not given any training data, and we are not told how many possible control states there are.","metadata":{}},{"cell_type":"markdown","source":"<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px\"> Load data and libraries </h2>","metadata":{}},{"cell_type":"code","source":"## This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom pathlib import Path\nfrom tqdm import tqdm\nfrom matplotlib import pyplot as plt\nimport seaborn as sns\nfrom sklearn.model_selection import train_test_split\nfrom lightgbm import LGBMClassifier\nfrom xgboost import XGBClassifier\n\nfrom sklearn.metrics import accuracy_score, recall_score, precision_score\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:06:49.927947Z","iopub.execute_input":"2022-07-28T15:06:49.928543Z","iopub.status.idle":"2022-07-28T15:06:54.651541Z","shell.execute_reply.started":"2022-07-28T15:06:49.928481Z","shell.execute_reply":"2022-07-28T15:06:54.650263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.read_csv(\"../input/tabular-playground-series-jul-2022/data.csv\", index_col='id')\nsubmission = pd.read_csv(\"../input/tabular-playground-series-jul-2022/sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:06:54.657505Z","iopub.execute_input":"2022-07-28T15:06:54.660479Z","iopub.status.idle":"2022-07-28T15:06:56.226416Z","shell.execute_reply.started":"2022-07-28T15:06:54.660436Z","shell.execute_reply":"2022-07-28T15:06:56.224813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px;font-family:sans-serif;\"> Data exploration </h2>\n\n* No NA data\n* 29 features\n* f07-f13 correlates with f22-28 and with each other","metadata":{}},{"cell_type":"code","source":"data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:58:00.583319Z","iopub.execute_input":"2022-07-27T06:58:00.583696Z","iopub.status.idle":"2022-07-27T06:58:00.607534Z","shell.execute_reply.started":"2022-07-27T06:58:00.583653Z","shell.execute_reply":"2022-07-27T06:58:00.606488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.columns","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:09:53.692532Z","iopub.execute_input":"2022-07-27T06:09:53.692955Z","iopub.status.idle":"2022-07-27T06:09:53.703334Z","shell.execute_reply.started":"2022-07-27T06:09:53.692926Z","shell.execute_reply":"2022-07-27T06:09:53.701855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_corr = data.corr()\nplt.subplots(figsize=(15,10))\nsns.heatmap(data_corr, annot= True, cmap=\"YlOrRd\", fmt = '0.1f', vmin=-0.6, vmax=0.6);","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:09:53.704870Z","iopub.execute_input":"2022-07-27T06:09:53.705294Z","iopub.status.idle":"2022-07-27T06:09:57.100742Z","shell.execute_reply.started":"2022-07-27T06:09:53.705239Z","shell.execute_reply":"2022-07-27T06:09:57.099331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px;font-family:sans-serif;\"> Data transformation </h2>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(18, 18))\nfor i, col in enumerate(list(data.columns)): \n    ax = plt.subplot(11,5, i+1) \n    sns.histplot(data,x=col,ax=ax)\nplt.suptitle('Data distribution of continuous variables')\nplt.tight_layout()     ","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:09:57.102221Z","iopub.execute_input":"2022-07-27T06:09:57.102605Z","iopub.status.idle":"2022-07-27T06:10:08.440100Z","shell.execute_reply.started":"2022-07-27T06:09:57.102546Z","shell.execute_reply":"2022-07-27T06:10:08.438984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<!-- Let's scale the data\n\nfrom sklearn.preprocessing import RobustScaler\nscaler = RobustScaler()\ndata_scaled = pd.DataFrame(scaler.fit_transform(data),columns = data.columns) -->","metadata":{}},{"cell_type":"markdown","source":"Let's power transform the data. We see that the variables f07-f13 follows more of power law distribution than a gaussian distribution. A power/log transform will make the probability distribution of the variables more Gaussian.\n\nhttps://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.PowerTransformer.html","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import PowerTransformer\npower = PowerTransformer(method='yeo-johnson', standardize=False)\ndata_scaled = pd.DataFrame( power.fit_transform(data),columns = data.columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:10:08.441554Z","iopub.execute_input":"2022-07-27T06:10:08.442310Z","iopub.status.idle":"2022-07-27T06:10:12.087812Z","shell.execute_reply.started":"2022-07-27T06:10:08.442266Z","shell.execute_reply":"2022-07-27T06:10:12.086706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(18, 18))\nfor i, col in enumerate(list(data.columns)): \n    ax = plt.subplot(11,5, i+1) \n    sns.histplot(data_scaled,x=col,ax=ax)\nplt.suptitle('Data distribution of continuous variables')\nplt.tight_layout()     ","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:10:12.089176Z","iopub.execute_input":"2022-07-27T06:10:12.089531Z","iopub.status.idle":"2022-07-27T06:10:23.183828Z","shell.execute_reply.started":"2022-07-27T06:10:12.089500Z","shell.execute_reply":"2022-07-27T06:10:23.182559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px;font-family:sans-serif;\"> K-Means clustering </h2>\n\nAdjusted Rank Index : 0.23671","metadata":{}},{"cell_type":"code","source":"from yellowbrick.cluster import KElbowVisualizer\nfrom sklearn.cluster import KMeans\nkmeans = KMeans(random_state=42)\nvisualizer = KElbowVisualizer(kmeans, k=(5,10), timings=True,locate_elbow=False)\nvisualizer.fit(data_scaled)\nvisualizer.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:10:23.185238Z","iopub.execute_input":"2022-07-27T06:10:23.186152Z","iopub.status.idle":"2022-07-27T06:11:00.524998Z","shell.execute_reply.started":"2022-07-27T06:10:23.186111Z","shell.execute_reply":"2022-07-27T06:11:00.523693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_scaled[\"Predicted\"] = KMeans(n_clusters=8, random_state=0).fit_predict(data_scaled)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:11:00.530344Z","iopub.execute_input":"2022-07-27T06:11:00.530704Z","iopub.status.idle":"2022-07-27T06:11:07.419408Z","shell.execute_reply.started":"2022-07-27T06:11:00.530672Z","shell.execute_reply":"2022-07-27T06:11:07.418298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for c in data_scaled[['f_07', 'f_08','f_09', 'f_10', 'f_11', 'f_12', 'f_13']]:\n    grid= sns.FacetGrid(data_scaled, col='Predicted') \n    grid.map(plt.hist, c)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:11:07.420991Z","iopub.execute_input":"2022-07-27T06:11:07.421465Z","iopub.status.idle":"2022-07-27T06:11:17.149504Z","shell.execute_reply.started":"2022-07-27T06:11:07.421419Z","shell.execute_reply":"2022-07-27T06:11:17.147986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clusters = []\n\nfor i in range(1, 20):\n    km = KMeans(n_clusters=i).fit(data_scaled)\n    clusters.append(km.inertia_)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:11:17.150730Z","iopub.execute_input":"2022-07-27T06:11:17.151316Z","iopub.status.idle":"2022-07-27T06:12:34.530131Z","shell.execute_reply.started":"2022-07-27T06:11:17.151284Z","shell.execute_reply":"2022-07-27T06:12:34.525751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(12, 8))\nsns.lineplot(x=list(range(1, 20)), y=clusters, ax=ax)\nax.set_title('Searching for Elbow')\nax.set_xlabel('Clusters')\nax.set_ylabel('Inertia')","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.532682Z","iopub.status.idle":"2022-07-27T06:12:34.533400Z","shell.execute_reply.started":"2022-07-27T06:12:34.533035Z","shell.execute_reply":"2022-07-27T06:12:34.533066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission[\"Predicted\"] = data_scaled[\"Predicted\"]\nsubmission.to_csv('submission_knn.csv', index=False)\nsubmission","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.535625Z","iopub.status.idle":"2022-07-27T06:12:34.536585Z","shell.execute_reply.started":"2022-07-27T06:12:34.536280Z","shell.execute_reply":"2022-07-27T06:12:34.536313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px;font-family:sans-serif;\">  Gaussian Mixture Models </h2>\n\nThough GMM is often categorized as a clustering algorithm, fundamentally it is an algorithm for density estimation. That is to say, the result of a GMM fit to some data is technically not a clustering model, but a generative probabilistic model describing the distribution of the data i.e what's the likelihood of the data given the model?\n\nThe optimal number of clusters is the value that minimizes the AIC or BIC, depending on which approximation we wish to use.\n\nhttps://jakevdp.github.io/PythonDataScienceHandbook/05.12-gaussian-mixtures.html\n\nAdjusted Rand Index = 0.59931 (based on power transform of data)","metadata":{}},{"cell_type":"markdown","source":"### What's the ideal number of components in GMM? Let's find out using BIC and AIC","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import PowerTransformer\npower = PowerTransformer(method='yeo-johnson', standardize=False)\ndata_scaled = pd.DataFrame( power.fit_transform(data),columns = data.columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.538750Z","iopub.status.idle":"2022-07-27T06:12:34.539608Z","shell.execute_reply.started":"2022-07-27T06:12:34.539305Z","shell.execute_reply":"2022-07-27T06:12:34.539336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.mixture import GaussianMixture","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.540946Z","iopub.status.idle":"2022-07-27T06:12:34.542175Z","shell.execute_reply.started":"2022-07-27T06:12:34.541860Z","shell.execute_reply":"2022-07-27T06:12:34.541892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_components = np.arange(3, 10)\nmodels = []\nfor n in tqdm(n_components):\n    models.append(GaussianMixture(n_components=n, covariance_type='full', random_state=1).fit(data_scaled))\nplt.plot(n_components, [m.bic(data_scaled) for m in models], label='BIC')\nplt.plot(n_components, [m.aic(data_scaled) for m in models], label='AIC')\nplt.legend(loc='best')\nplt.xlabel('n_components');","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.543697Z","iopub.status.idle":"2022-07-27T06:12:34.545125Z","shell.execute_reply.started":"2022-07-27T06:12:34.544826Z","shell.execute_reply":"2022-07-27T06:12:34.544856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After 7 clusters, we see BIC tapers off. Since BIC penalizes more complex models, 7 seems to be the ideal number.","metadata":{}},{"cell_type":"markdown","source":"### Let's compare BayesianGaussian and Gaussian","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import PowerTransformer\npower = PowerTransformer(method='yeo-johnson', standardize=False)\ndata_scaled = pd.DataFrame( power.fit_transform(data),columns = data.columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.546496Z","iopub.status.idle":"2022-07-27T06:12:34.547504Z","shell.execute_reply.started":"2022-07-27T06:12:34.547175Z","shell.execute_reply":"2022-07-27T06:12:34.547205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gmm = GaussianMixture(n_components=7, covariance_type='full',  random_state=1) \nsubmission[\"Predicted_GMM\"] =  gmm.fit_predict(data_scaled)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.548996Z","iopub.status.idle":"2022-07-27T06:12:34.549569Z","shell.execute_reply.started":"2022-07-27T06:12:34.549292Z","shell.execute_reply":"2022-07-27T06:12:34.549319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.mixture import BayesianGaussianMixture\ngmm = BayesianGaussianMixture(n_components=7, covariance_type='full', random_state=1)\nsubmission[\"Predicted_BGMM\"] =  gmm.fit_predict(data_scaled)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.551685Z","iopub.status.idle":"2022-07-27T06:12:34.552276Z","shell.execute_reply.started":"2022-07-27T06:12:34.551976Z","shell.execute_reply":"2022-07-27T06:12:34.552004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.mixture import BayesianGaussianMixture\ngmm = BayesianGaussianMixture(n_components=7, covariance_type='full', random_state=1)\nsubmission[\"Predicted_BGMM\"] =  gmm.fit_predict(data_scaled)\nsubmission[\"GMM_BGMM_match\"] = submission['Predicted_GMM'] == submission['Predicted_BGMM']","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.554115Z","iopub.status.idle":"2022-07-27T06:12:34.554696Z","shell.execute_reply.started":"2022-07-27T06:12:34.554410Z","shell.execute_reply":"2022-07-27T06:12:34.554436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission[\"GMM_BGMM_match\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.555916Z","iopub.status.idle":"2022-07-27T06:12:34.556472Z","shell.execute_reply.started":"2022-07-27T06:12:34.556180Z","shell.execute_reply":"2022-07-27T06:12:34.556206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_clusters_ = (np.round(gmm.weights_, 2) > 0).sum()\nprint('Estimated number of clusters: ' + str(n_clusters_))","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.559430Z","iopub.status.idle":"2022-07-27T06:12:34.559996Z","shell.execute_reply.started":"2022-07-27T06:12:34.559715Z","shell.execute_reply":"2022-07-27T06:12:34.559741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.scatter(data_scaled.iloc[:, 7], data_scaled.iloc[:, 8], c=submission[\"Predicted_GMM\"], cmap=\"viridis\")\nplt.xlabel(\"Feature 1\")\nplt.ylabel(\"Feature 2\")","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.561426Z","iopub.status.idle":"2022-07-27T06:12:34.561994Z","shell.execute_reply.started":"2022-07-27T06:12:34.561709Z","shell.execute_reply":"2022-07-27T06:12:34.561737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Final predictions using GMM","metadata":{}},{"cell_type":"code","source":"submission[\"Predicted\"] = submission[\"Predicted_GMM\"]\nsubmission_gmm = submission.drop(columns=[\"Predicted_GMM\",\"GMM_BGMM_match\",\"Predicted_BGMM\"])\nsubmission_gmm.to_csv('submission_gmm.csv', index=False)\nsubmission_gmm","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.563738Z","iopub.status.idle":"2022-07-27T06:12:34.564336Z","shell.execute_reply.started":"2022-07-27T06:12:34.564014Z","shell.execute_reply":"2022-07-27T06:12:34.564051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission[\"Predicted\"] = submission[\"Predicted_BGMM\"]\nsubmission_bgmm = submission.drop(columns=[\"Predicted_GMM\",\"GMM_BGMM_match\",\"Predicted_BGMM\"])\nsubmission_bgmm.to_csv('submission_bgmm.csv', index=False)\nsubmission_bgmm","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.566587Z","iopub.status.idle":"2022-07-27T06:12:34.567183Z","shell.execute_reply.started":"2022-07-27T06:12:34.566874Z","shell.execute_reply":"2022-07-27T06:12:34.566916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px;font-family:sans-serif;\"> Principal Component Analysis (PCA) </h2>\n\nDoes dimensionality reduction with PCA help?\n\nAdjusted Rand Index = 0.5123 ","metadata":{}},{"cell_type":"code","source":"from sklearn.decomposition import PCA, KernelPCA\nimport plotly.graph_objs as go\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\n#pca = KernelPCA(n_components=2, kernel='rbf').fit(data_scaled)\npca = PCA().fit(data_scaled)\npcaratio = pca.explained_variance_ratio_\ntrace = go.Scatter(x=np.arange(len(pcaratio)),y=np.cumsum(pcaratio))\ndata = [trace]\nlayout = dict(title=\"Manufacturing control data - PCA Explained Variance | 80% variance explained by 20 features\")\nfig = dict(data=data, layout=layout)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.569352Z","iopub.status.idle":"2022-07-27T06:12:34.569917Z","shell.execute_reply.started":"2022-07-27T06:12:34.569637Z","shell.execute_reply":"2022-07-27T06:12:34.569663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from plotly.offline import init_notebook_mode, iplot\niplot(fig)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.572272Z","iopub.status.idle":"2022-07-27T06:12:34.572832Z","shell.execute_reply.started":"2022-07-27T06:12:34.572550Z","shell.execute_reply":"2022-07-27T06:12:34.572578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_scaled_pca = PCA(n_components='mle').fit_transform(data_scaled)\n\n# Convert to dataframe\ncomponent_names = [f\"PC{i+1}\" for i in range(data_scaled_pca.shape[1])]\ndata_scaled_pca = pd.DataFrame(data_scaled_pca, columns=component_names)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.577805Z","iopub.status.idle":"2022-07-27T06:12:34.578558Z","shell.execute_reply.started":"2022-07-27T06:12:34.578214Z","shell.execute_reply":"2022-07-27T06:12:34.578260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_scaled_pca","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.580272Z","iopub.status.idle":"2022-07-27T06:12:34.580895Z","shell.execute_reply.started":"2022-07-27T06:12:34.580584Z","shell.execute_reply":"2022-07-27T06:12:34.580613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.mixture import GaussianMixture\ngmm = GaussianMixture(n_components=7, covariance_type='full',  random_state=1) \nsubmission[\"Predicted\"] =  gmm.fit_predict(data_scaled_pca)","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.582860Z","iopub.status.idle":"2022-07-27T06:12:34.583912Z","shell.execute_reply.started":"2022-07-27T06:12:34.583585Z","shell.execute_reply":"2022-07-27T06:12:34.583618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission[\"Predicted\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.585221Z","iopub.status.idle":"2022-07-27T06:12:34.585817Z","shell.execute_reply.started":"2022-07-27T06:12:34.585529Z","shell.execute_reply":"2022-07-27T06:12:34.585558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission_gmm_pca.csv', index=False)\nsubmission","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.587439Z","iopub.status.idle":"2022-07-27T06:12:34.588015Z","shell.execute_reply.started":"2022-07-27T06:12:34.587732Z","shell.execute_reply":"2022-07-27T06:12:34.587758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px;font-family:sans-serif;\"> Competition Metric Understanding </h2>\n\nUnderstanding Adjusted rand index - Comparing submission from GMM and from PCA + GMM","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import adjusted_rand_score, adjusted_mutual_info_score\nsubmission_gmm = pd.read_csv(\"submission_gmm.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.590441Z","iopub.status.idle":"2022-07-27T06:12:34.591553Z","shell.execute_reply.started":"2022-07-27T06:12:34.591219Z","shell.execute_reply":"2022-07-27T06:12:34.591269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"adjusted_mutual_info_score(submission[\"Predicted\"],submission_gmm[\"Predicted\"])","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.593493Z","iopub.status.idle":"2022-07-27T06:12:34.594082Z","shell.execute_reply.started":"2022-07-27T06:12:34.593792Z","shell.execute_reply":"2022-07-27T06:12:34.593820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"adjusted_rand_score(submission[\"Predicted\"],submission_gmm[\"Predicted\"]) ","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.596229Z","iopub.status.idle":"2022-07-27T06:12:34.596881Z","shell.execute_reply.started":"2022-07-27T06:12:34.596577Z","shell.execute_reply":"2022-07-27T06:12:34.596606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"adjusted_rand_score([1, 1, 2, 2, 3],[1,1,2,2,2]) #less penalty - even if less clusters but more equal distribution","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.598536Z","iopub.status.idle":"2022-07-27T06:12:34.599118Z","shell.execute_reply.started":"2022-07-27T06:12:34.598820Z","shell.execute_reply":"2022-07-27T06:12:34.598848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"adjusted_rand_score([1, 1, 2, 2, 3],[1,2,2,2,3]) #more penalty - when less equal distribution between clusters","metadata":{"execution":{"iopub.status.busy":"2022-07-27T06:12:34.601732Z","iopub.status.idle":"2022-07-27T06:12:34.602354Z","shell.execute_reply.started":"2022-07-27T06:12:34.602041Z","shell.execute_reply":"2022-07-27T06:12:34.602070Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px;font-family:sans-serif;\">  Gaussian Mixture Models + Boosting </h2>\n\nThanks to https://www.kaggle.com/code/ricopue/tps-jul22-clusters-and-lgb/notebook for the idea.\n\nUsing GMM, we can get the probability that each data point belongs to each of the 7 clusters. Then for each row, we get the highest probability. Using the rows with high confidence (say p > 0.7), we use boosting to predict the ones that GMM predicted with lower confidence.\n","metadata":{}},{"cell_type":"code","source":"data = pd.read_csv(\"../input/tabular-playground-series-jul-2022/data.csv\", index_col='id')\nsubmission = pd.read_csv(\"../input/tabular-playground-series-jul-2022/sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:06:57.775979Z","iopub.execute_input":"2022-07-28T15:06:57.776567Z","iopub.status.idle":"2022-07-28T15:06:58.597857Z","shell.execute_reply.started":"2022-07-28T15:06:57.776519Z","shell.execute_reply":"2022-07-28T15:06:58.596402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.preprocessing import PowerTransformer\npower = PowerTransformer(method='yeo-johnson', standardize=False)\ndata_scaled = pd.DataFrame( power.fit_transform(data),columns = data.columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:06:58.600897Z","iopub.execute_input":"2022-07-28T15:06:58.601783Z","iopub.status.idle":"2022-07-28T15:07:02.049462Z","shell.execute_reply.started":"2022-07-28T15:06:58.601742Z","shell.execute_reply":"2022-07-28T15:07:02.047980Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_scaled = data_scaled[['f_07','f_08', 'f_09', 'f_10','f_11', 'f_12', 'f_13', 'f_22','f_23', 'f_24', 'f_25','f_26','f_27', 'f_28']]","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:07:02.051372Z","iopub.execute_input":"2022-07-28T15:07:02.053218Z","iopub.status.idle":"2022-07-28T15:07:02.065502Z","shell.execute_reply.started":"2022-07-28T15:07:02.053171Z","shell.execute_reply":"2022-07-28T15:07:02.064181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_scaled","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:07:02.069131Z","iopub.execute_input":"2022-07-28T15:07:02.069596Z","iopub.status.idle":"2022-07-28T15:07:02.104298Z","shell.execute_reply.started":"2022-07-28T15:07:02.069556Z","shell.execute_reply":"2022-07-28T15:07:02.102956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.mixture import BayesianGaussianMixture\nbgmm = BayesianGaussianMixture(n_components=7, covariance_type='full',  random_state=1) \nbgmm.fit(data_scaled)\npredict=bgmm.predict(data_scaled)\nprobability=bgmm.predict_proba(data_scaled)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:07:02.105880Z","iopub.execute_input":"2022-07-28T15:07:02.106319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_scaled['predict']=predict\ndata_scaled['predict_proba']=np.max(probability, axis=1) \n\ntrain_index=data_scaled[data_scaled.predict_proba >= 0.8].index\ntest_index=data_scaled[data_scaled.predict_proba < 0.8].index","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_scaled","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(round(len(train_index)/len(data_scaled),3))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X=data_scaled.loc[train_index][['f_07','f_08', 'f_09', 'f_10','f_11', 'f_12', 'f_13', 'f_22','f_23', 'f_24', 'f_25','f_26','f_27', 'f_28']]\ny=data_scaled.loc[train_index]['predict']","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test=data_scaled.loc[test_index][['f_07','f_08', 'f_09', 'f_10','f_11', 'f_12', 'f_13', 'f_22','f_23', 'f_24', 'f_25','f_26','f_27', 'f_28']]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_valid, y_train, y_valid = train_test_split(X, y, random_state=1)\nmodel = LGBMClassifier(n_estimators=500, \n                       random_state=42)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.fit(X_train, y_train)\n\n# # Preprocessing of validation data, get predictions\npreds = model.predict(X_valid)\n\n# # Evaluate the model\nscore = accuracy_score(y_valid, preds)\nprint('accuracy_score:', score)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predict=model.predict(X_test)\nprobabilities = model.predict_proba(X_test)\nprobability=[max(probabilities[i]) for i in tqdm(range(len(X_test)))] ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_scaled.at[test_index,'predict_proba'] = probability\ndata_scaled.at[test_index,'predict'] = predict","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.491710Z","iopub.status.idle":"2022-07-28T15:03:11.492617Z","shell.execute_reply.started":"2022-07-28T15:03:11.492304Z","shell.execute_reply":"2022-07-28T15:03:11.492335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sum(data_scaled.predict_proba >= 0.8)/len(data_scaled)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.494205Z","iopub.status.idle":"2022-07-28T15:03:11.495133Z","shell.execute_reply.started":"2022-07-28T15:03:11.494794Z","shell.execute_reply":"2022-07-28T15:03:11.494823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission[\"Predicted\"] = data_scaled[\"predict\"]\nsubmission.to_csv('submission_bgmm_lgbm.csv', index=False)\nsubmission","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.616979Z","iopub.execute_input":"2022-07-28T15:03:11.618064Z","iopub.status.idle":"2022-07-28T15:03:11.636280Z","shell.execute_reply.started":"2022-07-28T15:03:11.618019Z","shell.execute_reply":"2022-07-28T15:03:11.633114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission[\"Predicted\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.637277Z","iopub.status.idle":"2022-07-28T15:03:11.637856Z","shell.execute_reply.started":"2022-07-28T15:03:11.637565Z","shell.execute_reply":"2022-07-28T15:03:11.637601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px;font-family:sans-serif;\">  Gaussian Mixture Models + NN </h2>","metadata":{}},{"cell_type":"code","source":"X.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.640306Z","iopub.status.idle":"2022-07-28T15:03:11.641362Z","shell.execute_reply.started":"2022-07-28T15:03:11.641004Z","shell.execute_reply":"2022-07-28T15:03:11.641035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow import keras\nfrom tensorflow.keras import layers, callbacks\n\nearly_stopping = callbacks.EarlyStopping(\n    min_delta=0.001, # minimium amount of change to count as an improvement\n    patience=20, # how many epochs to wait before stopping\n    restore_best_weights=True,\n)\n\nmodel = keras.Sequential([\n    layers.Dense(1024, activation='relu', input_shape=[14]),\n    layers.Dropout(0.3),\n    layers.BatchNormalization(),\n    layers.Dense(1024, activation='relu'),\n    layers.Dropout(0.3),\n    layers.BatchNormalization(),\n    layers.Dense(1024, activation='relu'),\n    layers.Dropout(0.3),\n    layers.BatchNormalization(),\n    layers.Dense(7, activation=\"softmax\"),\n])\n\nmodel.compile(optimizer=keras.optimizers.Adam(learning_rate=0.003),\n                loss=\"sparse_categorical_crossentropy\",\n                metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.643274Z","iopub.status.idle":"2022-07-28T15:03:11.644293Z","shell.execute_reply.started":"2022-07-28T15:03:11.643948Z","shell.execute_reply":"2022-07-28T15:03:11.643979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learning_rate = 0.003\nepochs = 50\nbatch_size = 1000\nvalidation_split = 0.2","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.646260Z","iopub.status.idle":"2022-07-28T15:03:11.647268Z","shell.execute_reply.started":"2022-07-28T15:03:11.646928Z","shell.execute_reply":"2022-07-28T15:03:11.646959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(x=X, y=y, batch_size=batch_size,\n                      epochs=epochs, shuffle=True, callbacks=[early_stopping], # put your callbacks in a list\n                      validation_split=validation_split)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.649271Z","iopub.status.idle":"2022-07-28T15:03:11.650416Z","shell.execute_reply.started":"2022-07-28T15:03:11.650020Z","shell.execute_reply":"2022-07-28T15:03:11.650053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history_df = pd.DataFrame(history.history)\nhistory_df.loc[:, ['accuracy', 'val_accuracy']].plot();\nprint(\"Minimum validation loss: {}\".format(history_df['val_accuracy'].min()))","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.652281Z","iopub.status.idle":"2022-07-28T15:03:11.653317Z","shell.execute_reply.started":"2022-07-28T15:03:11.652997Z","shell.execute_reply":"2022-07-28T15:03:11.653042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_results = model.predict(X_test)\npredict= [np.argmax(test_results[i], axis=0) for i in tqdm(range(len(X_test)))]\nprobability= [max(test_results[i]) for i in tqdm(range(len(X_test)))]","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.655055Z","iopub.status.idle":"2022-07-28T15:03:11.656067Z","shell.execute_reply.started":"2022-07-28T15:03:11.655699Z","shell.execute_reply":"2022-07-28T15:03:11.655727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_scaled.at[test_index,'predict_proba'] = probability\ndata_scaled.at[test_index,'predict'] = predict","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.657937Z","iopub.status.idle":"2022-07-28T15:03:11.658961Z","shell.execute_reply.started":"2022-07-28T15:03:11.658604Z","shell.execute_reply":"2022-07-28T15:03:11.658634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_scaled","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.660768Z","iopub.status.idle":"2022-07-28T15:03:11.661828Z","shell.execute_reply.started":"2022-07-28T15:03:11.661501Z","shell.execute_reply":"2022-07-28T15:03:11.661533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sum(data_scaled.predict_proba >= 0.8)/len(data_scaled)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.663528Z","iopub.status.idle":"2022-07-28T15:03:11.664539Z","shell.execute_reply.started":"2022-07-28T15:03:11.664212Z","shell.execute_reply":"2022-07-28T15:03:11.664242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission[\"Predicted\"] = data_scaled[\"predict\"]\nsubmission.to_csv('submission_bgmm_nn.csv', index=False)\nsubmission","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.666393Z","iopub.status.idle":"2022-07-28T15:03:11.667524Z","shell.execute_reply.started":"2022-07-28T15:03:11.667147Z","shell.execute_reply":"2022-07-28T15:03:11.667177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission[\"Predicted\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.669438Z","iopub.status.idle":"2022-07-28T15:03:11.670573Z","shell.execute_reply.started":"2022-07-28T15:03:11.670245Z","shell.execute_reply":"2022-07-28T15:03:11.670275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission","metadata":{"execution":{"iopub.status.busy":"2022-07-28T15:03:11.672436Z","iopub.status.idle":"2022-07-28T15:03:11.673512Z","shell.execute_reply.started":"2022-07-28T15:03:11.673202Z","shell.execute_reply":"2022-07-28T15:03:11.673231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px;font-family:sans-serif;\"> Resources </h2>\n\n1. https://www.kaggle.com/code/minc33/visualizing-high-dimensional-clusters/notebook\n2. https://www.kaggle.com/code/fazilbtopal/popular-unsupervised-clustering-algorithms/notebook\n3. https://www.kaggle.com/code/sabanasimbutt/clustering-visualization-of-clusters-using-pca\n4. https://www.kaggle.com/code/athews/tpsjune22-eda-xgboost-model\n5. https://www.kaggle.com/code/nikhilkhetan/setting-up-evaluation-metrics/notebook \n","metadata":{}},{"cell_type":"markdown","source":"<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px;font-family:sans-serif;\"> To do next </h2>\n\n\n1. https://www.kaggle.com/code/thedevastator/how-to-ensemble-clustering-algorithms-updated/comments - Clustering Ensemble","metadata":{}}]}