{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <span style=\"color:#ffd514;\">**TPS Jul 22**<span>\n\n### <span style=\"color:#ffd514;\">*Table of content*<span>\n<a id=\"table-of-contents\"></a>\n- [1. Introduction](#1)\n    - [1.1 Evaluation](#1.1)\n- [2. Preparations](#2)\n- [3. Dataset Overview](#3)\n    - [3.1 Distribution of features](#3.1)\n    - [3.2 Shapiro test](#3.2)\n    - [3.3 Correlation](#3.3)\n    - [3.4 Scatter of categorical features](#3.4)\n- [4. K-Means](#4)\n    - [4.1 PCA for best transformation](#4.1)\n    - [4.2 UMAP](#4.2)\n\n- [5. Gaussian Mixture Models](#5)    \n- [6. References](#6)    \n    \n    \n    \n    \n    \n### <span style=\"color:#C70039;\">**Upvote and give me feedback if helpful!**<span>\n#### <span style=\"color:#C70039;\">**I suggest to checkout my other kernels**<span> \n[Spaceship Titanic](https://www.kaggle.com/code/kartushovdanil/deep-eda-to-infinity-and-beyond)  \n[TPS FEB 22](https://www.kaggle.com/code/kartushovdanil/tps-feb-22-advanced-eda)  \n[TPS JAN 21 PART 2](https://www.kaggle.com/code/kartushovdanil/overfit-me-catboost-lgbm)  \n[TPS JAN 21 PART 1](https://www.kaggle.com/code/kartushovdanil/tps-jan-22-eda-atboost-prophet)  \n[TPS NOV 21](https://www.kaggle.com/code/kartushovdanil/tps-nov21-4-begginers-rus-links)  \n ","metadata":{}},{"cell_type":"code","source":"from IPython.core.display import display, HTML, Javascript\n\n# ----- Notebook Theme -----\ncolor_map = ['#f4a261', '#e8f6f3', '#d0ece7', '#a2d9ce', '#73c6b6', '#45b39d', \n                        '#16a085', '#138d75', '#117a65', '#0e6655', '#e76f51']\n\nprompt = color_map[-1]\nmain_color = color_map[0]\nstrong_main_color = color_map[1]\ncustom_colors = [strong_main_color, main_color]\n\ncss_file = ''' \n\ndiv #notebook {\nbackground-color: white;\nline-height: 20px;\n}\n\n#notebook-container {\n%s\nmargin-top: 2em;\npadding-top: 2em;\nborder-top: 4px solid %s; /* light orange */\n-webkit-box-shadow: 0px 0px 8px 2px rgba(224, 212, 226, 0.5); /* pink */\n    box-shadow: 0px 0px 8px 2px rgba(224, 212, 226, 0.5); /* pink */\n}\n\ndiv .input {\nmargin-bottom: 1em;\n}\n\n.rendered_html h1, .rendered_html h2, .rendered_html h3, .rendered_html h4, .rendered_html h5, .rendered_html h6 {\ncolor: %s; /* light orange */\nfont-weight: 600;\n}\n\ndiv.input_area {\nborder: none;\n    background-color: %s; /* rgba(229, 143, 101, 0.1); light orange [exactly #E58F65] */\n    border-top: 2px solid %s; /* light orange */\n}\n\ndiv.input_prompt {\ncolor: %s; /* light blue */\n}\n\ndiv.output_prompt {\ncolor: %s; /* strong orange */\n}\n\ndiv.cell.selected:before, div.cell.selected.jupyter-soft-selected:before {\nbackground: %s; /* light orange */\n}\n\ndiv.cell.selected, div.cell.selected.jupyter-soft-selected {\n    border-color: %s; /* light orange */\n}\n\n.edit_mode div.cell.selected:before {\nbackground: %s; /* light orange */\n}\n\n.edit_mode div.cell.selected {\nborder-color: %s; /* light orange */\n\n}\n'''\ndef to_rgb(h): \n    return tuple(int(h[i:i+2], 16) for i in [0, 2, 4])\n\nmain_color_rgba = 'rgba(%s, %s, %s, 0.1)' % (to_rgb(main_color[1:]))\nopen('notebook.css', 'w').write(css_file % ('width: 95%;', main_color, main_color, main_color_rgba, main_color,  main_color, prompt, main_color, main_color, main_color, main_color))\n\ndef nb(): \n    return HTML(\"<style>\" + open(\"notebook.css\", \"r\").read() + \"</style>\")\nnb()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"source_hidden":true}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[back to top](#table-of-contents)\n<a id=\"1\"></a>\n## **<span style=\"color:#ffd514;\">1. Introduction</span>**\nWelcome to Kaggle's first ever unsupervised clustering challenge!\n\nIn this challenge, you are given a dataset where each row belongs to a particular cluster. Your job is to predict the cluster each row belongs to. You are not given any training data, and you are not told how many clusters are found in the ground truth labels. \n\n[back to top](#table-of-contents)\n<a id=\"1.1\"></a>\n### **<span style=\"color:#ffd514;\"> 1.1 Evaluation </span>**\nGiven a set `S` of `n` elements, the division into classes $X=[X_1,X_2,\\ldots,X_r]$, and the resulting division into clusters $Y=[Y_1,Y_2,\\ldots,Y_s]$, matches between `X` and `Y` can be reflected in the contingency table $n_{ij}$, where each $n_{ij}$ denotes the number of objects in both $X_i$ and $Y_j$ : $n_{ij}=|Xi \\cap Yj|$.\n\n![](https://wikimedia.org/api/rest_v1/media/math/render/svg/fea9be4505450896e2386b9f2b88b0eb673a71cc)  \nThe base metrics is $\\text{Rand}$\nThe Rand index evaluates how many of those pairs of elements that were in the same class and those pairs of elements that were in different classes retained this state after clustering by the algorithm. It has a domain of definition from 0 to 1, where 1 is a complete match of clusters with given classes, and 0 is no match.  \n$\\text{Rand}=\\frac{TP+TN}{TP+TN+FP+FN}$  \n\n\nBut in this competition used `Adjusted Rand`, this metrics different `Rand` in that it can take negative values if $\\text{Index}<\\text{ExpectedIndex}$. \n\n$\\text{ARI}= \\frac{\\sum_{ij} \\begin{pmatrix} n_{ij}\\\\ 2 \\end{pmatrix} - [\\sum_{i} \\begin{pmatrix} a_{i}\\\\ 2 \\end{pmatrix} \\sum_{j} \\begin{pmatrix} b_{j}\\\\ 2 \\end{pmatrix}] / \\begin{pmatrix} n\\\\ 2 \\end{pmatrix}}{\\frac{1}{2} [\\sum_{i} \\begin{pmatrix} a_{i}\\\\ 2 \\end{pmatrix} + \\sum_{j} \\begin{pmatrix} b_{j}\\\\ 2 \\end{pmatrix}]- [\\sum_{i} \\begin{pmatrix} a_{i}\\\\ 2 \\end{pmatrix} \\sum_{j} \\begin{pmatrix} b_{j}\\\\ 2 \\end{pmatrix}] / \\begin{pmatrix} n\\\\ 2 \\end{pmatrix}}$\n\nwhere $n_{ij}$ , $a_{i}$ , $b_{j}$  are values from the contingency table. ","metadata":{}},{"cell_type":"markdown","source":"[back to top](#table-of-contents)\n<a id=\"2\"></a>\n### **<span style=\"color:#ffd514;\"> 2. Preparations </span>**\nPreparing packages and data that will be used in the analysis process. Packages that will be loaded are mainly for data manipulation, data visualization and modeling. There are datasets that are used in the analysis, they are data.csv and sample_sabmisions.csv dataset. The main use of data dataset is numerical features, that used in clasterisation. While sample submission file is used to informed participants on the expected submission for the competition. (to see the details, please expand)","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom matplotlib import pyplot as plt\nimport seaborn as sns\n\nfrom scipy.stats import norm\nfrom scipy.stats import shapiro\nfrom scipy.stats import skew\n\nfrom sklearn.cluster import KMeans\nfrom sklearn.mixture import GaussianMixture\nfrom sklearn.mixture import BayesianGaussianMixture\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nfrom termcolor import colored\n\n#!pip install umap-learn\nimport umap\n","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-05T07:30:45.787769Z","iopub.execute_input":"2022-07-05T07:30:45.78824Z","iopub.status.idle":"2022-07-05T07:30:45.796837Z","shell.execute_reply.started":"2022-07-05T07:30:45.788148Z","shell.execute_reply":"2022-07-05T07:30:45.79507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[back to top](#table-of-contents)\n<a id=\"3\"></a>\n### **<span style=\"color:#ffd514;\"> 3. Dataset Overview</span>**\n\n","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('../input/tabular-playground-series-jul-2022/data.csv', index_col='id')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-05T07:30:47.722584Z","iopub.execute_input":"2022-07-05T07:30:47.723252Z","iopub.status.idle":"2022-07-05T07:30:48.654749Z","shell.execute_reply.started":"2022-07-05T07:30:47.72321Z","shell.execute_reply":"2022-07-05T07:30:48.653618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-05T07:18:45.132877Z","iopub.execute_input":"2022-07-05T07:18:45.133289Z","iopub.status.idle":"2022-07-05T07:18:45.19129Z","shell.execute_reply.started":"2022-07-05T07:18:45.133252Z","shell.execute_reply":"2022-07-05T07:18:45.19017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[back to top](#table-of-contents)\n<a id=\"3.1\"></a>\n### **<span style=\"color:#ffd514;\"> 3.1 Distribution of features</span>**\n","metadata":{}},{"cell_type":"code","source":"plt.rcParams['figure.dpi'] = 600\nfig = plt.figure(figsize=(10, 10), facecolor='#f6f5f5')\ngs = fig.add_gridspec(5, 6)\ngs.update(wspace=0.3, hspace=0.3)\nbackground_color = '#f6f5f5'\nrun_no = 0\n\ncolormap = ['#1DBA94','#1C5ED2', '#FFC300', '#C70039']\n\nfor row in range(0, 5):\n    for col in range(0, 6):\n        locals()[\"ax\"+str(run_no)] = fig.add_subplot(gs[row, col])\n        locals()[\"ax\"+str(run_no)].set_facecolor(background_color)\n        for s in [\"top\",\"right\"]:\n            locals()[\"ax\"+str(run_no)].spines[s].set_visible(False)\n        run_no += 1  \n\n\nfeatures = list(df.columns)\n\nrun_no = 0\nfor col in features:\n\n    sns.kdeplot(ax=locals()[\"ax\"+str(run_no)], x=norm(loc=df[col].mean(), scale=np.std(df[col])).rvs(size=100000), zorder=2, alpha=0.9, fill=True, color='#0db887')\n    sns.kdeplot(ax=locals()[\"ax\"+str(run_no)], x=df[col], zorder=2, alpha=1, linewidth=1, color='#ffd514')    \n    \n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='x', zorder=0, color='#EEEEEE', linewidth=0.4)\n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='y', zorder=0, color='#EEEEEE', linewidth=0.4)\n    locals()[\"ax\"+str(run_no)].set_ylabel('')\n    locals()[\"ax\"+str(run_no)].set_xlabel(col, fontsize=4, fontweight='bold')\n    locals()[\"ax\"+str(run_no)].tick_params(labelsize=4, width=0.5)\n    locals()[\"ax\"+str(run_no)].xaxis.offsetText.set_fontsize(4)\n    locals()[\"ax\"+str(run_no)].yaxis.offsetText.set_fontsize(4)\n    #locals()[\"ax\"+str(run_no)].get_legend().remove()\n    \n    run_no += 1\n\nlocals()[\"ax\"+str(29)].remove()\nlocals()[\"ax\"+str(0)].legend(['Monte-Carlo', 'True'], fontsize=4, loc='lower center', facecolor=background_color, edgecolor=background_color)\n\nplt.text(-99, 1.8, 'Distribution of features', fontsize=10, fontweight='bold')\nplt.text(-99, 1.765, 'Modeling normal distribution with used Monte-Carlo method', fontsize=7)\n\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T12:59:32.459404Z","iopub.execute_input":"2022-07-03T12:59:32.459734Z","iopub.status.idle":"2022-07-03T13:00:05.111776Z","shell.execute_reply.started":"2022-07-03T12:59:32.459705Z","shell.execute_reply":"2022-07-03T13:00:05.109664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[back to top](#table-of-contents)\n<a id=\"3.2\"></a>\n### **<span style=\"color:#ffd514;\"> 3.2 Shapiro test</span>**\n","metadata":{}},{"cell_type":"code","source":"norm_distr_feat = []\nfor col in df.columns:\n    print(f'========== {col} ==========')\n    stat, p = shapiro(df[col]) # тест Шапиро-Уилк \n    print('Statistics=%.3f, p-value=%.3f' % (stat, p), )\n    alpha = 0.05\n    if p > alpha:\n        print(colored('Apply', 'green'), f' hypotesis that {col} has normal distribution')\n        norm_distr_feat.append(col);\n    else:\n        print(colored('Disapply', 'red'), f' hypotesis that {col} has normal distribution')  \n        \n        \n    print(f'M[xi]: {np.mean(df[col]):.4f} +- {np.std(df[col]):.4f}')\n    print(f'D[xi]: {np.var(df[col]):.4f}')\n    print(f'Min: {np.min(df[col]):.4f}')\n    print(f'Max: {np.max(df[col]):.4f}')\n    print(f'Skew: {skew(df[col]):.4f}')\n    print(' ')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T13:00:05.114503Z","iopub.execute_input":"2022-07-03T13:00:05.11488Z","iopub.status.idle":"2022-07-03T13:00:05.511042Z","shell.execute_reply.started":"2022-07-03T13:00:05.114847Z","shell.execute_reply":"2022-07-03T13:00:05.509815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[back to top](#table-of-contents)\n<a id=\"3.3\"></a>\n### **<span style=\"color:#ffd514;\"> 3.3 Correlation</span>**\n","metadata":{}},{"cell_type":"code","source":"train_corr = df.corr(method='pearson')\ntest_corr = df.corr(method='kendall')\ndiff_corr = df.corr(method='spearman')\n\nplt.rcParams['figure.dpi'] = 600\nfig = plt.figure(figsize=(10, 5), facecolor='#f6f5f5')\ngs = fig.add_gridspec(1, 3)\ngs.update(wspace=0.5, hspace=0)\n\nbackground_color = \"#f6f5f5\"\ncmap_train = sns.dark_palette('#fcd12a', as_cmap=True)\ncmap_test = sns.dark_palette('#287094', as_cmap=True)\ncmap_diff = sns.dark_palette('#ff69b4', as_cmap=True)\n\nmask = np.triu(np.ones_like(train_corr, dtype=bool))\n\nrun_no = 0\nfor row in range(0, 1):\n    for col in range(0, 3):\n        locals()[\"ax\"+str(run_no)] = fig.add_subplot(gs[row, col])\n        locals()[\"ax\"+str(run_no)].set_facecolor(background_color)\n        for s in [\"top\",\"right\"]:\n            locals()[\"ax\"+str(run_no)].spines[s].set_visible(False)\n        run_no += 1\n        \nsns.heatmap(train_corr, ax=ax0, cmap=cmap_train, square=True, mask=mask, linewidths=.5, linecolor='#f6f5f5', \n            cbar_kws={\"shrink\": .3}, annot_kws={\"fontsize\":4})\nax0.set_xlabel(col, fontsize=5, fontweight='bold')\nax0.tick_params(labelsize=5, width=0.5, length=1.5)\nax0.set_xlabel('Pearson correlation', fontsize=5, fontweight='bold')\ncax = plt.gcf().axes[-1]\ncax.tick_params(labelsize=5)\n\nsns.heatmap(test_corr, ax=ax1, cmap=cmap_test, square=True, mask=mask, linewidths=.5, linecolor='#f6f5f5', \n            cbar_kws={\"shrink\": .3}, annot_kws={\"fontsize\":4})\nax1.set_xlabel(col, fontsize=5, fontweight='bold')\nax1.tick_params(labelsize=5, width=0.5, length=1.5)\nax1.set_xlabel('Kendall correlations', fontsize=5, fontweight='bold')\ncax = plt.gcf().axes[-1]\ncax.tick_params(labelsize=5)\n\nsns.heatmap(diff_corr, ax=ax2, cmap=cmap_diff, square=True, mask=mask, linewidths=.5, linecolor='#f6f5f5', \n            cbar_kws={\"shrink\": .3}, annot_kws={\"fontsize\":4})\nax2.set_xlabel(col, fontsize=5, fontweight='bold')\nax2.tick_params(labelsize=5, width=0.5, length=1.5)\nax2.set_xlabel('Spearman correlations', fontsize=5, fontweight='bold')\ncax = plt.gcf().axes[-1]\ncax.tick_params(labelsize=5)\n\nax0.text(0, -3.4, 'Features Correlation', fontsize=10, fontweight='bold')\nax0.text(0, -1.4, 'Three different methods of correlation', fontsize=7)\n\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T13:00:05.512509Z","iopub.execute_input":"2022-07-03T13:00:05.512865Z","iopub.status.idle":"2022-07-03T13:00:19.934651Z","shell.execute_reply.started":"2022-07-03T13:00:05.512835Z","shell.execute_reply":"2022-07-03T13:00:19.932901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[back to top](#table-of-contents)\n<a id=\"3.4\"></a>\n### **<span style=\"color:#ffd514;\"> 3.4 Scatter of categorical features</span>**\n","metadata":{}},{"cell_type":"code","source":"features = [f'f_0{i}' for i in range(7,10)] + [f'f_{i}' for i in range(10,14)]\n\nplt.rcParams['figure.dpi'] = 600\nfig = plt.figure(figsize=(8, 5), facecolor='#f6f5f5')\ngs = fig.add_gridspec(2, 3)\ngs.update(wspace=0.2, hspace=0.25)\n\nbackground_color = \"#f6f5f5\"\ncmap = sns.light_palette('#fcd12a', as_cmap=True)\n\nrun_no = 0\nfor row in range(0, 2):\n    for col in range(0, 3):\n        locals()[\"ax\"+str(run_no)] = fig.add_subplot(gs[row, col])\n        locals()[\"ax\"+str(run_no)].set_facecolor(background_color)\n        for s in [\"top\",\"right\"]:\n            locals()[\"ax\"+str(run_no)].spines[s].set_visible(False)\n        run_no += 1\n\nrun_no = 0\nfor col in features:\n    locals()[\"ax\"+str(run_no)].hexbin(x=df[col], y=df['f_13'], gridsize=15, \n                                      cmap=cmap, zorder=2, facecolor='black', mincnt=1)\n    locals()[\"ax\"+str(run_no)].set_ylabel('f_13', fontsize=5, fontweight='bold')\n    locals()[\"ax\"+str(run_no)].set_xlabel(col, fontsize=5, fontweight='bold')\n    locals()[\"ax\"+str(run_no)].tick_params(labelsize=5, width=0.5, length=1.5)\n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='x', zorder=0, color='#EEEEEE', linewidth=0.7)\n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='y', zorder=0, color='#EEEEEE', linewidth=0.7)\n    run_no += 1\n    \n#ax7.remove()\n#ax8.remove()\nax0.text(0, 38, 'Scatter', fontsize=10, fontweight='bold')\nax0.text(0, 35, 'Hexabin plot between categorical features', fontsize=7)\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T13:00:19.93662Z","iopub.execute_input":"2022-07-03T13:00:19.937086Z","iopub.status.idle":"2022-07-03T13:00:21.643934Z","shell.execute_reply.started":"2022-07-03T13:00:19.937046Z","shell.execute_reply":"2022-07-03T13:00:21.642157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[back to top](#table-of-contents)\n<a id=\"4\"></a>\n### **<span style=\"color:#ffd514;\"> 4 K-Means</span>**\n","metadata":{}},{"cell_type":"code","source":"KM = KMeans(n_clusters=8)\npred_kmeans = KM.fit_predict(df)\npred_kmeans","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T16:49:03.364876Z","iopub.execute_input":"2022-07-03T16:49:03.365207Z","iopub.status.idle":"2022-07-03T16:49:08.695123Z","shell.execute_reply.started":"2022-07-03T16:49:03.365171Z","shell.execute_reply":"2022-07-03T16:49:08.694057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[back to top](#table-of-contents)\n<a id=\"4.1\"></a>\n### **<span style=\"color:#ffd514;\"> 4.1 PCA for best transformation</span>**\n","metadata":{}},{"cell_type":"code","source":"from sklearn.decomposition import PCA\nfrom sklearn.preprocessing import *\n\ndef apply_transformations(transformer=False):\n    data = df\n    if transformer:\n        transformer.fit(data)\n        data = transformer.transform(data)\n        \n    pcaData = PCA(n_components=20).fit(data)\n    transformedData = pcaData.fit_transform(data)\n    \n    varianceExplained = np.sum(29*pcaData.explained_variance_ratio_)\n    \n    fig,ax = plt.subplots(1,figsize=(20,8), facecolor='#f6f5f5')\n    ax.set_facecolor(background_color)\n    for s in [\"top\",\"right\"]:\n        ax.spines[s].set_visible(False)\n        \n    \n    ax.plot(100*pcaData.explained_variance_ratio_,'ms--')\n    ax.set_xlabel('Components', fontsize=14,fontweight='bold')\n    ax.set_ylabel('Percent variance explained', fontsize=14, fontweight='bold')\n    #ax.set_title('PCA scree plot', fontsize=24, loc='left', fontweight='bold')\n    ax.tick_params(labelsize=14, width=0.5, length=1.5)\n    ax.grid(which='major', axis='x', zorder=0, color='#EEEEEE', linewidth=0.7)\n    ax.grid(which='major', axis='y', zorder=0, color='#EEEEEE', linewidth=0.7)\n    #ax.text(-0.7, 29, 'PCA scree plot', fontsize=24, fontweight='bold')\n    ax.legend()\n    plt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-05T07:18:54.290228Z","iopub.execute_input":"2022-07-05T07:18:54.291013Z","iopub.status.idle":"2022-07-05T07:18:54.305231Z","shell.execute_reply.started":"2022-07-05T07:18:54.290973Z","shell.execute_reply":"2022-07-05T07:18:54.303946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **<span style=\"color:#ffd514;\"> No transformation</span>**\n","metadata":{}},{"cell_type":"code","source":"apply_transformations()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-05T07:18:55.709608Z","iopub.execute_input":"2022-07-05T07:18:55.710841Z","iopub.status.idle":"2022-07-05T07:18:58.770794Z","shell.execute_reply.started":"2022-07-05T07:18:55.710799Z","shell.execute_reply":"2022-07-05T07:18:58.769393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **<span style=\"color:#ffd514;\"> Standard Scaler</span>**\n","metadata":{}},{"cell_type":"code","source":"apply_transformations(StandardScaler())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-05T07:18:58.771691Z","iopub.status.idle":"2022-07-05T07:18:58.772386Z","shell.execute_reply.started":"2022-07-05T07:18:58.772166Z","shell.execute_reply":"2022-07-05T07:18:58.772191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **<span style=\"color:#ffd514;\"> Power Transformer</span>**\n","metadata":{}},{"cell_type":"code","source":"apply_transformations(PowerTransformer())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-05T07:18:58.773638Z","iopub.status.idle":"2022-07-05T07:18:58.774028Z","shell.execute_reply.started":"2022-07-05T07:18:58.77383Z","shell.execute_reply":"2022-07-05T07:18:58.773847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **<span style=\"color:#ffd514;\"> MinMax Scaler</span>**\n","metadata":{}},{"cell_type":"code","source":"apply_transformations(MinMaxScaler())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-05T07:18:58.775893Z","iopub.status.idle":"2022-07-05T07:18:58.776858Z","shell.execute_reply.started":"2022-07-05T07:18:58.776632Z","shell.execute_reply":"2022-07-05T07:18:58.776655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **<span style=\"color:#ffd514;\"> Binarizer</span>**\n","metadata":{}},{"cell_type":"code","source":"apply_transformations(Binarizer())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-05T07:18:58.841784Z","iopub.execute_input":"2022-07-05T07:18:58.842197Z","iopub.status.idle":"2022-07-05T07:19:01.343204Z","shell.execute_reply.started":"2022-07-05T07:18:58.842169Z","shell.execute_reply":"2022-07-05T07:19:01.341436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **<span style=\"color:#ffd514;\"> Max Abs Scaler</span>**\n","metadata":{}},{"cell_type":"code","source":"apply_transformations(MaxAbsScaler())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-05T07:19:01.344613Z","iopub.status.idle":"2022-07-05T07:19:01.345976Z","shell.execute_reply.started":"2022-07-05T07:19:01.345625Z","shell.execute_reply":"2022-07-05T07:19:01.345656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **<span style=\"color:#ffd514;\"> Quantile Transformation</span>**\n","metadata":{}},{"cell_type":"code","source":"apply_transformations(QuantileTransformer())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-05T07:19:01.347851Z","iopub.status.idle":"2022-07-05T07:19:01.348826Z","shell.execute_reply.started":"2022-07-05T07:19:01.348522Z","shell.execute_reply":"2022-07-05T07:19:01.348553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **<span style=\"color:#ffd514;\"> Ordinal Encoder</span>**\n","metadata":{}},{"cell_type":"code","source":"apply_transformations(OrdinalEncoder())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T13:01:07.904813Z","iopub.execute_input":"2022-07-03T13:01:07.905253Z","iopub.status.idle":"2022-07-03T13:01:16.47512Z","shell.execute_reply.started":"2022-07-03T13:01:07.905219Z","shell.execute_reply":"2022-07-03T13:01:16.473725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **<span style=\"color:#ffd514;\"> Normalizer</span>**\n","metadata":{}},{"cell_type":"code","source":"apply_transformations(Normalizer())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T13:01:16.476915Z","iopub.execute_input":"2022-07-03T13:01:16.477408Z","iopub.status.idle":"2022-07-03T13:01:21.997597Z","shell.execute_reply.started":"2022-07-03T13:01:16.477362Z","shell.execute_reply":"2022-07-03T13:01:21.995619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[back to top](#table-of-contents)\n<a id=\"4.2\"></a>\n### **<span style=\"color:#ffd514;\"> 4.2 UMAP</span>**\n","metadata":{}},{"cell_type":"markdown","source":"\n### **<span style=\"color:#ffd514;\"> K-Means</span>**\n","metadata":{}},{"cell_type":"code","source":"df_with_cluster = df.copy()\ndf_with_cluster['clusters'] = pred_kmeans\nfeatures = [col for col in df_with_cluster.columns if col not in ['clusters']]\ncluster_list = []\nfor cluster in set(pred_kmeans):\n    X = df_with_cluster[df_with_cluster['clusters'] == cluster]\n    reducer = umap.UMAP(random_state=42, n_components=2, low_memory=True)\n    cluster_list.append(reducer.fit_transform(X[features]))\n    ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T13:01:21.999494Z","iopub.execute_input":"2022-07-03T13:01:22.000505Z","iopub.status.idle":"2022-07-03T13:03:28.682665Z","shell.execute_reply.started":"2022-07-03T13:01:22.00046Z","shell.execute_reply":"2022-07-03T13:03:28.681333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.rcParams['figure.dpi'] = 600\nfig = plt.figure(figsize=(10, 5), facecolor='#f6f5f5')\ngs = fig.add_gridspec(1, 1)\ngs.update(wspace=0.5, hspace=0)\n\nbackground_color = \"#f6f5f5\"\n\nax = fig.add_subplot(gs[0, 0])\nax.set_facecolor(background_color)\nfor s in [\"top\",\"right\"]:\n    ax.spines[s].set_visible(False)\n\ncolormap = ['#1DBA94','#1C5ED2', '#FFC300', '#C70039', '#42f5e0', '#54f542', '#f5e342', '#a742f5']\nrun_no = 0\nfor cluster in cluster_list:\n    x = [val[0] for val in cluster]\n    y = [val[1] for val in cluster]\n    sns.scatterplot(ax=ax, x=x, y=y, zorder=2, linewidth=0, color=colormap[run_no])\n    run_no += 1\n\nax.set_ylabel('PC1', fontsize=5, fontweight='bold')\nax.set_xlabel('PC2', fontsize=5, fontweight='bold')\nax.tick_params(labelsize=5, width=0.5, length=1.5)\nax.grid(which='major', axis='x', zorder=0, color='#EEEEEE', linewidth=0.7)\nax.grid(which='major', axis='y', zorder=0, color='#EEEEEE', linewidth=0.7)\n\nax.legend([f'Cluster {i}' for i in range(len(cluster_list))], ncol=4, facecolor=background_color, edgecolor=background_color, loc='lower center');","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T13:03:28.684558Z","iopub.execute_input":"2022-07-03T13:03:28.685033Z","iopub.status.idle":"2022-07-03T13:03:31.944948Z","shell.execute_reply.started":"2022-07-03T13:03:28.684988Z","shell.execute_reply":"2022-07-03T13:03:31.94308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from cycler import cycler\n\nplt.rcParams['figure.dpi'] = 600\nfig = plt.figure(figsize=(10, 10), facecolor='#f6f5f5')\ngs = fig.add_gridspec(5, 6)\ngs.update(wspace=0.3, hspace=0.3)\nbackground_color = '#f6f5f5'\nrun_no = 0\n\ncolormap = ['#1DBA94','#1C5ED2', '#FFC300', '#C70039', '#42f5e0', '#54f542', '#f5e342', '#a742f5']\nplt.rc('axes', prop_cycle=(cycler('color', colormap)))\n\nfor row in range(0, 5):\n    for col in range(0, 6):\n        locals()[\"ax\"+str(run_no)] = fig.add_subplot(gs[row, col])\n        locals()[\"ax\"+str(run_no)].set_facecolor(background_color)\n        for s in [\"top\",\"right\"]:\n            locals()[\"ax\"+str(run_no)].spines[s].set_visible(False)\n        run_no += 1  \n\n\nfeatures = list(df.columns)\n\nrun_no = 0\nfor col in features:\n\n    #sns.kdeplot(ax=locals()[\"ax\"+str(run_no)], x=norm(loc=df[col].mean(), scale=np.std(df[col])).rvs(size=100000), zorder=2, alpha=0.9, fill=True, color='#0db887')\n    sns.kdeplot(ax=locals()[\"ax\"+str(run_no)], x=df[col], zorder=2, alpha=1, linewidth=1, palette=colormap, hue=pred_kmeans)  \n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='x', zorder=0, color='#EEEEEE', linewidth=0.4)\n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='y', zorder=0, color='#EEEEEE', linewidth=0.4)\n    locals()[\"ax\"+str(run_no)].set_ylabel('')\n    locals()[\"ax\"+str(run_no)].set_xlabel(col, fontsize=4, fontweight='bold')\n    locals()[\"ax\"+str(run_no)].tick_params(labelsize=4, width=0.5)\n    locals()[\"ax\"+str(run_no)].xaxis.offsetText.set_fontsize(4)\n    locals()[\"ax\"+str(run_no)].yaxis.offsetText.set_fontsize(4)\n    locals()[\"ax\"+str(run_no)].get_legend().remove()\n    \n    run_no += 1\n\nlocals()[\"ax\"+str(29)].remove()\nlocals()[\"ax\"+str(0)].legend([f'Cluster {i}' for i in set(pred_kmeans)], ncol=2, fontsize=2, loc='lower center', facecolor=background_color, edgecolor=background_color)\n\n#plt.text(-99, 1.8, 'Distribution of features', fontsize=10, fontweight='bold')\n#plt.text(-99, 1.765, 'Modeling normal distribution with used Monte-Carlo method', fontsize=7)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-03T13:11:38.598894Z","iopub.execute_input":"2022-07-03T13:11:38.59929Z","iopub.status.idle":"2022-07-03T13:12:04.670184Z","shell.execute_reply.started":"2022-07-03T13:11:38.599258Z","shell.execute_reply":"2022-07-03T13:12:04.668279Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style=\"color:#ffd514;\"> K-means with Standart Scaler</span>**\n","metadata":{}},{"cell_type":"code","source":"transformer = StandardScaler().fit(df)\nss_df = transformer.transform(df)\nKM = KMeans(n_clusters=8)\nss_pred_kmeans = KM.fit_predict(ss_df)\nss_pred_kmeans","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T16:50:49.231102Z","iopub.execute_input":"2022-07-03T16:50:49.231564Z","iopub.status.idle":"2022-07-03T16:50:59.784603Z","shell.execute_reply.started":"2022-07-03T16:50:49.231512Z","shell.execute_reply":"2022-07-03T16:50:59.783308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_with_cluster = df.copy()\ndf_with_cluster['clusters'] = ss_pred_kmeans\nfeatures = [col for col in df_with_cluster.columns if col not in ['clusters']]\ncluster_list = []\nfor cluster in set(ss_pred_kmeans):\n    X = df_with_cluster[df_with_cluster['clusters'] == cluster]\n    reducer = umap.UMAP(random_state=42, n_components=2, low_memory=True)\n    cluster_list.append(reducer.fit_transform(X[features]))\n    ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T13:16:59.586548Z","iopub.execute_input":"2022-07-03T13:16:59.586943Z","iopub.status.idle":"2022-07-03T13:18:25.691747Z","shell.execute_reply.started":"2022-07-03T13:16:59.586911Z","shell.execute_reply":"2022-07-03T13:18:25.690742Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.rcParams['figure.dpi'] = 600\nfig = plt.figure(figsize=(10, 5), facecolor='#f6f5f5')\ngs = fig.add_gridspec(1, 1)\ngs.update(wspace=0.5, hspace=0)\n\nbackground_color = \"#f6f5f5\"\n\nax = fig.add_subplot(gs[0, 0])\nax.set_facecolor(background_color)\nfor s in [\"top\",\"right\"]:\n    ax.spines[s].set_visible(False)\n\ncolormap = ['#1DBA94','#1C5ED2', '#FFC300', '#C70039', '#42f5e0', '#54f542', '#f5e342', '#a742f5']\nrun_no = 0\nfor cluster in cluster_list:\n    x = [val[0] for val in cluster]\n    y = [val[1] for val in cluster]\n    sns.scatterplot(ax=ax, x=x, y=y, zorder=2, linewidth=0, color=colormap[run_no])\n    run_no += 1\n\nax.set_ylabel('PC1', fontsize=5, fontweight='bold')\nax.set_xlabel('PC2', fontsize=5, fontweight='bold')\nax.tick_params(labelsize=5, width=0.5, length=1.5)\nax.grid(which='major', axis='x', zorder=0, color='#EEEEEE', linewidth=0.7)\nax.grid(which='major', axis='y', zorder=0, color='#EEEEEE', linewidth=0.7)\nax.legend([f'Cluster {i}' for i in range(len(cluster_list))], ncol=4, facecolor=background_color, edgecolor=background_color, loc='lower center');","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T13:18:25.693279Z","iopub.execute_input":"2022-07-03T13:18:25.693912Z","iopub.status.idle":"2022-07-03T13:18:28.928833Z","shell.execute_reply.started":"2022-07-03T13:18:25.693869Z","shell.execute_reply":"2022-07-03T13:18:28.927369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ss_df = pd.DataFrame(ss_df, columns=df.columns)\n\nplt.rcParams['figure.dpi'] = 600\nfig = plt.figure(figsize=(10, 10), facecolor='#f6f5f5')\ngs = fig.add_gridspec(5, 6)\ngs.update(wspace=0.3, hspace=0.3)\nbackground_color = '#f6f5f5'\nrun_no = 0\n\ncolormap = ['#1DBA94','#1C5ED2', '#FFC300', '#C70039', '#42f5e0', '#54f542', '#f5e342', '#a742f5']\nplt.rc('axes', prop_cycle=(cycler('color', colormap)))\n\nfor row in range(0, 5):\n    for col in range(0, 6):\n        locals()[\"ax\"+str(run_no)] = fig.add_subplot(gs[row, col])\n        locals()[\"ax\"+str(run_no)].set_facecolor(background_color)\n        for s in [\"top\",\"right\"]:\n            locals()[\"ax\"+str(run_no)].spines[s].set_visible(False)\n        run_no += 1  \n\n\nfeatures = list(df.columns)\n\nrun_no = 0\nfor col in features:\n\n    #sns.kdeplot(ax=locals()[\"ax\"+str(run_no)], x=norm(loc=df[col].mean(), scale=np.std(df[col])).rvs(size=100000), zorder=2, alpha=0.9, fill=True, color='#0db887')\n    sns.kdeplot(ax=locals()[\"ax\"+str(run_no)], x=ss_df[col], zorder=2, alpha=1, linewidth=1, palette=colormap, hue=ss_pred_kmeans)  \n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='x', zorder=0, color='#EEEEEE', linewidth=0.4)\n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='y', zorder=0, color='#EEEEEE', linewidth=0.4)\n    locals()[\"ax\"+str(run_no)].set_ylabel('')\n    locals()[\"ax\"+str(run_no)].set_xlabel(col, fontsize=4, fontweight='bold')\n    locals()[\"ax\"+str(run_no)].tick_params(labelsize=4, width=0.5)\n    locals()[\"ax\"+str(run_no)].xaxis.offsetText.set_fontsize(4)\n    locals()[\"ax\"+str(run_no)].yaxis.offsetText.set_fontsize(4)\n    locals()[\"ax\"+str(run_no)].get_legend().remove()\n    \n    run_no += 1\n\nlocals()[\"ax\"+str(29)].remove()\nlocals()[\"ax\"+str(0)].legend([f'Cluster {i}' for i in set(ss_pred_kmeans)], ncol=2, fontsize=2, loc='lower center', facecolor=background_color, edgecolor=background_color)\n\n#plt.text(-99, 1.8, 'Distribution of features', fontsize=10, fontweight='bold')\n#plt.text(-99, 1.765, 'Modeling normal distribution with used Monte-Carlo method', fontsize=7)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-03T13:20:25.690766Z","iopub.execute_input":"2022-07-03T13:20:25.691168Z","iopub.status.idle":"2022-07-03T13:20:51.641516Z","shell.execute_reply.started":"2022-07-03T13:20:25.691139Z","shell.execute_reply":"2022-07-03T13:20:51.640382Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **<span style=\"color:#ffd514;\"> K-means with Max Abs Scaler</span>**\n","metadata":{}},{"cell_type":"code","source":"transformer = MaxAbsScaler().fit(df)\nmas_df = transformer.transform(df)\nKM = KMeans(n_clusters=8)\nmas_pred_kmeans = KM.fit_predict(mas_df)\nmas_pred_kmeans","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-04T09:35:16.472682Z","iopub.execute_input":"2022-07-04T09:35:16.473185Z","iopub.status.idle":"2022-07-04T09:35:16.551283Z","shell.execute_reply.started":"2022-07-04T09:35:16.473144Z","shell.execute_reply":"2022-07-04T09:35:16.549933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_with_cluster = df.copy()\ndf_with_cluster['clusters'] = mas_pred_kmeans\nfeatures = [col for col in df_with_cluster.columns if col not in ['clusters']]\ncluster_list = []\nfor cluster in set(ss_pred_kmeans):\n    X = df_with_cluster[df_with_cluster['clusters'] == cluster]\n    reducer = umap.UMAP(random_state=42, n_components=2, low_memory=True)\n    cluster_list.append(reducer.fit_transform(X[features]))\n    ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T13:26:39.68535Z","iopub.execute_input":"2022-07-03T13:26:39.685902Z","iopub.status.idle":"2022-07-03T13:28:13.967209Z","shell.execute_reply.started":"2022-07-03T13:26:39.685851Z","shell.execute_reply":"2022-07-03T13:28:13.965665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.rcParams['figure.dpi'] = 600\nfig = plt.figure(figsize=(10, 5), facecolor='#f6f5f5')\ngs = fig.add_gridspec(1, 1)\ngs.update(wspace=0.5, hspace=0)\n\nbackground_color = \"#f6f5f5\"\n\nax = fig.add_subplot(gs[0, 0])\nax.set_facecolor(background_color)\nfor s in [\"top\",\"right\"]:\n    ax.spines[s].set_visible(False)\n\ncolormap = ['#1DBA94','#1C5ED2', '#FFC300', '#C70039', '#42f5e0', '#54f542', '#f5e342', '#a742f5']\nrun_no = 0\nfor cluster in cluster_list:\n    x = [val[0] for val in cluster]\n    y = [val[1] for val in cluster]\n    sns.scatterplot(ax=ax, x=x, y=y, zorder=2, linewidth=0, color=colormap[run_no])\n    run_no += 1\n\nax.set_ylabel('PC1', fontsize=5, fontweight='bold')\nax.set_xlabel('PC2', fontsize=5, fontweight='bold')\nax.tick_params(labelsize=5, width=0.5, length=1.5)\nax.grid(which='major', axis='x', zorder=0, color='#EEEEEE', linewidth=0.7)\nax.grid(which='major', axis='y', zorder=0, color='#EEEEEE', linewidth=0.7)\nax.legend([f'Cluster {i}' for i in range(len(cluster_list))], ncol=4, facecolor=background_color, edgecolor=background_color, loc='lower center');","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T13:28:13.968844Z","iopub.execute_input":"2022-07-03T13:28:13.969171Z","iopub.status.idle":"2022-07-03T13:28:17.189102Z","shell.execute_reply.started":"2022-07-03T13:28:13.969142Z","shell.execute_reply":"2022-07-03T13:28:17.187657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mas_df = pd.DataFrame(mas_df, columns=df.columns)\n\nplt.rcParams['figure.dpi'] = 600\nfig = plt.figure(figsize=(10, 10), facecolor='#f6f5f5')\ngs = fig.add_gridspec(5, 6)\ngs.update(wspace=0.3, hspace=0.3)\nbackground_color = '#f6f5f5'\nrun_no = 0\n\ncolormap = ['#1DBA94','#1C5ED2', '#FFC300', '#C70039', '#42f5e0', '#54f542', '#f5e342', '#a742f5']\nplt.rc('axes', prop_cycle=(cycler('color', colormap)))\n\nfor row in range(0, 5):\n    for col in range(0, 6):\n        locals()[\"ax\"+str(run_no)] = fig.add_subplot(gs[row, col])\n        locals()[\"ax\"+str(run_no)].set_facecolor(background_color)\n        for s in [\"top\",\"right\"]:\n            locals()[\"ax\"+str(run_no)].spines[s].set_visible(False)\n        run_no += 1  \n\n\nfeatures = list(df.columns)\n\nrun_no = 0\nfor col in features:\n\n    #sns.kdeplot(ax=locals()[\"ax\"+str(run_no)], x=norm(loc=df[col].mean(), scale=np.std(df[col])).rvs(size=100000), zorder=2, alpha=0.9, fill=True, color='#0db887')\n    sns.kdeplot(ax=locals()[\"ax\"+str(run_no)], x=mas_df[col], zorder=2, alpha=1, linewidth=1, palette=colormap, hue=mas_pred_kmeans)  \n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='x', zorder=0, color='#EEEEEE', linewidth=0.4)\n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='y', zorder=0, color='#EEEEEE', linewidth=0.4)\n    locals()[\"ax\"+str(run_no)].set_ylabel('')\n    locals()[\"ax\"+str(run_no)].set_xlabel(col, fontsize=4, fontweight='bold')\n    locals()[\"ax\"+str(run_no)].tick_params(labelsize=4, width=0.5)\n    locals()[\"ax\"+str(run_no)].xaxis.offsetText.set_fontsize(4)\n    locals()[\"ax\"+str(run_no)].yaxis.offsetText.set_fontsize(4)\n    locals()[\"ax\"+str(run_no)].get_legend().remove()\n    \n    run_no += 1\n\nlocals()[\"ax\"+str(29)].remove()\nlocals()[\"ax\"+str(0)].legend([f'Cluster {i}' for i in set(mas_pred_kmeans)], ncol=2, fontsize=2, loc='lower center', facecolor=background_color, edgecolor=background_color)\n\n#plt.text(-99, 1.8, 'Distribution of features', fontsize=10, fontweight='bold')\n#plt.text(-99, 1.765, 'Modeling normal distribution with used Monte-Carlo method', fontsize=7)\n\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2022-07-03T13:28:17.192098Z","iopub.execute_input":"2022-07-03T13:28:17.192636Z","iopub.status.idle":"2022-07-03T13:28:44.057828Z","shell.execute_reply.started":"2022-07-03T13:28:17.192588Z","shell.execute_reply":"2022-07-03T13:28:44.055663Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[back to top](#table-of-contents)\n<a id=\"5\"></a>\n### **<span style=\"color:#ffd514;\"> 5 Gaussian Mixture Models</span>**\n#### **<span style=\"color:#ffd514;\"> With out transformation</span>**\n","metadata":{}},{"cell_type":"code","source":"n_components = np.arange(1, 21)\nmodels = [GaussianMixture(n, covariance_type='full', random_state=0).fit(df) for n in n_components]\n","metadata":{"execution":{"iopub.status.busy":"2022-07-03T13:54:47.633932Z","iopub.execute_input":"2022-07-03T13:54:47.634481Z","iopub.status.idle":"2022-07-03T14:25:32.992975Z","shell.execute_reply.started":"2022-07-03T13:54:47.634434Z","shell.execute_reply":"2022-07-03T14:25:32.99117Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.rcParams['figure.dpi'] = 600\nfig = plt.figure(figsize=(10, 4), facecolor='#f6f5f5')\ngs = fig.add_gridspec(1, 1)\ngs.update(wspace=0.3, hspace=0.3)\nbackground_color = '#f6f5f5'\n\nax = fig.add_subplot(gs[0, 0])\nax.set_facecolor(background_color)\nfor s in [\"top\",\"right\"]:\n    ax.spines[s].set_visible(False)\n\nax.plot(n_components, [m.bic(df) for m in models], label='BIC')\nax.plot(n_components, [m.aic(df) for m in models], label='AIC')\n\nax.grid(which='major', axis='x', zorder=0, color='#EEEEEE', linewidth=0.4)\nax.grid(which='major', axis='y', zorder=0, color='#EEEEEE', linewidth=0.4)\nax.set_ylabel('BIC/AIC', fontsize=6, fontweight='bold')\nax.set_xlabel('n components', fontsize=6, fontweight='bold')\nax.tick_params(labelsize=6, width=0.5)\nax.xaxis.offsetText.set_fontsize(6)\nax.yaxis.offsetText.set_fontsize(6)\n\n\nplt.legend(ncol=2, fontsize=6, loc='upper center', facecolor=background_color, edgecolor=background_color)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-03T14:43:03.05732Z","iopub.execute_input":"2022-07-03T14:43:03.057825Z","iopub.status.idle":"2022-07-03T14:43:56.646586Z","shell.execute_reply.started":"2022-07-03T14:43:03.057784Z","shell.execute_reply":"2022-07-03T14:43:56.64533Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gmm = GaussianMixture(n_components=7)\npred_gmm = gmm.fit_predict(df)\npred_gmm\n\n\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T16:49:19.962066Z","iopub.execute_input":"2022-07-03T16:49:19.963377Z","iopub.status.idle":"2022-07-03T16:50:05.29464Z","shell.execute_reply.started":"2022-07-03T16:49:19.963326Z","shell.execute_reply":"2022-07-03T16:50:05.293508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_with_cluster = df.copy()\ndf_with_cluster['clusters'] = pred_gmm\nfeatures = [col for col in df_with_cluster.columns if col not in ['clusters']]\ncluster_list = []\nfor cluster in set(pred_gmm):\n    X = df_with_cluster[df_with_cluster['clusters'] == cluster]\n    reducer = umap.UMAP(random_state=42, n_components=2, low_memory=True)\n    cluster_list.append(reducer.fit_transform(X[features]))\n    ","metadata":{"execution":{"iopub.status.busy":"2022-07-03T14:48:20.391078Z","iopub.execute_input":"2022-07-03T14:48:20.391522Z","iopub.status.idle":"2022-07-03T14:52:02.324013Z","shell.execute_reply.started":"2022-07-03T14:48:20.391487Z","shell.execute_reply":"2022-07-03T14:52:02.32277Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.rcParams['figure.dpi'] = 600\nfig = plt.figure(figsize=(10, 5), facecolor='#f6f5f5')\ngs = fig.add_gridspec(1, 1)\ngs.update(wspace=0.5, hspace=0)\n\nbackground_color = \"#f6f5f5\"\n\nax = fig.add_subplot(gs[0, 0])\nax.set_facecolor(background_color)\nfor s in [\"top\",\"right\"]:\n    ax.spines[s].set_visible(False)\n\ncolormap = ['#0634e1', '#6bae9d', '#763085', '#e1bd4c', '#467acf', '#f6242b', '#0d7cdd']\nrun_no = 0\nfor cluster in cluster_list:\n    x = [val[0] for val in cluster]\n    y = [val[1] for val in cluster]\n    sns.scatterplot(ax=ax, x=x, y=y, zorder=2, linewidth=0, color=colormap[run_no])\n    run_no += 1\n\nax.set_ylabel('PC1', fontsize=5, fontweight='bold')\nax.set_xlabel('PC2', fontsize=5, fontweight='bold')\nax.tick_params(labelsize=5, width=0.5, length=1.5)\nax.grid(which='major', axis='x', zorder=0, color='#EEEEEE', linewidth=0.7)\nax.grid(which='major', axis='y', zorder=0, color='#EEEEEE', linewidth=0.7)\nax.legend([f'Cluster {i}' for i in range(len(cluster_list))], ncol=4, facecolor=background_color, edgecolor=background_color, loc='lower center');","metadata":{"execution":{"iopub.status.busy":"2022-07-03T14:55:51.328856Z","iopub.execute_input":"2022-07-03T14:55:51.329367Z","iopub.status.idle":"2022-07-03T14:55:54.954608Z","shell.execute_reply.started":"2022-07-03T14:55:51.329323Z","shell.execute_reply":"2022-07-03T14:55:54.95333Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.rcParams['figure.dpi'] = 600\nfig = plt.figure(figsize=(10, 10), facecolor='#f6f5f5')\ngs = fig.add_gridspec(5, 6)\ngs.update(wspace=0.3, hspace=0.3)\nbackground_color = '#f6f5f5'\nrun_no = 0\n\nplt.rc('axes', prop_cycle=(cycler('color', colormap)))\n\nfor row in range(0, 5):\n    for col in range(0, 6):\n        locals()[\"ax\"+str(run_no)] = fig.add_subplot(gs[row, col])\n        locals()[\"ax\"+str(run_no)].set_facecolor(background_color)\n        for s in [\"top\",\"right\"]:\n            locals()[\"ax\"+str(run_no)].spines[s].set_visible(False)\n        run_no += 1  \n\n\nfeatures = list(df.columns)\n\nrun_no = 0\nfor col in features:\n\n    #sns.kdeplot(ax=locals()[\"ax\"+str(run_no)], x=norm(loc=df[col].mean(), scale=np.std(df[col])).rvs(size=100000), zorder=2, alpha=0.9, fill=True, color='#0db887')\n    sns.kdeplot(ax=locals()[\"ax\"+str(run_no)], x=df[col], zorder=2, alpha=1, linewidth=1, palette=colormap, hue=pred_gmm)  \n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='x', zorder=0, color='#EEEEEE', linewidth=0.4)\n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='y', zorder=0, color='#EEEEEE', linewidth=0.4)\n    locals()[\"ax\"+str(run_no)].set_ylabel('')\n    locals()[\"ax\"+str(run_no)].set_xlabel(col, fontsize=4, fontweight='bold')\n    locals()[\"ax\"+str(run_no)].tick_params(labelsize=4, width=0.5)\n    locals()[\"ax\"+str(run_no)].xaxis.offsetText.set_fontsize(4)\n    locals()[\"ax\"+str(run_no)].yaxis.offsetText.set_fontsize(4)\n    locals()[\"ax\"+str(run_no)].get_legend().remove()\n    \n    run_no += 1\n\nlocals()[\"ax\"+str(29)].remove()\nlocals()[\"ax\"+str(0)].legend([f'Cluster {i}' for i in set(pred_gmm)], ncol=2, fontsize=2, loc='lower center', facecolor=background_color, edgecolor=background_color)\n\n#plt.text(-99, 1.8, 'Distribution of features', fontsize=10, fontweight='bold')\n#plt.text(-99, 1.765, 'Modeling normal distribution with used Monte-Carlo method', fontsize=7)\n\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2022-07-03T14:56:08.834413Z","iopub.execute_input":"2022-07-03T14:56:08.835132Z","iopub.status.idle":"2022-07-03T14:56:34.176397Z","shell.execute_reply.started":"2022-07-03T14:56:08.835092Z","shell.execute_reply":"2022-07-03T14:56:34.175049Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **<span style=\"color:#ffd514;\"> Standart Scaler</span>**\n","metadata":{}},{"cell_type":"code","source":"#features = [col for col in df.columns if col not in ['clusters']]\n#ss_df = pd.DataFrame(ss_df, columns=df.columns)\ngmm = GaussianMixture(n_components=7)\nss_pred_gmm = gmm.fit_predict(ss_df[df.columns])\nss_pred_gmm\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T17:02:44.045559Z","iopub.execute_input":"2022-07-03T17:02:44.046007Z","iopub.status.idle":"2022-07-03T17:02:58.830251Z","shell.execute_reply.started":"2022-07-03T17:02:44.045971Z","shell.execute_reply":"2022-07-03T17:02:58.828913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_with_cluster = df.copy()\ndf_with_cluster['clusters'] = ss_pred_gmm\nfeatures = [col for col in df_with_cluster.columns if col not in ['clusters']]\ncluster_list = []\nfor cluster in set(ss_pred_gmm):\n    X = df_with_cluster[df_with_cluster['clusters'] == cluster]\n    reducer = umap.UMAP(random_state=42, n_components=2, low_memory=True)\n    cluster_list.append(reducer.fit_transform(X[features]))\n    ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T16:51:30.24477Z","iopub.execute_input":"2022-07-03T16:51:30.245197Z","iopub.status.idle":"2022-07-03T16:51:30.282362Z","shell.execute_reply.started":"2022-07-03T16:51:30.245163Z","shell.execute_reply":"2022-07-03T16:51:30.280903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.rcParams['figure.dpi'] = 600\nfig = plt.figure(figsize=(10, 5), facecolor='#f6f5f5')\ngs = fig.add_gridspec(1, 1)\ngs.update(wspace=0.5, hspace=0)\n\nbackground_color = \"#f6f5f5\"\n\nax = fig.add_subplot(gs[0, 0])\nax.set_facecolor(background_color)\nfor s in [\"top\",\"right\"]:\n    ax.spines[s].set_visible(False)\n\ncolormap = ['#0634e1', '#6bae9d', '#763085', '#e1bd4c', '#467acf', '#f6242b', '#0d7cdd']\nrun_no = 0\nfor cluster in cluster_list:\n    x = [val[0] for val in cluster]\n    y = [val[1] for val in cluster]\n    sns.scatterplot(ax=ax, x=x, y=y, zorder=2, linewidth=0, color=colormap[run_no])\n    run_no += 1\n\nax.set_ylabel('PC1', fontsize=5, fontweight='bold')\nax.set_xlabel('PC2', fontsize=5, fontweight='bold')\nax.tick_params(labelsize=5, width=0.5, length=1.5)\nax.grid(which='major', axis='x', zorder=0, color='#EEEEEE', linewidth=0.7)\nax.grid(which='major', axis='y', zorder=0, color='#EEEEEE', linewidth=0.7)\nax.legend([f'Cluster {i}' for i in range(len(cluster_list))], ncol=4, facecolor=background_color, edgecolor=background_color, loc='lower center');","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T15:10:37.357023Z","iopub.execute_input":"2022-07-03T15:10:37.358204Z","iopub.status.idle":"2022-07-03T15:10:41.08752Z","shell.execute_reply.started":"2022-07-03T15:10:37.358165Z","shell.execute_reply":"2022-07-03T15:10:41.086134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ss_df = pd.DataFrame(ss_df, columns=df.columns)\n\nplt.rcParams['figure.dpi'] = 600\nfig = plt.figure(figsize=(10, 10), facecolor='#f6f5f5')\ngs = fig.add_gridspec(5, 6)\ngs.update(wspace=0.3, hspace=0.3)\nbackground_color = '#f6f5f5'\nrun_no = 0\n\nplt.rc('axes', prop_cycle=(cycler('color', colormap)))\n\nfor row in range(0, 5):\n    for col in range(0, 6):\n        locals()[\"ax\"+str(run_no)] = fig.add_subplot(gs[row, col])\n        locals()[\"ax\"+str(run_no)].set_facecolor(background_color)\n        for s in [\"top\",\"right\"]:\n            locals()[\"ax\"+str(run_no)].spines[s].set_visible(False)\n        run_no += 1  \n\n\nfeatures = list(df.columns)\n\nrun_no = 0\nfor col in features:\n\n    #sns.kdeplot(ax=locals()[\"ax\"+str(run_no)], x=norm(loc=df[col].mean(), scale=np.std(df[col])).rvs(size=100000), zorder=2, alpha=0.9, fill=True, color='#0db887')\n    sns.kdeplot(ax=locals()[\"ax\"+str(run_no)], x=ss_df[col], zorder=2, alpha=1, linewidth=1, palette=colormap, hue=ss_pred_gmm)  \n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='x', zorder=0, color='#EEEEEE', linewidth=0.4)\n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='y', zorder=0, color='#EEEEEE', linewidth=0.4)\n    locals()[\"ax\"+str(run_no)].set_ylabel('')\n    locals()[\"ax\"+str(run_no)].set_xlabel(col, fontsize=4, fontweight='bold')\n    locals()[\"ax\"+str(run_no)].tick_params(labelsize=4, width=0.5)\n    locals()[\"ax\"+str(run_no)].xaxis.offsetText.set_fontsize(4)\n    locals()[\"ax\"+str(run_no)].yaxis.offsetText.set_fontsize(4)\n    locals()[\"ax\"+str(run_no)].get_legend().remove()\n    \n    run_no += 1\n\nlocals()[\"ax\"+str(29)].remove()\nlocals()[\"ax\"+str(0)].legend([f'Cluster {i}' for i in set(ss_pred_gmm)], ncol=2, fontsize=2, loc='lower center', facecolor=background_color, edgecolor=background_color)\n\n#plt.text(-99, 1.8, 'Distribution of features', fontsize=10, fontweight='bold')\n#plt.text(-99, 1.765, 'Modeling normal distribution with used Monte-Carlo method', fontsize=7)\n\nplt.show();","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-03T15:10:41.09106Z","iopub.execute_input":"2022-07-03T15:10:41.09172Z","iopub.status.idle":"2022-07-03T15:11:03.195329Z","shell.execute_reply.started":"2022-07-03T15:10:41.091671Z","shell.execute_reply":"2022-07-03T15:11:03.194074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transformer = MaxAbsScaler().fit(df)\ndf_mas = transformer.transform(df)\n#rob_scaler = RobustScaler().fit(df_mas)\n#data_standard = rob_scaler.transform(df_mas)\n\n# Transforming data with Yeo jhonson\npower_transformer = PowerTransformer().fit(df_mas)\ndata_transformed = power_transformer.transform(df_mas)\n\ndf_transformed = pd.DataFrame(data_transformed, columns=df.columns)\n\ngmm = BayesianGaussianMixture(n_components=7)\npred_gmm = gmm.fit_predict(data_transformed)\npred_gmm\n\nssub = pd.read_csv('../input/tabular-playground-series-jul-2022/sample_submission.csv', index_col='Id')\nssub['Predicted'] = pred_gmm\nssub.to_csv('submission.csv')\n\nssub","metadata":{"execution":{"iopub.status.busy":"2022-07-05T07:33:28.81422Z","iopub.execute_input":"2022-07-05T07:33:28.815395Z","iopub.status.idle":"2022-07-05T07:34:39.499155Z","shell.execute_reply.started":"2022-07-05T07:33:28.815334Z","shell.execute_reply":"2022-07-05T07:34:39.497901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **<span style=\"color:#ffd514;\">Good luck! </span>**\n","metadata":{}}]}