{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Cluster Analysis\nThere are many differing methods of clustering, as such it is a general term that refers to the finding subgroups or clusters in a dataset. When clustering these observations, the data is partioned into distinct groups, the observations within these groups are similar to each other and as such, observations in other groups are dissmilar from others in other groups. In other words, the resulting clusters should indicate high internal homogenity (in the cluster) and high external (among different clusters) heterogenity.\n\nThe cluster that the observations are assigned to are a mathematical representation of the observations of the objects similarities. What is similar or dissimilar is left up to the user/researcher and their domain specific criteria, this is a crucial step in cluster analysis.\n\nCluster analysis can also be used to form descriptive statistics to determine whether or not the data consists of distinct subgroups. To acheive this, the extent of the difference is assessed between the observations assigned to their respective clusters.\n\nCluster analysis is also referred to as data segmentation.\n\n### Uses of Clustering:\n* Customer Segmentation\n* Preliminary Data Analysis\n* Anomaly Detection (Outlier Detection)\n* Image Segmentation\n* Reduce the size of the data set (see also, PCA)","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport plotly.express as px\nimport seaborn as sns \nimport matplotlib.pyplot as plt\nfrom plotly.subplots import make_subplots\nimport plotly.graph_objects as go\nimport math\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\nfrom plotly.offline import plot, iplot, init_notebook_mode\ninit_notebook_mode(connected=True)\n\nRANDOM_STATE = 115","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:13.076261Z","iopub.execute_input":"2022-07-19T02:44:13.076767Z","iopub.status.idle":"2022-07-19T02:44:13.086166Z","shell.execute_reply.started":"2022-07-19T02:44:13.076725Z","shell.execute_reply":"2022-07-19T02:44:13.085101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set(rc={'figure.figsize':(11.7,8.27)})","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:13.120686Z","iopub.execute_input":"2022-07-19T02:44:13.121851Z","iopub.status.idle":"2022-07-19T02:44:13.128636Z","shell.execute_reply.started":"2022-07-19T02:44:13.121795Z","shell.execute_reply":"2022-07-19T02:44:13.127430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_original = pd.read_csv(\"/kaggle/input/tabular-playground-series-jul-2022/data.csv\")\ndf_sample_submission = pd.read_csv(\"/kaggle/input/tabular-playground-series-jul-2022/sample_submission.csv\")\nid_col = df_original[\"id\"]\ndf_original = df_original.drop(\"id\",axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:13.167831Z","iopub.execute_input":"2022-07-19T02:44:13.168516Z","iopub.status.idle":"2022-07-19T02:44:13.942353Z","shell.execute_reply.started":"2022-07-19T02:44:13.168460Z","shell.execute_reply":"2022-07-19T02:44:13.941130Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_original.describe().T","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:13.944566Z","iopub.execute_input":"2022-07-19T02:44:13.945037Z","iopub.status.idle":"2022-07-19T02:44:14.146073Z","shell.execute_reply.started":"2022-07-19T02:44:13.944994Z","shell.execute_reply":"2022-07-19T02:44:14.144971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Correleation Matrix \n\nFor All Features.","metadata":{}},{"cell_type":"code","source":"fig = px.imshow(df_original.corr(),color_continuous_scale=\"agsunset\")\nfig.update_layout(title = \"Correlation Matrix for all features\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:14.147704Z","iopub.execute_input":"2022-07-19T02:44:14.148045Z","iopub.status.idle":"2022-07-19T02:44:14.435682Z","shell.execute_reply.started":"2022-07-19T02:44:14.148013Z","shell.execute_reply":"2022-07-19T02:44:14.434628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Top 5 Correlated Features\n\nFinding out the top 5 Correleated Features for any particular feature.","metadata":{}},{"cell_type":"code","source":"from plotly.subplots import make_subplots\nimport plotly.graph_objects as go\namt_of_features = 6 \ntitles = []\nn = 5\nsuffix = None\n\nfor i in range(0,amt_of_features):\n    if i < 10:\n        suffix = \"0\"\n    else: \n        suffix = \"\"\n        \n    target = \"f_\"+suffix+str(i)\n    titles.append(\"Top \"+str(n)+\" features correlated with \"+target)\n    \nfig = make_subplots(rows=2, cols=3,subplot_titles=tuple(titles)) # change accordingly wiht the number of features\nsuffix = None\nrow:int = 1\ncounter:int = 0\nfor i in range(0,amt_of_features):\n    counter+=1\n    if counter % 4 == 0:\n        row+=1\n        counter = 1\n#     print(row,counter)\n\n    if i < 10:\n        suffix = \"0\"\n    else: \n        suffix = \"\"\n        \n    target = \"f_\"+suffix+str(i)\n\n    corr = df_original.corr()\n    x = corr.nlargest(n+1,target).index\n    corr_df =  df_original[list(x)]\n    corr = corr_df.corr()\n    fig.add_trace(go.Heatmap(x=list(x),y=list(x),z=corr,colorscale=\"picnic\",coloraxis=\"coloraxis\"),row=row,col=counter)\n\nfig.update_layout(title_text=\"Top 5 Correlated Features\", height=1000) # change accordingly wiht the number of features\nfig.update_layout(coloraxis = {'colorscale':'picnic'}) \nfig.show()\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:14.439059Z","iopub.execute_input":"2022-07-19T02:44:14.440001Z","iopub.status.idle":"2022-07-19T02:44:16.064555Z","shell.execute_reply.started":"2022-07-19T02:44:14.439956Z","shell.execute_reply":"2022-07-19T02:44:16.063428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Tests For Normality\n\n[Come on guys, is normal](https://www.youtube.com/watch?v=WO-E7zOHHp0&ab_channel=MontyBurns)\n\n* Skewness assesses the extent to which a variable’s distribution is symmetrical. If the distribution of responses for a variable stretches toward the right or left tail of the distribution, then the distribution is referred to as skewed.\n* Kurtosis is a measure of whether the distribution is too peaked (a very narrow distribution with most of the responses in the center).\n\nWhen both skewness and kurtosis are zero (an unlikely scenario), the distribution is considered a normal distribution.\n\nIf skewness is greater than +1 or less than -1, this is usually a good indication of a significantly skewed distribution. With respect to kurtosis, if it is greater than +1, then the distribution is too peaked, whereas a kurtosis is less than -1, the distribution is too flat. For distributions that exhibit skewness and or kurtosis that go beyond these guidelines are considered nonnormal. \n\nWe can test for normality using the two following tests:\n\n1. Shapiro-Wilk Test\n2. Kolmogorov-Smirnov Test\n3. Anderson-Darling Test","metadata":{}},{"cell_type":"code","source":"from scipy.stats import shapiro \nfrom scipy.stats import kstest\nfrom scipy.stats import anderson","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:16.066003Z","iopub.execute_input":"2022-07-19T02:44:16.066305Z","iopub.status.idle":"2022-07-19T02:44:16.071740Z","shell.execute_reply.started":"2022-07-19T02:44:16.066276Z","shell.execute_reply":"2022-07-19T02:44:16.070640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"normality_summary_df = pd.DataFrame(columns=[\"col_name\",\n                                             \"skewness\",\n                                             \"kurtosis\",\n                                             \"shapiro_p_value\",\n                                             \"ks_p_value\",\n                                             \"anderson_p_value\",\n                                             \"shapiro_normal_status\",\n                                             \"ks_test_normal_status\",\n                                             \"anderson_normal_status\"])\nskewness_list = []\nkurtosis_list = []\n\nshapiro_p_value_list = []\nks_p_value_list = []\nanderson_p_value_list =[]\n\nshapiro_normal_status_list = []\nks_normal_status_list = []\nanderson_normal_status_list=[]","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:16.073206Z","iopub.execute_input":"2022-07-19T02:44:16.073517Z","iopub.status.idle":"2022-07-19T02:44:16.086797Z","shell.execute_reply.started":"2022-07-19T02:44:16.073486Z","shell.execute_reply":"2022-07-19T02:44:16.085638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def es_normal(val:float):\n    if val <= 0.05:\n        return 1 #non normal\n    elif val > 0.05:\n        return 0# normal\n    \ndef p_value_ad(AD):\n    if AD <=0.2:\n        return 1 - math.exp(-13.436+(101.14*AD)-(223.73*(AD**2)))\n    elif 0.2 < AD <= 0.34:\n        return 1 - math.exp(-8.318+(42.796*AD)-(59.938*(AD**2)))\n    elif 0.34 < AD < 0.6:\n        return math.exp(0.91770-(4.279*AD)-(1.38*(AD**2)))\n    elif AD >= 0.6:\n        return math.exp(1.2937-(5.709*AD)-(0.0186*(AD**2)))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:16.088540Z","iopub.execute_input":"2022-07-19T02:44:16.089000Z","iopub.status.idle":"2022-07-19T02:44:16.104115Z","shell.execute_reply.started":"2022-07-19T02:44:16.088957Z","shell.execute_reply":"2022-07-19T02:44:16.103047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in list(df_original.columns):\n    skewness_list.append(df_original[col].skew())\n    kurtosis_list.append(df_original[col].kurtosis())\n    \n    shapiro_p_value_list.append(shapiro(df_original[col]).pvalue)\n    ks_p_value_list.append(kstest(df_original[col],\"norm\").pvalue)\n    anderson_p_value_list.append(p_value_ad(anderson(df_original[col],\"norm\").statistic))\n    \n    shapiro_normal_status_list.append(es_normal(shapiro(df_original[col]).pvalue))\n    ks_normal_status_list.append(es_normal(kstest(df_original[col],\"norm\").pvalue))\n    anderson_normal_status_list.append(es_normal(p_value_ad(anderson(df_original[col],\"norm\").statistic)))","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:16.105275Z","iopub.execute_input":"2022-07-19T02:44:16.105623Z","iopub.status.idle":"2022-07-19T02:44:18.614749Z","shell.execute_reply.started":"2022-07-19T02:44:16.105591Z","shell.execute_reply":"2022-07-19T02:44:18.613654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"normality_summary_df[\"col_name\"] = list(df_original.columns)\nnormality_summary_df[\"skewness\"] = skewness_list\nnormality_summary_df[\"kurtosis\"] = kurtosis_list\n\nnormality_summary_df[\"shapiro_p_value\"] = shapiro_p_value_list\nnormality_summary_df[\"ks_p_value\"] = ks_p_value_list\nnormality_summary_df[\"anderson_p_value\"] = anderson_p_value_list\n\nnormality_summary_df[\"shapiro_normal_status\"] = shapiro_normal_status_list\nnormality_summary_df[\"ks_test_normal_status\"] = ks_normal_status_list\nnormality_summary_df[\"anderson_normal_status\"] = anderson_normal_status_list\nnormality_summary_df[\"score\"] = normality_summary_df[\"shapiro_normal_status\"]+normality_summary_df[\"ks_test_normal_status\"]+normality_summary_df[\"anderson_normal_status\"]","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:18.616333Z","iopub.execute_input":"2022-07-19T02:44:18.616836Z","iopub.status.idle":"2022-07-19T02:44:18.630457Z","shell.execute_reply.started":"2022-07-19T02:44:18.616789Z","shell.execute_reply":"2022-07-19T02:44:18.629419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"normality_summary_df","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:18.634557Z","iopub.execute_input":"2022-07-19T02:44:18.634967Z","iopub.status.idle":"2022-07-19T02:44:18.665841Z","shell.execute_reply.started":"2022-07-19T02:44:18.634916Z","shell.execute_reply":"2022-07-19T02:44:18.664586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ua_normal = normality_summary_df.query(\"score == 0\")\nua_normal = ua_normal.reset_index(drop=True)\n\nua_non_normal = normality_summary_df.query(\"score == 3\")\nua_non_normal = ua_non_normal.reset_index(drop=True)\n\nunsure_normal = normality_summary_df.query(\"score >= 1 & score <= 2\")\nunsure_normal = unsure_normal.reset_index(drop=True)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:18.667514Z","iopub.execute_input":"2022-07-19T02:44:18.668658Z","iopub.status.idle":"2022-07-19T02:44:18.685551Z","shell.execute_reply.started":"2022-07-19T02:44:18.668622Z","shell.execute_reply":"2022-07-19T02:44:18.684518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Distributions Of Normal Featrues","metadata":{}},{"cell_type":"code","source":"dist_subplots= plt.figure(figsize=(20,16))\n\nfor idx, feature_name in enumerate(ua_normal[\"col_name\"]):\n    plt.subplot(6,5,idx+1)\n    sns.histplot(x = df_original[feature_name])\n    plt.title(f'Feature Name: {feature_name}')\n    plt.xlabel('')\n\ndist_subplots.suptitle('Distribution of Normal Features',  size=20)\ndist_subplots.tight_layout() \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:18.687144Z","iopub.execute_input":"2022-07-19T02:44:18.687461Z","iopub.status.idle":"2022-07-19T02:44:24.555145Z","shell.execute_reply.started":"2022-07-19T02:44:18.687431Z","shell.execute_reply":"2022-07-19T02:44:24.553794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Distributions Of Non-Normal Featrues","metadata":{}},{"cell_type":"code","source":"dist_subplots= plt.figure(figsize=(20,16))\n\nfor idx, feature_name in enumerate(ua_non_normal[\"col_name\"]):\n    plt.subplot(5,5,idx+1)\n    sns.histplot(x = df_original[feature_name])\n    plt.title(f'Feature Name: {feature_name}')\n    plt.xlabel('')\n\ndist_subplots.suptitle('Distribution of Non Normal Features',  size=20)\ndist_subplots.tight_layout() \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:24.556631Z","iopub.execute_input":"2022-07-19T02:44:24.557650Z","iopub.status.idle":"2022-07-19T02:44:31.458050Z","shell.execute_reply.started":"2022-07-19T02:44:24.557610Z","shell.execute_reply":"2022-07-19T02:44:31.457191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Distributions Of Possibly Non-Normal Features","metadata":{}},{"cell_type":"code","source":"dist_subplots= plt.figure(figsize=(16,8))\n\nfor idx, feature_name in enumerate(unsure_normal[\"col_name\"]):\n    plt.subplot(1,3,idx+1)\n    sns.histplot(x = df_original[feature_name])\n    plt.title(f'Feature Name: {feature_name}')\n    plt.xlabel('')\n\ndist_subplots.suptitle('Distribution of Features (at least one test indicated departure from normality)',  size=20)\ndist_subplots.tight_layout() \nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:31.459449Z","iopub.execute_input":"2022-07-19T02:44:31.460033Z","iopub.status.idle":"2022-07-19T02:44:33.063630Z","shell.execute_reply.started":"2022-07-19T02:44:31.459997Z","shell.execute_reply":"2022-07-19T02:44:33.062777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def scaled_df(X_scaled,list_of_cols,y_target=None):\n    scaled_df = pd.DataFrame(X_scaled,columns=list_of_cols)\n    if y_target is not None:\n        scaled_df = pd.concat([scaled_df,pd.Series(y_target)],axis=1)\n    return scaled_df","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:33.065086Z","iopub.execute_input":"2022-07-19T02:44:33.065655Z","iopub.status.idle":"2022-07-19T02:44:33.070787Z","shell.execute_reply.started":"2022-07-19T02:44:33.065619Z","shell.execute_reply":"2022-07-19T02:44:33.070009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def scale_comparision(scaled_df,orig_df,x_var,y_var,x_title,y_title,scaler_name):\n    fig = make_subplots(\n        rows=1, cols=2,\n        subplot_titles=(\"Without Scaling\", \"With \"+str(scaler_name)))\n\n    fig.add_trace(go.Scattergl(x=orig_df[x_var],\n                             y=orig_df[y_var],\n                            mode=\"markers\"),\n                  row=1, col=1)\n\n    fig.add_trace(go.Scattergl(x=scaled_df[x_var],\n                             y=scaled_df[y_var],\n                            mode=\"markers\"),\n                  row=1, col=2)\n\n\n\n    fig.update_layout(title_text=\"Comparision of Unscaled and Scaled Data \")\n    fig.update_xaxes(title=x_title)\n    fig.update_yaxes(title=y_title)\n    fig.update_layout(showlegend=False)\n    fig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:33.072312Z","iopub.execute_input":"2022-07-19T02:44:33.073060Z","iopub.status.idle":"2022-07-19T02:44:33.085636Z","shell.execute_reply.started":"2022-07-19T02:44:33.073025Z","shell.execute_reply":"2022-07-19T02:44:33.084418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# x = \"f_10\",f_28,f_12,f27,f_13\n# y = \"f_01\",f_10,f_26,f10,f_12\nx = \"f_28\"\ny = \"f_10\"\n# cols = list(df_original.columns)\n# x = np.random.choice(cols)\n# cols.remove(x)\n# y = np.random.choice(cols)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:33.087413Z","iopub.execute_input":"2022-07-19T02:44:33.088005Z","iopub.status.idle":"2022-07-19T02:44:33.095921Z","shell.execute_reply.started":"2022-07-19T02:44:33.087956Z","shell.execute_reply":"2022-07-19T02:44:33.094837Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Scaling\n\n### Original data\nNo scaling is done to the data, the data is left as is in it's original form.\n\n### StandardScaler\nThe StandardScaler subtracts the mean (standardized values have zero mean) and divides it by the unit variance. Therefore the resulting distribution has unit variance. It should be noted that outliers affects feature standardization using this method, the mean is influenced by outliers and thus each feature may not always have similar scales. Therefore StandardScaler cannot always gurantee balanced feature scales when outliers are in the data.","metadata":{"execution":{"iopub.status.busy":"2022-07-15T16:14:24.754687Z","iopub.execute_input":"2022-07-15T16:14:24.755067Z","iopub.status.idle":"2022-07-15T16:14:24.761647Z","shell.execute_reply.started":"2022-07-15T16:14:24.755038Z","shell.execute_reply":"2022-07-15T16:14:24.760601Z"}}},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\n\nstd_scaler = StandardScaler()\nX_std_scaled = std_scaler.fit_transform(df_original)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:33.097292Z","iopub.execute_input":"2022-07-19T02:44:33.097878Z","iopub.status.idle":"2022-07-19T02:44:33.151115Z","shell.execute_reply.started":"2022-07-19T02:44:33.097844Z","shell.execute_reply":"2022-07-19T02:44:33.149937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# scale_comparision(scaled_df(X_std_scaled,list(df_original.columns),None),\n#                  df_original,x,y,x,y,\n#                  \"StandardScaler\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:33.152866Z","iopub.execute_input":"2022-07-19T02:44:33.153506Z","iopub.status.idle":"2022-07-19T02:44:33.159722Z","shell.execute_reply.started":"2022-07-19T02:44:33.153458Z","shell.execute_reply":"2022-07-19T02:44:33.158509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### MinMaxScaler\nThis scaling method shifts and rescales the data so it's range is $[0,1]$.\nThis is done by subtracting the minimum value from the feature and then dividing it by the max value of the feature minus the minmum value of the feature. It is then standardized by multiplying it by the $(max-min)+min$ where $max$ and $min$ are the maximum and minimum values of the feature range. This method of scaling, like StandardScaler is very sensitive to outliers.","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import MinMaxScaler\n\nmin_max_scaler = MinMaxScaler()\nX_min_max_scaled = min_max_scaler.fit_transform(df_original)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:33.161525Z","iopub.execute_input":"2022-07-19T02:44:33.161968Z","iopub.status.idle":"2022-07-19T02:44:33.205854Z","shell.execute_reply.started":"2022-07-19T02:44:33.161925Z","shell.execute_reply":"2022-07-19T02:44:33.204604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# scale_comparision(scaled_df(X_min_max_scaled,list(df_original.columns),None),\n#                  df_original,x,y,x,y,\n#                  \"MinMaxScaler\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:33.207326Z","iopub.execute_input":"2022-07-19T02:44:33.208308Z","iopub.status.idle":"2022-07-19T02:44:33.212658Z","shell.execute_reply.started":"2022-07-19T02:44:33.208257Z","shell.execute_reply":"2022-07-19T02:44:33.211840Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### MaxAbsScaler\nThis scaling method behaves in a similar way to MinMaxScaler. In the case where:\n* Only Positive values are in the feature, the range is $[0,1]$, \n* Only Negative values, the range is from $[-1,0]$ \n* Both negative and positive values are present, the range is from $[-1,1]$. \n\nAs with the previous scalers, this method is also sensitive to outliers. ","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import MaxAbsScaler\n\nmax_abs_scaler = MaxAbsScaler()\nX_max_abs_scaled = max_abs_scaler.fit_transform(df_original)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:33.214000Z","iopub.execute_input":"2022-07-19T02:44:33.214518Z","iopub.status.idle":"2022-07-19T02:44:33.256998Z","shell.execute_reply.started":"2022-07-19T02:44:33.214484Z","shell.execute_reply":"2022-07-19T02:44:33.256109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# scale_comparision(scaled_df(X_max_abs_scaled,list(df_original.columns),None),\n#                  df_original,x,y,x,y,\n#                  \"MaxAbsScaler\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:33.258311Z","iopub.execute_input":"2022-07-19T02:44:33.259223Z","iopub.status.idle":"2022-07-19T02:44:33.262737Z","shell.execute_reply.started":"2022-07-19T02:44:33.259184Z","shell.execute_reply":"2022-07-19T02:44:33.262023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### RobustScaler\nThis scaling method scales the data by removing the median and scaling accoriding to the quantile range where the default is the IQR (interquartile range, between the 25th and 75th quantile). RobustScaler is therefore not influenced by outliers. The transformed feature value ranges are larger compared to the MinMax, MaxAbs and Standard scalers. The transformed features are on approximately similar scales as well. However, it should be noted that the outliers themselves are present in the transformed data. ","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import RobustScaler\n\nrobust_scaler = RobustScaler()\nX_robust_scaled = robust_scaler.fit_transform(df_original)","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-07-19T02:44:33.263978Z","iopub.execute_input":"2022-07-19T02:44:33.264477Z","iopub.status.idle":"2022-07-19T02:44:33.409243Z","shell.execute_reply.started":"2022-07-19T02:44:33.264446Z","shell.execute_reply":"2022-07-19T02:44:33.408092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scale_comparision(scaled_df(X_robust_scaled,list(df_original.columns),None),\n                 df_original,x,y,x,y,\n                 \"RobustScaler\")","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:33.411672Z","iopub.execute_input":"2022-07-19T02:44:33.414031Z","iopub.status.idle":"2022-07-19T02:44:34.016640Z","shell.execute_reply.started":"2022-07-19T02:44:33.413979Z","shell.execute_reply":"2022-07-19T02:44:34.015783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### PowerTransformer\nA power transformation is applied to each feature to make the data have a gaussian/normal like distribution, this is done to stabilize variance and to minimize skewness (Skewness is the measure of asymmetery in a distribution).","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import PowerTransformer\n\n# the box-cox method can only be applied strictly to positive data\npow_transformer = PowerTransformer(method = \"yeo-johnson\")\nX_pow_transform = pow_transformer.fit_transform(df_original)\nX_pow_transform.shape","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:34.017888Z","iopub.execute_input":"2022-07-19T02:44:34.018828Z","iopub.status.idle":"2022-07-19T02:44:37.718690Z","shell.execute_reply.started":"2022-07-19T02:44:34.018789Z","shell.execute_reply":"2022-07-19T02:44:37.717551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# scale_comparision(scaled_df(X_pow_transform,list(df_original.columns),None),\n#                  df_original,x,y,x,y,\n#                  \"PowerTransformer\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:37.720584Z","iopub.execute_input":"2022-07-19T02:44:37.721371Z","iopub.status.idle":"2022-07-19T02:44:37.726143Z","shell.execute_reply.started":"2022-07-19T02:44:37.721324Z","shell.execute_reply":"2022-07-19T02:44:37.724976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### QuantileTransformer (Uniform output/Gaussian output)\nA non linear transfromation is applied s.t the PDF (probabilty density function) for each feature will be mapped to either a Uniform or Gaussian (Normal) distribution. This method of transformation, like RobustScaler are not sensitive to outliers, however, QuantileTransformer will collapse the outliers into the defined range. The transformation is done on each feature independently. ","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import QuantileTransformer","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:37.735428Z","iopub.execute_input":"2022-07-19T02:44:37.735767Z","iopub.status.idle":"2022-07-19T02:44:37.740240Z","shell.execute_reply.started":"2022-07-19T02:44:37.735737Z","shell.execute_reply":"2022-07-19T02:44:37.739014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"qt_transformer = QuantileTransformer(output_distribution = \"uniform\",\n                                     random_state = RANDOM_STATE)\nX_qt_transform_uni = qt_transformer.fit_transform(df_original)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:37.742063Z","iopub.execute_input":"2022-07-19T02:44:37.742759Z","iopub.status.idle":"2022-07-19T02:44:38.494362Z","shell.execute_reply.started":"2022-07-19T02:44:37.742686Z","shell.execute_reply":"2022-07-19T02:44:38.493155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# scale_comparision(scaled_df(X_pow_transform,list(df_original.columns),None),\n#                  df_original,x,y,x,y,\n#                  \"QuantileTransformer(Uniform)\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:38.495598Z","iopub.execute_input":"2022-07-19T02:44:38.495948Z","iopub.status.idle":"2022-07-19T02:44:38.502580Z","shell.execute_reply.started":"2022-07-19T02:44:38.495916Z","shell.execute_reply":"2022-07-19T02:44:38.501172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"qt_transformer = QuantileTransformer(output_distribution = \"normal\",\n                                     random_state = RANDOM_STATE)\nX_qt_transform_normal = qt_transformer.fit_transform(df_original)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:38.504655Z","iopub.execute_input":"2022-07-19T02:44:38.505772Z","iopub.status.idle":"2022-07-19T02:44:39.442454Z","shell.execute_reply.started":"2022-07-19T02:44:38.505719Z","shell.execute_reply":"2022-07-19T02:44:39.441060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# scale_comparision(scaled_df(X_qt_transform_normal,list(df_original.columns),None),\n#                  df_original,x,y,x,y,\n#                  \"QuantileTransformer(Normal)\")","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:39.444208Z","iopub.execute_input":"2022-07-19T02:44:39.444886Z","iopub.status.idle":"2022-07-19T02:44:39.450276Z","shell.execute_reply.started":"2022-07-19T02:44:39.444835Z","shell.execute_reply":"2022-07-19T02:44:39.449284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Normalizer\nUnlike the previous scalers where features (columns) are scaled, for normalization the samples (rows) are scaled to have unit norm indpendent of the distribution of the samples. \n\nTypes of unit norm:\n* $l^1$ $normalization$: Summing the absolute values of the normalized vector is equal to 1. \n* $l^2$ $normalization$: If each element in the normalized vector is squared and then summed,it would be equal to 1. This is the most commonly used vector normalization method\n* $l^{\\infty}$ $normalization$: also called Vector Max Norm. Gives the largest magnitude among each element of a vector","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import Normalizer","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:39.451482Z","iopub.execute_input":"2022-07-19T02:44:39.452487Z","iopub.status.idle":"2022-07-19T02:44:39.471144Z","shell.execute_reply.started":"2022-07-19T02:44:39.452442Z","shell.execute_reply":"2022-07-19T02:44:39.470248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"normalizer_l1 = Normalizer(norm = \"l1\")\nX_normalize_l1 = normalizer_l1.fit_transform(df_original)","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-07-19T02:44:39.472672Z","iopub.execute_input":"2022-07-19T02:44:39.473789Z","iopub.status.idle":"2022-07-19T02:44:39.518832Z","shell.execute_reply.started":"2022-07-19T02:44:39.473738Z","shell.execute_reply":"2022-07-19T02:44:39.517953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scale_comparision(scaled_df(X_normalize_l1,list(df_original.columns),None),\n                 df_original,x,y,x,y,\n                 \"l1 Normalization\")","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-07-19T02:44:39.520501Z","iopub.execute_input":"2022-07-19T02:44:39.521325Z","iopub.status.idle":"2022-07-19T02:44:40.161288Z","shell.execute_reply.started":"2022-07-19T02:44:39.521275Z","shell.execute_reply":"2022-07-19T02:44:40.159155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"normalizer_l2 = Normalizer(norm = \"l2\")\nX_normalize_l2 = normalizer_l2.fit_transform(df_original)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:40.163891Z","iopub.execute_input":"2022-07-19T02:44:40.164238Z","iopub.status.idle":"2022-07-19T02:44:40.200529Z","shell.execute_reply.started":"2022-07-19T02:44:40.164205Z","shell.execute_reply":"2022-07-19T02:44:40.199342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scale_comparision(scaled_df(X_normalize_l2,list(df_original.columns),None),\n                 df_original,x,y,x,y,\n                 \"l2 Normalization\")","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:40.202009Z","iopub.execute_input":"2022-07-19T02:44:40.202337Z","iopub.status.idle":"2022-07-19T02:44:40.844684Z","shell.execute_reply.started":"2022-07-19T02:44:40.202306Z","shell.execute_reply":"2022-07-19T02:44:40.842959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"normalizer_max = Normalizer(norm = \"max\")\nX_normalize_max = normalizer_max.fit_transform(df_original)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:40.846347Z","iopub.execute_input":"2022-07-19T02:44:40.846861Z","iopub.status.idle":"2022-07-19T02:44:40.892055Z","shell.execute_reply.started":"2022-07-19T02:44:40.846825Z","shell.execute_reply":"2022-07-19T02:44:40.890882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scale_comparision(scaled_df(X_normalize_max,list(df_original.columns),None),\n                 df_original,x,y,x,y,\n                 \"max Normalization\")","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:40.893402Z","iopub.execute_input":"2022-07-19T02:44:40.894437Z","iopub.status.idle":"2022-07-19T02:44:41.495365Z","shell.execute_reply.started":"2022-07-19T02:44:40.894397Z","shell.execute_reply":"2022-07-19T02:44:41.493106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explained Variance","metadata":{}},{"cell_type":"code","source":"from sklearn.decomposition import PCA","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:41.497518Z","iopub.execute_input":"2022-07-19T02:44:41.498128Z","iopub.status.idle":"2022-07-19T02:44:41.503626Z","shell.execute_reply.started":"2022-07-19T02:44:41.498080Z","shell.execute_reply":"2022-07-19T02:44:41.502159Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def disp_explained_variance_ratio(X,title,n_components=None):\n    \n    if n_components is not None:\n        pca = PCA(n_components=n_components)\n    else:\n        pca = PCA()\n    pca.fit(X)\n    exp_var_cumul = np.cumsum(pca.explained_variance_ratio_)\n    \n    explained_variance = px.area(\n        x=range(1, exp_var_cumul.shape[0] + 1),\n        y=exp_var_cumul,\n        color_discrete_sequence=[\"deepskyblue\"],\n        labels={\"x\": \"# Components\", \"y\": \"Explained Variance\"}\n    )\n    \n    explained_variance.update_layout(title=title)\n    explained_variance.show()\n    \n    dimensions = px.bar(x=range(pca.n_components_), y=pca.explained_variance_ratio_,\n                        color_discrete_sequence=[\"deeppink\"],\n                        labels={\"x\":\"Component #\",\"y\":\"Explained Variance\"})\n    dimensions.update_yaxes(range=[0.0, 1.0])\n    dimensions.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:41.505675Z","iopub.execute_input":"2022-07-19T02:44:41.506124Z","iopub.status.idle":"2022-07-19T02:44:41.518990Z","shell.execute_reply.started":"2022-07-19T02:44:41.506080Z","shell.execute_reply":"2022-07-19T02:44:41.517993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Explained Variance using Unscaled Data","metadata":{}},{"cell_type":"code","source":"# disp_explained_variance_ratio(df_original,title=\"PCA 99.99% explained variance using Unscaled data\",n_components = None)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:41.520223Z","iopub.execute_input":"2022-07-19T02:44:41.521232Z","iopub.status.idle":"2022-07-19T02:44:41.531194Z","shell.execute_reply.started":"2022-07-19T02:44:41.521193Z","shell.execute_reply":"2022-07-19T02:44:41.530269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Explained Variance using MinMax Scaled Data","metadata":{}},{"cell_type":"code","source":"# disp_explained_variance_ratio(X_min_max_scaled,title=\"PCA 99.99% explained variance using MinMax data\",n_components = None)","metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2022-07-19T02:44:41.532627Z","iopub.execute_input":"2022-07-19T02:44:41.533286Z","iopub.status.idle":"2022-07-19T02:44:41.546126Z","shell.execute_reply.started":"2022-07-19T02:44:41.533253Z","shell.execute_reply":"2022-07-19T02:44:41.545102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Explained Variance using MaxAbs Scaled Data","metadata":{}},{"cell_type":"code","source":"# disp_explained_variance_ratio(X_max_abs_scaled,title=\"PCA 99.99% explained variance using MaxAbs data\",n_components = None)","metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2022-07-19T02:44:41.547262Z","iopub.execute_input":"2022-07-19T02:44:41.548044Z","iopub.status.idle":"2022-07-19T02:44:41.558295Z","shell.execute_reply.started":"2022-07-19T02:44:41.547995Z","shell.execute_reply":"2022-07-19T02:44:41.557310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Explained Variance using Robust Scaled Data","metadata":{}},{"cell_type":"code","source":"disp_explained_variance_ratio(X_robust_scaled,title=\"PCA 99.99% explained variance using Robust Scaled data\",n_components = None)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:41.559803Z","iopub.execute_input":"2022-07-19T02:44:41.560434Z","iopub.status.idle":"2022-07-19T02:44:41.847765Z","shell.execute_reply.started":"2022-07-19T02:44:41.560392Z","shell.execute_reply":"2022-07-19T02:44:41.846861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Explained Variance using Power Transformed Data","metadata":{}},{"cell_type":"code","source":"# disp_explained_variance_ratio(X_pow_transform,title=\"PCA 99.99% explained variance using Power Transformed data\",n_components = None)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:41.849161Z","iopub.execute_input":"2022-07-19T02:44:41.849734Z","iopub.status.idle":"2022-07-19T02:44:41.853963Z","shell.execute_reply.started":"2022-07-19T02:44:41.849684Z","shell.execute_reply":"2022-07-19T02:44:41.853115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Explained Variance using Quantile Transformed (Uniform) Data","metadata":{}},{"cell_type":"code","source":"# disp_explained_variance_ratio(X_qt_transform_uni,title=\"PCA 99.99% explained variance using Quantile Transformer(Uniform) data\",n_components = None)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:41.855257Z","iopub.execute_input":"2022-07-19T02:44:41.856234Z","iopub.status.idle":"2022-07-19T02:44:41.864499Z","shell.execute_reply.started":"2022-07-19T02:44:41.856197Z","shell.execute_reply":"2022-07-19T02:44:41.863602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Explained Variance using Quantile Transformed (Normal) Data","metadata":{}},{"cell_type":"code","source":"# disp_explained_variance_ratio(X_qt_transform_normal,title=\"PCA 99.99% explained variance using Quantile Transformer (Normal) data\",n_components = None)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:41.865889Z","iopub.execute_input":"2022-07-19T02:44:41.867068Z","iopub.status.idle":"2022-07-19T02:44:41.875736Z","shell.execute_reply.started":"2022-07-19T02:44:41.867025Z","shell.execute_reply":"2022-07-19T02:44:41.874778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Explained Variance using l1 Normalized Data","metadata":{}},{"cell_type":"code","source":"# disp_explained_variance_ratio(X_normalize_l1,title=\"PCA 99.99% explained variance using l1 data\",n_components = None)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:41.877131Z","iopub.execute_input":"2022-07-19T02:44:41.878082Z","iopub.status.idle":"2022-07-19T02:44:41.885975Z","shell.execute_reply.started":"2022-07-19T02:44:41.878044Z","shell.execute_reply":"2022-07-19T02:44:41.885154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Explained Variance using l2 Normalized Data","metadata":{}},{"cell_type":"code","source":"disp_explained_variance_ratio(X_normalize_l2,title=\"PCA 99.99% explained variance using l2 data\",n_components = None)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:41.887273Z","iopub.execute_input":"2022-07-19T02:44:41.888156Z","iopub.status.idle":"2022-07-19T02:44:42.258665Z","shell.execute_reply.started":"2022-07-19T02:44:41.888118Z","shell.execute_reply":"2022-07-19T02:44:42.257799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Explained Variance using Max Normalized Data","metadata":{}},{"cell_type":"code","source":"# disp_explained_variance_ratio(X_normalize_max,title=\"PCA 99.99% explained variance using norm max data\",n_components = None)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:42.260134Z","iopub.execute_input":"2022-07-19T02:44:42.260716Z","iopub.status.idle":"2022-07-19T02:44:42.264784Z","shell.execute_reply.started":"2022-07-19T02:44:42.260668Z","shell.execute_reply":"2022-07-19T02:44:42.263887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Making Different Combinations Of Scaled Data to try.","metadata":{}},{"cell_type":"code","source":"orig_cols = list(df_original.columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:42.266200Z","iopub.execute_input":"2022-07-19T02:44:42.266493Z","iopub.status.idle":"2022-07-19T02:44:42.279204Z","shell.execute_reply.started":"2022-07-19T02:44:42.266464Z","shell.execute_reply":"2022-07-19T02:44:42.278095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scales_list = [StandardScaler(),MaxAbsScaler(),MinMaxScaler(),RobustScaler()]\n\ntransformer_list = [PowerTransformer(),\n                    QuantileTransformer(output_distribution=\"uniform\",random_state=RANDOM_STATE),\n                    QuantileTransformer(output_distribution=\"normal\",random_state=RANDOM_STATE)]\n\nnormalizer_list = [Normalizer(norm=\"l1\"),Normalizer(norm=\"l2\"),Normalizer(norm=\"max\")]","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:42.280666Z","iopub.execute_input":"2022-07-19T02:44:42.281261Z","iopub.status.idle":"2022-07-19T02:44:42.291158Z","shell.execute_reply.started":"2022-07-19T02:44:42.281226Z","shell.execute_reply":"2022-07-19T02:44:42.290040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.pipeline import Pipeline\n\ndef scale_data(df,scaler,do_pca:bool=True,n_components = None):\n    if not do_pca:\n        X_scaled = scaler.fit_transform(df)\n        return pd.DataFrame(X_scaled,columns=list(df.columns))\n        pass\n    elif do_pca:\n        if n_components is None:\n            pca = PCA()\n        else:\n            pca = PCA(n_components)\n        pipeline = Pipeline([(\"scaler\",scaler),(\"PCA\",pca)])\n        X_scale_pca = pipeline.fit_transform(df)\n        return pd.DataFrame(X_scale_pca,columns=[\"pca_\"+str(i) for i in range(0,len(pca.components_))])","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:42.292484Z","iopub.execute_input":"2022-07-19T02:44:42.292958Z","iopub.status.idle":"2022-07-19T02:44:42.303645Z","shell.execute_reply.started":"2022-07-19T02:44:42.292919Z","shell.execute_reply":"2022-07-19T02:44:42.302507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scaled = []\nscaled_pca = []\nfor scales in scales_list:\n    scaled.append((str(scales),scale_data(df=df_original,scaler=scales,do_pca=False),False))\n    scaled_pca.append((str(scales)+\"_pca\",scale_data(df=df_original,scaler=scales,do_pca=True),True))\n    \ntransformed = []\ntransformed_pca = []\nfor transformer in transformer_list:\n    transformed.append((str(transformer),scale_data(df=df_original,scaler=scales,do_pca=False),False))\n    transformed_pca.append((str(transformer)+\"_pca\",scale_data(df=df_original,scaler=scales,do_pca=True),True))\n    \nnormalized = []\nnormalized_pca = []\nfor normalizer in normalizer_list:\n    normalized.append((str(normalizer),scale_data(df=df_original,scaler=scales,do_pca=False),False))\n    normalized_pca.append((str(normalizer)+\"_pca\",scale_data(df=df_original,scaler=scales,do_pca=True),True))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:42.305410Z","iopub.execute_input":"2022-07-19T02:44:42.305792Z","iopub.status.idle":"2022-07-19T02:44:46.021401Z","shell.execute_reply.started":"2022-07-19T02:44:42.305748Z","shell.execute_reply":"2022-07-19T02:44:46.020064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_normal_features_only = df_original[ua_normal[\"col_name\"].reset_index(drop=True).tolist()]\nrob  = scaled[3][1]\nrob_pca_data  = scaled_pca[3][1]\n\npower_data = transformed[0][1]\nqt_uni_data = transformed[1][1]\nqt_norm_data = transformed[2][1]\n\npower_pca_data = transformed_pca[0][1]\nqt_uni_pca_data = transformed_pca[1][1]\nqt_norm_pca_data = transformed_pca[2][1]\n\nl1_data = normalized[0][1]\nl2_data = normalized[1][1]\nmax_data = normalized[1][1]\n\nl1_pca_data = normalized_pca[0][1]\nl2_pca_data = normalized_pca[1][1]\nmax_pca_data = normalized_pca[1][1]","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.023530Z","iopub.execute_input":"2022-07-19T02:44:46.024278Z","iopub.status.idle":"2022-07-19T02:44:46.047029Z","shell.execute_reply.started":"2022-07-19T02:44:46.024221Z","shell.execute_reply":"2022-07-19T02:44:46.043968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# K-Means Clustering\n\nThis approach partitions the data into $K$ distinct clusters. We need to specify the number of clusters beforehand, that is, $K$ must be specified. \n\nThe K-Means clustering algorithm originates from set theory, if we were to define sets from $C_1$ to $C_K$, the union of these sets would give all the observations. \n$$C_1 \\cup C_2 \\cup C_3 \\cup ... \\cup C_K =\\{1,...,n\\}$$\nThis means, each observation belongs to at least one of the $K$ clusters.\n\nThe clusters are also non overlapping/distinct, no observation is in more than one cluster.\nSay we have some clutser $C_b$, then the intersection of that cluster and some other cluster $C_{b'}$, is the empty set, \n$$C_b \\cap C_{b'}=\\emptyset$$\n\nWe want to reduce the *within cluster variation* as much as possible, this can be thought of like a loss function for supervised learning problems.\n\nThis can be written as:\n\n$$\\underset{C_1,...,C_K}{\\min}\\Biggr[\\sum\\limits_{k=1}^{K} W(C_K)\\Biggr]$$\n\nThe within cluster variation; $W(C_K)$ can be interpted as the amount by which the observations within each cluster differs.\n\nMost common approach in minimizing this within cluster variation is to use the **squared Euclidean distance**.\n\n$$W(C_k) = \\frac{1}{|C_K|}\\Biggr[\\sum\\limits_{i,i^{'} \\in C_K}\\sum\\limits_{j=1}^{p} (x_{ij}-x_{i^{'}j})^2\\Biggr]$$\n\nwhere $|C_K|$ is the number of observations in the cluster, $p$ is the number of features, $i$ and $i^{'}$ are some observations in the cluster, $i$ $\\neq$ $i^{'}$. In essence, for every pair of observation in the cluster, it's squared distance is calculated and then divided by the number of observations in that respective cluster. \n\nI suspect this may be computationlly expensive.\n\nThis can be finally written as: \n\n$$\\underset{C_1,...,C_K}{\\min}\\Biggr[\\sum\\limits_{k=1}^{K}\\frac{1}{|C_K|}\\sum\\limits_{i,i^{'} \\in C_K}\\sum\\limits_{j=1}^{p} (x_{ij}-x_{i^{'}j})^2\\Biggr]$$\n\nThe KMeans Clustering generally works as follows: \n\nYou first randomly initialize a number from $1$ to $K$ for the respective observations.\n\nThen, we repeat the following until the cluster assignments stop changing. \n* Compute the cluster centroid (the centroid is the vector of the $p$ feature means (the means in K-means) for the observations in the $k^{th}$ cluster. Each observation is assgned to the cluster whose centroid is closet (closest wrt to Euclidean distance).\n\nThe algorithm is guranteed to converge in a finite number of steps, that is the clustering will keep improving until it no longer changes meaning a local optimum has been reached. The local optimum that is reached depends on the initial random assigmenet of clusters. ","metadata":{}},{"cell_type":"code","source":"#Plotly \ndef clusters_sub_plots(x,\n                       y,\n                       x_lab,\n                       y_lab,\n                       clusters_list,\n                       model_name,\n                       cluster_range,\n                       data_name,\n                       colorscale=None,\n                       title=None):\n    titles = []\n    for cluster_list in clusters_list:\n        titles.append(str(cluster_list[0])+\" Clusters\")\n    if colorscale is None:\n        colorscale = \"Jet\"#options:picnic,Jet\n    fig = make_subplots(rows=3, cols=3,subplot_titles = tuple(titles))\n    counter:int = 0\n    row:int = 1\n    for cluster_list in clusters_list:\n        counter+=1\n        if counter % 4 == 0:\n            row+=1\n            counter = 1\n        fig.add_trace(go.Scattergl(\n        x = x,\n        y = y,\n        mode='markers',\n        marker=dict(\n            color=cluster_list[1],\n            colorscale=colorscale,\n            line_width=1,\n            showscale=not True,\n            )\n    ),row=row,col=counter)\n    \n    if x_lab is None: \n        x_lab = str(x.name)\n    if y_lab is None:\n        y_lab = str(y.name)\n    fig.update_xaxes(title = x_lab)\n    fig.update_yaxes(title = y_lab)\n    if title is None:\n        title = \"Visualizing \"+str(cluster_range[0])+\" to \"+str(cluster_range[-1])+\" clusters for \"+model_name+\" using \"+data_name+\" data\"\n    fig.update_layout(title=title)\n    fig.update_layout(height = 1000)\n    fig.update_layout(showlegend=False)\n    fig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:46.048670Z","iopub.execute_input":"2022-07-19T02:44:46.049366Z","iopub.status.idle":"2022-07-19T02:44:46.068412Z","shell.execute_reply.started":"2022-07-19T02:44:46.049319Z","shell.execute_reply":"2022-07-19T02:44:46.067164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#seaborn\ndef clusters_sub_plots_seaborn(x,\n                       y,\n                       x_lab,\n                       y_lab,\n                       clusters_list,\n                       model_name,\n                       cluster_range,\n                       data_name,\n                       colorscale=None,\n                       title=None):\n    \n    fig, axes = plt.subplots(3, 3)\n    sns.set(rc={'figure.figsize':(20,15)})    \n    row = 0\n    col = 0\n    for cluster_list in clusters_list:\n        sns.scatterplot(x=x, y=y,hue=cluster_list[1] ,ax=axes[row,col]).set_title(str(cluster_list[0])+\" Clusters\")\n        col +=1\n        if col == 3:\n            col = 0\n            row += 1\n    plt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:46.070404Z","iopub.execute_input":"2022-07-19T02:44:46.071210Z","iopub.status.idle":"2022-07-19T02:44:46.086019Z","shell.execute_reply.started":"2022-07-19T02:44:46.071163Z","shell.execute_reply":"2022-07-19T02:44:46.084755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Defining the cluster range\n\nRange: 2 clusters to 10 clusters.","metadata":{}},{"cell_type":"code","source":"cluster_range = np.arange(2,10+1)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.088266Z","iopub.execute_input":"2022-07-19T02:44:46.089303Z","iopub.status.idle":"2022-07-19T02:44:46.101290Z","shell.execute_reply.started":"2022-07-19T02:44:46.089253Z","shell.execute_reply":"2022-07-19T02:44:46.100031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.cluster import KMeans","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.103570Z","iopub.execute_input":"2022-07-19T02:44:46.104608Z","iopub.status.idle":"2022-07-19T02:44:46.111881Z","shell.execute_reply.started":"2022-07-19T02:44:46.104560Z","shell.execute_reply":"2022-07-19T02:44:46.110631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import calinski_harabasz_score as chs\nfrom sklearn.metrics import davies_bouldin_score as dbs\nfrom sklearn.metrics import silhouette_score as shs","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.113925Z","iopub.execute_input":"2022-07-19T02:44:46.114894Z","iopub.status.idle":"2022-07-19T02:44:46.122174Z","shell.execute_reply.started":"2022-07-19T02:44:46.114849Z","shell.execute_reply":"2022-07-19T02:44:46.121402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def clusters_info_kmeans(df,cluster_range):\n    summary_df = pd.DataFrame()\n    clusters_list = []\n    inertia_list = []\n    chs_scores = []\n    dbs_scores = []\n    silo_scores = []\n    labels = []\n    for cluster in cluster_range:\n        kmeans = KMeans(n_clusters = cluster)\n        kmeans.fit(df)\n        clusters_list.append((cluster,kmeans.labels_))\n        inertia_list.append(kmeans.inertia_)\n        chs_scores.append(chs(df,kmeans.labels_))\n        dbs_scores.append(dbs(df,kmeans.labels_))\n        labels.append(kmeans.labels_)\n        silo_scores.append(shs(df,kmeans.labels_))\n    \n    summary_df[\"clusters\"] = cluster_range\n    summary_df[\"labels\"] = labels\n    summary_df[\"inertia\"] = inertia_list\n    summary_df[\"calinski_harabasz_score\"] = chs_scores\n    summary_df[\"davies_bouldin_score\"] = dbs_scores\n    summary_df[\"silo_score\"] = silo_scores\n    return clusters_list,summary_df","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.123934Z","iopub.execute_input":"2022-07-19T02:44:46.124688Z","iopub.status.idle":"2022-07-19T02:44:46.135181Z","shell.execute_reply.started":"2022-07-19T02:44:46.124641Z","shell.execute_reply":"2022-07-19T02:44:46.134373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def disp_kmeans_scores(summary_df):\n    \n    fig, axes = plt.subplots(2, 2)\n    sns.set(rc={'figure.figsize':(30,20)})    \n    sns.lineplot(data = summary_df,x=\"clusters\", y=\"inertia\" ,ax=axes[0,0]).set(title='Inertia vs Amt of Clusters')\n    sns.lineplot(data = summary_df,x=\"clusters\", y=\"calinski_harabasz_score\" ,ax=axes[0,1]).set(title='Calinski Harabasaz Score vs Amt of Clusters')\n    sns.lineplot(data = summary_df,x=\"clusters\", y=\"davies_bouldin_score\" ,ax=axes[1,0]).set(title='Davies Bouldin vs Amt of Clusters')\n    sns.lineplot(data = summary_df,x=\"clusters\", y=\"silo_score\" ,ax=axes[1,1]).set(title='Silohouette Score vs Amt of Clusters')\n    plt.show()\n    ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-19T02:44:46.136975Z","iopub.execute_input":"2022-07-19T02:44:46.137801Z","iopub.status.idle":"2022-07-19T02:44:46.147750Z","shell.execute_reply.started":"2022-07-19T02:44:46.137752Z","shell.execute_reply":"2022-07-19T02:44:46.146876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Scoring K-Means\n\n- Inertia\n- Silohouette Score\n- Calinski Harahasz Score\n- Davies Bouldin Score","metadata":{}},{"cell_type":"markdown","source":"### K-Means with Robust Scaled PCA Data","metadata":{}},{"cell_type":"code","source":"clusters_list,summary_df_k_means_rob_scaled = clusters_info_kmeans(rob_pca_data,cluster_range)\ndisp_kmeans_scores(summary_df_k_means_rob_scaled)\nclusters_sub_plots_seaborn(rob_pca_data[\"pca_1\"],rob_pca_data[\"pca_2\"],None,None,clusters_list,\"KMeans\",cluster_range,\"Robust Scaled with PCA\")","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.149156Z","iopub.execute_input":"2022-07-19T02:44:46.149731Z","iopub.status.idle":"2022-07-19T02:44:46.163907Z","shell.execute_reply.started":"2022-07-19T02:44:46.149672Z","shell.execute_reply":"2022-07-19T02:44:46.162804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### K-Means with Quantile Transformed (Normal) PCA Data","metadata":{}},{"cell_type":"code","source":"clusters_list,summary_df_k_means_qt_normal_pca = clusters_info_kmeans(qt_norm_pca_data,cluster_range)\ndisp_kmeans_scores(summary_df_k_means_qt_normal_pca)\nclusters_sub_plots_seaborn(qt_norm_pca_data[\"pca_1\"],qt_norm_pca_data[\"pca_2\"],None,None,clusters_list,\"KMeans\",cluster_range,\"Quantile Transform (Normal) with PCA\")","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.165394Z","iopub.execute_input":"2022-07-19T02:44:46.166012Z","iopub.status.idle":"2022-07-19T02:44:46.175396Z","shell.execute_reply.started":"2022-07-19T02:44:46.165975Z","shell.execute_reply":"2022-07-19T02:44:46.174218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### K-Means with L2 Normalized PCA Data","metadata":{}},{"cell_type":"code","source":"clusters_list,summary_df_k_means_l2_pca_data = clusters_info_kmeans(l2_pca_data,cluster_range)\ndisp_kmeans_scores(summary_df_k_means_l2_pca_data)\nclusters_sub_plots_seaborn(l2_pca_data[\"pca_1\"],l2_pca_data[\"pca_2\"],None,None,clusters_list,\"KMeans\",cluster_range,\"l2 Normalized with PCA\")","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.176880Z","iopub.execute_input":"2022-07-19T02:44:46.177203Z","iopub.status.idle":"2022-07-19T02:44:46.186301Z","shell.execute_reply.started":"2022-07-19T02:44:46.177172Z","shell.execute_reply":"2022-07-19T02:44:46.185067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Mixture Models\n\n1. Gaussian Mixture Models\n2. Bayesian Gaussian Mixture Models","metadata":{}},{"cell_type":"markdown","source":"## Gaussian Mixture Models","metadata":{}},{"cell_type":"code","source":"from sklearn.mixture import GaussianMixture\n\ndef cluster_info_gaussian_mixture(df,cluster_range):\n    \n    summary_df = pd.DataFrame()\n    clusters_list = []\n    bic_score = []\n    aic_score = []\n    did_converge=[]\n    chs_scores = []\n    dbs_scores = []\n    labels = []\n    \n    for cluster in cluster_range:\n        gm = GaussianMixture(n_components=cluster)\n        gm.fit(df)\n        y_pred = gm.predict(df)\n        labels.append(y_pred)\n        clusters_list.append((cluster,y_pred))\n        did_converge.append(gm.converged_)\n        chs_scores.append(chs(df,y_pred))\n        dbs_scores.append(dbs(df,y_pred))\n        bic_score.append(gm.bic(df))\n        aic_score.append(gm.aic(df))\n        \n        \n    summary_df[\"clusters\"] = cluster_range\n    summary_df[\"labels\"] = labels\n    summary_df[\"did_converge\"] = did_converge\n    summary_df[\"calinski_harabasz_score\"] = chs_scores\n    summary_df[\"davies_bouldin_score\"] = dbs_scores\n    summary_df[\"aic\"] = aic_score\n    summary_df[\"bic\"] = bic_score\n    \n    return clusters_list,summary_df\n        ","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.187638Z","iopub.execute_input":"2022-07-19T02:44:46.187995Z","iopub.status.idle":"2022-07-19T02:44:46.200066Z","shell.execute_reply.started":"2022-07-19T02:44:46.187964Z","shell.execute_reply":"2022-07-19T02:44:46.198828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Scoring Gaussian Mixture\n\n- Calinski Harabasz Score\n- Davies Bouldin Score\n- AIC - Akaike Information Criterion\n- BIC - Bayesian Infomration Criterion","metadata":{}},{"cell_type":"code","source":"def disp_gm_scores(summary_df):\n    \n    fig, axes = plt.subplots(2, 2)\n    sns.set(rc={'figure.figsize':(30,20)})    \n    \n    sns.lineplot(data = summary_df,x=\"clusters\", y=\"calinski_harabasz_score\" ,ax=axes[0,0]).set(title='Calinski Harabasaz Score vs Amt of Clusters')\n    sns.lineplot(data = summary_df,x=\"clusters\", y=\"davies_bouldin_score\" ,ax=axes[0,1]).set(title='Davies Bouldin vs Amt of Clusters')\n    sns.lineplot(data = summary_df,x=\"clusters\", y=\"aic\" ,ax=axes[1,0]).set(title='AIC vs Amt of Clusters')\n    sns.lineplot(data = summary_df,x=\"clusters\", y=\"bic\" ,ax=axes[1,1]).set(title='BIC vs Amt of Clusters')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.201475Z","iopub.execute_input":"2022-07-19T02:44:46.202317Z","iopub.status.idle":"2022-07-19T02:44:46.215169Z","shell.execute_reply.started":"2022-07-19T02:44:46.202284Z","shell.execute_reply":"2022-07-19T02:44:46.214187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### GM with L2 Normalized with PCA Data","metadata":{}},{"cell_type":"code","source":"clusters_list,summary_gm_l2_pca_data = cluster_info_gaussian_mixture(l2_pca_data,cluster_range)\ndisp_gm_scores(summary_gm_l2_pca_data)\nclusters_sub_plots_seaborn(l2_pca_data[\"pca_1\"],l2_pca_data[\"pca_2\"],None,None,clusters_list,\"Gaussian Mixture\",cluster_range,\"l2 Normalized with PCA\")","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.216610Z","iopub.execute_input":"2022-07-19T02:44:46.217217Z","iopub.status.idle":"2022-07-19T02:44:46.225314Z","shell.execute_reply.started":"2022-07-19T02:44:46.217182Z","shell.execute_reply":"2022-07-19T02:44:46.224533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### GM with Power Transformed with PCA Data","metadata":{}},{"cell_type":"code","source":"clusters_list,summary_gm_power_pca_data = cluster_info_gaussian_mixture(power_pca_data,cluster_range)\ndisp_gm_scores(summary_gm_power_pca_data)\nclusters_sub_plots_seaborn(power_pca_data[\"pca_1\"],power_pca_data[\"pca_2\"],None,None,clusters_list,\"Gaussian Mixture\",cluster_range,\"l2 Normalized with PCA\")","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.226500Z","iopub.execute_input":"2022-07-19T02:44:46.227276Z","iopub.status.idle":"2022-07-19T02:44:46.237725Z","shell.execute_reply.started":"2022-07-19T02:44:46.227226Z","shell.execute_reply":"2022-07-19T02:44:46.236296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Bayesian Gaussian Mixture Model","metadata":{}},{"cell_type":"code","source":"from sklearn.mixture import BayesianGaussianMixture\n\ndef cluster_info_bgm(df,cluster_range):\n    \n    summary_df = pd.DataFrame()\n    clusters_list = []\n    did_converge=[]\n    chs_scores = []\n    dbs_scores = []\n    labels = []\n    \n    for cluster in cluster_range:\n        bgm = BayesianGaussianMixture(n_components=cluster)\n        bgm.fit(df)\n        y_pred = bgm.predict(df)\n        labels.append(y_pred)\n        clusters_list.append((cluster,y_pred))\n        did_converge.append(bgm.converged_)\n        chs_scores.append(chs(df,y_pred))\n        dbs_scores.append(dbs(df,y_pred))\n        \n        \n    summary_df[\"clusters\"] = cluster_range\n    summary_df[\"labels\"] = labels\n    summary_df[\"did_converge\"] = did_converge\n    summary_df[\"calinski_harabasz_score\"] = chs_scores\n    summary_df[\"davies_bouldin_score\"] = dbs_scores\n    \n    return clusters_list,summary_df","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.239150Z","iopub.execute_input":"2022-07-19T02:44:46.240098Z","iopub.status.idle":"2022-07-19T02:44:46.252519Z","shell.execute_reply.started":"2022-07-19T02:44:46.240062Z","shell.execute_reply":"2022-07-19T02:44:46.251740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Scoring Bayesian Gaussian Mixtures\n\n- Calinski Harabasz Score\n- Davies Bouldin Score","metadata":{}},{"cell_type":"code","source":"def disp_bgm_scores(summary_df):\n    \n    fig, axes = plt.subplots(1, 2)\n    sns.set(rc={'figure.figsize':(20,15)})    \n    sns.lineplot(data = summary_df,x=\"clusters\", y=\"calinski_harabasz_score\" ,ax=axes[0]).set(title='Calinski Harabasaz Score vs Amt of Clusters')\n    sns.lineplot(data = summary_df,x=\"clusters\", y=\"davies_bouldin_score\" ,ax=axes[1]).set(title='Davies Bouldin vs Amt of Clusters')\n    \n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.253796Z","iopub.execute_input":"2022-07-19T02:44:46.254657Z","iopub.status.idle":"2022-07-19T02:44:46.263129Z","shell.execute_reply.started":"2022-07-19T02:44:46.254623Z","shell.execute_reply":"2022-07-19T02:44:46.262025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### BGM with L2 Normalized with PCA Data","metadata":{}},{"cell_type":"code","source":"clusters_list,summary_bgm_l2_pca_data = cluster_info_bgm(l2_pca_data,cluster_range)\ndisp_bgm_scores(summary_bgm_l2_pca_data)\nclusters_sub_plots_seaborn(l2_pca_data[\"pca_1\"],l2_pca_data[\"pca_2\"],None,None,clusters_list,\"Bayesian Gaussian Mixture\",cluster_range,\"l2 Normalized with PCA\")","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.264970Z","iopub.execute_input":"2022-07-19T02:44:46.265304Z","iopub.status.idle":"2022-07-19T02:44:46.274428Z","shell.execute_reply.started":"2022-07-19T02:44:46.265273Z","shell.execute_reply":"2022-07-19T02:44:46.273316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### BGM with Quantile Transformed (Normal) with PCA Data","metadata":{}},{"cell_type":"code","source":"clusters_list,summary_bgm_qt_normal_pca_data = cluster_info_bgm(qt_norm_pca_data,cluster_range)\ndisp_bgm_scores(summary_bgm_qt_normal_pca_data)\nclusters_sub_plots_seaborn(qt_norm_pca_data[\"pca_1\"],qt_norm_pca_data[\"pca_2\"],None,None,clusters_list,\" Bayesian Gaussian Mixture\",cluster_range,\"QT Normal with PCA\")","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.275793Z","iopub.execute_input":"2022-07-19T02:44:46.276205Z","iopub.status.idle":"2022-07-19T02:44:46.284621Z","shell.execute_reply.started":"2022-07-19T02:44:46.276174Z","shell.execute_reply":"2022-07-19T02:44:46.283464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# clust_df = summary_bgm_l2_pca_data.query(\"clusters==7\")\n# output = pd.DataFrame()\n# output[\"id\"] = id_col\n# output[\"predicted\"] =clust_df[\"labels\"].reset_index(drop=True)[0].tolist()\n# output.to_csv(\"submission.csv\",index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-19T02:44:46.286890Z","iopub.execute_input":"2022-07-19T02:44:46.287420Z","iopub.status.idle":"2022-07-19T02:44:46.298413Z","shell.execute_reply.started":"2022-07-19T02:44:46.287285Z","shell.execute_reply":"2022-07-19T02:44:46.297579Z"},"trusted":true},"execution_count":null,"outputs":[]}]}