{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# TPS-JUN22; Unsupervised Clustering with Keras 🍇\n\n\n\nIn this Notebook I will be using Tensorflow to identify the Clusters in the Dataset...\nFirst I will build a simple model using K-Means as a Baseline and then move to Tensorflow, to implement the model in TF I will be building a \n\nI'm just getting started so give me a few hours to update everything...\nI plan to follow this strategy on the Notebook...\n\n**Notebook Strategy**\n\nThis is more a list of logical steps that I ussually like to implement on my Notebooks; the idea is to describe to the audience how I will aproach this type of problem\n* Loading all the requiered Libraries\n* Setup the Notebook for Optimal Usage\n* Read all the Information into Pandas DataFrames\n* Understand the Information Loaded\n* Complete Some Exploratory Data Analysis\n* Build Features and Preprocess some of the Information\n* Build A Simple K-Means Model; Select optimal Clusters based on the Elbow Methodology\n* Implement a Tensor Flow Model for Clustering, Probably an Ecoder Decoder Model\n* Estimate the Clusters and Upload to Kaggle for Evaluation\n\n\n**Credits and References**\n\nBelow some of the blog post and other sources of informatio that I have used to produce this Notebook.\n\n**Website Articles**\n\n1. https://www.dlology.com/blog/how-to-do-unsupervised-clustering-with-keras/\n2. https://www.geeksforgeeks.org/elbow-method-for-optimal-value-of-k-in-kmeans/\n\n**Other Notebooks**\n\n1. https://www.kaggle.com/code/thedevastator/bruteforce-clustering/notebook?scriptVersionId=100001658\n2. ...\n\n---\n","metadata":{}},{"cell_type":"markdown","source":"## Work in Progress...\n## Still working on the Baseline, Notebook has not been refactored yet.","metadata":{}},{"cell_type":"markdown","source":"# 1. Loading All the Requiered Libraries...","metadata":{}},{"cell_type":"markdown","source":"## 1.1. Default Kaggle Libraries","metadata":{}},{"cell_type":"code","source":"%%time\n# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-05T00:27:42.214159Z","iopub.execute_input":"2022-07-05T00:27:42.214456Z","iopub.status.idle":"2022-07-05T00:27:42.222786Z","shell.execute_reply.started":"2022-07-05T00:27:42.214427Z","shell.execute_reply":"2022-07-05T00:27:42.221683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1.2. Extra Libraries Needed","metadata":{}},{"cell_type":"code","source":"%%time\n# Model Libraries for K-Means\nfrom sklearn.cluster import KMeans\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn import metrics\nfrom sklearn.decomposition import PCA\nfrom sklearn.mixture import GaussianMixture, BayesianGaussianMixture\nfrom sklearn.preprocessing import StandardScaler, RobustScaler, PowerTransformer\n\n# Scientific Libraries for Calculations\nfrom scipy.spatial.distance import cdist\n\n# Visualization Libraries\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:42.240705Z","iopub.execute_input":"2022-07-05T00:27:42.241310Z","iopub.status.idle":"2022-07-05T00:27:42.248754Z","shell.execute_reply.started":"2022-07-05T00:27:42.241278Z","shell.execute_reply":"2022-07-05T00:27:42.247815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 2. Configurating the Notebook","metadata":{}},{"cell_type":"markdown","source":"## 2.1. Setting the Amount of Decimals, Rows, Columns and Warnings","metadata":{}},{"cell_type":"code","source":"%%time\n# I like to disable my Notebook Warnings To Reduce Noice.\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:42.273593Z","iopub.execute_input":"2022-07-05T00:27:42.274168Z","iopub.status.idle":"2022-07-05T00:27:42.280424Z","shell.execute_reply.started":"2022-07-05T00:27:42.274142Z","shell.execute_reply":"2022-07-05T00:27:42.279381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Notebook Configuration...\n\n# Amount of data we want to load into the Model...\nDATA_ROWS = None\n# Dataframe, the amount of rows and cols to visualize...\nNROWS = 25\nNCOLS = 10\n# Main data location path...\nBASE_PATH = '...'","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:42.282217Z","iopub.execute_input":"2022-07-05T00:27:42.283129Z","iopub.status.idle":"2022-07-05T00:27:42.292019Z","shell.execute_reply.started":"2022-07-05T00:27:42.283060Z","shell.execute_reply":"2022-07-05T00:27:42.290927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Configure notebook display settings to only use 2 decimal places, tables look nicer.\npd.options.display.float_format = '{:,.2f}'.format\npd.set_option('display.max_columns', NCOLS) \npd.set_option('display.max_rows', NROWS)","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:42.295805Z","iopub.execute_input":"2022-07-05T00:27:42.296617Z","iopub.status.idle":"2022-07-05T00:27:42.305996Z","shell.execute_reply.started":"2022-07-05T00:27:42.296566Z","shell.execute_reply":"2022-07-05T00:27:42.304945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 3. Reading the Datasets","metadata":{}},{"cell_type":"markdown","source":"## 3.1. Load the Information into Pandas","metadata":{}},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\ndataset = pd.read_csv('/kaggle/input/tabular-playground-series-jul-2022/data.csv')\nsubmission = pd.read_csv(\"../input/tabular-playground-series-jul-2022/sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:42.307767Z","iopub.execute_input":"2022-07-05T00:27:42.308077Z","iopub.status.idle":"2022-07-05T00:27:42.747933Z","shell.execute_reply.started":"2022-07-05T00:27:42.308050Z","shell.execute_reply":"2022-07-05T00:27:42.746979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 4. Exploring the Information Loaded","metadata":{}},{"cell_type":"markdown","source":"## 4.1. DataFrame Information","metadata":{}},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\ndataset.info(verbose = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:42.749870Z","iopub.execute_input":"2022-07-05T00:27:42.750196Z","iopub.status.idle":"2022-07-05T00:27:42.761884Z","shell.execute_reply.started":"2022-07-05T00:27:42.750163Z","shell.execute_reply":"2022-07-05T00:27:42.760716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.2. Visulization of the First Records","metadata":{}},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\ndataset.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:42.763125Z","iopub.execute_input":"2022-07-05T00:27:42.763414Z","iopub.status.idle":"2022-07-05T00:27:42.785203Z","shell.execute_reply.started":"2022-07-05T00:27:42.763390Z","shell.execute_reply":"2022-07-05T00:27:42.784009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.3. Description of the DataFrame","metadata":{}},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\ndataset.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:42.787601Z","iopub.execute_input":"2022-07-05T00:27:42.787984Z","iopub.status.idle":"2022-07-05T00:27:42.933767Z","shell.execute_reply.started":"2022-07-05T00:27:42.787933Z","shell.execute_reply":"2022-07-05T00:27:42.933158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.4. Identifying Unique Values","metadata":{}},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\ndataset.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:42.934748Z","iopub.execute_input":"2022-07-05T00:27:42.935428Z","iopub.status.idle":"2022-07-05T00:27:43.043807Z","shell.execute_reply.started":"2022-07-05T00:27:42.935404Z","shell.execute_reply":"2022-07-05T00:27:43.043004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.5. Identifying Null in the DataFrame","metadata":{}},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\ndataset.isnull().sum().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:43.044885Z","iopub.execute_input":"2022-07-05T00:27:43.045927Z","iopub.status.idle":"2022-07-05T00:27:43.057827Z","shell.execute_reply.started":"2022-07-05T00:27:43.045895Z","shell.execute_reply":"2022-07-05T00:27:43.056947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\ndataset.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:43.060301Z","iopub.execute_input":"2022-07-05T00:27:43.060676Z","iopub.status.idle":"2022-07-05T00:27:43.076050Z","shell.execute_reply.started":"2022-07-05T00:27:43.060645Z","shell.execute_reply":"2022-07-05T00:27:43.075012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 5. Pre-Processing the Dataset","metadata":{}},{"cell_type":"markdown","source":"## 5.1. Separating Categorical from Numeric and Others","metadata":{}},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\ncategorical_cols = ['f_07','f_08','f_09','f_10','f_11','f_12','f_13']","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:43.077636Z","iopub.execute_input":"2022-07-05T00:27:43.077984Z","iopub.status.idle":"2022-07-05T00:27:43.083290Z","shell.execute_reply.started":"2022-07-05T00:27:43.077924Z","shell.execute_reply":"2022-07-05T00:27:43.082123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\ndataset[categorical_cols].describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:43.084835Z","iopub.execute_input":"2022-07-05T00:27:43.085133Z","iopub.status.idle":"2022-07-05T00:27:43.134362Z","shell.execute_reply.started":"2022-07-05T00:27:43.085103Z","shell.execute_reply":"2022-07-05T00:27:43.133482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## 5.2. Removing Outliers Function, Based on Quantiles","metadata":{}},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\ndef quantile_outlier(df, variables, lower = 0.05, upper = 0.95):\n    '''\n    '''\n    \n    for col in df[variables].columns:\n        lower_bound = df[col].quantile(lower)\n        upper_bound = df[col].quantile(upper)\n        df[col + '_outlier'] = np.where((df[col] > upper_bound) |  (df[col] < lower_bound), 1, 0)\n    \n    outlier_cols = [feat for feat in df.columns if '_outlier' in feat]\n    df['total_outlier'] = df[outlier_cols].sum(axis=1)\n    \n    \n    df = df.drop(columns = outlier_cols, axis = 1)\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:43.139384Z","iopub.execute_input":"2022-07-05T00:27:43.139612Z","iopub.status.idle":"2022-07-05T00:27:43.146120Z","shell.execute_reply.started":"2022-07-05T00:27:43.139591Z","shell.execute_reply":"2022-07-05T00:27:43.145412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## 5.3. OneHot Encoding Categorical Variables","metadata":{}},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\ndef one_hot_encode(df, categ_feat = categorical_cols):\n    '''\n    Convert the selected variables into a onehot encoded field...\n    \n    '''\n    df = pd.get_dummies(df, columns = categ_feat)\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:43.147025Z","iopub.execute_input":"2022-07-05T00:27:43.147256Z","iopub.status.idle":"2022-07-05T00:27:43.158539Z","shell.execute_reply.started":"2022-07-05T00:27:43.147233Z","shell.execute_reply":"2022-07-05T00:27:43.157573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## 5.4. Creating Some Features","metadata":{}},{"cell_type":"code","source":"%%time\n# Create multiple features...\ndef create_features(df):\n    '''\n    '''\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:43.159548Z","iopub.execute_input":"2022-07-05T00:27:43.159842Z","iopub.status.idle":"2022-07-05T00:27:43.176722Z","shell.execute_reply.started":"2022-07-05T00:27:43.159810Z","shell.execute_reply":"2022-07-05T00:27:43.175699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.5. Selecting and Applying Data-Preprocessing Steps","metadata":{}},{"cell_type":"code","source":"%%time\n# Apply all the transformations nesesary to process the dataset...\nskip = ['id']\nfeatures = [feat for feat in dataset.columns if feat not in skip]\n\n# Apply all the Transformation and Build the Selected Features...\n\n#dataset = one_hot_encode(dataset, categ_feat = categorical_cols)\n#dataset = quantile_outlier(dataset, variables = features)\n#dataset = create_features(dataset)","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:43.178070Z","iopub.execute_input":"2022-07-05T00:27:43.178296Z","iopub.status.idle":"2022-07-05T00:27:43.189598Z","shell.execute_reply.started":"2022-07-05T00:27:43.178267Z","shell.execute_reply":"2022-07-05T00:27:43.188528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.6. Verifying the Results of the Data Pre-Processing Steps","metadata":{}},{"cell_type":"code","source":"%%time\ndataset.sample(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:43.190976Z","iopub.execute_input":"2022-07-05T00:27:43.191488Z","iopub.status.idle":"2022-07-05T00:27:43.221592Z","shell.execute_reply.started":"2022-07-05T00:27:43.191448Z","shell.execute_reply":"2022-07-05T00:27:43.220513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndataset.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:43.223439Z","iopub.execute_input":"2022-07-05T00:27:43.223808Z","iopub.status.idle":"2022-07-05T00:27:43.368669Z","shell.execute_reply.started":"2022-07-05T00:27:43.223773Z","shell.execute_reply":"2022-07-05T00:27:43.367745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n","metadata":{}},{"cell_type":"markdown","source":"# 6. Exploratory Data Analysis","metadata":{}},{"cell_type":"markdown","source":"## 6.1. Data Visualizations","metadata":{}},{"cell_type":"code","source":"%%time\n# Seaborn visualization library\nimport seaborn as sns\n\n# Increase the size of the heatmap.\nplt.figure(figsize = (5,5))\n\n# Create the default pairplot\nsns.pairplot(dataset[categorical_cols], size = 1.5)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:27:43.369845Z","iopub.execute_input":"2022-07-05T00:27:43.370127Z","iopub.status.idle":"2022-07-05T00:28:01.561274Z","shell.execute_reply.started":"2022-07-05T00:27:43.370100Z","shell.execute_reply":"2022-07-05T00:28:01.560248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\nplt.hist(dataset['f_18'], bins = 30)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:28:01.562335Z","iopub.execute_input":"2022-07-05T00:28:01.562559Z","iopub.status.idle":"2022-07-05T00:28:01.716553Z","shell.execute_reply.started":"2022-07-05T00:28:01.562537Z","shell.execute_reply":"2022-07-05T00:28:01.715999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\nplt.scatter(dataset['f_12'], dataset['f_27'], alpha = 0.2)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:28:01.717450Z","iopub.execute_input":"2022-07-05T00:28:01.718198Z","iopub.status.idle":"2022-07-05T00:28:02.116328Z","shell.execute_reply.started":"2022-07-05T00:28:01.718169Z","shell.execute_reply":"2022-07-05T00:28:02.115447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## 6.2. Understanding the Correlations","metadata":{}},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\ncorrelation = dataset[features].corr()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:28:02.117642Z","iopub.execute_input":"2022-07-05T00:28:02.117886Z","iopub.status.idle":"2022-07-05T00:28:02.291508Z","shell.execute_reply.started":"2022-07-05T00:28:02.117864Z","shell.execute_reply":"2022-07-05T00:28:02.290458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\nimport seaborn as sns\n# Increase the size of the heatmap.\nplt.figure(figsize=(12, 8))\n\n# Store heatmap object in a variable to easily access it when you want to include more features (such as title).\n# Set the range of values to be displayed on the colormap from -1 to 1, and set the annotation to True to display the correlation values on the heatmap.\n\nheatmap = sns.heatmap(correlation, \n                      vmin=-1, \n                      vmax=1, \n                      annot=False, \n                      fmt='.2f', \n                      cmap = 'viridis')\n\n# Give a title to the heatmap. Pad defines the distance of the title from the top of the heatmap.\nheatmap.set_title('Correlation Heatmap', fontdict={'fontsize':16}, pad=12);","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:28:02.292891Z","iopub.execute_input":"2022-07-05T00:28:02.293189Z","iopub.status.idle":"2022-07-05T00:28:02.844046Z","shell.execute_reply.started":"2022-07-05T00:28:02.293162Z","shell.execute_reply":"2022-07-05T00:28:02.842914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g = sns.heatmap(correlation, \n                vmin=-1,\n                vmax=1,\n                center=0,\n                square=True,\n                linewidths=1,\n                cbar_kws={\"shrink\": 0.82},\n                fmt='.2f',\n                cmap='viridis')\n\nsns.despine()\ng.figure.set_size_inches(9,9)\n    \n# Give a title to the heatmap. Pad defines the distance of the title from the top of the heatmap.\ng.set_title('Correlation Heatmap', fontdict={'fontsize':16}, pad=12);\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T02:07:52.625964Z","iopub.execute_input":"2022-07-05T02:07:52.626329Z","iopub.status.idle":"2022-07-05T02:07:52.969012Z","shell.execute_reply.started":"2022-07-05T02:07:52.626300Z","shell.execute_reply":"2022-07-05T02:07:52.968006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 7. Data Transformation, Scaler & Normalization","metadata":{}},{"cell_type":"markdown","source":"## 7.1. Configuring the Baseline Model Data","metadata":{"execution":{"iopub.status.busy":"2022-07-03T02:37:08.511239Z","iopub.execute_input":"2022-07-03T02:37:08.511736Z","iopub.status.idle":"2022-07-03T02:37:08.534285Z","shell.execute_reply.started":"2022-07-03T02:37:08.511624Z","shell.execute_reply":"2022-07-03T02:37:08.5333Z"}}},{"cell_type":"code","source":"dataset.columns","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:28:02.845071Z","iopub.execute_input":"2022-07-05T00:28:02.845296Z","iopub.status.idle":"2022-07-05T00:28:02.852059Z","shell.execute_reply.started":"2022-07-05T00:28:02.845267Z","shell.execute_reply":"2022-07-05T00:28:02.851137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\nX = dataset.drop(\"id\", axis = 1).values\n\nX = StandardScaler().fit_transform(X)\nX = PowerTransformer().fit(X).transform(X)","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:28:02.853408Z","iopub.execute_input":"2022-07-05T00:28:02.853894Z","iopub.status.idle":"2022-07-05T00:28:02.889039Z","shell.execute_reply.started":"2022-07-05T00:28:02.853863Z","shell.execute_reply":"2022-07-05T00:28:02.888179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 8. Simple K-Means Using Elbow Technique Insights","metadata":{}},{"cell_type":"code","source":"%%time\n# Set a baseline with K-Means\n# 10 clusters\nn_clusters = 7\n# Runs in parallel All CPUs\n# Train K-Means.\nkmeans = KMeans(n_clusters = n_clusters, \n                n_init = 20,\n                algorithm = 'elkan',\n                max_iter = 300,\n                random_state = 1).fit(X)\n\nkmeans_predictions = kmeans.predict(X)","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:28:02.890532Z","iopub.execute_input":"2022-07-05T00:28:02.892151Z","iopub.status.idle":"2022-07-05T00:28:10.047286Z","shell.execute_reply.started":"2022-07-05T00:28:02.892119Z","shell.execute_reply":"2022-07-05T00:28:10.046500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\nkmeans.labels_","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:28:10.050838Z","iopub.execute_input":"2022-07-05T00:28:10.053118Z","iopub.status.idle":"2022-07-05T00:28:10.063586Z","shell.execute_reply.started":"2022-07-05T00:28:10.053091Z","shell.execute_reply":"2022-07-05T00:28:10.062007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\nkmeans.labels_","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:28:10.065881Z","iopub.execute_input":"2022-07-05T00:28:10.066266Z","iopub.status.idle":"2022-07-05T00:28:10.076086Z","shell.execute_reply.started":"2022-07-05T00:28:10.066235Z","shell.execute_reply":"2022-07-05T00:28:10.074900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 9. Baseline Model, A Simple GMM","metadata":{}},{"cell_type":"code","source":"%%time\n# Set a baseline with K-Means\n# 10 clusters\nn_components = 7\n# Runs in parallel All CPUs\n# Train Gaussian Mixture.\n\ngmm = GaussianMixture(n_components = n_components,\n                      covariance_type = 'full',\n                      tol = 0.001,\n                      reg_covar = 1e-06,\n                      max_iter = 100,\n                      n_init = 20, \n                      init_params = 'kmeans',\n                      weights_init = None,\n                      means_init = None,\n                      precisions_init = None,\n                      random_state = 1,\n                      warm_start = False,\n                      verbose = 0,\n                      verbose_interval = 10)\n\ngmm_predictions = gmm.fit_predict(X)","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:28:10.078007Z","iopub.execute_input":"2022-07-05T00:28:10.078637Z","iopub.status.idle":"2022-07-05T00:30:05.016509Z","shell.execute_reply.started":"2022-07-05T00:28:10.078614Z","shell.execute_reply":"2022-07-05T00:30:05.015713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 10. Baseline Model, A Simple BGM","metadata":{}},{"cell_type":"code","source":"%%time\n# Set a baseline with K-Means\n# 10 clusters\nn_components = 7\n# Runs in parallel All CPUs\n# Train Gaussian Mixture.\n\nbgm = BayesianGaussianMixture(n_components = n_components, \n                              covariance_type = 'full', \n                              tol = 0.001, \n                              reg_covar = 1e-06, \n                              max_iter = 100, \n                              n_init = 1, \n                              init_params = 'kmeans', \n                              weight_concentration_prior_type = 'dirichlet_process', \n                              weight_concentration_prior = None, \n                              mean_precision_prior = None, \n                              mean_prior = None, \n                              degrees_of_freedom_prior = None, \n                              covariance_prior = None, \n                              random_state = 1, \n                              warm_start = False, \n                              verbose = 0, \n                              verbose_interval = 10)\n\nbgm_predictions = bgm.fit_predict(X)","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:30:05.020039Z","iopub.execute_input":"2022-07-05T00:30:05.021970Z","iopub.status.idle":"2022-07-05T00:30:05.550806Z","shell.execute_reply.started":"2022-07-05T00:30:05.021910Z","shell.execute_reply":"2022-07-05T00:30:05.549993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 11. Unsupervised Machine Learning Example in Keras","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 12. Identifying the Optimal Number of Clusters","metadata":{}},{"cell_type":"markdown","source":"## 12.1. Creating a Multiple Models\nI have disable the code in the Cells using **%%script false --no-raise-error** the code takes 16 mins to run...","metadata":{}},{"cell_type":"code","source":"%%script false --no-raise-error\n# Placeholder, Describe Code...","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:30:05.560525Z","iopub.execute_input":"2022-07-05T00:30:05.561685Z","iopub.status.idle":"2022-07-05T00:30:05.585924Z","shell.execute_reply.started":"2022-07-05T00:30:05.561654Z","shell.execute_reply":"2022-07-05T00:30:05.584568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%script false --no-raise-error\n%%time\n# Placeholder, Describe Code...\n\ndistortions = []\ninertias = []\n\nmapping1 = {}\nmapping2 = {}\n\nmax_centroids = 30\n\nK = range(1, max_centroids)\n\n\nfor k in K:\n    # Building and fitting the model\n    kmeanModel = KMeans(n_clusters = k).fit(X)\n    kmeanModel.fit(X)\n  \n    distortions.append(sum(np.min(cdist(X, kmeanModel.cluster_centers_,\n                                        'euclidean'), axis=1)) / X.shape[0])\n    inertias.append(kmeanModel.inertia_)\n  \n    mapping1[k] = sum(np.min(cdist(X, kmeanModel.cluster_centers_,\n                                   'euclidean'), axis=1)) / X.shape[0]\n    mapping2[k] = kmeanModel.inertia_","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:30:05.587211Z","iopub.execute_input":"2022-07-05T00:30:05.587457Z","iopub.status.idle":"2022-07-05T00:30:05.607528Z","shell.execute_reply.started":"2022-07-05T00:30:05.587432Z","shell.execute_reply":"2022-07-05T00:30:05.606222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## 12.2. Visualizing the Results, Distortion Aproach","metadata":{}},{"cell_type":"code","source":"%%script false --no-raise-error\n%%time\n# Placeholder, Describe Code...\nfor key, val in mapping1.items():\n    print(f'{key} : {val}')","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:30:05.608815Z","iopub.execute_input":"2022-07-05T00:30:05.609289Z","iopub.status.idle":"2022-07-05T00:30:05.629366Z","shell.execute_reply.started":"2022-07-05T00:30:05.609247Z","shell.execute_reply":"2022-07-05T00:30:05.628415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%script false --no-raise-error\n%%time\n# Placeholder, Describe Code...\nplt.plot(K, distortions, 'bx-')\nplt.xlabel('Values of K')\nplt.ylabel('Distortion')\nplt.title('The Elbow Method using Distortion')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:30:05.630542Z","iopub.execute_input":"2022-07-05T00:30:05.631251Z","iopub.status.idle":"2022-07-05T00:30:05.652153Z","shell.execute_reply.started":"2022-07-05T00:30:05.631223Z","shell.execute_reply":"2022-07-05T00:30:05.651049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## 12.3. Visualizing the Results, Inertia Aproach","metadata":{}},{"cell_type":"code","source":"%%script false --no-raise-error\n%%time\n# Placeholder, Describe Code...\nfor key, val in mapping2.items():\n    print(f'{key} : {val}')","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:30:05.654600Z","iopub.execute_input":"2022-07-05T00:30:05.655016Z","iopub.status.idle":"2022-07-05T00:30:05.676723Z","shell.execute_reply.started":"2022-07-05T00:30:05.654975Z","shell.execute_reply":"2022-07-05T00:30:05.675441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%script false --no-raise-error\n%%time\n# Placeholder, Describe Code...\nplt.plot(K, inertias, 'bx-')\nplt.xlabel('Values of K')\nplt.ylabel('Inertia')\nplt.title('The Elbow Method using Inertia')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:30:05.678477Z","iopub.execute_input":"2022-07-05T00:30:05.679570Z","iopub.status.idle":"2022-07-05T00:30:05.698360Z","shell.execute_reply.started":"2022-07-05T00:30:05.679532Z","shell.execute_reply":"2022-07-05T00:30:05.697289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 13. Creating a Submission Dataset","metadata":{}},{"cell_type":"markdown","source":"## 13.1. Generating a CSV to Upload for Evaluation","metadata":{}},{"cell_type":"code","source":"%%time\n%%time\n# Placeholder, Describe Code...\nsubmission[\"Predicted\"] = bgm_predictions\nsubmission.to_csv(\"submission.csv\", index = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:30:05.699894Z","iopub.execute_input":"2022-07-05T00:30:05.700238Z","iopub.status.idle":"2022-07-05T00:30:05.839756Z","shell.execute_reply.started":"2022-07-05T00:30:05.700201Z","shell.execute_reply":"2022-07-05T00:30:05.838465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n%%time\n# Placeholder, Describe Code...\nsubmission.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-05T00:30:05.843092Z","iopub.execute_input":"2022-07-05T00:30:05.844495Z","iopub.status.idle":"2022-07-05T00:30:05.855242Z","shell.execute_reply.started":"2022-07-05T00:30:05.844455Z","shell.execute_reply":"2022-07-05T00:30:05.853905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}}]}