{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# TPS-JUN22 Unsupervised Clustering Using PyCaret ⚡ \n## Tutorial Level Beginner\n\n**Created using: PyCaret 3.0.0.rc3** <br>\n**Date Updated: July 17, 2022**\n\n\nHello Kaggle In this Notebook, I developed an End to End model using one of my favorites **PyCaret** The strategy that I will follow is super simple, I took most of my inspiration from the online tutorials available at PyCaret\n\nhttps://github.com/pycaret/pycaret/blob/master/tutorials/Clustering%20Tutorial%20Level%20Beginner%20-%20CLU101.ipynb\n\n# 1.0 Notebook Objective\n\n**In this Notebook you will learn:**\n\n* Getting Data: How to import data from PyCaret repository\n* Setting up Environment: How to setup an experiment in PyCaret and get started with building multiclass models\n* Create Model: How to create a model and assign cluster labels to the original dataset for analysis\n* Plot Model: How to analyze model performance using various plots\n* Predict Model: How to assign cluster labels to new and unseen datasets based on a trained model\n* Save / Load Model: How to save / load model for future use\n\n\nRead Time : Approx. 25 Minutes\n\n**Credits and References:**\n* https://github.com/pycaret/pycaret/blob/master/tutorials/Clustering%20Tutorial%20Level%20Beginner%20-%20CLU101.ipynb","metadata":{}},{"cell_type":"markdown","source":"## 1.1 Installing PyCaret","metadata":{}},{"cell_type":"code","source":"%%capture\n!pip install --pre pycaret","metadata":{"execution":{"iopub.status.busy":"2022-07-18T01:48:56.707284Z","iopub.execute_input":"2022-07-18T01:48:56.707876Z","iopub.status.idle":"2022-07-18T01:49:37.768554Z","shell.execute_reply.started":"2022-07-18T01:48:56.707793Z","shell.execute_reply":"2022-07-18T01:49:37.767146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1.2 Pre-Requisites\n* PyCaret 3.0 or greater\n* Python 3.6 or greater\n* Internet connection to load data from pycaret's repository\n* Basic Knowledge of Clustering","metadata":{}},{"cell_type":"markdown","source":"## 1.3 Importing Libraries","metadata":{}},{"cell_type":"code","source":"%%time\n# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-18T01:49:37.771073Z","iopub.execute_input":"2022-07-18T01:49:37.771531Z","iopub.status.idle":"2022-07-18T01:49:37.783206Z","shell.execute_reply.started":"2022-07-18T01:49:37.771487Z","shell.execute_reply":"2022-07-18T01:49:37.782382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport pycaret\nfrom pycaret.clustering import *","metadata":{"execution":{"iopub.status.busy":"2022-07-18T01:49:37.784772Z","iopub.execute_input":"2022-07-18T01:49:37.785537Z","iopub.status.idle":"2022-07-18T01:49:42.755883Z","shell.execute_reply.started":"2022-07-18T01:49:37.785491Z","shell.execute_reply":"2022-07-18T01:49:42.754422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1.4 Configuring the Environment","metadata":{}},{"cell_type":"code","source":"%%time\n# I like to disable my Notebook Warnings To Reduce Noice.\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-07-18T01:49:42.758619Z","iopub.execute_input":"2022-07-18T01:49:42.758957Z","iopub.status.idle":"2022-07-18T01:49:42.766059Z","shell.execute_reply.started":"2022-07-18T01:49:42.758918Z","shell.execute_reply":"2022-07-18T01:49:42.764919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport plotly.express as px\nimport plotly.io as pio\n\npx.defaults.width = 600\npx.defaults.height = 200","metadata":{"execution":{"iopub.status.busy":"2022-07-18T01:49:42.767531Z","iopub.execute_input":"2022-07-18T01:49:42.768393Z","iopub.status.idle":"2022-07-18T01:49:42.781503Z","shell.execute_reply.started":"2022-07-18T01:49:42.768348Z","shell.execute_reply":"2022-07-18T01:49:42.780156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Notebook Configuration...\n\n# Amount of data we want to load into the Model...\nDATA_ROWS = None\n# Dataframe, the amount of rows and cols to visualize...\nNROWS = 25\nNCOLS = 20\n# Main data location path...\nBASE_PATH = '...'","metadata":{"execution":{"iopub.status.busy":"2022-07-18T01:49:42.782481Z","iopub.execute_input":"2022-07-18T01:49:42.782783Z","iopub.status.idle":"2022-07-18T01:49:42.792231Z","shell.execute_reply.started":"2022-07-18T01:49:42.782756Z","shell.execute_reply":"2022-07-18T01:49:42.791143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Configure notebook display settings to only use 2 decimal places, tables look nicer.\npd.options.display.float_format = '{:,.2f}'.format\npd.set_option('display.max_columns', NCOLS) \npd.set_option('display.max_rows', NROWS)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T01:49:42.793521Z","iopub.execute_input":"2022-07-18T01:49:42.793940Z","iopub.status.idle":"2022-07-18T01:49:42.802998Z","shell.execute_reply.started":"2022-07-18T01:49:42.793906Z","shell.execute_reply":"2022-07-18T01:49:42.802261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 2.0 Getting the Data","metadata":{}},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\ndataset = pd.read_csv('/kaggle/input/tabular-playground-series-jul-2022/data.csv')\nsubmission = pd.read_csv(\"../input/tabular-playground-series-jul-2022/sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-18T01:49:42.804343Z","iopub.execute_input":"2022-07-18T01:49:42.804946Z","iopub.status.idle":"2022-07-18T01:49:44.082044Z","shell.execute_reply.started":"2022-07-18T01:49:42.804914Z","shell.execute_reply":"2022-07-18T01:49:44.080802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 3.0 Exploring the Datasets","metadata":{}},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\ndataset.info(verbose = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T01:49:44.083196Z","iopub.execute_input":"2022-07-18T01:49:44.083493Z","iopub.status.idle":"2022-07-18T01:49:44.102920Z","shell.execute_reply.started":"2022-07-18T01:49:44.083466Z","shell.execute_reply":"2022-07-18T01:49:44.101769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Placeholder, Describe Code...\ndataset.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T01:49:44.107124Z","iopub.execute_input":"2022-07-18T01:49:44.107533Z","iopub.status.idle":"2022-07-18T01:49:44.137449Z","shell.execute_reply.started":"2022-07-18T01:49:44.107501Z","shell.execute_reply":"2022-07-18T01:49:44.136592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 4.0 Setting up Environment in PyCaret","metadata":{}},{"cell_type":"code","source":"%%time\ncluster_01 = setup(dataset, normalize = True, ignore_features = ['id'], session_id = 123)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T01:49:44.138645Z","iopub.execute_input":"2022-07-18T01:49:44.139210Z","iopub.status.idle":"2022-07-18T01:49:45.924313Z","shell.execute_reply.started":"2022-07-18T01:49:44.139178Z","shell.execute_reply":"2022-07-18T01:49:45.922984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 5.0 Create a Model","metadata":{}},{"cell_type":"markdown","source":"## 5.1 K-means","metadata":{}},{"cell_type":"code","source":"%%time\nkmeans = create_model('kmeans', num_clusters = 5)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T01:49:45.925996Z","iopub.execute_input":"2022-07-18T01:49:45.926612Z","iopub.status.idle":"2022-07-18T01:52:27.010767Z","shell.execute_reply.started":"2022-07-18T01:49:45.926564Z","shell.execute_reply":"2022-07-18T01:52:27.009608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nprint(kmeans)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T01:52:27.012171Z","iopub.execute_input":"2022-07-18T01:52:27.012615Z","iopub.status.idle":"2022-07-18T01:52:27.019798Z","shell.execute_reply.started":"2022-07-18T01:52:27.012571Z","shell.execute_reply":"2022-07-18T01:52:27.018671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.2 Meanshift","metadata":{}},{"cell_type":"code","source":"%%time\nmeanshift = create_model('meanshift')","metadata":{"execution":{"iopub.status.busy":"2022-07-18T01:52:27.021114Z","iopub.execute_input":"2022-07-18T01:52:27.021484Z","iopub.status.idle":"2022-07-18T02:59:20.618461Z","shell.execute_reply.started":"2022-07-18T01:52:27.021451Z","shell.execute_reply":"2022-07-18T02:59:20.617120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nprint(meanshift)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T02:59:20.620063Z","iopub.execute_input":"2022-07-18T02:59:20.620869Z","iopub.status.idle":"2022-07-18T02:59:20.628882Z","shell.execute_reply.started":"2022-07-18T02:59:20.620833Z","shell.execute_reply":"2022-07-18T02:59:20.627766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nmodels()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T02:59:20.630190Z","iopub.execute_input":"2022-07-18T02:59:20.630532Z","iopub.status.idle":"2022-07-18T02:59:20.652305Z","shell.execute_reply.started":"2022-07-18T02:59:20.630503Z","shell.execute_reply":"2022-07-18T02:59:20.651245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 6.0 Assign a Model","metadata":{}},{"cell_type":"code","source":"%%time\nkmean_results = assign_model(kmeans)\nkmean_results.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T02:59:20.653802Z","iopub.execute_input":"2022-07-18T02:59:20.654265Z","iopub.status.idle":"2022-07-18T02:59:20.722293Z","shell.execute_reply.started":"2022-07-18T02:59:20.654222Z","shell.execute_reply":"2022-07-18T02:59:20.721066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 7.0 Plot a Model","metadata":{}},{"cell_type":"markdown","source":"## 7.1 Cluster PCA Plot","metadata":{}},{"cell_type":"code","source":"%%time\nplot_model(kmeans, scale = 0.7)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T02:59:20.723686Z","iopub.execute_input":"2022-07-18T02:59:20.724007Z","iopub.status.idle":"2022-07-18T02:59:22.813463Z","shell.execute_reply.started":"2022-07-18T02:59:20.723978Z","shell.execute_reply":"2022-07-18T02:59:22.810554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 7.2 Distribution Plot","metadata":{}},{"cell_type":"code","source":"%%time\nplot_model(kmeans, plot = 'distribution') #to see size of clusters","metadata":{"execution":{"iopub.status.busy":"2022-07-18T02:59:22.815263Z","iopub.execute_input":"2022-07-18T02:59:22.816327Z","iopub.status.idle":"2022-07-18T02:59:24.594294Z","shell.execute_reply.started":"2022-07-18T02:59:22.816285Z","shell.execute_reply":"2022-07-18T02:59:24.592969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 7.3 Elbow Plot","metadata":{}},{"cell_type":"code","source":"%%time\nplot_model(kmeans, plot = 'elbow', scale = 1.2)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T02:59:24.596265Z","iopub.execute_input":"2022-07-18T02:59:24.596795Z","iopub.status.idle":"2022-07-18T03:00:26.351098Z","shell.execute_reply.started":"2022-07-18T02:59:24.596732Z","shell.execute_reply":"2022-07-18T03:00:26.349879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 7.4 Silhouette Plot","metadata":{}},{"cell_type":"code","source":"%%time\nplot_model(kmeans, plot = 'silhouette', scale = 1.2)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T03:00:26.352589Z","iopub.execute_input":"2022-07-18T03:00:26.353330Z","iopub.status.idle":"2022-07-18T03:05:42.185104Z","shell.execute_reply.started":"2022-07-18T03:00:26.353292Z","shell.execute_reply":"2022-07-18T03:05:42.183884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 8.0 Predict on unseen data","metadata":{}},{"cell_type":"code","source":"%%time\nunseen_predictions = predict_model(kmeans, data = dataset)\nunseen_predictions.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-18T03:05:42.186695Z","iopub.execute_input":"2022-07-18T03:05:42.187010Z","iopub.status.idle":"2022-07-18T03:05:42.420356Z","shell.execute_reply.started":"2022-07-18T03:05:42.186979Z","shell.execute_reply":"2022-07-18T03:05:42.419423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nsubmission_pycaret = pd.read_csv(\"../input/tabular-playground-series-jul-2022/sample_submission.csv\")\nsubmission_pycaret['Predicted'] = unseen_predictions['Cluster'].str[-1:].astype('int')\nsubmission_pycaret.to_csv(\"submission_lgbm.csv\",index = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-18T03:05:42.421890Z","iopub.execute_input":"2022-07-18T03:05:42.422213Z","iopub.status.idle":"2022-07-18T03:05:42.590252Z","shell.execute_reply.started":"2022-07-18T03:05:42.422185Z","shell.execute_reply":"2022-07-18T03:05:42.589193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nsubmission_pycaret","metadata":{"execution":{"iopub.status.busy":"2022-07-18T03:05:42.591403Z","iopub.execute_input":"2022-07-18T03:05:42.592406Z","iopub.status.idle":"2022-07-18T03:05:42.606652Z","shell.execute_reply.started":"2022-07-18T03:05:42.592368Z","shell.execute_reply":"2022-07-18T03:05:42.605426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 9.0 Saving the model","metadata":{}},{"cell_type":"code","source":"%%time\nsave_model(kmeans,'Final KMeans Model 17Jul2022')","metadata":{"execution":{"iopub.status.busy":"2022-07-18T03:05:42.608662Z","iopub.execute_input":"2022-07-18T03:05:42.609595Z","iopub.status.idle":"2022-07-18T03:05:42.849936Z","shell.execute_reply.started":"2022-07-18T03:05:42.609545Z","shell.execute_reply":"2022-07-18T03:05:42.848757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# 10.0 Loading the saved model","metadata":{}},{"cell_type":"code","source":"%%time\nsaved_kmeans = load_model('Final KMeans Model 17Jul2022')","metadata":{"execution":{"iopub.status.busy":"2022-07-18T03:05:42.851922Z","iopub.execute_input":"2022-07-18T03:05:42.852388Z","iopub.status.idle":"2022-07-18T03:05:42.861595Z","shell.execute_reply.started":"2022-07-18T03:05:42.852345Z","shell.execute_reply":"2022-07-18T03:05:42.860162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}}]}