{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Tabular Playground July 2022 Series - Simple explanation\n\nHello Everyone! As you might have seen on other notebooks of this competition, many have used only 14 out of 30 features, and considered number of clusters as 7. But noone has explained why. Well this notebook is aimed at explaining those things to new learners. \n\nThis is an unsupervised learning type problem. But I have converted it into supervised learning type. I have first done clustering of some of the samples for which I am most sure of (probability > 70%), and then used that as a training dataset to predict the clusters of other samples using a classification model. \n\nI have also performed a simple hyperparameter tuning and model selection for 4 different types of classification models (i.e. XGBoost, kNN, SVC, and Naive Bayes). Since this notebook is only for explanatory purposes, I have tuned only some of the parameters and that too within a small range, otherwise the code would run for hours! ","metadata":{}},{"cell_type":"markdown","source":"**Contents**\n\n* Importing the dataset\n* Selecting important features\n* The Elbow Method\n* Bayesian Gaussian Mixture model\n* Model Selection (Classification)\n  * XGBoost Classification\n  * k-Nearest Neighbors (kNN) Classification\n  * Support Vector Classification (SVC)\n  * Naive Bayes Classification\n* Final predictions\n* Submission","metadata":{}},{"cell_type":"markdown","source":"**Importing the dataset**","metadata":{}},{"cell_type":"code","source":"data = pd.read_csv('/kaggle/input/tabular-playground-series-jul-2022/data.csv')\nId = data.loc[:, 'id'].values\ndata = data.drop(columns = 'id')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Selecting important features**\n\nUsed PCA to determine the \"importance\" of all features (in determining the output). A feature is considered \"important\" only if its emperical mean is greater than 0.01. These important features are stored in \"data_best\" dataframe.","metadata":{}},{"cell_type":"code","source":"from sklearn.decomposition import PCA\n\npca = PCA(n_components = None)\npca.fit(data)\n\ndata_best = data.loc[:, abs(pca.mean_) > 0.01]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Turns out that PCA took 14 features to be important (from f_07-f_13 and f_22-f_28). In the following models, only these 14 features are used to determine the output.","metadata":{}},{"cell_type":"markdown","source":"**The Elbow method**\n\nSince we don't know the number of clusters, I have used the Elbow method to determine the optimum number of clusters that would be suitable for this dataset.","metadata":{}},{"cell_type":"code","source":"# import matplotlib.pyplot as plt\n# from sklearn.cluster import KMeans\n\n# wcss = []\n# for n in range(10):\n#     kmeans = KMeans(n_clusters = n+1,\n#                     init = 'k-means++',\n#                     tol = 1e-2,\n#                     random_state = 6)\n#     kmeans.fit(data_best)\n#     wcss.append(kmeans.inertia_)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Plotting the values of WCSS (Within-Cluster Sum of Squares) for different number of clusters to find the optimum number of clusters","metadata":{}},{"cell_type":"code","source":"# plt.figure(figsize = (16,9))\n# plt.title('WCSS for different number of clusters')\n# plt.xlabel('Number of clusters (0 indicates 1)')\n# plt.ylabel('WCSS')\n# plt.plot(wcss)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen from the plot, the optimum number of clusters is 7.","metadata":{}},{"cell_type":"code","source":"n_clusters = 7","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Bayesian Gaussian Mixture Model**\n\nUsing BGM model to predict the clusters of the samples that we are most sure of. Then making this prediction as a training set to train our classifier model. Then using this classifier model to classify other samples.","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import PowerTransformer\nfrom sklearn.mixture import BayesianGaussianMixture\n\npt = PowerTransformer()\ndata_best_scaled = pd.DataFrame(pt.fit_transform(data_best), columns = data_best.columns)\n\nbgm = BayesianGaussianMixture(n_components = n_clusters,\n                              random_state = 6)\nbgm.fit(data_best_scaled)","metadata":{"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now predicting the probabilities (that a particular sample belongs to a particular cluster) for all samples for all clusters. Then, for a sample, considering the cluster with maximum probability.","metadata":{}},{"cell_type":"code","source":"bgm_probs = bgm.predict_proba(data_best_scaled)\nbgm_pred = np.argmax(bgm_probs, axis = 1)\nbgm_max_prob = pd.Series([max(row) for row in bgm_probs])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now only selecting the samples that have probability (of belonging to a particular cluster) more than 70%. That means that we are sure that these samples definately belong to this cluster label.","metadata":{}},{"cell_type":"code","source":"threshold = 0.7\nindexes = np.array([])\n\nfor cluster in range(n_clusters):\n    indexes = np.concatenate((indexes, data_best_scaled[(bgm_pred == cluster) & (bgm_max_prob >= threshold)].index)).astype(np.int32)\n\nX = data_best_scaled.loc[indexes].reset_index(drop = True)\ny = pd.Series(bgm_pred)[indexes]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Turns out that there are 72,687 (out of 98,000) samples for which we are sure about the cluster they belong to. Now these samples are used as training set to train the classification model and classify other samples.","metadata":{}},{"cell_type":"markdown","source":"**Model Selection (Classification)**\n\nI have used GridSearchCV for hyperparameter tuning and determining the accuracy. But wherever hyperparameter tuning was not required (XGBoost and NaiveBayes), I simply used k-Fold cross-validation to determine the accuracy.","metadata":{}},{"cell_type":"code","source":"# from sklearn.model_selection import cross_val_score, GridSearchCV\n\n# accuracies_of_models = {}","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**XGBoost Classification**","metadata":{}},{"cell_type":"code","source":"# from xgboost import XGBClassifier\n\n# accuracies = cross_val_score(estimator = XGBClassifier(),\n#                              X = X,\n#                              y = y,\n#                              scoring = 'accuracy',\n#                              cv = 5)\n# accuracies_of_models['XGB'] = accuracies.mean()\n# del accuracies","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**k-Nearest Neighbors (kNN) Classification**","metadata":{}},{"cell_type":"code","source":"# from sklearn.neighbors import KNeighborsClassifier\n\n# params_knn = [{'n_neighbors': [10, 15, 30, 50],\n#                'weights': ['uniform', 'distance']}]\n# gs_knn = GridSearchCV(estimator = KNeighborsClassifier(),\n#                       param_grid = params_knn,\n#                       n_jobs = -1,\n#                       cv = 5)\n# gs_knn.fit(X, y)\n# print(gs_knn.best_params_)\n# accuracies_of_models['kNN'] = gs_knn.best_score_\n# del params_knn, gs_knn","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Best parameters for kNN:\n* n_neighbors = 30\n* weights = 'distance'","metadata":{}},{"cell_type":"markdown","source":"**Support Vector Classification (SVC)**","metadata":{}},{"cell_type":"code","source":"# from sklearn.svm import SVC\n\n# params_svc = [{'C': [1, 10, 100],\n#                'gamma': [0.1, 'scale', 0.01],\n#                'kernel': ['rbf', 'poly'],\n#                'degree': [3, 4, 5]}]\n# gs_svc = GridSearchCV(estimator = SVC(tol = 0.1),\n#                       param_grid = params_svc,\n#                       scoring = 'accuracy',\n#                       n_jobs = -1,\n#                       cv = 5)\n# gs_svc.fit(X, y)\n# print(gs_svc.best_params_)\n# accuracies_of_models['SVC'] = gs_svc.best_score_\n# del params_svc, gs_svc","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Best parameters for SVC:\n* C = 100\n* gamma = 0.01\n* kernel = 'rbf'","metadata":{}},{"cell_type":"markdown","source":"**Naive Bayes Classification**","metadata":{}},{"cell_type":"code","source":"# from sklearn.naive_bayes import GaussianNB\n\n# accuracies = cross_val_score(estimator = GaussianNB(),\n#                              X = X,\n#                              y = y,\n#                              scoring = 'accuracy',\n#                              cv = 5)\n# accuracies_of_models['NB'] = accuracies.mean()\n# del accuracies","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now comparing the accuracies of all models","metadata":{}},{"cell_type":"code","source":"# accuracies_of_models","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* XGB: 98.62%\n* kNN: 97.78%\n* SVC: 99.89%\n* NB: 90.82%\n\nThus, I haves used SVC for final predictions.","metadata":{}},{"cell_type":"markdown","source":"**Final predictions**","metadata":{}},{"cell_type":"code","source":"from sklearn.svm import SVC\n\nsvc = SVC(C = 100,\n          gamma = 0.01,\n          kernel = 'rbf',\n          tol = 0.1)\nsvc.fit(X, y)\n\nfinal_predictions = svc.predict(data_best)\nfinal_predictions[indexes] = y","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Submission**","metadata":{}},{"cell_type":"code","source":"submission = pd.DataFrame({'Id': Id, 'Predicted': final_predictions}).set_index('Id')\nsubmission.to_csv('submission.csv')","metadata":{},"execution_count":null,"outputs":[]}]}