{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🚀 Spaceship Titanic - 📊 EDA | ML 17 models | DL 🤖\n#### This is my first Kaggle Notebook which I will be using solely for practice.\n#### If you found this Kernel insightful, please consider upvoting it 👍\n#### I am all ears if you have any advice :)","metadata":{}},{"cell_type":"markdown","source":"<img src= \"https://img.freepik.com/free-vector/spaceship-cockpit-interior-space-planets-view_33099-2159.jpg?t=st=1658011031~exp=1658011631~hmac=4d584b8e56c9a53634137f256d197440fb35cf906a2f4b561c206a2856cfc9c0&w=996\" alt =\"Titanic\" style='width: 75%; margin-left: 12.5%\n'>","metadata":{}},{"cell_type":"markdown","source":"<a href=\"https://www.freepik.com/vectors/cockpit\">Cockpit vector created by vectorpouch - www.freepik.com</a>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           background-color:#F4E06D;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n           flex-direction: row;\">\n\n<h1 style=\"padding: 2rem;\n          color:black;\n          text-align:center;\n          margin:0 auto;\n          font-size:3rem;\">\n   THE NOTEBOOK CONTAINS THE FOLLOWING SECTIONS:\n</h1>\n \n</div>\n<div style=\"\n           display:fill;\n           background-color:#F4E06D;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n           flex-direction: row;\n           justify-content: center\"\n     >\n<ul style=\"\n           background-color:#F4E06D;\n           margin-left: 2rem;\n               \">\n\n<li style=\"color: black;\n           font-size:1.75rem;\">\n<text >\n Problem Description 📄\n</text>\n</li>\n<li style=\"color: black;\n           font-size:1.75rem;\">\n\n<text >\n Exploratory Data Analysis 🔍📁\n</text>\n</li>\n\n    \n<li style=\"color: black;\n           font-size:1.75rem;\">\n<text >\n  Data Visualization 📊📈\n</text>\n</li>\n\n<li style=\"color: black;\n           font-size:1.75rem;\">\n<text >\n  Data Preprocessing 🏭\n</text>\n</li>\n\n <li style=\"color: black;\n           font-size:1.75rem;\">\n<text >\n  ML and Model Selection 🤖👨‍💻 (16 Models + 1 DL Model)\n  \n</text>\n</li> \n\n <li style=\"\n           font-size:1.75rem;\">\n<text >\n  DL (Work In Progress)\n  \n</text>\n</li> \n    \n</ul>\n</div>","metadata":{}},{"cell_type":"markdown","source":"## Table of Contents:\n[A) PROBLEM DESCRIPTION](#1)\n* [1. OBJECTIVE](#2)\n* [2. ABOUT THE GIVEN DATASETS](#3)\n* [3. EVALUATION METRICS USED](#4)\n\n[B) WORK](#5)\n* [4. IMPORTING RELEVANT LIBRARIES](#6)\n* [5. READING DATA](#7)\n* [6. EXPLORATORY DATA ANALYSIS AND VISUALIZATION](#8)\n    - [6.1 View Data Types of Predictors and Target Variables](#9)\n    - [6.2 Analysis of Missing Values](#10)\n    - [6.3 Data Analysis of Individual Predictors](#11)\n* [7. DATA PREPROCESSING](#12)\n* [8. MACHINE LEARNING](#13)\n* [9. DEEP LEARNING](#14)\n\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"<div id=1 style=\"color:white;\n           display:fill;\n           border-radius:10px;\n           background-color:#36AE7C;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n           justify-content:center;\">\n\n<h1 style=\"padding: 2.5rem;\n          color:white;\n          text-align:center;\n          margin:0 auto;\n          font-size:3rem;\">\n   Problem Description\n</h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div id = 2 style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#5642C5;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\nObjective ⛳\n</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"The Spaceship Titanic Competition wants the participants to predict whether a passanger was successfuly transported to an alternate dimension. In order to make these predictions, we are provided with both training and testing datasets for which we will appy data exploration and preprocessing techniques in order to reach our end goal.\n","metadata":{}},{"cell_type":"markdown","source":"<div id = 3 style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#5642C5;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\n    About the given Datasets 📁📁\n</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"* PassengerId - A unique Id for each passenger. Each Id takes the form gggg_pp where gggg indicates a group the passenger is travelling with and pp is their number within the group. People in a group are often family members, but not always.\n* HomePlanet - The planet the passenger departed from, typically their planet of permanent residence.\n* CryoSleep - Indicates whether the passenger elected to be put into suspended animation for the duration of the voyage. Passengers in cryosleep are confined to their cabins.\n* Cabin - The cabin number where the passenger is staying. Takes the form deck/num/side, where side can be either P for Port or S for Starboard.\nDestination - The planet the passenger will be debarking to.\n* Age - The age of the passenger.\n* VIP - Whether the passenger has paid for special VIP service during the voyage.\n* RoomService, FoodCourt, ShoppingMall, Spa, VRDeck - Amount the passenger has billed at each of the Spaceship Titanic's many luxury amenities.\n* Name - The first and last names of the passenger.\n\n* Transported - Whether the passenger was transported to another dimension. This is the target, the column you are trying to predict","metadata":{}},{"cell_type":"markdown","source":"<div id = 4 style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#5642C5;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\n    Evaluation Metrics Used 📏\n</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### Based on the problem, we will use the Accuracy metric, perfect for classification.\n$Accuracy = \\frac{Number of Correct Predictions}{Total Number of Predictions Made} = \\frac{TP + TN}{ TP + FP + TN + FN}$\n\nTP: True Positives\n\nFP: False Positivies\n\nTN: True Negatives\n\nFN: False Negatives\n\n#### If you want to learn more about which metrics to use in your Machine Learning Problem, visit: \nhttps://towardsdatascience.com/metrics-to-evaluate-your-machine-learning-algorithm-f10ba6e38234","metadata":{}},{"cell_type":"markdown","source":"\n<div id = 5 style=\"color:white;\n           display:fill;\n           border-radius:10px;\n           background-color:#36AE7C;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n           justify-content:center;\">\n\n<h1 style=\"padding: 2.5rem;\n          color:white;\n          text-align:center;\n          margin:0 auto;\n          font-size:3rem;\">\n   Work 🔨\n</h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div id=6 style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#5642C5;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\nImporting relevant libraries 📚📚📚\n</h2>\n</div>","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os\nfrom termcolor import colored\nimport warnings\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import LabelEncoder, MinMaxScaler, StandardScaler\nfrom sklearn.linear_model import LogisticRegression, SGDClassifier\nfrom sklearn.metrics import accuracy_score, f1_score, confusion_matrix, plot_confusion_matrix\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.discriminant_analysis import LinearDiscriminantAnalysis\nfrom sklearn.svm import SVC\nfrom sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifier\nfrom xgboost import XGBClassifier\nfrom sklearn.model_selection import GridSearchCV as gscv\nfrom sklearn.gaussian_process import GaussianProcessClassifier\nfrom sklearn.gaussian_process.kernels import RBF\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import AdaBoostClassifier, ExtraTreesClassifier\nfrom lightgbm import LGBMClassifier\nfrom sklearn.ensemble import VotingClassifier\nfrom scipy.stats import expon, uniform\n\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import Sequential, Input, Model\nfrom tensorflow.keras.layers import Dense, Flatten, Dropout, BatchNormalization\nwarnings.simplefilter('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:35:58.796851Z","iopub.execute_input":"2022-08-07T15:35:58.797719Z","iopub.status.idle":"2022-08-07T15:35:58.807090Z","shell.execute_reply.started":"2022-08-07T15:35:58.797679Z","shell.execute_reply":"2022-08-07T15:35:58.806165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div id=7 style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#5642C5;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\nReading Data 📖\n</h2>\n</div>","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv(\"../input/spaceship-titanic/train.csv\")\ntest_data = pd.read_csv(\"../input/spaceship-titanic/test.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:38.319754Z","iopub.execute_input":"2022-08-07T15:44:38.320091Z","iopub.status.idle":"2022-08-07T15:44:38.374185Z","shell.execute_reply.started":"2022-08-07T15:44:38.320057Z","shell.execute_reply":"2022-08-07T15:44:38.373423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Train Data","metadata":{}},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:38.598075Z","iopub.execute_input":"2022-08-07T15:44:38.598383Z","iopub.status.idle":"2022-08-07T15:44:38.619572Z","shell.execute_reply.started":"2022-08-07T15:44:38.598350Z","shell.execute_reply":"2022-08-07T15:44:38.618657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Test Data","metadata":{}},{"cell_type":"code","source":"test_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:38.826346Z","iopub.execute_input":"2022-08-07T15:44:38.826656Z","iopub.status.idle":"2022-08-07T15:44:38.844966Z","shell.execute_reply.started":"2022-08-07T15:44:38.826625Z","shell.execute_reply":"2022-08-07T15:44:38.844342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('TRAIN DATA')\nprint(colored(f'Number of rows in train data: {train_data.shape[0]}', 'cyan'))\nprint(colored(f'Number of columns in train data: {train_data.shape[1]}\\n', 'cyan'))\nprint(\"TEST DATA\")\nprint(colored(f'Number of rows in test data: {test_data.shape[0]}', 'green'))\nprint(colored(f'Number of columns in test data: {test_data.shape[1]}', 'green'))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:39.091308Z","iopub.execute_input":"2022-08-07T15:44:39.091976Z","iopub.status.idle":"2022-08-07T15:44:39.099361Z","shell.execute_reply.started":"2022-08-07T15:44:39.091940Z","shell.execute_reply":"2022-08-07T15:44:39.098771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div id=8 style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#5642C5;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\n Exploratory Data Analysis and Visualization 🔍 👀\n</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"#### Performing Data Analysis and Visualization is a necessary step to solve Data Science problems. It allows you to gain insights that are relevant to the problem, which potentially will be useful down the line for reaching a solution. ","metadata":{}},{"cell_type":"markdown","source":"\n<div id=9 style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           background-color:#14C38E;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"text-align:center;\n          margin:0 auto;\n          color:white;\n           padding:1rem\n          \">\nView Data Types of Predictors and Target Variables\n</h3>\n</div>","metadata":{}},{"cell_type":"code","source":"train_data.dtypes","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:39.223812Z","iopub.execute_input":"2022-08-07T15:44:39.224616Z","iopub.status.idle":"2022-08-07T15:44:39.232327Z","shell.execute_reply.started":"2022-08-07T15:44:39.224566Z","shell.execute_reply":"2022-08-07T15:44:39.231329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div id=10 style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           background-color:#14C38E;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"text-align:center;\n          margin:0 auto;\n          color:white;\n           padding:1rem\n          \">\nAnalysis of Missing Values ⚠\n</h3>\n</div>","metadata":{}},{"cell_type":"markdown","source":"##### Let's start by having a look into missing vlaues. Deciding how to handle this inconvenience will affect our final result whether it is by removing the missing rows or by using the mode, median, or mean value. Below I provide a count plot of the missing values by column","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          text-decoration: underline\n          \">\nMissing Values Train Data\n</h4>\n</div>","metadata":{}},{"cell_type":"code","source":"nan_cols = train_data.columns[train_data.isna().any()].tolist()\nplt.figure(figsize=(19,8))\nnan_count_cols = train_data[nan_cols].isna().sum()\nprint(\"MISSING VALS IN THE TRAINING SET:\")\nprint(colored(nan_count_cols, \"green\"))\nsns.barplot(y=nan_count_cols, x=nan_cols, palette='mako')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:39.365656Z","iopub.execute_input":"2022-08-07T15:44:39.366587Z","iopub.status.idle":"2022-08-07T15:44:39.650209Z","shell.execute_reply.started":"2022-08-07T15:44:39.366545Z","shell.execute_reply":"2022-08-07T15:44:39.649562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          text-decoration: underline\n          \">\nMissing Values Test Data\n</h4>\n</div>","metadata":{}},{"cell_type":"code","source":"nan_cols = test_data.columns[test_data.isna().any()].tolist()\nplt.figure(figsize=(30,8))\nnan_count_cols = test_data[nan_cols].isna().sum()\nsns.barplot(y=nan_count_cols, x=nan_cols, palette='mako')\nplt.show()\nprint(colored(\"MISSING VALS IN THE TESTING SET:\\n\", 'magenta', attrs=['bold', 'underline']))\nprint(colored(nan_count_cols, \"cyan\", attrs=['bold']),'\\n')","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:39.651539Z","iopub.execute_input":"2022-08-07T15:44:39.651863Z","iopub.status.idle":"2022-08-07T15:44:39.951131Z","shell.execute_reply.started":"2022-08-07T15:44:39.651835Z","shell.execute_reply":"2022-08-07T15:44:39.950428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\nConclusions 📍\n</h4>\n</div>\n\n* Both the Training and Test datasets contain missing values, they must be handled properly.\n* In the data preprocessing sections I will replace missing values with the mode for categorical predictors and the mean when it comes to interval predictors.\n","metadata":{}},{"cell_type":"markdown","source":"<div id=11 style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#14C38E;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"padding: 1rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\n Data Analysis of Individual Predictors 🔮\n</h3>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\n    Passenger ID's 👨👩\n</h3>\n</div>\n\n\n##### It must be noted that even though the Passenger ID's are unique based on the passenger, useful information can still be extracted since the PassengerId predictor has the following format: gggg_pp.\n\n* gggg indicates the group the passenger is part of.\n\n* pp is the unique number of each passenger of a specific group (the highest number is the size of the group)\n","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h5 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\n    Use pp from PassengerId for both the training and test sets\n</h5>\n</div>\n","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          text-decoration: underline\n          \">\nGroup Size train set\n</h4>\n</div>","metadata":{}},{"cell_type":"code","source":"MAX_GROUP_SIZE = 8\ngggg_pp = train_data['PassengerId'].apply(lambda x: x.split('_')).values\ngggg = list(map(lambda x: x[0], gggg_pp))\npp = list(map(lambda x: x[1], gggg_pp))\ntrain_data['gggg'] = gggg\ntrain_data['pp'] = pp\n\ntrain_data['pp'] = train_data['pp'].astype('int64')\ntrain_data['group_size'] = 0\nfor i in range(MAX_GROUP_SIZE):\n    curr_gggg = train_data[train_data['pp'] == i + 1]['gggg'].to_numpy()\n    train_data.loc[train_data['gggg'].isin(curr_gggg), ['group_size']] = i + 1\n\nplt.figure(figsize=(19,8))\nprint(colored(\"Value Counts based on the group size:\\n\", 'green', attrs=['underline', 'bold']))\nprint(colored(\"Group Size, Count\", 'magenta', attrs=['bold']))\nprint(colored(train_data['group_size'].value_counts(), \"cyan\", attrs=['bold']))\nsns.barplot(y=train_data['group_size'].value_counts(), x=np.unique(train_data['pp']), palette='mako')\nplt.show()\nsns.catplot(x=\"group_size\",  kind=\"count\", hue='Transported', data=train_data, palette='mako').set(title='Group Size and Transported Count')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:39.952388Z","iopub.execute_input":"2022-08-07T15:44:39.952851Z","iopub.status.idle":"2022-08-07T15:44:40.612602Z","shell.execute_reply.started":"2022-08-07T15:44:39.952815Z","shell.execute_reply":"2022-08-07T15:44:40.611418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          text-decoration: underline\n          \">\nGroup Size test set\n</h4>\n</div>","metadata":{}},{"cell_type":"code","source":"gggg_pp = test_data['PassengerId'].apply(lambda x: x.split('_')).values\ngggg = list(map(lambda x: x[0], gggg_pp))\npp = list(map(lambda x: x[1], gggg_pp))\n   \nplt.figure(figsize=(19,8))\ntest_data['pp'] = pp\ntest_data['gggg'] = gggg\n\ntest_data['pp'] = test_data['pp'].astype('int64')\ntest_data['group_size'] = 0\n\nfor i in range(MAX_GROUP_SIZE):\n    curr_gggg = test_data[test_data['pp'] == i + 1]['gggg'].to_numpy()\n    test_data.loc[test_data['gggg'].isin(curr_gggg), ['group_size']] = i + 1\n\nprint(colored(\"Value Counts based on the group size:\\n\", 'green', attrs=['underline', 'bold']))\nprint(colored(\"Group Size, Count\", 'magenta', attrs=['bold']))\nprint(colored(test_data['group_size'].value_counts(), \"cyan\", attrs=['bold']))\nsns.barplot(y=test_data['group_size'].value_counts(), x=np.unique(test_data['group_size']), palette='mako')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:40.615473Z","iopub.execute_input":"2022-08-07T15:44:40.615806Z","iopub.status.idle":"2022-08-07T15:44:40.835265Z","shell.execute_reply.started":"2022-08-07T15:44:40.615763Z","shell.execute_reply":"2022-08-07T15:44:40.834373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Now for each group we must assign the final group size, which is max(pp)","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\nConclusions 📍\n</h4>\n</div>","metadata":{}},{"cell_type":"markdown","source":"* Based on the data visualizations centered on the pp predictor (number of people/groupsize) It can be concluded that It has an impact whether the Passenger is transported or not.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\n    HomePlanet Visualization 🌍🚀\n</h3>\n</div>\n\n#### In this section, we will visualize:\n* The HomePlanet count for both the Train and Test sets.\n* The HomePlanet predictor with respect to the Transported target variable. ","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          text-decoration: underline\n          \">\nHomePlanet count visualization\n</h4>\n</div>","metadata":{}},{"cell_type":"code","source":"# The possible values present in the \"HomePlanet\" predictor\nHOME_PLANET_UNIQUE_VALS = ['Earth', 'Europa', 'Mars']\nsns.barplot(y=train_data['HomePlanet'].value_counts(), x=HOME_PLANET_UNIQUE_VALS, palette='mako')\nplt.show()\nprint(colored(\"HomePlanet Count train data:\\n\", 'magenta', attrs=['underline', 'bold']))\nprint(colored(train_data['HomePlanet'].value_counts(), 'cyan', attrs=['bold']))\nsns.barplot(y=test_data['HomePlanet'].value_counts(), x=HOME_PLANET_UNIQUE_VALS, palette='mako')\nplt.show()\nprint(colored(\"HomePlanet Count test data:\\n\", 'magenta', attrs=['underline', 'bold']))\nprint(colored(test_data['HomePlanet'].value_counts(), 'cyan', attrs=['bold']))\n","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:40.837058Z","iopub.execute_input":"2022-08-07T15:44:40.837671Z","iopub.status.idle":"2022-08-07T15:44:41.066205Z","shell.execute_reply.started":"2022-08-07T15:44:40.837625Z","shell.execute_reply":"2022-08-07T15:44:41.065275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          text-decoration: underline\n          \">\nHomePlanet with respect to Transported\n</h4>\n</div>","metadata":{}},{"cell_type":"code","source":"homeplanet_transported_count = train_data[['HomePlanet', 'Transported']].value_counts()\nsns.catplot(x=\"HomePlanet\",  kind=\"count\", hue='Transported', data=train_data, palette='mako')\nplt.show()\nprint(colored(\"COUNT STATISTICS:\\n\", 'magenta', attrs=['underline', 'bold']))\nprint(colored(homeplanet_transported_count, 'cyan', attrs=['bold']))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:41.067970Z","iopub.execute_input":"2022-08-07T15:44:41.068616Z","iopub.status.idle":"2022-08-07T15:44:41.331127Z","shell.execute_reply.started":"2022-08-07T15:44:41.068567Z","shell.execute_reply":"2022-08-07T15:44:41.330227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\nConclusions 📍\n</h4>\n</div>\n\n##### Based on the visualization:\n* Europa and Mars seem to indicate a higher chance of the passanger being transported.\n* Earth seems to imply lower chances of being transported.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\n    CryoSleep Visualization 💤💤 \n</h3>\n</div>\n\n\n##### In this section, we will visualize:\n* The CryoSleep count for both the Train and Test sets.\n* The CryoSleep predictor with respect to the Transported target variable (only the train set obviously). ","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          text-decoration: underline\n          \">\nCryoSleep Count\n</h4>\n</div>","metadata":{}},{"cell_type":"code","source":"sns.barplot(y=train_data['CryoSleep'].value_counts(), x=[True, False], palette='mako').set(title=\"Train Set CryoSleep value count\")\nplt.show()\nprint(colored(\"CryoSleep Count Train data:\\n\", 'magenta', attrs=['underline', 'bold']))\nprint(colored(train_data['CryoSleep'].value_counts(), 'cyan', attrs=['bold']), '\\n')\nsns.barplot(y=test_data['CryoSleep'].value_counts(), x=[True, False], palette='mako').set(title=\"Test Set CryoSleep value count\")\nplt.show()\nprint(colored(\"CryoSleep Count Test data:\\n\", 'magenta', attrs=['underline', 'bold']))\nprint(colored(test_data['CryoSleep'].value_counts(), 'cyan', attrs=['bold']))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:41.332629Z","iopub.execute_input":"2022-08-07T15:44:41.332943Z","iopub.status.idle":"2022-08-07T15:44:41.680339Z","shell.execute_reply.started":"2022-08-07T15:44:41.332901Z","shell.execute_reply":"2022-08-07T15:44:41.679520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          text-decoration: underline\n          \">\nCryoSleep with respect to Transported\n</h4>\n</div>","metadata":{}},{"cell_type":"code","source":"cryosleep_transported_count = train_data[['CryoSleep', 'Transported']].value_counts()\nsns.catplot(x=\"CryoSleep\",  kind=\"count\", hue='Transported', data=train_data, palette='mako')\nplt.show()\nprint(colored(\"COUNT STATISTICS:\\n\", 'magenta', attrs=['underline', 'bold']))\nprint(colored(cryosleep_transported_count, 'cyan', attrs=['bold']))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:41.681741Z","iopub.execute_input":"2022-08-07T15:44:41.681957Z","iopub.status.idle":"2022-08-07T15:44:42.056670Z","shell.execute_reply.started":"2022-08-07T15:44:41.681931Z","shell.execute_reply":"2022-08-07T15:44:42.056025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\nConclusions 📍\n</h4>\n</div>\n\n* Based on the visualization, CryoSleep and Transported are positively correlated.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\n    Cabin Visualization 📊\n</h3>\n</div>\n\n##### The following must be taken into account for the Cabin predictor\n* It takes the following form: deck/num/side\n* Side can be either P (Port) or S (Starboard)\n\n#####  Below I perform feature engineering to when it comes to the side feature to see if useful information can be extracted","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          text-decoration: underline\n          \">\nCabin Visualization\n</h4>\n</div>","metadata":{}},{"cell_type":"code","source":"train_data['Cabin'].fillna('N/N/N', inplace=True)\ntest_data['Cabin'].fillna('N/N/N', inplace=True)\n\ndeck_num_side = train_data['Cabin'].apply(lambda x: x.split('/'))\nside = list(map(lambda x: x[-1], deck_num_side))\nside_order_vals = ['S', 'P', 'N']\ntrain_data['side'] = side\ntest_data['side'] = list(map(lambda x: x[-1], test_data['Cabin'].apply(lambda x: x.split('/'))))\n\nplt.figure(figsize=(19,8))\nsns.countplot(x='side', data=train_data, order=side_order_vals, palette='mako')\nplt.show()\nprint(colored(\"Value Counts based on the Cabin side:\\n\", 'green', attrs=['underline', 'bold']))\nprint(colored(\"Cabin Side, Count\", 'magenta', attrs=['bold']))\nprint(colored(train_data['side'].value_counts(), \"cyan\", attrs=['bold']))\nsns.catplot(x=\"side\",  kind=\"count\", hue='Transported', data=train_data, palette='mako').set(title='Cabin Side and Transported Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:42.059194Z","iopub.execute_input":"2022-08-07T15:44:42.060086Z","iopub.status.idle":"2022-08-07T15:44:42.608339Z","shell.execute_reply.started":"2022-08-07T15:44:42.060047Z","shell.execute_reply":"2022-08-07T15:44:42.607708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          text-decoration: underline\n          \">\nDeck Visualization\n</h4>\n</div>","metadata":{}},{"cell_type":"code","source":"deck =  list(map(lambda x: x[0], deck_num_side))\n\ntrain_data['deck'] = deck\ntest_data['deck'] = list(map(lambda x: x[0], test_data['Cabin'].apply(lambda x: x.split('/'))))\n\ndeck_order_vals = ['A','B','C','D','E','F','G','T','N']\nplt.figure(figsize=(19,8))\nsns.countplot(x='deck', data=train_data, order=deck_order_vals, palette='mako')\nplt.show()\nprint(colored(\"Value Counts based on the Cabin side:\\n\", 'green', attrs=['underline', 'bold']))\nprint(colored(\"Cabin Side, Count\", 'magenta', attrs=['bold']))\nprint(colored(train_data['deck'].value_counts(), \"cyan\", attrs=['bold']))\nsns.catplot(x=\"deck\",  kind=\"count\", hue='Transported', data=train_data, palette='mako').set(title='Cabin Deck and Transported Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:42.609448Z","iopub.execute_input":"2022-08-07T15:44:42.610284Z","iopub.status.idle":"2022-08-07T15:44:43.289916Z","shell.execute_reply.started":"2022-08-07T15:44:42.610237Z","shell.execute_reply":"2022-08-07T15:44:43.288999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\nConclusions 📍\n</h4>\n</div>\n\n* Based on the visualization, the side of the Cabin predictor seems to have an impact on whether the passenger is transported or not.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\n    Destination Visualization 📊\n</h3>\n</div>\n\n#####  I will simply perform a count based on all the values in Destination and also compare it with the target\n","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(19,8))\nDESTINATION_UNIQUE_VALS = train_data['Destination'].unique()\nsns.countplot(x='Destination', data=train_data, order=DESTINATION_UNIQUE_VALS, palette='mako')\nplt.show()\nprint(colored(\"Value Counts based on the Destination:\\n\", 'green', attrs=['underline', 'bold']))\nprint(colored(\"Destination, Count\", 'magenta', attrs=['bold']))\nprint(colored(train_data['Destination'].value_counts(), \"cyan\", attrs=['bold']))\nsns.catplot(x=\"Destination\",  kind=\"count\", hue='Transported', data=train_data, palette='mako').set(title='Cabin Side and Transported Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:43.291665Z","iopub.execute_input":"2022-08-07T15:44:43.292190Z","iopub.status.idle":"2022-08-07T15:44:43.881613Z","shell.execute_reply.started":"2022-08-07T15:44:43.292146Z","shell.execute_reply":"2022-08-07T15:44:43.877803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\nConclusions 📍\n</h4>\n</div>\n\n* Based on the visualization, the destination \"TRAPPIST-1e\", exhibits a negative correlation with Transported, while the other two destinations have a positive correlation with the passangers being successfuly transported to a certain degree, \"55 Cancri e\" being the strongest one.\n","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\n    Age Visualization 📊\n</h3>\n</div>\n\n#####  Below I provide information about the min, mean and max age values for both Target scenarios (Transported==1 and Transported==0). Lastly I also provide visualization of the data with a Box Plot and a Histogram.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 2, figsize=(19,8))\nfig.suptitle('Age Histogram and Boxplot', size=22)\nsns.histplot(x='Age', data=train_data, hue='Transported', palette='mako', kde=True, element='step', ax=ax[0])\nsns.boxplot(x='Transported', y='Age', data=train_data, palette='mako', ax=ax[1])\nplt.show()\n\nprint(colored(\"Age Min, Mean and Max:\", 'green', attrs=['underline', 'bold']))\n\ntransported_1 = train_data[train_data['Transported']==True]['Age']\ntransported_0 = train_data[train_data['Transported']==False]['Age']\n\nprint(colored(\"\\tTransported == 1\", 'magenta', attrs=['bold']))\nprint('\\tAge Minimum: ', colored(transported_1.describe()['min'], \"cyan\", attrs=['bold']))\nprint('\\tAge Mean:', colored(transported_1.describe()['mean'], \"cyan\", attrs=['bold']))\nprint('\\tAge Maximum:', colored(transported_1.describe()['max'], \"cyan\", attrs=['bold']), '\\n')\n\nprint(colored(\"\\tTransported == 0\", 'magenta', attrs=['bold']))\nprint('\\tAge Minimum: ', colored(transported_0.describe()['min'], \"cyan\", attrs=['bold']))\nprint('\\tAge Mean:', colored(transported_0.describe()['mean'], \"cyan\", attrs=['bold']))\nprint('\\tAge Maximum:', colored(transported_0.describe()['max'], \"cyan\", attrs=['bold']))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:43.883616Z","iopub.execute_input":"2022-08-07T15:44:43.884012Z","iopub.status.idle":"2022-08-07T15:44:44.377246Z","shell.execute_reply.started":"2022-08-07T15:44:43.883973Z","shell.execute_reply":"2022-08-07T15:44:44.376355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\nConclusions 📍\n</h4>\n</div>\n\n* Based on the visualization, passengers with lower age have a higher chance of being Transported, which is reflected in the histogram.\n","metadata":{}},{"cell_type":"markdown","source":"\"","metadata":{}},{"cell_type":"code","source":"VIP_transported_count = train_data[['VIP', 'Transported']].value_counts()\nsns.catplot(x=\"VIP\",  kind=\"count\", hue='Transported', data=train_data, palette='mako')\nplt.show()\nprint(colored(\"COUNT STATISTICS:\\n\", 'magenta', attrs=['underline', 'bold']))\nprint(colored(VIP_transported_count, 'cyan', attrs=['bold']))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:44.378882Z","iopub.execute_input":"2022-08-07T15:44:44.379186Z","iopub.status.idle":"2022-08-07T15:44:44.756309Z","shell.execute_reply.started":"2022-08-07T15:44:44.379144Z","shell.execute_reply":"2022-08-07T15:44:44.755345Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\nConclusions 📍\n</h4>\n</div>\n\n* Based on the visualization, VIP exhibits a negative correlation with Transported. Essentialy, VIP passengers have a higher chance of Transported being False.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\n    RoomService, FoodCourt, ShoppingMall, Spa, VRDeck 📊\n</h3>\n</div>\n\n#####  Below I provide a summary regarding the following predictors: RoomService, FoodCourt, ShoppingMall, Spa, VRDeck","metadata":{}},{"cell_type":"code","source":"fig,ax = plt.subplots(1, 2, figsize=(22,7))\nfig.suptitle('RoomService Histogram and Boxplot', size=22)\nsns.histplot(x='RoomService', data=train_data, hue='Transported', palette='mako', kde=True, element='step', ax=ax[0])\nsns.boxplot(x='Transported', y='Age', data=train_data, palette='mako', ax=ax[1])\nplt.show()\n\nprint(colored(\"Age Min, Mean and Max:\", 'green', attrs=['underline', 'bold']))\n\ntransported_1 = train_data[train_data['Transported']==True]['RoomService']\ntransported_0 = train_data[train_data['Transported']==False]['RoomService']\n\nprint(colored(\"\\tTransported == 1\", 'magenta', attrs=['bold']))\nprint('\\tRoomService Minimum: ', colored(transported_1.describe()['min'], \"cyan\", attrs=['bold']))\nprint('\\tRoomService Mean:', colored(transported_1.describe()['mean'], \"cyan\", attrs=['bold']))\nprint('\\tRoomService Maximum:', colored(transported_1.describe()['max'], \"cyan\", attrs=['bold']), '\\n')\n\nprint(colored(\"\\tTransported == 0\", 'magenta', attrs=['bold']))\nprint('\\tRoomService Minimum: ', colored(transported_0.describe()['min'], \"cyan\", attrs=['bold']))\nprint('\\tRoomService Mean:', colored(transported_0.describe()['mean'], \"cyan\", attrs=['bold']))\nprint('\\tRoomService Maximum:', colored(transported_0.describe()['max'], \"cyan\", attrs=['bold']))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:44.757845Z","iopub.execute_input":"2022-08-07T15:44:44.758081Z","iopub.status.idle":"2022-08-07T15:44:45.600330Z","shell.execute_reply.started":"2022-08-07T15:44:44.758053Z","shell.execute_reply":"2022-08-07T15:44:45.599456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig,ax = plt.subplots(1, 2, figsize=(18,5))\nfig.suptitle('FoodCourt Histogram and Boxplot', size=22)\ncurrent_var = 'FoodCourt'\nsns.histplot(x='FoodCourt', data=train_data, hue='Transported', palette='mako', kde=True, element='step', ax=ax[0])\nsns.boxplot(x='Transported', y='FoodCourt', data=train_data, palette='mako', ax=ax[1])\nplt.show()\n\nprint(colored(\"FoodCourt Min, Mean and Max:\", 'green', attrs=['underline', 'bold']))\n\ntransported_1 = train_data[train_data['Transported']==True]['FoodCourt']\ntransported_0 = train_data[train_data['Transported']==False]['FoodCourt']\n\nprint(colored(\"\\tTransported == 1\", 'magenta', attrs=['bold']))\nprint('\\t FoodCourt Minimum: ', colored(transported_1.describe()['min'], \"cyan\", attrs=['bold']))\nprint('\\t FoodCourt Mean:', colored(transported_1.describe()['mean'], \"cyan\", attrs=['bold']))\nprint('\\t FoodCourt Maximum:', colored(transported_1.describe()['max'], \"cyan\", attrs=['bold']), '\\n')\n\nprint(colored(\"\\tTransported == 0\", 'magenta', attrs=['bold']))\nprint('\\t FoodCourt Minimum: ', colored(transported_0.describe()['min'], \"cyan\", attrs=['bold']))\nprint('\\t FoodCourt Mean:', colored(transported_0.describe()['mean'], \"cyan\", attrs=['bold']))\nprint('\\t FoodCourt Maximum:', colored(transported_0.describe()['max'], \"cyan\", attrs=['bold']))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:45.601832Z","iopub.execute_input":"2022-08-07T15:44:45.602050Z","iopub.status.idle":"2022-08-07T15:44:46.077153Z","shell.execute_reply.started":"2022-08-07T15:44:45.602024Z","shell.execute_reply":"2022-08-07T15:44:46.076193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig,ax = plt.subplots(1, 2, figsize=(18,5))\nfig.suptitle('ShoppingMall Histogram and Boxplot', size=22)\ncurrent_var = 'ShoppingMall'\nsns.histplot(x='ShoppingMall', data=train_data, hue='Transported', palette='mako', kde=True, element='step', ax=ax[0])\nsns.boxplot(x='Transported', y='ShoppingMall', data=train_data, palette='mako', ax=ax[1])\nplt.show()\n\nprint(colored(\"ShoppingMall Min, Mean and Max:\", 'green', attrs=['underline', 'bold']))\n\ntransported_1 = train_data[train_data['Transported']==True]['ShoppingMall']\ntransported_0 = train_data[train_data['Transported']==False]['ShoppingMall']\n\nprint(colored(\"\\t Transported == 1\", 'magenta', attrs=['bold']))\nprint('\\t ShoppingMall Minimum: ', colored(transported_1.describe()['min'], \"cyan\", attrs=['bold']))\nprint('\\t ShoppingMall Mean:', colored(transported_1.describe()['mean'], \"cyan\", attrs=['bold']))\nprint('\\t ShoppingMall Maximum:', colored(transported_1.describe()['max'], \"cyan\", attrs=['bold']), '\\n')\n\nprint(colored(\"\\t Transported == 0\", 'magenta', attrs=['bold']))\nprint('\\t ShoppingMall Minimum: ', colored(transported_0.describe()['min'], \"cyan\", attrs=['bold']))\nprint('\\t ShoppingMall Mean:', colored(transported_0.describe()['mean'], \"cyan\", attrs=['bold']))\nprint('\\t ShoppingMall Maximum:', colored(transported_0.describe()['max'], \"cyan\", attrs=['bold']))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:46.078636Z","iopub.execute_input":"2022-08-07T15:44:46.078961Z","iopub.status.idle":"2022-08-07T15:44:46.561675Z","shell.execute_reply.started":"2022-08-07T15:44:46.078918Z","shell.execute_reply":"2022-08-07T15:44:46.560718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig,ax = plt.subplots(1, 2, figsize=(18,5))\nfig.suptitle('Spa Histogram and Boxplot', size=22)\n\nsns.histplot(x='Spa', data=train_data, hue='Transported', palette='mako', kde=True, element='step', ax=ax[0])\nsns.boxplot(x='Transported', y='Spa', data=train_data, palette='mako', ax=ax[1])\nplt.show()\n\nprint(colored(\"Spa Min, Mean and Max:\", 'green', attrs=['underline', 'bold']))\n\ntransported_1 = train_data[train_data['Transported']==True]['Spa']\ntransported_0 = train_data[train_data['Transported']==False]['Spa']\n\nprint(colored(\"\\t Transported == 1\", 'magenta', attrs=['bold']))\nprint('\\t Spa Minimum: ', colored(transported_1.describe()['min'], \"cyan\", attrs=['bold']))\nprint('\\t Spa Mean:', colored(transported_1.describe()['mean'], \"cyan\", attrs=['bold']))\nprint('\\t Spa Maximum:', colored(transported_1.describe()['max'], \"cyan\", attrs=['bold']), '\\n')\n\nprint(colored(\"\\t Transported == 0\", 'magenta', attrs=['bold']))\nprint('\\t Spa Minimum: ', colored(transported_0.describe()['min'], \"cyan\", attrs=['bold']))\nprint('\\t Spa Mean:', colored(transported_0.describe()['mean'], \"cyan\", attrs=['bold']))\nprint('\\t Spa Maximum:', colored(transported_0.describe()['max'], \"cyan\", attrs=['bold']))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:46.562959Z","iopub.execute_input":"2022-08-07T15:44:46.563235Z","iopub.status.idle":"2022-08-07T15:44:47.025362Z","shell.execute_reply.started":"2022-08-07T15:44:46.563203Z","shell.execute_reply":"2022-08-07T15:44:47.024471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig,ax = plt.subplots(1, 2, figsize=(18,5))\nfig.suptitle('VRDeck Histogram and Boxplot', size=22)\n\nsns.histplot(x='VRDeck', data=train_data, hue='Transported', palette='mako', kde=True, element='step', ax=ax[0])\nsns.boxplot(x='Transported', y='VRDeck', data=train_data, palette='mako', ax=ax[1])\nplt.show()\n\nprint(colored(\"VRDeck Min, Mean and Max:\", 'green', attrs=['underline', 'bold']))\n\ntransported_1 = train_data[train_data['Transported']==True]['VRDeck']\ntransported_0 = train_data[train_data['Transported']==False]['VRDeck']\n\nprint(colored(\"\\t Transported == 1\", 'magenta', attrs=['bold']))\nprint('\\t VRDeck Minimum: ', colored(transported_1.describe()['min'], \"cyan\", attrs=['bold']))\nprint('\\t VRDeck Mean:', colored(transported_1.describe()['mean'], \"cyan\", attrs=['bold']))\nprint('\\t VRDeck Maximum:', colored(transported_1.describe()['max'], \"cyan\", attrs=['bold']), '\\n')\n\nprint(colored(\"\\t Transported == 0\", 'magenta', attrs=['bold']))\nprint('\\t VRDeck Minimum: ', colored(transported_0.describe()['min'], \"cyan\", attrs=['bold']))\nprint('\\t VRDeck Mean:', colored(transported_0.describe()['mean'], \"cyan\", attrs=['bold']))\nprint('\\t VRDeck Maximum:', colored(transported_0.describe()['max'], \"cyan\", attrs=['bold']))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:47.026848Z","iopub.execute_input":"2022-08-07T15:44:47.027105Z","iopub.status.idle":"2022-08-07T15:44:47.497118Z","shell.execute_reply.started":"2022-08-07T15:44:47.027073Z","shell.execute_reply":"2022-08-07T15:44:47.496173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px;\n           display:flex;\n            justify-content:center;\">\n\n<h4 style=\"text-align:center;\n          margin:0 auto;\n          color:black;\n          \">\nConclusions 📍\n</h4>\n</div>\n\n* By looking each of these individual cases, it can be concluded that passengers that spent more money on ameneties had a higher Transported rate","metadata":{}},{"cell_type":"markdown","source":"####  Visualizing Age vs Money Spent on Amenities","metadata":{}},{"cell_type":"code","source":"AMENITIES = ['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\nfig, ax = plt.subplots(3, 2, figsize=(18,5))\n\nfor i, amenity in enumerate(AMENITIES):\n    sns.scatterplot(x='Age', y=amenity, data=train_data, hue='Transported', palette='mako', ax=fig.axes[i])\n    fig.axes[i].set_title(f'{amenity} vs Age', weight='bold')\n    \n    \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:47.498538Z","iopub.execute_input":"2022-08-07T15:44:47.498840Z","iopub.status.idle":"2022-08-07T15:44:49.572783Z","shell.execute_reply.started":"2022-08-07T15:44:47.498800Z","shell.execute_reply":"2022-08-07T15:44:49.571932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Create a new column \"Amenities\" which is the sum of all the amenities","metadata":{}},{"cell_type":"code","source":"train_data['Amenities'] = train_data[AMENITIES].sum(axis=1)\ntest_data['Amenities'] = test_data[AMENITIES].sum(axis=1)\ntrain_data['NoAmenities'] = train_data['Amenities']==0\ntest_data['NoAmenities'] = test_data['Amenities']==0","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:49.574248Z","iopub.execute_input":"2022-08-07T15:44:49.574653Z","iopub.status.idle":"2022-08-07T15:44:49.592763Z","shell.execute_reply.started":"2022-08-07T15:44:49.574592Z","shell.execute_reply":"2022-08-07T15:44:49.591446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Visualizing the correlation matrix of the dataset\n#### If features are highly correlated then they need to be handled accordingly","metadata":{}},{"cell_type":"code","source":"plt.matshow(train_data.corr())\nplt.colorbar()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:49.593967Z","iopub.execute_input":"2022-08-07T15:44:49.594202Z","iopub.status.idle":"2022-08-07T15:44:49.959450Z","shell.execute_reply.started":"2022-08-07T15:44:49.594175Z","shell.execute_reply":"2022-08-07T15:44:49.957998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr_matrix = train_data.corr().abs()\nupper = corr_matrix.where(np.triu(np.ones(corr_matrix.shape), k=1).astype(np.bool))\nto_drop = [column for column in upper.columns if any(upper[column] > 0.9)]\nprint(to_drop)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:49.963213Z","iopub.execute_input":"2022-08-07T15:44:49.963662Z","iopub.status.idle":"2022-08-07T15:44:49.984799Z","shell.execute_reply.started":"2022-08-07T15:44:49.963605Z","shell.execute_reply":"2022-08-07T15:44:49.983909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:49.986063Z","iopub.execute_input":"2022-08-07T15:44:49.986721Z","iopub.status.idle":"2022-08-07T15:44:50.031480Z","shell.execute_reply.started":"2022-08-07T15:44:49.986685Z","shell.execute_reply":"2022-08-07T15:44:50.030859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Since no predictor exhibits a correlation above abs(0.9), no columns need to be dropped from the dataset","metadata":{}},{"cell_type":"markdown","source":"### We will use all of the information obtained in the Data Analysis and Visualization Section to our advantage and preprocess the data which will later on be used to train models","metadata":{}},{"cell_type":"markdown","source":"<div id = 12 style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#5642C5;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\n Data Preprocessing 🔄📁\n</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#14C38E;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"padding: 1rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\n Because we are now familiar with the data and we have created additional predictors that will be helpful (deck, pp, side, Amenities, NoAmenities), we will focus on processing and formatting the data so that it can be used to train a model.\n</h3>\n</div>","metadata":{}},{"cell_type":"markdown","source":"#### We can remove PassengerId, Cabin, and the Name predictors, since they will not result in improved predictions","metadata":{}},{"cell_type":"code","source":"# SAVE PASSENGERID's for submission\nPASSENGER_ID = test_data[['PassengerId']]\ntrain_data.drop(['PassengerId', 'Cabin', 'Name', 'Amenities', 'pp', 'gggg'], axis=1, inplace=True)\ntest_data.drop(['PassengerId', 'Cabin', 'Name', 'Amenities', 'pp', 'gggg'], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:50.034290Z","iopub.execute_input":"2022-08-07T15:44:50.035005Z","iopub.status.idle":"2022-08-07T15:44:50.047059Z","shell.execute_reply.started":"2022-08-07T15:44:50.034969Z","shell.execute_reply":"2022-08-07T15:44:50.046310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### The remaining predictors are:\n* HomePlanet: **object**\n* CryoSleep: **object**\n* Destination: **object**\n* Age: **float64**\n* VIP: **bool**\n* RoomService: **float64**\n* FoodCourt: **float64**\n* ShoppingMall: **float64**\n* Spa: **float64**\n* VRDeck: **float64**\n* group_size: **int64**\n* side: **object**\n* deck: **object**\n* Amenities: **float64**\n* NoAmenities: **bool**","metadata":{}},{"cell_type":"code","source":"train_data","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:50.048695Z","iopub.execute_input":"2022-08-07T15:44:50.049947Z","iopub.status.idle":"2022-08-07T15:44:50.081282Z","shell.execute_reply.started":"2022-08-07T15:44:50.049898Z","shell.execute_reply":"2022-08-07T15:44:50.080624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#14C38E;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"padding: 1rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\n Filling Missing Values, Encoding and Scaling 👩‍💻\n</h3>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### Filling Missing values ","metadata":{}},{"cell_type":"code","source":"print(colored(f\"Total NaN values before:\", 'green', attrs=['underline', 'bold']))\nprint(colored(train_data.isna().sum().sum(), attrs=['bold']))\nLABELS = test_data.columns\nfor col in LABELS:\n    if col in ['Age', 'RoomService',\n       'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']:\n        train_data[col].fillna(train_data[col].median(), inplace=True)\n        test_data[col].fillna(train_data[col].median(), inplace=True)\n    else:\n        train_data[col].fillna(train_data[col].mode()[0], inplace=True)\n        test_data[col].fillna(train_data[col].mode()[0], inplace=True)\n\nprint(colored(f\"Total NaN values after:\", 'green', attrs=['underline', 'bold']))\nprint(colored(train_data.isna().sum().sum(), attrs=['bold']))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:50.082455Z","iopub.execute_input":"2022-08-07T15:44:50.083168Z","iopub.status.idle":"2022-08-07T15:44:50.128704Z","shell.execute_reply.started":"2022-08-07T15:44:50.083133Z","shell.execute_reply":"2022-08-07T15:44:50.127546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Encode labels and Scale\n* ##### Objects will use scikit-learn's LabelEncoder\n* ##### Booleans will be converted to int\n* ##### Floats will be scaled with MinMaxScaler","metadata":{}},{"cell_type":"code","source":"for col in LABELS:\n    # Check if object\n    if train_data[col].dtype == 'O':\n        encoder = LabelEncoder()\n        train_data[col] = encoder.fit_transform(train_data[col])\n        test_data[col] = encoder.transform(test_data[col])\n        \n    elif train_data[col].dtype == 'bool':\n        train_data[col] = train_data[col].astype('int')\n        test_data[col] = test_data[col].astype('int')\n\nencoder = LabelEncoder()\ntrain_data['Transported'] = train_data['Transported'].astype('int')\nLABELS_MM = ['Age']\nLABELS_SS = ['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\nmm_scaler = MinMaxScaler()\nss_scaler = StandardScaler()\n# Apply Min-Max Scaling\ntrain_data[LABELS_MM] = mm_scaler.fit_transform(train_data[LABELS_MM])\ntest_data[LABELS_MM] = mm_scaler.transform(test_data[LABELS_MM])\n# Apply Standard Scaling\ntrain_data[LABELS_SS] = ss_scaler.fit_transform(train_data[LABELS_SS])\ntest_data[LABELS_SS] = ss_scaler.transform(test_data[LABELS_SS])","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:50.130040Z","iopub.execute_input":"2022-08-07T15:44:50.130282Z","iopub.status.idle":"2022-08-07T15:44:50.176250Z","shell.execute_reply.started":"2022-08-07T15:44:50.130250Z","shell.execute_reply":"2022-08-07T15:44:50.175204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#14C38E;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"padding: 1rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\n Splitting Data into Training and Validation Sets 📁📂\n</h3>\n</div>","metadata":{}},{"cell_type":"code","source":"X, y = train_data.drop('Transported', axis=1), train_data[['Transported']]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:50.177596Z","iopub.execute_input":"2022-08-07T15:44:50.177852Z","iopub.status.idle":"2022-08-07T15:44:50.184724Z","shell.execute_reply.started":"2022-08-07T15:44:50.177821Z","shell.execute_reply":"2022-08-07T15:44:50.183717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.0499, random_state=43, stratify=y)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:51.672389Z","iopub.execute_input":"2022-08-07T15:44:51.673240Z","iopub.status.idle":"2022-08-07T15:44:51.721300Z","shell.execute_reply.started":"2022-08-07T15:44:51.673199Z","shell.execute_reply":"2022-08-07T15:44:51.720620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div id=13 style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#5642C5;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\n Machine Learning 🤖\n</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#14C38E;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"padding: 1rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\n Model Selection \n</h3>\n</div>","metadata":{}},{"cell_type":"markdown","source":"#### We will now focus on determining the most accurate model which will be determined by using the validation set. Below I include the ML Algorithms that we will work with:\n* Naive Bayes\n* Linear Discriminant Analysis\n* Logistic Regression\n* Support Vector Classifier\n* K-Neighbors Classifier\n* Stochastic Gradient Descent Classifier\n* Random Forest Classifier\n* Gradient Boosting Classifier\n* XGBoost Classifier\n* AdaBoost Classifier\n* Light GBM Classifier\n* Extra Trees Classifier","metadata":{}},{"cell_type":"markdown","source":"### We will keep track of the accuracies of all the models in the dictionary","metadata":{}},{"cell_type":"code","source":"# Define dictionary with model accuracies\nmodel_dict = {}","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:53.974872Z","iopub.execute_input":"2022-08-07T15:44:53.975383Z","iopub.status.idle":"2022-08-07T15:44:53.979894Z","shell.execute_reply.started":"2022-08-07T15:44:53.975326Z","shell.execute_reply":"2022-08-07T15:44:53.978968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Naive Bayes \n##### Model 1","metadata":{}},{"cell_type":"code","source":"classifer = GaussianNB()\npredictor = classifer.fit(X_train, y_train)\ny_pred = predictor.predict(X_val)\naccuracy_naive_bayes = accuracy_score(y_val, y_pred)\nmodel_dict['naive_bayes'] = accuracy_naive_bayes\nprint(accuracy_naive_bayes)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:54.261812Z","iopub.execute_input":"2022-08-07T15:44:54.262336Z","iopub.status.idle":"2022-08-07T15:44:54.276666Z","shell.execute_reply.started":"2022-08-07T15:44:54.262260Z","shell.execute_reply":"2022-08-07T15:44:54.275536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Linear Discriminant Analysis\n##### Model 2","metadata":{}},{"cell_type":"code","source":"classifer = LinearDiscriminantAnalysis()\npredictor = classifer.fit(X_train, y_train)\ny_pred = predictor.predict(X_val)\naccuracy_lda = accuracy_score(y_val, y_pred)\nmodel_dict['linear_discriminant_analysis'] = accuracy_lda\nprint(accuracy_lda)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:54.807701Z","iopub.execute_input":"2022-08-07T15:44:54.807977Z","iopub.status.idle":"2022-08-07T15:44:54.864856Z","shell.execute_reply.started":"2022-08-07T15:44:54.807948Z","shell.execute_reply":"2022-08-07T15:44:54.863318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Logistic Regression\n##### Model 3","metadata":{}},{"cell_type":"code","source":"classifier = LogisticRegression(random_state=42)\npredictor = classifier.fit(X_train, y_train)\ny_pred = predictor.predict(X_val)\naccuracy_log_reg = accuracy_score(y_val, y_pred)\nmodel_dict['logistic_regression'] = accuracy_log_reg\nprint(accuracy_log_reg)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:55.067939Z","iopub.execute_input":"2022-08-07T15:44:55.068549Z","iopub.status.idle":"2022-08-07T15:44:55.186430Z","shell.execute_reply.started":"2022-08-07T15:44:55.068509Z","shell.execute_reply":"2022-08-07T15:44:55.185223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Support Vector Classifier\n##### Model 4","metadata":{}},{"cell_type":"code","source":"classifier = SVC(random_state=42)\npredictor_svc = classifier.fit(X_train, y_train)\ny_pred = predictor_svc.predict(X_val)\naccuracy_svc = accuracy_score(y_val, y_pred)\nmodel_dict['SVC'] = accuracy_svc\nprint(accuracy_svc)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:55.290685Z","iopub.execute_input":"2022-08-07T15:44:55.291394Z","iopub.status.idle":"2022-08-07T15:44:57.896696Z","shell.execute_reply.started":"2022-08-07T15:44:55.291327Z","shell.execute_reply":"2022-08-07T15:44:57.895572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### K-Neighbors Classifier\n##### Model 5","metadata":{}},{"cell_type":"code","source":"from sklearn.neighbors import KNeighborsClassifier\nclassifier = KNeighborsClassifier()\npredictor = classifier.fit(X_train, y_train)\ny_pred = predictor.predict(X_val)\naccuracy_knn = accuracy_score(y_val, y_pred)\nmodel_dict['kneighbors_classifier'] = accuracy_knn\nprint(accuracy_knn)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:57.898343Z","iopub.execute_input":"2022-08-07T15:44:57.898613Z","iopub.status.idle":"2022-08-07T15:44:57.976086Z","shell.execute_reply.started":"2022-08-07T15:44:57.898582Z","shell.execute_reply":"2022-08-07T15:44:57.975344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Stochastic Gradient Descent Classifier\n##### Model 6","metadata":{}},{"cell_type":"code","source":"classifier = SGDClassifier(random_state=42)\npredictor = classifier.fit(X_train, y_train)\ny_pred = predictor.predict(X_val)\naccuracy_sgdc = accuracy_score(y_val, y_pred)\nmodel_dict['sgd_classifier'] = accuracy_sgdc\nprint(accuracy_sgdc)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:57.977238Z","iopub.execute_input":"2022-08-07T15:44:57.977688Z","iopub.status.idle":"2022-08-07T15:44:58.048011Z","shell.execute_reply.started":"2022-08-07T15:44:57.977650Z","shell.execute_reply":"2022-08-07T15:44:58.047099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Random Forest Classifier\n##### Model 7","metadata":{}},{"cell_type":"code","source":"classifier = RandomForestClassifier(random_state=42)\npredictor = classifier.fit(X_train, y_train)\ny_pred = predictor.predict(X_val)\naccuracy_rfc = accuracy_score(y_val, y_pred)\nmodel_dict['random_forest_classifier'] = accuracy_rfc\nprint(accuracy_rfc)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:58.050174Z","iopub.execute_input":"2022-08-07T15:44:58.050967Z","iopub.status.idle":"2022-08-07T15:44:59.159029Z","shell.execute_reply.started":"2022-08-07T15:44:58.050918Z","shell.execute_reply":"2022-08-07T15:44:59.158133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Gradient Boosting Classifier\n##### Model 8","metadata":{}},{"cell_type":"code","source":"classifier = GradientBoostingClassifier(random_state=42)\npredictor_gbc = classifier.fit(X_train, y_train)\ny_pred = predictor_gbc.predict(X_val)\naccuracy_gbc = accuracy_score(y_val, y_pred)\nmodel_dict['gradient_boosting_classifier'] = accuracy_gbc\nprint(accuracy_gbc)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:44:59.160319Z","iopub.execute_input":"2022-08-07T15:44:59.160961Z","iopub.status.idle":"2022-08-07T15:45:00.179133Z","shell.execute_reply.started":"2022-08-07T15:44:59.160927Z","shell.execute_reply":"2022-08-07T15:45:00.178093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### XGBoost Classifier\n##### Model 9","metadata":{}},{"cell_type":"code","source":"classifier = XGBClassifier(random_state=42, eval_metric='logloss')\npredictor_xgb = classifier.fit(X_train, y_train)\ny_pred = predictor_xgb.predict(X_val)\naccuracy_xgb = accuracy_score(y_val, y_pred)\nmodel_dict['xgboost_classifier'] = accuracy_xgb\nprint(accuracy_xgb)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:00.180691Z","iopub.execute_input":"2022-08-07T15:45:00.180958Z","iopub.status.idle":"2022-08-07T15:45:01.053047Z","shell.execute_reply.started":"2022-08-07T15:45:00.180907Z","shell.execute_reply":"2022-08-07T15:45:01.052035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### AdaBoost Classifier\n##### Model 10","metadata":{}},{"cell_type":"code","source":"dtc=DecisionTreeClassifier(criterion='entropy', random_state=42)\nclassifier = AdaBoostClassifier(dtc,random_state=42)\npredictor = classifier.fit(X_train, y_train)\ny_pred = predictor.predict(X_val)\naccuracy_ada = accuracy_score(y_val, y_pred)\nmodel_dict['adaboost_classifier'] = accuracy_ada\nprint(accuracy_ada)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:01.054483Z","iopub.execute_input":"2022-08-07T15:45:01.054839Z","iopub.status.idle":"2022-08-07T15:45:02.744508Z","shell.execute_reply.started":"2022-08-07T15:45:01.054779Z","shell.execute_reply":"2022-08-07T15:45:02.743353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### LGBM Classifier\n##### Model 11\n##### If you want to learn more about the LGBM classifier I suggest looking at this kernel: https://www.kaggle.com/code/prashant111/lightgbm-classifier-in-python/notebook","metadata":{}},{"cell_type":"code","source":"classifier = LGBMClassifier(random_state=42)\npredictor_lgbm = classifier.fit(X_train, y_train)\ny_pred = predictor_lgbm.predict(X_val)\naccuracy_lgbm = accuracy_score(y_val, y_pred)\nmodel_dict['lgbm_classifier'] = accuracy_lgbm\nprint(accuracy_lgbm)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:02.745796Z","iopub.execute_input":"2022-08-07T15:45:02.746029Z","iopub.status.idle":"2022-08-07T15:45:02.973525Z","shell.execute_reply.started":"2022-08-07T15:45:02.746001Z","shell.execute_reply":"2022-08-07T15:45:02.972499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Extra Trees Classifier\n##### Model 12","metadata":{}},{"cell_type":"code","source":"classifier = ExtraTreesClassifier(random_state=42)\npredictor_etc = classifier.fit(X_train, y_train)\ny_pred = predictor_etc.predict(X_val)\naccuracy_etc = accuracy_score(y_val, y_pred)\nmodel_dict['etc_classifier'] = accuracy_etc\nprint(accuracy_etc)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:02.975804Z","iopub.execute_input":"2022-08-07T15:45:02.976112Z","iopub.status.idle":"2022-08-07T15:45:03.930175Z","shell.execute_reply.started":"2022-08-07T15:45:02.976079Z","shell.execute_reply":"2022-08-07T15:45:03.929125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#14C38E;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"padding: 1rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\n Visualize model accuracies\n</h3>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### Function to visualizing model accuracies","metadata":{}},{"cell_type":"code","source":"def visualize_model_accuracies(model_dict):\n    model_accuracies_df = pd.DataFrame(columns=['Model', 'Accuracy'])\n    model_accuracies_df['Model'] = model_dict.keys()\n    model_accuracies_df['Accuracy'] = model_dict.values()\n    model_accuracies_df.sort_values('Accuracy', inplace=True, ascending=False)\n\n    plt.figure(figsize=(28,8),)\n    plt.ylabel(\"Models\", fontsize=16)\n    plt.xlabel(\"Accuracy\", fontsize=16)\n    plt.title(\"Model Accuracies\", fontsize=22)\n    sns.barplot(y = pd.to_numeric(model_accuracies_df['Accuracy']), x = model_accuracies_df['Model'], palette='mako')\n    plt.margins(x=0.005)\n    plt.show()\n\n    print(colored(\"The 3 models with the highest accuracies are:\"))\n    print(f\"{model_accuracies_df.iloc[:3, ]}\")","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:03.931343Z","iopub.execute_input":"2022-08-07T15:45:03.931638Z","iopub.status.idle":"2022-08-07T15:45:03.939266Z","shell.execute_reply.started":"2022-08-07T15:45:03.931605Z","shell.execute_reply":"2022-08-07T15:45:03.938424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"visualize_model_accuracies(model_dict)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:03.940485Z","iopub.execute_input":"2022-08-07T15:45:03.940727Z","iopub.status.idle":"2022-08-07T15:45:04.227610Z","shell.execute_reply.started":"2022-08-07T15:45:03.940697Z","shell.execute_reply":"2022-08-07T15:45:04.226727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### The three most accurate models include: XGBoost Classifer, Gradient Boosting Classifier, and SVC.\n##### We will focus on tuning their hyperparameters and once again compare their accuracies.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#14C38E;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"padding: 1rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\n Tuning Hyperparameters 🔧\n</h3>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### I will now perform Grid Search on the three candidate models\n#### Below I define yet another dictionary to keep track of the models\nNote:\nXGBoost Classifier is a very similar model to GB classifier, therefore I will use SVC as the 3rd option despite it being the 4th best performing model","metadata":{}},{"cell_type":"code","source":"gs_model_dict = {}","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:04.228690Z","iopub.execute_input":"2022-08-07T15:45:04.228901Z","iopub.status.idle":"2022-08-07T15:45:04.232808Z","shell.execute_reply.started":"2022-08-07T15:45:04.228874Z","shell.execute_reply":"2022-08-07T15:45:04.231790Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Grid Search for XGBoost Classifer\n##### Model 13","metadata":{}},{"cell_type":"code","source":"# GRID SEARCH PARAMETERS USED\n# parameters_xgbc = {\n#     'learning_rate': [0.1, 0.09 ,0.08, 0.05, 0.025, 0.01],\n#     'n_estimators': range(60, 220, 20),\n#     'max_depth':  range (2, 8, 1),\n# }\n\n# Results after performing grid search\nfinal_parameters_xgbc = {'learning_rate': 0.09, 'max_depth': 6, 'n_estimators': 100}\nclf_xgbc = XGBClassifier(random_state=42, eval_metric='logloss', base_score=0.5, booster='gbtree', gamma=0.45, **final_parameters_xgbc)\nclf_xgbc.fit(X_train, y_train)\ny_pred_xgb = clf_xgbc.predict(X_val)\naccuracy_xgb = accuracy_score(y_val, y_pred_xgb)\ngs_model_dict['xgboost_classifier'] = accuracy_xgb\nprint(f'\\n\\n Accuracy: {accuracy_xgb}')\nplot_confusion_matrix(clf_xgbc, X_val, y_val)  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:04.234154Z","iopub.execute_input":"2022-08-07T15:45:04.234390Z","iopub.status.idle":"2022-08-07T15:45:05.165262Z","shell.execute_reply.started":"2022-08-07T15:45:04.234361Z","shell.execute_reply":"2022-08-07T15:45:05.164334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Grid Search LGBM Classifier\n##### Model 14","metadata":{}},{"cell_type":"code","source":"# GRID SEARCH PARAMETERS USED\n# parameters_lgbm = {\n#     'learning_rate': [0.09, 0.085, 0.05, 0.02],\n#     'n_estimators': [100, 90],\n#     'min_child_samples': [25, 30, 40, 50, 60, 70],\n#     'num_leaves': [34, 37, 40, 43, 46, 49, 52]\n# }\n\n# Results after performing grid search\nfinal_parameters_lgbm = {\n 'learning_rate': 0.09,\n 'min_child_samples': 70,\n 'n_estimators': 113, # Adjusted by trial and error\n 'num_leaves': 40\n}\n\nclf_lgbm = LGBMClassifier(random_state=42, max_depth=8, subsample=0.6, **final_parameters_lgbm)\nclf_lgbm.fit(X_train, y_train)\ny_pred_lgbm = clf_lgbm.predict(X_val)\naccuracy_lgbm = accuracy_score(y_val, y_pred_lgbm)\ngs_model_dict['lgbm_classifier'] = accuracy_lgbm\nprint(f'Accuracy: {accuracy_lgbm}')\nplot_confusion_matrix(clf_lgbm, X_val, y_val)  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:05.167888Z","iopub.execute_input":"2022-08-07T15:45:05.169050Z","iopub.status.idle":"2022-08-07T15:45:05.535058Z","shell.execute_reply.started":"2022-08-07T15:45:05.169003Z","shell.execute_reply":"2022-08-07T15:45:05.532602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Grid Search SVC \n##### Model 15","metadata":{}},{"cell_type":"code","source":"# GRID SEARCH PARAMETERS USED\n# parameters_svc = {\n#     'C':range(1, 101, 5),\n#     'kernel':['rbf'],\n#     'class_weight':['balanced', None]\n# }\n\n# Results after performing grid search\nfinal_parameters_svc = {'C': 91, 'class_weight': None, 'kernel': 'rbf'}\nclf_svc = SVC(random_state=42, **final_parameters_svc)\nclf_svc.fit(X_train, y_train)\ny_pred_svc = clf_svc.predict(X_val)\naccuracy_svc = accuracy_score(y_val, y_pred_svc)\ngs_model_dict['SVC'] = accuracy_svc\nprint(f'Accuracy: {accuracy_svc}')\nplot_confusion_matrix(clf_svc, X_val, y_val) \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:05.536521Z","iopub.execute_input":"2022-08-07T15:45:05.537019Z","iopub.status.idle":"2022-08-07T15:45:11.190098Z","shell.execute_reply.started":"2022-08-07T15:45:05.536971Z","shell.execute_reply":"2022-08-07T15:45:11.189185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Visualize accuracies","metadata":{}},{"cell_type":"code","source":"visualize_model_accuracies(gs_model_dict)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:11.192629Z","iopub.execute_input":"2022-08-07T15:45:11.193154Z","iopub.status.idle":"2022-08-07T15:45:11.389763Z","shell.execute_reply.started":"2022-08-07T15:45:11.193107Z","shell.execute_reply":"2022-08-07T15:45:11.388847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#14C38E;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h3 style=\"padding: 1rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\n Voting 🗳\n</h3>\n</div>","metadata":{}},{"cell_type":"markdown","source":"#### Voting Ensemble\n##### Model 16","metadata":{}},{"cell_type":"code","source":"voting_ensemble = VotingClassifier(estimators= \n                                   [('SVC', clf_svc),\n                                    ('XBG', clf_xgbc),\n                                    ('LGBM', clf_lgbm)],\n                              voting = 'hard')\n\nensemble_predictor = voting_ensemble.fit(X_train, y_train)\n\ny_pred = ensemble_predictor.predict(X_val)\nprint(f'Accuracy score for the first ensemble: {accuracy_score(y_val, y_pred)}')\n","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:11.391370Z","iopub.execute_input":"2022-08-07T15:45:11.392860Z","iopub.status.idle":"2022-08-07T15:45:17.811047Z","shell.execute_reply.started":"2022-08-07T15:45:11.392796Z","shell.execute_reply":"2022-08-07T15:45:17.810332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Confusion matrix for ensemble composed of SVC, XGBC, and LGBM","metadata":{}},{"cell_type":"code","source":"print(\"Performance on validation data:\", f1_score(y_val, y_pred, average='micro'))\nplot_confusion_matrix(ensemble_predictor, X_val, y_val)  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:17.815372Z","iopub.execute_input":"2022-08-07T15:45:17.817597Z","iopub.status.idle":"2022-08-07T15:45:18.110948Z","shell.execute_reply.started":"2022-08-07T15:45:17.817546Z","shell.execute_reply":"2022-08-07T15:45:18.109178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Confusion matrix for ensemble composed of XGBC and LGBM","metadata":{}},{"cell_type":"code","source":"# print(\"Performance on validation data:\", f1_score(y_val, y_pred2, average='micro'))\n# plot_confusion_matrix(ensemble_predictor2, X_val, y_val)  \n# plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:18.112628Z","iopub.execute_input":"2022-08-07T15:45:18.113149Z","iopub.status.idle":"2022-08-07T15:45:18.117577Z","shell.execute_reply.started":"2022-08-07T15:45:18.113087Z","shell.execute_reply":"2022-08-07T15:45:18.116393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Final Classifier\n##### The model with the higher accuracy isn't that good","metadata":{}},{"cell_type":"code","source":"classifier = ensemble_predictor","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:18.119388Z","iopub.execute_input":"2022-08-07T15:45:18.120011Z","iopub.status.idle":"2022-08-07T15:45:18.130279Z","shell.execute_reply.started":"2022-08-07T15:45:18.119968Z","shell.execute_reply":"2022-08-07T15:45:18.129325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Note:\n##### If you improve your model selection, I am sure that the hyperparameters for XGBoost will/can be better, allowing for an improved accuracy","metadata":{}},{"cell_type":"markdown","source":"<div id=14 style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#5642C5;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n            justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:white;\n          text-align:center;\n          margin:0 auto;\n          \">\n Deep Learning 🤖\n</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"### Change data to numpy format","metadata":{}},{"cell_type":"code","source":"X_train = X_train.to_numpy()\ny_train = y_train.to_numpy()\nX_val = X_val.to_numpy()\ny_val = y_val.to_numpy()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:18.133453Z","iopub.execute_input":"2022-08-07T15:45:18.134077Z","iopub.status.idle":"2022-08-07T15:45:18.142953Z","shell.execute_reply.started":"2022-08-07T15:45:18.134031Z","shell.execute_reply":"2022-08-07T15:45:18.142347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model Architecture\n##### Model 17","metadata":{}},{"cell_type":"code","source":"LOSS_CONST = 30e-6\nSHAPE = X_train.shape[1]\nDROPOUT = 0.05\ninputs = Input(shape=(SHAPE))\n\n# Hidden Layers\nx = Dense(1024, kernel_regularizer=tf.keras.regularizers.l2(LOSS_CONST), activation='relu', name='dense_1')(inputs)\nx = BatchNormalization(name='batch_normalization_1')(x)\nx = Dropout(DROPOUT, name='dropout_1')(x)\n\n# x = Dense(128, kernel_regularizer=tf.keras.regularizers.l2(LOSS_CONST), activation='swish', name='dense_2')(inputs)\n# x = BatchNormalization(name='batch_normalization_2')(x)\n# x = Dropout(0.025, name='dropout_2')(x)\n\nx = Dense(512, kernel_regularizer=tf.keras.regularizers.l2(LOSS_CONST), activation='relu', name='dense_3')(inputs)\nx = BatchNormalization(name='batch_normalization_3')(x)\nx = Dropout(DROPOUT, name='dropout_3')(x)\n\n# x = Dense(128, kernel_regularizer=tf.keras.regularizers.l2(LOSS_CONST), activation='swish', name='dense_2')(inputs)\n# x = BatchNormalization(name='batch_normalization_2')(x)\n# x = Dropout(0.025, name='dropout_2')(x)\n\nx = Dense(128, kernel_regularizer=tf.keras.regularizers.l2(LOSS_CONST), activation='relu', name='dense_3')(inputs)\nx = BatchNormalization(name='batch_normalization_3')(x)\nx = Dropout(DROPOUT, name='dropout_3')(x)\n\n# x = Dense(32, kernel_regularizer=tf.keras.regularizers.l2(LOSS_CONST), activation='swish', name='dense_2')(inputs)\n# x = BatchNormalization(name='batch_normalization_2')(x)\n# x = Dropout(0.025, name='dropout_2')(x)\n\nx = Dense(32, kernel_regularizer=tf.keras.regularizers.l2(LOSS_CONST), activation='relu', name='dense_3')(inputs)\nx = BatchNormalization(name='batch_normalization_3')(x)\nx = Dropout(DROPOUT, name='dropout_3')(x)\n\nx = Dense(16, kernel_regularizer=tf.keras.regularizers.l2(LOSS_CONST), activation='relu', name='dense_3')(inputs)\nx = BatchNormalization(name='batch_normalization_3')(x)\nx = Dropout(DROPOUT, name='dropout_3')(x)\n\n\n# x = Dense(8, kernel_regularizer=tf.keras.regularizers.l2(LOSS_CONST), activation='relu', name='dense_3')(inputs)\n# x = BatchNormalization(name='batch_normalization_3')(x)\n# x = Dropout(DROPOUT, name='dropout_3')(x)\n\n# x = Dense(4, kernel_regularizer=tf.keras.regularizers.l2(LOSS_CONST), activation='relu', name='dense_3')(inputs)\n# x = BatchNormalization(name='batch_normalization_3')(x)\n# x = Dropout(DROPOUT, name='dropout_3')(x)\n\n\n\ny = Dense(1, activation='sigmoid', name='output_layer')(x)\n\nmodel = Model(inputs, y)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:18.144428Z","iopub.execute_input":"2022-08-07T15:45:18.144937Z","iopub.status.idle":"2022-08-07T15:45:18.275222Z","shell.execute_reply.started":"2022-08-07T15:45:18.144900Z","shell.execute_reply":"2022-08-07T15:45:18.274588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.compile(loss='binary_crossentropy', optimizer=keras.optimizers.Adam(learning_rate=0.0001), metrics=[tf.keras.metrics.AUC()])\nhistory = model.fit(X_train, \n                    y_train, \n                    validation_data=(X_val, y_val),\n                    epochs=100,\n                    batch_size=64,\n                   )","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:18.276384Z","iopub.execute_input":"2022-08-07T15:45:18.276739Z","iopub.status.idle":"2022-08-07T15:45:44.004493Z","shell.execute_reply.started":"2022-08-07T15:45:18.276711Z","shell.execute_reply":"2022-08-07T15:45:44.003491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = model.predict(X_val).round()\naccuracy_score(y_val, y_pred)\nprint(\"Performance on validation data:\", f1_score(y_val, y_pred, average='micro'))\nprint(confusion_matrix(y_val, y_pred))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:44.005953Z","iopub.execute_input":"2022-08-07T15:45:44.006274Z","iopub.status.idle":"2022-08-07T15:45:44.134198Z","shell.execute_reply.started":"2022-08-07T15:45:44.006241Z","shell.execute_reply":"2022-08-07T15:45:44.133593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Submission","metadata":{}},{"cell_type":"markdown","source":"#### For now we submit the ML voting classifier, since the DL model needs a lot of work","metadata":{}},{"cell_type":"code","source":"submission = classifier.predict(test_data)\n\ndata = {\n    'PassengerId': np.array(PASSENGER_ID).reshape(len(PASSENGER_ID)),\n    'Transported': np.array(submission).astype('bool')\n}\n\nsubmission_df = pd.DataFrame(data=data).reset_index()\nsubmission_df.drop(columns=['index'], inplace=True, axis=1)\nsubmission_df.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T15:45:44.135219Z","iopub.execute_input":"2022-08-07T15:45:44.135507Z","iopub.status.idle":"2022-08-07T15:45:45.271287Z","shell.execute_reply.started":"2022-08-07T15:45:44.135476Z","shell.execute_reply":"2022-08-07T15:45:45.270023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Upvote if you thought the kernel was insightful 👍\n### If you have any suggestions feel free to leave a comment, I am willing to listen and learn :)","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           background-color:#F4E06D;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px;\n           display:flex;\n           flex-direction: row;\">\n\n<h1 style=\"padding: 2rem;\n          color:black;\n          text-align:center;\n          margin:0 auto;\n          font-size:3rem;\">\n   WORK IN PROGRESS ⚠\n</h1>\n \n</div>","metadata":{}}]}