{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"#Table of Contents\n<a id=\"toc\"></a>\n\n1. [Importing libraries](#1)\n\n2. [Loading data](#2)\n\n    2.1 Train Data\n    \n    2.2 Test Data\n    \n3. [Exploratory data analysis](#3)\n\n    3.1 Missing values\n    \n    3.2 Categorical features\n\n    3.3 `Transported`\n   \n    3.4 `Age`\n    \n    3.5 Numerical features\n    \n4. [Feature Engineering](#4)\n    \n    4.1 `Age`\n    \n    4.2 Expenditure\n    \n    4.3 Passenger group extraction\n\n    4.4 `Cabin`\n    \n    4.5 `FamilySize`\n\n    4.6 Missing values\n    \n      4.6.1 `HomePlanet`\n    \n      4.6.2 `Destination`\n      \n      4.6.3 `Surname`\n      \n      4.6.4 `CabinSide`\n      \n      4.6.5 `CabinDeck`\n      \n      4.6.6 `CabinNumber`\n      \n      4.6.7 `VIP`\n      \n      4.6.8 `Age`\n      \n      4.6.9 `CryoSleep`\n      \n      4.6.10 Expenditure\n\n5. [Data Pre-Processing](#5)\n\n      5.1 Dropping columns\n      \n      5.2 Log transform\n      \n      5.3 Encoding and Scaling\n      \n      5.4 Validation split\n      \n6. [Models](#6)\n    \n    6.1 Classifiers\n\n    6.2 Catboost Classifier\n    \n7. [Submission](#7)","metadata":{}},{"cell_type":"markdown","source":"Let's start with importing necessary libraries!\n# 1. Importing libraries  <a id=\"1\"></a>\n","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport warnings\nimport ggplot\nfrom ggplot import aes\nfrom sklearn.preprocessing import OneHotEncoder\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.model_selection import train_test_split, GridSearchCV\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.pipeline import Pipeline\nwarnings.simplefilter(\"ignore\")\n\n\nfrom xgboost import XGBClassifier\nfrom lightgbm import LGBMClassifier\nfrom catboost import CatBoostClassifier\nfrom sklearn.metrics import confusion_matrix,plot_confusion_matrix\n\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\n\n","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:17.959223Z","iopub.execute_input":"2022-08-11T13:14:17.959728Z","iopub.status.idle":"2022-08-11T13:14:28.124888Z","shell.execute_reply.started":"2022-08-11T13:14:17.959624Z","shell.execute_reply":"2022-08-11T13:14:28.123698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Loading data <a id=\"2\"></a>","metadata":{}},{"cell_type":"code","source":"df_train = pd.read_csv('../input/spaceship-titanic/train.csv')\ndf_test = pd.read_csv('../input/spaceship-titanic/test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:28.127059Z","iopub.execute_input":"2022-08-11T13:14:28.128294Z","iopub.status.idle":"2022-08-11T13:14:28.199528Z","shell.execute_reply.started":"2022-08-11T13:14:28.128242Z","shell.execute_reply":"2022-08-11T13:14:28.198399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* `PassengerId` - A unique Id for each passenger. Each Id takes the form gggg_pp where gggg indicates a group the passenger is travelling with and pp is their number within the group. People in a group are often family members, but not always.\n* `HomePlanet` - The planet the passenger departed from, typically their planet of permanent residence.\n* `CryoSleep` - Indicates whether the passenger elected to be put into suspended animation for the duration of the voyage. Passengers in cryosleep are confined to their cabins.\n* `Cabin` - The cabin number where the passenger is staying. Takes the form deck/num/side, where side can be either 'P' for Port or 'S' for Starboard.\n* `Destination` - The planet the passenger will be debarking to.\n* `Age` - The age of the passenger.\n* `VIP` - Whether the passenger has paid for special VIP service during the voyage.\n* `RoomService`, `FoodCourt`, `ShoppingMall, Spa`, `VRDeck` - Amount the passenger has billed at each of the Spaceship Titanic's many luxury amenities.\n* `Name` - The first and last names of the passenger.\n* `Transported` - Whether the passenger was transported to another dimension. This is the target, the column you are trying to predict.","metadata":{}},{"cell_type":"markdown","source":"## 2.1 Train data","metadata":{}},{"cell_type":"code","source":"df_train.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:28.201354Z","iopub.execute_input":"2022-08-11T13:14:28.201811Z","iopub.status.idle":"2022-08-11T13:14:28.239527Z","shell.execute_reply.started":"2022-08-11T13:14:28.201767Z","shell.execute_reply":"2022-08-11T13:14:28.238276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:28.242332Z","iopub.execute_input":"2022-08-11T13:14:28.242746Z","iopub.status.idle":"2022-08-11T13:14:28.274787Z","shell.execute_reply.started":"2022-08-11T13:14:28.242711Z","shell.execute_reply":"2022-08-11T13:14:28.272867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_train.shape)\ndf_train.columns","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:28.276877Z","iopub.execute_input":"2022-08-11T13:14:28.277712Z","iopub.status.idle":"2022-08-11T13:14:28.286131Z","shell.execute_reply.started":"2022-08-11T13:14:28.277676Z","shell.execute_reply":"2022-08-11T13:14:28.284906Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have 14 columns and 8693 rows which means we have around 8700 passengers. `Transported` is our target data we will try to predict this value.","metadata":{}},{"cell_type":"markdown","source":"## 2.2 Test data","metadata":{}},{"cell_type":"code","source":"df_test.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:28.287375Z","iopub.execute_input":"2022-08-11T13:14:28.288449Z","iopub.status.idle":"2022-08-11T13:14:28.315775Z","shell.execute_reply.started":"2022-08-11T13:14:28.288386Z","shell.execute_reply":"2022-08-11T13:14:28.314500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_test.shape)\ndf_test.columns","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:28.317808Z","iopub.execute_input":"2022-08-11T13:14:28.318245Z","iopub.status.idle":"2022-08-11T13:14:28.327192Z","shell.execute_reply.started":"2022-08-11T13:14:28.318208Z","shell.execute_reply":"2022-08-11T13:14:28.326139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Exploratory data analysis <a id=\"3\"></a>","metadata":{}},{"cell_type":"code","source":"df_train.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:28.328715Z","iopub.execute_input":"2022-08-11T13:14:28.329321Z","iopub.status.idle":"2022-08-11T13:14:28.375332Z","shell.execute_reply.started":"2022-08-11T13:14:28.329286Z","shell.execute_reply":"2022-08-11T13:14:28.373649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The average age of approximately 8700 passengers in our train data is 28,82.","metadata":{}},{"cell_type":"markdown","source":"## 3.1 Missing values","metadata":{}},{"cell_type":"markdown","source":"Visualizing the number of missing values in `df_train` and `df_test`","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (20,5))\nplt.bar(df_train.columns, df_train.isna().sum(),color='orange')\nplt.xlabel(\"Columns name\")\nplt.ylabel(\"Number of missing values in data\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:28.377540Z","iopub.execute_input":"2022-08-11T13:14:28.378504Z","iopub.status.idle":"2022-08-11T13:14:28.684734Z","shell.execute_reply.started":"2022-08-11T13:14:28.378455Z","shell.execute_reply":"2022-08-11T13:14:28.683505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_train.isna().sum().sort_values(ascending=False))","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:28.690215Z","iopub.execute_input":"2022-08-11T13:14:28.690644Z","iopub.status.idle":"2022-08-11T13:14:28.707863Z","shell.execute_reply.started":"2022-08-11T13:14:28.690609Z","shell.execute_reply":"2022-08-11T13:14:28.706670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (20,5))\nplt.bar(df_test.columns, df_test.isna().sum(),color='pink')\nplt.xlabel(\"Columns name\")\nplt.ylabel(\"Number of missing values in data\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:28.709257Z","iopub.execute_input":"2022-08-11T13:14:28.710210Z","iopub.status.idle":"2022-08-11T13:14:28.979131Z","shell.execute_reply.started":"2022-08-11T13:14:28.710171Z","shell.execute_reply":"2022-08-11T13:14:28.977628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_test.isna().sum().sort_values(ascending=False))","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:28.980888Z","iopub.execute_input":"2022-08-11T13:14:28.981294Z","iopub.status.idle":"2022-08-11T13:14:28.992842Z","shell.execute_reply.started":"2022-08-11T13:14:28.981260Z","shell.execute_reply":"2022-08-11T13:14:28.991416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Almost all features have missing values.","metadata":{}},{"cell_type":"markdown","source":"## 3.2 Categorical features","metadata":{}},{"cell_type":"markdown","source":"Let's see categorical columns of `df_train` so we have 8 categorical columns.","metadata":{}},{"cell_type":"code","source":"df_train.select_dtypes(include=['object','bool']).columns.tolist()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:28.994501Z","iopub.execute_input":"2022-08-11T13:14:28.994942Z","iopub.status.idle":"2022-08-11T13:14:29.008668Z","shell.execute_reply.started":"2022-08-11T13:14:28.994904Z","shell.execute_reply":"2022-08-11T13:14:29.007135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's see numerical columns of `df_train` so we have 6 numerical columns.","metadata":{}},{"cell_type":"code","source":"df_train.select_dtypes(include=['float64']).columns.tolist()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:29.010676Z","iopub.execute_input":"2022-08-11T13:14:29.011256Z","iopub.status.idle":"2022-08-11T13:14:29.021355Z","shell.execute_reply.started":"2022-08-11T13:14:29.011222Z","shell.execute_reply":"2022-08-11T13:14:29.020525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_features=['HomePlanet', 'CryoSleep', 'Destination', 'VIP']\n\nfig=plt.figure(figsize=(10,16))\nfor i, var_name in enumerate(cat_features):\n    ax=fig.add_subplot(4,1,i+1)\n    sns.countplot(data = df_train, x=var_name, axes=ax, hue='Transported',color='green')\n    ax.set_title(var_name)\nfig.tight_layout()  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:29.023027Z","iopub.execute_input":"2022-08-11T13:14:29.023387Z","iopub.status.idle":"2022-08-11T13:14:29.787476Z","shell.execute_reply.started":"2022-08-11T13:14:29.023355Z","shell.execute_reply":"2022-08-11T13:14:29.786155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* `VIP` probably isn’t  a useful feature.\n* `CryoSleep` appears the be a very helpful feature.","metadata":{}},{"cell_type":"code","source":"#other categorical features\nfeatures=['PassengerId', 'Cabin' ,'Name']\n\n# Preview qualitative features\ndf_train[features].head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:29.789117Z","iopub.execute_input":"2022-08-11T13:14:29.789630Z","iopub.status.idle":"2022-08-11T13:14:29.805126Z","shell.execute_reply.started":"2022-08-11T13:14:29.789584Z","shell.execute_reply":"2022-08-11T13:14:29.804142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* We can extract the group and group size from the `PassengerId` feature.\n* We can extract the deck, number and side from the `Cabin` feature.\n* We could extract the surname from the `Name` feature to identify families.","metadata":{}},{"cell_type":"markdown","source":"## 3.3 `Transported`","metadata":{}},{"cell_type":"code","source":"ax = sns.countplot(x=\"Transported\", data=df_train,color='red')\nfor p in ax.patches:\n   ax.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0.25, p.get_height()+0.01))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:29.806666Z","iopub.execute_input":"2022-08-11T13:14:29.807465Z","iopub.status.idle":"2022-08-11T13:14:29.990520Z","shell.execute_reply.started":"2022-08-11T13:14:29.807430Z","shell.execute_reply":"2022-08-11T13:14:29.989261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The target is balanced!","metadata":{}},{"cell_type":"markdown","source":"## 3.4 `Age`","metadata":{}},{"cell_type":"code","source":"sns.histplot(data=df_train, x=\"Age\", hue=\"Transported\")\nplt.title('Age Histogram')\nplt.xlabel('Age (years)')\nplt.ylabel('Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:29.992156Z","iopub.execute_input":"2022-08-11T13:14:29.992517Z","iopub.status.idle":"2022-08-11T13:14:30.640481Z","shell.execute_reply.started":"2022-08-11T13:14:29.992484Z","shell.execute_reply":"2022-08-11T13:14:30.639328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* The ages between 0 - 18 are more likely to be transported. \n* The ages between 18 - 32 are more likely to be transported. \n* Those over 32 years old is equally likely to be transported. \n\nI think we should create a new age-related feature, maybe we can call the 0-18 age range a child, the 18-32 age range young, and the rest of the age range an adult.","metadata":{}},{"cell_type":"markdown","source":"## 3.5 Numerical features\n\nSuch as `RoomService`, `FoodCourt`, `ShoppingMall`, `Spa`, `VRDeck` . These are the amount the passenger has billed at each of the Spaceship Titanic's many luxury amenities.","metadata":{}},{"cell_type":"code","source":"df_train.RoomService.value_counts().sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:30.642714Z","iopub.execute_input":"2022-08-11T13:14:30.643216Z","iopub.status.idle":"2022-08-11T13:14:30.658999Z","shell.execute_reply.started":"2022-08-11T13:14:30.643166Z","shell.execute_reply":"2022-08-11T13:14:30.657466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.Spa.value_counts().sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:30.660401Z","iopub.execute_input":"2022-08-11T13:14:30.660791Z","iopub.status.idle":"2022-08-11T13:14:30.673184Z","shell.execute_reply.started":"2022-08-11T13:14:30.660757Z","shell.execute_reply":"2022-08-11T13:14:30.671714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Mostly zero expenditure.","metadata":{}},{"cell_type":"code","source":"num_features = ['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n\nfig=plt.figure(figsize=(10,20))\nfor i, var_name in enumerate(num_features):\n    # Left plot\n    ax=fig.add_subplot(5,2,2*i+1)\n    sns.histplot(data = df_train, x=var_name, axes=ax, bins=30, hue='Transported')\n    ax.set_title(var_name)\n    \n    # Right plot (truncated)\n    ax=fig.add_subplot(5,2,2*i+2)\n    sns.histplot(data = df_train, x=var_name, axes=ax, bins=30, hue='Transported')\n    plt.ylim([0,1000])\n    plt.xlim([0,10000])\n    ax.set_title(var_name)\nfig.tight_layout()  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:30.676126Z","iopub.execute_input":"2022-08-11T13:14:30.676514Z","iopub.status.idle":"2022-08-11T13:14:33.825880Z","shell.execute_reply.started":"2022-08-11T13:14:30.676462Z","shell.execute_reply":"2022-08-11T13:14:33.824918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Many people don't spend any money.\n* There are a small number of outliers.\n* People who were transported tended to spend less.\n* `RoomService`, `Spa` and `VRDeck` have different distributions to `FoodCourt` and `ShoppingMall`.\n\nWe should create a new feature that tracks the total expenditure. Create a binary feature to indicate if the person has not spent anything also take the log transform to reduce skewness.","metadata":{}},{"cell_type":"markdown","source":"# 4. Feature engineering <a id=\"4\"></a>","metadata":{}},{"cell_type":"markdown","source":"## 4.1 Age\n\nLet's arrange the age ranges by adding new column named as `AgeGroup` to `df_train` and `df_test` .","metadata":{}},{"cell_type":"code","source":"df_train['AgeGroup']=np.nan\ndf_train.loc[df_train['Age']<=14,'AgeGroup']='Children'\ndf_train.loc[(df_train['Age']>14) & (df_train['Age']<24),'AgeGroup']='Youth'\ndf_train.loc[(df_train['Age']>=24) & (df_train['Age']<=64),'AgeGroup']='Adult'\ndf_train.loc[df_train['Age']>64,'AgeGroup']='Seniors'\n\ndf_test['AgeGroup']=np.nan\ndf_test.loc[df_test['Age']<=14,'AgeGroup']='Children'\ndf_test.loc[(df_test['Age']>14) & (df_test['Age']<24),'AgeGroup']='Youth'\ndf_test.loc[(df_test['Age']>=24) & (df_test['Age']<=64),'AgeGroup']='Adult'\ndf_test.loc[df_test['Age']>64,'AgeGroup']='Seniors'","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:33.827164Z","iopub.execute_input":"2022-08-11T13:14:33.827649Z","iopub.status.idle":"2022-08-11T13:14:33.846613Z","shell.execute_reply.started":"2022-08-11T13:14:33.827618Z","shell.execute_reply":"2022-08-11T13:14:33.845253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10,4))\ng=sns.countplot(data=df_train, x='AgeGroup', hue='Transported', order=['Children','Youth','Adult','Seniors'],color='yellow')\nplt.title('Age group distribution')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:33.851222Z","iopub.execute_input":"2022-08-11T13:14:33.851606Z","iopub.status.idle":"2022-08-11T13:14:34.090822Z","shell.execute_reply.started":"2022-08-11T13:14:33.851550Z","shell.execute_reply":"2022-08-11T13:14:34.089504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.2 Expenditure\n\nLet's add new feature named `TotalSpending` to determine if the passenger spend any.","metadata":{}},{"cell_type":"code","source":"df_train['Expenditure'] = df_train[num_features].sum(axis=1)\ndf_train['TotalSpending'] = (df_train['Expenditure']==0).astype(int)\n\ndf_test['Expenditure'] = df_test[num_features].sum(axis=1)\ndf_test['TotalSpending'] = (df_test['Expenditure']==0).astype(int)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:34.092311Z","iopub.execute_input":"2022-08-11T13:14:34.093397Z","iopub.status.idle":"2022-08-11T13:14:34.109778Z","shell.execute_reply.started":"2022-08-11T13:14:34.093358Z","shell.execute_reply":"2022-08-11T13:14:34.108614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(data=df_train, x='TotalSpending', hue='Transported')\nplt.title('No spending indicator')\nfig.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:34.114775Z","iopub.execute_input":"2022-08-11T13:14:34.115150Z","iopub.status.idle":"2022-08-11T13:14:34.631592Z","shell.execute_reply.started":"2022-08-11T13:14:34.115117Z","shell.execute_reply":"2022-08-11T13:14:34.630522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.3 Passenger group extraction\n\nA unique Id for each passenger. Each Id takes the form gggg_pp where gggg indicates a group the passenger is travelling with and pp is their number within the group. So we seperate the gggg and pp from each other to create new features named `Group` and `GroupSize`","metadata":{}},{"cell_type":"code","source":"df_train.PassengerId.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:34.632833Z","iopub.execute_input":"2022-08-11T13:14:34.633141Z","iopub.status.idle":"2022-08-11T13:14:34.641819Z","shell.execute_reply.started":"2022-08-11T13:14:34.633111Z","shell.execute_reply":"2022-08-11T13:14:34.640584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['Group'] = df_train['PassengerId'].apply(lambda x: x.split('_')[0]).astype(int)\ndf_test['Group'] = df_test['PassengerId'].apply(lambda x: x.split('_')[0]).astype(int)\n\ndf_train['GroupSize']=df_train['Group'].map(lambda x: pd.concat([df_train['Group'], df_test['Group']]).value_counts()[x])\ndf_test['GroupSize']=df_test['Group'].map(lambda x: pd.concat([df_train['Group'], df_test['Group']]).value_counts()[x])","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:34.643087Z","iopub.execute_input":"2022-08-11T13:14:34.643408Z","iopub.status.idle":"2022-08-11T13:14:52.816803Z","shell.execute_reply.started":"2022-08-11T13:14:34.643372Z","shell.execute_reply":"2022-08-11T13:14:52.815532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(data=df_train, x='GroupSize', hue='Transported',color='purple')\nplt.title('Group size')\nfig.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:52.826170Z","iopub.execute_input":"2022-08-11T13:14:52.826636Z","iopub.status.idle":"2022-08-11T13:14:53.328781Z","shell.execute_reply.started":"2022-08-11T13:14:52.826588Z","shell.execute_reply":"2022-08-11T13:14:53.327504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.Group.value_counts().sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:53.330385Z","iopub.execute_input":"2022-08-11T13:14:53.331445Z","iopub.status.idle":"2022-08-11T13:14:53.343047Z","shell.execute_reply.started":"2022-08-11T13:14:53.331407Z","shell.execute_reply":"2022-08-11T13:14:53.341746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Also, we can add a feature called `IsAlone` to identify the passenger is alone or with group.","metadata":{}},{"cell_type":"code","source":"df_train['IsAlone']=(df_train['GroupSize']==1).astype(int)\ndf_test['IsAlone']=(df_test['GroupSize']==1).astype(int)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:53.344878Z","iopub.execute_input":"2022-08-11T13:14:53.345647Z","iopub.status.idle":"2022-08-11T13:14:53.353135Z","shell.execute_reply.started":"2022-08-11T13:14:53.345607Z","shell.execute_reply":"2022-08-11T13:14:53.352080Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10,4))\nsns.countplot(data=df_train, x='IsAlone', hue='Transported',color='blue')\nplt.title('Passenger is alone or not?')\nplt.ylim([0,3000])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:53.354902Z","iopub.execute_input":"2022-08-11T13:14:53.355605Z","iopub.status.idle":"2022-08-11T13:14:53.568763Z","shell.execute_reply.started":"2022-08-11T13:14:53.355545Z","shell.execute_reply":"2022-08-11T13:14:53.567345Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.4 `Cabin`\n\nWe should extract deck, number and side from `Cabin` feature.\n\n* Takes the form deck/num/side, where side can be either P for Port or S for Starboard.\n\nLet's add new features named as `CabinDeck`, `CabinNumber`, `CabinSide` .","metadata":{}},{"cell_type":"code","source":"df_train.Cabin.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:53.570194Z","iopub.execute_input":"2022-08-11T13:14:53.570817Z","iopub.status.idle":"2022-08-11T13:14:53.580980Z","shell.execute_reply.started":"2022-08-11T13:14:53.570782Z","shell.execute_reply":"2022-08-11T13:14:53.579658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['Cabin'].fillna('Z/9999/Z', inplace=True)\ndf_test['Cabin'].fillna('Z/9999/Z', inplace=True)\n\n\ndf_train['CabinDeck'] = df_train['Cabin'].apply(lambda x: x.split('/')[0])\ndf_train['CabinNumber'] = df_train['Cabin'].apply(lambda x: x.split('/')[1]).astype(int)\ndf_train['CabinSide'] = df_train['Cabin'].apply(lambda x: x.split('/')[2])\n\n\ndf_test['CabinDeck'] = df_test['Cabin'].apply(lambda x: x.split('/')[0])\ndf_test['CabinNumber'] = df_test['Cabin'].apply(lambda x: x.split('/')[1]).astype(int)\ndf_test['CabinSide'] = df_test['Cabin'].apply(lambda x: x.split('/')[2])\n\ndf_train.loc[df_train['CabinDeck']=='Z', 'CabinDeck']=np.nan\ndf_train.loc[df_train['CabinNumber']==9999, 'CabinNumber']=np.nan\ndf_train.loc[df_train['CabinSide']=='Z', 'CabinSide']=np.nan\ndf_test.loc[df_test['CabinDeck']=='Z', 'CabinDeck']=np.nan\ndf_test.loc[df_test['CabinNumber']==9999, 'CabinNumber']=np.nan\ndf_test.loc[df_test['CabinSide']=='Z', 'CabinSide']=np.nan","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:53.583232Z","iopub.execute_input":"2022-08-11T13:14:53.583780Z","iopub.status.idle":"2022-08-11T13:14:53.631909Z","shell.execute_reply.started":"2022-08-11T13:14:53.583732Z","shell.execute_reply":"2022-08-11T13:14:53.630475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Drop Cabin\ndf_train.drop('Cabin', axis=1, inplace=True)\ndf_test.drop('Cabin', axis=1, inplace=True)\n\nfig=plt.figure(figsize=(10,12))\nplt.subplot(3,1,1)\nsns.countplot(data=df_train, x='CabinDeck', hue='Transported', order=['A','B','C','D','E','F','G','T'])\nplt.title('Cabin Deck')\n\nplt.subplot(3,1,2)\nsns.histplot(data=df_train, x='CabinNumber', hue='Transported',binwidth=20)\nplt.title('Cabin Number')\nplt.xlim([0,2000])\n\nplt.subplot(3,1,3)\nsns.countplot(data=df_train, x='CabinSide', hue='Transported')\nplt.title('Cabin side')\nfig.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:53.633880Z","iopub.execute_input":"2022-08-11T13:14:53.634819Z","iopub.status.idle":"2022-08-11T13:14:54.825256Z","shell.execute_reply.started":"2022-08-11T13:14:53.634767Z","shell.execute_reply":"2022-08-11T13:14:54.823996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It looks that `CabinNumber` is grouped into chunks. This means we can convert this into a categorical feature, which indicates which chunk each passenger is in.\n\n* The cabin deck 'T' seems to be an outlier.","metadata":{}},{"cell_type":"code","source":"df_train['CabinRegion1']=(df_train['CabinNumber']<300).astype(int)  \ndf_train['CabinRegion2']=((df_train['CabinNumber']>=300) & (df_train['CabinNumber']<600)).astype(int)\ndf_train['CabinRegion3']=((df_train['CabinNumber']>=600) & (df_train['CabinNumber']<1200)).astype(int)\ndf_train['CabinRegion4']=(df_train['CabinNumber']>=1200).astype(int)\n\ndf_test['CabinRegion1']=(df_test['CabinNumber']<300).astype(int)  \ndf_test['CabinRegion2']=((df_test['CabinNumber']>=300) & (df_test['CabinNumber']<600)).astype(int)\ndf_test['CabinRegion3']=((df_test['CabinNumber']>=600) & (df_test['CabinNumber']<1200)).astype(int)\ndf_test['CabinRegion4']=(df_test['CabinNumber']>=1200).astype(int)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:54.826875Z","iopub.execute_input":"2022-08-11T13:14:54.827271Z","iopub.status.idle":"2022-08-11T13:14:54.846030Z","shell.execute_reply.started":"2022-08-11T13:14:54.827236Z","shell.execute_reply":"2022-08-11T13:14:54.844647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10,4))\ndf_train['CabinRegionsPlot']=(df_train['CabinRegion1']+2*df_train['CabinRegion2']+3*df_train['CabinRegion3']+4*df_train['CabinRegion4']).astype(int)\nsns.countplot(data=df_train, x='CabinRegionsPlot', hue='Transported')\nplt.title('Cabin regions')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:54.847390Z","iopub.execute_input":"2022-08-11T13:14:54.848548Z","iopub.status.idle":"2022-08-11T13:14:55.099432Z","shell.execute_reply.started":"2022-08-11T13:14:54.848506Z","shell.execute_reply":"2022-08-11T13:14:55.098260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.drop('CabinRegionsPlot', axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:55.101286Z","iopub.execute_input":"2022-08-11T13:14:55.102001Z","iopub.status.idle":"2022-08-11T13:14:55.110722Z","shell.execute_reply.started":"2022-08-11T13:14:55.101956Z","shell.execute_reply":"2022-08-11T13:14:55.109806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.5 `FamilySize`\n\nWe will calculate family size from last name. We will add the new features named `Surname` and `FamilySize` and then drop the `Name` column.","metadata":{}},{"cell_type":"code","source":"df_train.Name.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:55.112358Z","iopub.execute_input":"2022-08-11T13:14:55.113345Z","iopub.status.idle":"2022-08-11T13:14:55.124109Z","shell.execute_reply.started":"2022-08-11T13:14:55.113309Z","shell.execute_reply":"2022-08-11T13:14:55.122837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.Name.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:55.125808Z","iopub.execute_input":"2022-08-11T13:14:55.127052Z","iopub.status.idle":"2022-08-11T13:14:55.135945Z","shell.execute_reply.started":"2022-08-11T13:14:55.126992Z","shell.execute_reply":"2022-08-11T13:14:55.134893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['Name'].fillna('Unknown Unknown', inplace=True)\ndf_test['Name'].fillna('Unknown Unknown', inplace=True)\n\ndf_train['Surname']=df_train['Name'].str.split().str[-1]\ndf_test['Surname']=df_test['Name'].str.split().str[-1]\n\ndf_train['FamilySize']=df_train['Surname'].map(lambda x: pd.concat([df_train['Surname'],df_test['Surname']]).value_counts()[x])\ndf_test['FamilySize']=df_test['Surname'].map(lambda x: pd.concat([df_train['Surname'],df_test['Surname']]).value_counts()[x])\n","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:14:55.137256Z","iopub.execute_input":"2022-08-11T13:14:55.138418Z","iopub.status.idle":"2022-08-11T13:15:31.774687Z","shell.execute_reply.started":"2022-08-11T13:14:55.138370Z","shell.execute_reply":"2022-08-11T13:15:31.773456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.loc[df_train['Surname']=='Unknown','Surname']=np.nan\ndf_train.loc[df_train['FamilySize']>100,'FamilySize']=np.nan\ndf_test.loc[df_test['Surname']=='Unknown','Surname']=np.nan\ndf_test.loc[df_test['FamilySize']>100,'FamilySize']=np.nan\n\n# Drop Name\ndf_train.drop('Name', axis=1, inplace=True)\ndf_test.drop('Name', axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:31.776311Z","iopub.execute_input":"2022-08-11T13:15:31.776726Z","iopub.status.idle":"2022-08-11T13:15:31.796367Z","shell.execute_reply.started":"2022-08-11T13:15:31.776690Z","shell.execute_reply":"2022-08-11T13:15:31.795330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12,4))\nsns.countplot(data=df_train, x='FamilySize', hue='Transported',color='pink')\nplt.title('Size of the family')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:31.797623Z","iopub.execute_input":"2022-08-11T13:15:31.798404Z","iopub.status.idle":"2022-08-11T13:15:32.178279Z","shell.execute_reply.started":"2022-08-11T13:15:31.798371Z","shell.execute_reply":"2022-08-11T13:15:32.177109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.6 Missing values\n\nNow, we deal with missing values!\n\n* We should concat training and testing dataframes to deal with missing values all at once.","metadata":{}},{"cell_type":"code","source":"y = df_train['Transported'].copy().astype(int)\nX = df_train.drop('Transported', axis=1).copy()\n\ndf = pd.concat([X, df_test], axis=0).reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:32.179854Z","iopub.execute_input":"2022-08-11T13:15:32.180201Z","iopub.status.idle":"2022-08-11T13:15:32.205284Z","shell.execute_reply.started":"2022-08-11T13:15:32.180170Z","shell.execute_reply":"2022-08-11T13:15:32.204045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols = df.columns[df.isna().any()].tolist()\n\nMissingValues = pd.DataFrame(df[cols].isna().sum(), columns=['NumberMissing'])","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:32.207261Z","iopub.execute_input":"2022-08-11T13:15:32.208344Z","iopub.status.idle":"2022-08-11T13:15:32.232555Z","shell.execute_reply.started":"2022-08-11T13:15:32.208296Z","shell.execute_reply":"2022-08-11T13:15:32.231661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"MissingValues","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:32.234133Z","iopub.execute_input":"2022-08-11T13:15:32.234581Z","iopub.status.idle":"2022-08-11T13:15:32.247472Z","shell.execute_reply.started":"2022-08-11T13:15:32.234522Z","shell.execute_reply":"2022-08-11T13:15:32.246271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can find a relationship between PassengerId and other attributes and try to fill in the missing values.\n\n* The way to deal with missing values is to use the median for numerical features and the mode for categorical features.","metadata":{}},{"cell_type":"markdown","source":"### 4.6.1 `HomePlanet`\n\nLet's fill in the missing HomePlanet values by the same `Group`.","metadata":{}},{"cell_type":"code","source":"GroupHomePlanet = df.groupby(['Group','HomePlanet'])['HomePlanet'].size().unstack().fillna(0)\nGroupHomePlanet.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:32.248978Z","iopub.execute_input":"2022-08-11T13:15:32.249740Z","iopub.status.idle":"2022-08-11T13:15:32.279791Z","shell.execute_reply.started":"2022-08-11T13:15:32.249696Z","shell.execute_reply":"2022-08-11T13:15:32.278541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"GroupHomePlanetIndex = df[df['HomePlanet'].isna()][(df[df['HomePlanet'].isna()]['Group']).isin(GroupHomePlanet.index)].index\n\ndf.loc[GroupHomePlanetIndex,'HomePlanet']=df.iloc[GroupHomePlanetIndex,:]['Group'].map(lambda x: GroupHomePlanet.idxmax(axis=1)[x])","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:32.281307Z","iopub.execute_input":"2022-08-11T13:15:32.283528Z","iopub.status.idle":"2022-08-11T13:15:33.856257Z","shell.execute_reply.started":"2022-08-11T13:15:32.283480Z","shell.execute_reply":"2022-08-11T13:15:33.854973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('HomePlanet missing values left:',df['HomePlanet'].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:33.857845Z","iopub.execute_input":"2022-08-11T13:15:33.858922Z","iopub.status.idle":"2022-08-11T13:15:33.865691Z","shell.execute_reply.started":"2022-08-11T13:15:33.858872Z","shell.execute_reply":"2022-08-11T13:15:33.864732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's fill in the missing `HomePlanet` values by looking at `CabinDeck` feature.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12,5))\nax = sns.countplot('CabinDeck',hue='HomePlanet',data=df)\nfor p in ax.patches:\n   ax.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0.25, p.get_height()+0.01))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:33.866865Z","iopub.execute_input":"2022-08-11T13:15:33.867818Z","iopub.status.idle":"2022-08-11T13:15:34.275343Z","shell.execute_reply.started":"2022-08-11T13:15:33.867783Z","shell.execute_reply":"2022-08-11T13:15:34.273976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* All of the passengers on cabin deck A, B, C and T came from Europa. \n* All of the passengers on cabin deck G came from Europa.\n* Passengers on decks D, E or F came from multiple planets.\n\nWe can fill the other missing values according to these.","metadata":{}},{"cell_type":"code","source":"# CabinDecks A, B, C or T came from Europa\ndf.loc[(df['HomePlanet'].isna()) & (df['CabinDeck'].isin(['A', 'B', 'C', 'T'])), 'HomePlanet'] = 'Europa'\n\n# CabinDecks G came from Earth\ndf.loc[(df['HomePlanet'].isna()) & (df['CabinDeck']=='G'), 'HomePlanet'] = 'Earth'","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:34.277210Z","iopub.execute_input":"2022-08-11T13:15:34.277840Z","iopub.status.idle":"2022-08-11T13:15:34.293150Z","shell.execute_reply.started":"2022-08-11T13:15:34.277791Z","shell.execute_reply":"2022-08-11T13:15:34.291884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('#HomePlanet missing values left:',df['HomePlanet'].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:34.294615Z","iopub.execute_input":"2022-08-11T13:15:34.295646Z","iopub.status.idle":"2022-08-11T13:15:34.302778Z","shell.execute_reply.started":"2022-08-11T13:15:34.295602Z","shell.execute_reply":"2022-08-11T13:15:34.301638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's fill in the missing `HomePlanet` values by looking at `Surname` feature.","metadata":{}},{"cell_type":"code","source":"SurnameHomePlanet = df.groupby(['Surname','HomePlanet'])['HomePlanet'].size().unstack().fillna(0)\nSurnameHomePlanet.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:34.304326Z","iopub.execute_input":"2022-08-11T13:15:34.305006Z","iopub.status.idle":"2022-08-11T13:15:34.336272Z","shell.execute_reply.started":"2022-08-11T13:15:34.304964Z","shell.execute_reply":"2022-08-11T13:15:34.334908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The same surnames came from the same homeplanet, let's fill in the remaining missing values accordingly.","metadata":{}},{"cell_type":"code","source":"SurnameHomePlanetIndex = df[df['HomePlanet'].isna()][(df[df['HomePlanet'].isna()]['Surname']).isin(SurnameHomePlanet.index)].index\n\ndf.loc[SurnameHomePlanetIndex,'HomePlanet']=df.iloc[SurnameHomePlanetIndex,:]['Surname'].map(lambda x: SurnameHomePlanet.idxmax(axis=1)[x])","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:34.337642Z","iopub.execute_input":"2022-08-11T13:15:34.338392Z","iopub.status.idle":"2022-08-11T13:15:34.667021Z","shell.execute_reply.started":"2022-08-11T13:15:34.338342Z","shell.execute_reply":"2022-08-11T13:15:34.665394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('HomePlanet missing values left:',df['HomePlanet'].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:34.668484Z","iopub.execute_input":"2022-08-11T13:15:34.669017Z","iopub.status.idle":"2022-08-11T13:15:34.677378Z","shell.execute_reply.started":"2022-08-11T13:15:34.668981Z","shell.execute_reply":"2022-08-11T13:15:34.675851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's look at other 10 missing values.","metadata":{}},{"cell_type":"code","source":"df[df['HomePlanet'].isna()][['PassengerId','HomePlanet','Destination']]","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:34.679053Z","iopub.execute_input":"2022-08-11T13:15:34.680438Z","iopub.status.idle":"2022-08-11T13:15:34.696019Z","shell.execute_reply.started":"2022-08-11T13:15:34.680339Z","shell.execute_reply":"2022-08-11T13:15:34.694607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These 10 values have the same destination. We can fill in these last 10 values by looking at the `HomePlanet` of passengers with the same destination.","metadata":{}},{"cell_type":"code","source":"df.loc[(df['HomePlanet'].isna()) & ~(df['CabinDeck']=='D'), 'HomePlanet'] = 'Earth'\ndf.loc[(df['HomePlanet'].isna()) & (df['CabinDeck']=='D'), 'HomePlanet'] = 'Mars'","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:34.697720Z","iopub.execute_input":"2022-08-11T13:15:34.698439Z","iopub.status.idle":"2022-08-11T13:15:34.712987Z","shell.execute_reply.started":"2022-08-11T13:15:34.698395Z","shell.execute_reply":"2022-08-11T13:15:34.711660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('HomePlanet missing values left:',df['HomePlanet'].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:34.714858Z","iopub.execute_input":"2022-08-11T13:15:34.716003Z","iopub.status.idle":"2022-08-11T13:15:34.727001Z","shell.execute_reply.started":"2022-08-11T13:15:34.715950Z","shell.execute_reply":"2022-08-11T13:15:34.725637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.6.2 `Destination`\n\nSince the majority of passengers are goes towards TRAPPIST-1e we can fill missing values with TRAPPIST-1e intuitively.","metadata":{}},{"cell_type":"code","source":"df.loc[(df['Destination'].isna()), 'Destination'] = 'TRAPPIST-1e'","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:34.728651Z","iopub.execute_input":"2022-08-11T13:15:34.729836Z","iopub.status.idle":"2022-08-11T13:15:34.743002Z","shell.execute_reply.started":"2022-08-11T13:15:34.729799Z","shell.execute_reply":"2022-08-11T13:15:34.741458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Destination missing values left:',df['Destination'].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:34.745249Z","iopub.execute_input":"2022-08-11T13:15:34.746058Z","iopub.status.idle":"2022-08-11T13:15:34.755819Z","shell.execute_reply.started":"2022-08-11T13:15:34.746000Z","shell.execute_reply":"2022-08-11T13:15:34.754464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.6.3 `Surname`\n\nWe can fill the missing values from Surname with the group.","metadata":{}},{"cell_type":"code","source":"GroupSurname = df[df['GroupSize']>1].groupby(['Group','Surname'])['Surname'].size().unstack().fillna(0)\n\nGroupSurnameIndex=df[df['Surname'].isna()][(df[df['Surname'].isna()]['Group']).isin(GroupSurname.index)].index\n\ndf.loc[GroupSurnameIndex,'Surname']=df.iloc[GroupSurnameIndex,:]['Group'].map(lambda x: GroupSurname.idxmax(axis=1)[x])","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:34.757695Z","iopub.execute_input":"2022-08-11T13:15:34.758173Z","iopub.status.idle":"2022-08-11T13:15:38.488503Z","shell.execute_reply.started":"2022-08-11T13:15:34.758130Z","shell.execute_reply":"2022-08-11T13:15:38.487232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Surname missing values left:',df['Surname'].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:38.490205Z","iopub.execute_input":"2022-08-11T13:15:38.490699Z","iopub.status.idle":"2022-08-11T13:15:38.498946Z","shell.execute_reply.started":"2022-08-11T13:15:38.490653Z","shell.execute_reply":"2022-08-11T13:15:38.498004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can update `FamilySize` feature.","metadata":{}},{"cell_type":"code","source":"df['Surname'].fillna('Unknown', inplace=True)\n\n# Update FamilySize \ndf['FamilySize'] = df['Surname'].map(lambda x: df['Surname'].value_counts()[x])\n\n# Put NaN's back in place of outliers\ndf.loc[df['Surname'] == 'Unknown','Surname'] = np.nan\n\n# Unknown = no family\ndf.loc[df['FamilySize']>100,'FamilySize']=0","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:15:38.500090Z","iopub.execute_input":"2022-08-11T13:15:38.500864Z","iopub.status.idle":"2022-08-11T13:16:09.468836Z","shell.execute_reply.started":"2022-08-11T13:15:38.500829Z","shell.execute_reply":"2022-08-11T13:16:09.467596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.6.4 `CabinSide`","metadata":{}},{"cell_type":"code","source":"GroupCabinDeck =df[df['GroupSize']>1].groupby(['Group','CabinDeck'])['CabinDeck'].size().unstack().fillna(0)\nGroupCabinNumber =df[df['GroupSize']>1].groupby(['Group','CabinNumber'])['CabinNumber'].size().unstack().fillna(0)\nGroupCabinSide =df[df['GroupSize']>1].groupby(['Group','CabinSide'])['CabinSide'].size().unstack().fillna(0)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:09.470433Z","iopub.execute_input":"2022-08-11T13:16:09.470956Z","iopub.status.idle":"2022-08-11T13:16:09.537771Z","shell.execute_reply.started":"2022-08-11T13:16:09.470917Z","shell.execute_reply":"2022-08-11T13:16:09.536642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot((GroupCabinSide>0).sum(axis=1))\nplt.title('Unique CabinSide per group')\nfig.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:09.539209Z","iopub.execute_input":"2022-08-11T13:16:09.539539Z","iopub.status.idle":"2022-08-11T13:16:09.890885Z","shell.execute_reply.started":"2022-08-11T13:16:09.539509Z","shell.execute_reply":"2022-08-11T13:16:09.889524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that everyone in the same group is also on the same `CabinSide`.","metadata":{}},{"cell_type":"code","source":"GroupCabinSideIndex=df[df['CabinSide'].isna()][(df[df['CabinSide'].isna()]['Group']).isin(GroupCabinSide.index)].index\n\ndf.loc[GroupCabinSideIndex,'CabinSide']=df.iloc[GroupCabinSideIndex,:]['Group'].map(lambda x: GroupCabinSide.idxmax(axis=1)[x])","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:09.892877Z","iopub.execute_input":"2022-08-11T13:16:09.893322Z","iopub.status.idle":"2022-08-11T13:16:10.361624Z","shell.execute_reply.started":"2022-08-11T13:16:09.893281Z","shell.execute_reply":"2022-08-11T13:16:10.360374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('CabinSide missing values left:',df['CabinSide'].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:10.363235Z","iopub.execute_input":"2022-08-11T13:16:10.364043Z","iopub.status.idle":"2022-08-11T13:16:10.370953Z","shell.execute_reply.started":"2022-08-11T13:16:10.363998Z","shell.execute_reply":"2022-08-11T13:16:10.369843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.loc[df['CabinSide'].isna(),'CabinSide'] = 'Z' #filling with outlier\n\nprint('CabinSide missing values left:',df['CabinSide'].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:10.372459Z","iopub.execute_input":"2022-08-11T13:16:10.372893Z","iopub.status.idle":"2022-08-11T13:16:10.387454Z","shell.execute_reply.started":"2022-08-11T13:16:10.372857Z","shell.execute_reply":"2022-08-11T13:16:10.385948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.6.5 `CabinDeck`\n\nGroups tend to be on the same CabinDeck.","metadata":{}},{"cell_type":"code","source":"sns.countplot((GroupCabinDeck>0).sum(axis=1))\nplt.title('Unique CabiDeck per group')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:10.389400Z","iopub.execute_input":"2022-08-11T13:16:10.390038Z","iopub.status.idle":"2022-08-11T13:16:10.588826Z","shell.execute_reply.started":"2022-08-11T13:16:10.390002Z","shell.execute_reply":"2022-08-11T13:16:10.587585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"GroupCabinDeckIndex=df[df['CabinDeck'].isna()][(df[df['CabinDeck'].isna()]['Group']).isin(GroupCabinDeck.index)].index\n\n# Filling missing values\ndf.loc[GroupCabinDeckIndex,'CabinDeck']=df.iloc[GroupCabinDeckIndex,:]['Group'].map(lambda x: GroupCabinDeck.idxmax(axis=1)[x])","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:10.590699Z","iopub.execute_input":"2022-08-11T13:16:10.591479Z","iopub.status.idle":"2022-08-11T13:16:11.064451Z","shell.execute_reply.started":"2022-08-11T13:16:10.591423Z","shell.execute_reply":"2022-08-11T13:16:11.062960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('CabinDeck missing values left:',df['CabinDeck'].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.065992Z","iopub.execute_input":"2022-08-11T13:16:11.066364Z","iopub.status.idle":"2022-08-11T13:16:11.074232Z","shell.execute_reply.started":"2022-08-11T13:16:11.066330Z","shell.execute_reply":"2022-08-11T13:16:11.072844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's look at the relationship between `CabinDeck` and `HomePlanet` .","metadata":{}},{"cell_type":"code","source":"df.groupby(['HomePlanet','Destination','CabinDeck'])['CabinDeck'].size().unstack().fillna(0)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.075870Z","iopub.execute_input":"2022-08-11T13:16:11.076341Z","iopub.status.idle":"2022-08-11T13:16:11.112290Z","shell.execute_reply.started":"2022-08-11T13:16:11.076293Z","shell.execute_reply":"2022-08-11T13:16:11.110974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Passengers from Mars are most likely in deck F.\n* Passengers from Europa are most likely in deck C.\n* Passengers from Earth are most likely in deck G.","metadata":{}},{"cell_type":"code","source":"NACabinDeck = df.loc[df['CabinDeck'].isna(),'CabinDeck'].index\ndf.loc[df['CabinDeck'].isna(),'CabinDeck']=df.groupby(['HomePlanet','Destination'])['CabinDeck'].transform(lambda x: x.fillna(pd.Series.mode(x)[0]))[NACabinDeck]","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.113927Z","iopub.execute_input":"2022-08-11T13:16:11.114407Z","iopub.status.idle":"2022-08-11T13:16:11.142812Z","shell.execute_reply.started":"2022-08-11T13:16:11.114371Z","shell.execute_reply":"2022-08-11T13:16:11.141211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('CabinDeck missing values left:',df['CabinDeck'].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.144476Z","iopub.execute_input":"2022-08-11T13:16:11.144906Z","iopub.status.idle":"2022-08-11T13:16:11.152554Z","shell.execute_reply.started":"2022-08-11T13:16:11.144872Z","shell.execute_reply":"2022-08-11T13:16:11.151192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.6.6 `CabinNumber`","metadata":{}},{"cell_type":"code","source":"corr = df.corr()\ncorr['CabinNumber'].sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.154151Z","iopub.execute_input":"2022-08-11T13:16:11.154834Z","iopub.status.idle":"2022-08-11T13:16:11.181849Z","shell.execute_reply.started":"2022-08-11T13:16:11.154781Z","shell.execute_reply":"2022-08-11T13:16:11.180666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['CabinNumber'] = df['CabinNumber'].fillna(0)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.183440Z","iopub.execute_input":"2022-08-11T13:16:11.183812Z","iopub.status.idle":"2022-08-11T13:16:11.189822Z","shell.execute_reply.started":"2022-08-11T13:16:11.183780Z","shell.execute_reply":"2022-08-11T13:16:11.188617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('CabinNumber missing values left:',df['CabinNumber'].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.191628Z","iopub.execute_input":"2022-08-11T13:16:11.192019Z","iopub.status.idle":"2022-08-11T13:16:11.204047Z","shell.execute_reply.started":"2022-08-11T13:16:11.191982Z","shell.execute_reply":"2022-08-11T13:16:11.202646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's update the cabin regions with the new data.","metadata":{}},{"cell_type":"code","source":"df['CabinRegion1']=(df['CabinNumber']<300).astype(int)  \ndf['CabinRegion2']=((df['CabinNumber']>=300) & (df['CabinNumber']<600)).astype(int)\ndf['CabinRegion3']=((df['CabinNumber']>=600) & (df['CabinNumber']<1200)).astype(int)\ndf['CabinRegion4']=(df['CabinNumber']>=1200).astype(int)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.205395Z","iopub.execute_input":"2022-08-11T13:16:11.206510Z","iopub.status.idle":"2022-08-11T13:16:11.216541Z","shell.execute_reply.started":"2022-08-11T13:16:11.206464Z","shell.execute_reply":"2022-08-11T13:16:11.215454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.6.7 `VIP`","metadata":{}},{"cell_type":"code","source":"df.loc[df['VIP'].isna(),'VIP']=False\n\nprint('VIP missing values left:',df['VIP'].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.217984Z","iopub.execute_input":"2022-08-11T13:16:11.219002Z","iopub.status.idle":"2022-08-11T13:16:11.230430Z","shell.execute_reply.started":"2022-08-11T13:16:11.218966Z","shell.execute_reply":"2022-08-11T13:16:11.228933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.6.8 `Age`\n\nWe will fill the value with median.","metadata":{}},{"cell_type":"code","source":"NAAge = df.loc[df['Age'].isna(),'Age'].index\ndf.loc[df['Age'].isna(),'Age']=df.groupby(['HomePlanet','TotalSpending','IsAlone','CabinDeck'])['Age'].transform(lambda x: x.fillna(x.median()))[NAAge]\n\nprint('Age missing values left:',df['Age'].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.231964Z","iopub.execute_input":"2022-08-11T13:16:11.232855Z","iopub.status.idle":"2022-08-11T13:16:11.281059Z","shell.execute_reply.started":"2022-08-11T13:16:11.232798Z","shell.execute_reply":"2022-08-11T13:16:11.279671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's update the `AgeGroup` values.","metadata":{}},{"cell_type":"code","source":"df.loc[df['Age']<=14,'AgeGroup']='Children'\ndf.loc[(df['Age']>14) & (df['Age']<24),'AgeGroup']='Youth'\ndf.loc[(df['Age']>=24) & (df['Age']<=64),'AgeGroup']='Adult'\ndf.loc[df['Age']>64,'AgeGroup']='Seniors'","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.285873Z","iopub.execute_input":"2022-08-11T13:16:11.286294Z","iopub.status.idle":"2022-08-11T13:16:11.300047Z","shell.execute_reply.started":"2022-08-11T13:16:11.286260Z","shell.execute_reply":"2022-08-11T13:16:11.298627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.6.9 `CryoSleep`\n\nWe can look at the relationship between `CryoSleep` and `TotalSpending` .","metadata":{}},{"cell_type":"code","source":"df.groupby(['TotalSpending','CryoSleep'])['CryoSleep'].size().unstack().fillna(0)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.301797Z","iopub.execute_input":"2022-08-11T13:16:11.302705Z","iopub.status.idle":"2022-08-11T13:16:11.319595Z","shell.execute_reply.started":"2022-08-11T13:16:11.302664Z","shell.execute_reply":"2022-08-11T13:16:11.318620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"NA_CryoSleep = df.loc[df['CryoSleep'].isna(),'CryoSleep'].index\ndf.loc[df['CryoSleep'].isna(),'CryoSleep']= df.groupby(['TotalSpending'])['CryoSleep'].transform(lambda x: x.fillna(pd.Series.mode(x)[0]))[NA_CryoSleep]","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.330742Z","iopub.execute_input":"2022-08-11T13:16:11.331479Z","iopub.status.idle":"2022-08-11T13:16:11.349442Z","shell.execute_reply.started":"2022-08-11T13:16:11.331440Z","shell.execute_reply":"2022-08-11T13:16:11.347924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('CryoSleep missing values left:',df['CryoSleep'].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.351030Z","iopub.execute_input":"2022-08-11T13:16:11.351416Z","iopub.status.idle":"2022-08-11T13:16:11.358320Z","shell.execute_reply.started":"2022-08-11T13:16:11.351379Z","shell.execute_reply":"2022-08-11T13:16:11.357474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4.6.10 Expenditure","metadata":{}},{"cell_type":"code","source":"for col in num_features:\n    df.loc[(df[col].isna()) & (df['CryoSleep']==True), col]=0\n\nprint('Expenditure missing values left:',df[num_features].isna().sum().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.359479Z","iopub.execute_input":"2022-08-11T13:16:11.360468Z","iopub.status.idle":"2022-08-11T13:16:11.385419Z","shell.execute_reply.started":"2022-08-11T13:16:11.360431Z","shell.execute_reply":"2022-08-11T13:16:11.383826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[num_features] = df[num_features] .fillna(0)\nprint('Expenditure missing values left:',df[num_features].isna().sum().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.387588Z","iopub.execute_input":"2022-08-11T13:16:11.388419Z","iopub.status.idle":"2022-08-11T13:16:11.401560Z","shell.execute_reply.started":"2022-08-11T13:16:11.388369Z","shell.execute_reply":"2022-08-11T13:16:11.400423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Data pre-processing <a id=\"5\"></a>","metadata":{}},{"cell_type":"code","source":"X = df[df['PassengerId'].isin(df_train['PassengerId'].values)].copy()\nX_test = df[df['PassengerId'].isin(df_test['PassengerId'].values)].copy()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.402975Z","iopub.execute_input":"2022-08-11T13:16:11.403498Z","iopub.status.idle":"2022-08-11T13:16:11.418313Z","shell.execute_reply.started":"2022-08-11T13:16:11.403467Z","shell.execute_reply":"2022-08-11T13:16:11.416671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.1 Dropping columns\nLet's drop some unuseful columns.","metadata":{}},{"cell_type":"code","source":"X.drop(['PassengerId', 'Group', 'GroupSize', 'AgeGroup', 'CabinNumber'], axis=1, inplace=True)\nX_test.drop(['PassengerId', 'Group', 'GroupSize', 'AgeGroup', 'CabinNumber'], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.419852Z","iopub.execute_input":"2022-08-11T13:16:11.420750Z","iopub.status.idle":"2022-08-11T13:16:11.430920Z","shell.execute_reply.started":"2022-08-11T13:16:11.420707Z","shell.execute_reply":"2022-08-11T13:16:11.429518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.2 Log transform","metadata":{}},{"cell_type":"code","source":"for col in ['RoomService','FoodCourt','ShoppingMall','Spa','VRDeck','Expenditure']:\n    X[col]=np.log(1+X[col])\n    X_test[col]=np.log(1+X_test[col])","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.432473Z","iopub.execute_input":"2022-08-11T13:16:11.433095Z","iopub.status.idle":"2022-08-11T13:16:11.448445Z","shell.execute_reply.started":"2022-08-11T13:16:11.433040Z","shell.execute_reply":"2022-08-11T13:16:11.447003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.3 Encoding and scaling\n\nEncoding with `OneHotEncoder()` Scaling with `StandartScaler()`","metadata":{}},{"cell_type":"code","source":"# Indentify numerical and categorical columns\nnum_cols = X.select_dtypes(include=['float64','int64']).columns.tolist()\ncat_cols = X.select_dtypes(include=['object']).columns.tolist()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.449975Z","iopub.execute_input":"2022-08-11T13:16:11.450876Z","iopub.status.idle":"2022-08-11T13:16:11.461178Z","shell.execute_reply.started":"2022-08-11T13:16:11.450824Z","shell.execute_reply":"2022-08-11T13:16:11.459996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numerical_transformer = Pipeline(steps=[('scaler', StandardScaler())])\n\ncategorical_transformer = Pipeline(steps=[('onehot', OneHotEncoder(drop='if_binary', handle_unknown='ignore',sparse=False))])\n\n# Combine preprocessing\nct = ColumnTransformer(\n        transformers=[\n        ('num', numerical_transformer, num_cols),\n        ('cat', categorical_transformer, cat_cols)],\n        remainder='passthrough')\n\n# Apply preprocessing\nX = ct.fit_transform(X)\nX_test = ct.transform(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.462851Z","iopub.execute_input":"2022-08-11T13:16:11.463489Z","iopub.status.idle":"2022-08-11T13:16:11.883243Z","shell.execute_reply.started":"2022-08-11T13:16:11.463452Z","shell.execute_reply":"2022-08-11T13:16:11.882244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.4 Validation split","metadata":{}},{"cell_type":"code","source":"X_train, X_val, y_train, y_val = train_test_split(X,y,stratify=y,test_size=0.2,random_state=0)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.884798Z","iopub.execute_input":"2022-08-11T13:16:11.885428Z","iopub.status.idle":"2022-08-11T13:16:11.970256Z","shell.execute_reply.started":"2022-08-11T13:16:11.885384Z","shell.execute_reply":"2022-08-11T13:16:11.968908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Models <a id=\"6\"></a>\n\n## 6.1 Classifiers","metadata":{}},{"cell_type":"code","source":"from xgboost import XGBClassifier\n\n\nlgbm_params = {'num_leaves':15,'objective':'binary','learning_rate': 0.05, 'max_depth': 8, 'n_estimators': 100}\n\nlgbm = LGBMClassifier(**lgbm_params,verbose=0)\n\ncatb_params = {'learning_rate': 0.17, 'max_depth': 4, 'n_estimators': 100}\n\ncatboost = CatBoostClassifier(**catb_params)\n\nxgbc = XGBClassifier(gamma = 1.5,\n                           subsample = 1.0,\n                           max_depth = 5,\n                           colsample_bytree = 1.0,\n                           n_estimators = 100)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.971831Z","iopub.execute_input":"2022-08-11T13:16:11.972231Z","iopub.status.idle":"2022-08-11T13:16:11.985300Z","shell.execute_reply.started":"2022-08-11T13:16:11.972194Z","shell.execute_reply":"2022-08-11T13:16:11.984063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the list classifiers\nclassifiers = [\n    (\"LGBM\" , lgbm),\n    (\"CatBoost\" , catboost),\n    (\"XGBoost\" , xgbc) ]","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:11.986872Z","iopub.execute_input":"2022-08-11T13:16:11.987618Z","iopub.status.idle":"2022-08-11T13:16:11.996633Z","shell.execute_reply.started":"2022-08-11T13:16:11.987547Z","shell.execute_reply":"2022-08-11T13:16:11.995409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import accuracy_score\nfrom sklearn import metrics\n\n# Iterate over the pre-defined list of classifiers\nfor clf_name, clf in classifiers:\n    # Fit clf to the training set\n    clf.fit(X_train, y_train)\n    \n    # Predict y_pred\n    y_pred = clf.predict(X_val)\n    acc = accuracy_score(y_val, y_pred)\n\n    # Evaluate clf's accuracy on the test set\n    print('{:s} score : {:.3f}'.format(clf_name, acc))","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:12.000928Z","iopub.execute_input":"2022-08-11T13:16:12.001378Z","iopub.status.idle":"2022-08-11T13:16:44.199159Z","shell.execute_reply.started":"2022-08-11T13:16:12.001342Z","shell.execute_reply":"2022-08-11T13:16:44.198210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's do Catboost training and predict the labels!","metadata":{}},{"cell_type":"markdown","source":"## 6.2 Catboost Classifier\n\nCatBoost is a recently open-sourced machine learning algorithm from Yandex. It can easily integrate with deep learning frameworks like Google’s TensorFlow and Apple’s Core ML. It can work with diverse data types to help solve a wide range of problems that businesses face today. To top it up, it provides best-in-class accuracy.\n\nIt is especially powerful in two ways:\n\n* It yields state-of-the-art results without extensive data training typically required by other machine learning methods, and\n* Provides powerful out-of-the-box support for the more descriptive data formats that accompany many business problems. [X](https://www.analyticsvidhya.com/blog/2017/08/catboost-automated-categorical-data/)","metadata":{}},{"cell_type":"code","source":"X_train, X_val, y_train, y_val = train_test_split(X,y,test_size=0.3,random_state=0)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:44.200479Z","iopub.execute_input":"2022-08-11T13:16:44.201179Z","iopub.status.idle":"2022-08-11T13:16:44.263629Z","shell.execute_reply.started":"2022-08-11T13:16:44.201142Z","shell.execute_reply":"2022-08-11T13:16:44.262451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"catboost.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:44.265336Z","iopub.execute_input":"2022-08-11T13:16:44.265751Z","iopub.status.idle":"2022-08-11T13:16:49.767058Z","shell.execute_reply.started":"2022-08-11T13:16:44.265713Z","shell.execute_reply":"2022-08-11T13:16:49.764730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_confusion_matrix(catboost, X_val, y_val)  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:16:49.769172Z","iopub.execute_input":"2022-08-11T13:16:49.769676Z","iopub.status.idle":"2022-08-11T13:16:52.073723Z","shell.execute_reply.started":"2022-08-11T13:16:49.769628Z","shell.execute_reply":"2022-08-11T13:16:52.072624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 7. Submission <a id=\"7\"></a>","metadata":{}},{"cell_type":"code","source":"y_pred = catboost.predict(X_test)\ny_pred = (y_pred > 0.5).astype(int)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:17:10.454342Z","iopub.execute_input":"2022-08-11T13:17:10.454909Z","iopub.status.idle":"2022-08-11T13:17:14.250128Z","shell.execute_reply.started":"2022-08-11T13:17:10.454874Z","shell.execute_reply":"2022-08-11T13:17:14.248714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub = pd.read_csv('/kaggle/input/spaceship-titanic/sample_submission.csv')\nsub['Transported'] = list(map(int, y_pred))\nsub['Transported'] = sub['Transported'].replace({0.0:'False',1.0:'True'})\nsub.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:17:16.366340Z","iopub.execute_input":"2022-08-11T13:17:16.366821Z","iopub.status.idle":"2022-08-11T13:17:16.393924Z","shell.execute_reply.started":"2022-08-11T13:17:16.366783Z","shell.execute_reply":"2022-08-11T13:17:16.392845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thank you for reading my notebook. \n* I hope you find it useful!\n* Feel free to contact me for any feedback good or bad 😌 🚀","metadata":{}}]}