{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style='font-size= 20pt;'><center>Titanic SpaceShip</center></h1>\n<img src=\"https://img3.goodfon.com/wallpaper/nbig/3/78/art-sci-fi-spaceship.jpg\">\n\n### <span style=\"color:#F28000;\">*This is a work in progress, so I'll update this and try to improve the score!*<span>\n<a id=\"sections\"></a>\n- [1. Load the Data](#1)\n    - [1.1. Import Libararies](#1.1)\n    - [1.2. Load Datasets](#1.2)\n- [2. Summarize the Dataset](#2)\n    - [2.1. Dimensions of the Dataset](#2.1)\n    - [2.2. Peek at the Data](#2.2)\n    - [2.3. Statistical Summary](#2.3)\n-[3. Data Visualization](#3)\n    - [3.1. Univariate Plots for Numerical Features](#3.1)\n    - [3.2. Univariate Plots for Categorical Features](#3.2)\n    - [3.3. Multivariate Plots for Numerical Features](#3.3)\n    - [3.4. Multivariate Plots for Categorical Features](#3.4)\n    \n- [4. Handling Missing Values and Outliers](#4)\n    - [4.1 Outliers](#4.1)\n    - [4.2 Missing and Duplicates](#4.2)\n- [5. Exploring Features](#5)\n    - [5.1 Names](#5.1)\n    - [5.2 Cabin](#5.2)\n    - [5.3 PassengerId](#5.3)\n    - [5.4 Expenses](#5.4)\n    - [5.5 Age](#5.5)\n    - [5.6 Homeplanet-Destination](#5.6)\n    - [5.7 Chi2 test](#5.7)\n    - [5.8 Dropping Columns and Defining The Target Column](#5.8)\n- [6. PipeLine Implementation](#6)\n    - [6.1 Models and Scoring](#6.1)\n- [7. Learning Curve](#7)\n- [8. Model Tuning using GridSearchCV](#8)\n    - [8.1 Gradient Boost](#8.1)\n    - [8.2 LGBM](#8.2)\n    - [8.3 XGBClassifier](#8.3)\n    - [8.4 CatBoost](#8.4)\n    \n- [9. Transforming and Predicting Test Set](#9)\n- [10. Submission](#10)\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"1\"></a>\n## **<span style=\"color:#F26300;\">1. Load The Data</span>**\n\n<a id='1.1'></a>\n## **<span style=\"color:#F26300;\">1.1 Import Libraries</span>**","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\nfrom pandas import read_csv\nfrom pandas.plotting import scatter_matrix\nfrom matplotlib import pyplot as plt\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import KFold\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.metrics import classification_report\nfrom sklearn.metrics import confusion_matrix\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.discriminant_analysis import LinearDiscriminantAnalysis\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.svm import SVC\nimport seaborn as sns\nimport missingno as mno","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:05.629972Z","iopub.execute_input":"2022-07-16T11:45:05.630354Z","iopub.status.idle":"2022-07-16T11:45:05.649058Z","shell.execute_reply.started":"2022-07-16T11:45:05.630328Z","shell.execute_reply":"2022-07-16T11:45:05.648104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id='1.2'></a>\n## **<span style=\"color:#F26300;\">1.2 Load Datasets</span>**","metadata":{}},{"cell_type":"code","source":"train=pd.read_csv(\"/kaggle/input/spaceship-titanic/train.csv\")\ntest=pd.read_csv(\"/kaggle/input/spaceship-titanic/test.csv\")\nsample=pd.read_csv(\"/kaggle/input/spaceship-titanic/sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:05.839027Z","iopub.execute_input":"2022-07-16T11:45:05.839626Z","iopub.status.idle":"2022-07-16T11:45:05.879161Z","shell.execute_reply.started":"2022-07-16T11:45:05.839592Z","shell.execute_reply":"2022-07-16T11:45:05.878428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_set=train.copy()\ntest_set=test.copy()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:06.021056Z","iopub.execute_input":"2022-07-16T11:45:06.022296Z","iopub.status.idle":"2022-07-16T11:45:06.027461Z","shell.execute_reply.started":"2022-07-16T11:45:06.022246Z","shell.execute_reply":"2022-07-16T11:45:06.026675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id='2'></a>\n## **<span style=\"color:#F26300;\">2. Summarize the Dataset</span>**\n<a id='2.1'></a>\n## **<span style=\"color:#F26300;\">2.1 Dimensions of the Dataset</span>**\n","metadata":{}},{"cell_type":"code","source":"train_set.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:06.383419Z","iopub.execute_input":"2022-07-16T11:45:06.384005Z","iopub.status.idle":"2022-07-16T11:45:06.391083Z","shell.execute_reply.started":"2022-07-16T11:45:06.383977Z","shell.execute_reply":"2022-07-16T11:45:06.390402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_set.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:06.664317Z","iopub.execute_input":"2022-07-16T11:45:06.664853Z","iopub.status.idle":"2022-07-16T11:45:06.673132Z","shell.execute_reply.started":"2022-07-16T11:45:06.664807Z","shell.execute_reply":"2022-07-16T11:45:06.672119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id='2.2'></a>\n## **<span style=\"color:#F26300;\">2.2 Peek at the Data</span>**","metadata":{}},{"cell_type":"code","source":"train_set.head(10) ","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:07.119116Z","iopub.execute_input":"2022-07-16T11:45:07.119483Z","iopub.status.idle":"2022-07-16T11:45:07.141765Z","shell.execute_reply.started":"2022-07-16T11:45:07.119458Z","shell.execute_reply":"2022-07-16T11:45:07.140763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_set.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:07.305732Z","iopub.execute_input":"2022-07-16T11:45:07.306534Z","iopub.status.idle":"2022-07-16T11:45:07.322099Z","shell.execute_reply.started":"2022-07-16T11:45:07.306506Z","shell.execute_reply":"2022-07-16T11:45:07.321341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id='2.3'></a>\n## **<span style=\"color:#F26300;\">2.3 Statistical Summary</span>**\n","metadata":{}},{"cell_type":"code","source":"train_set.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:07.602821Z","iopub.execute_input":"2022-07-16T11:45:07.603150Z","iopub.status.idle":"2022-07-16T11:45:07.631611Z","shell.execute_reply.started":"2022-07-16T11:45:07.603126Z","shell.execute_reply":"2022-07-16T11:45:07.630954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that:\n* the numeric features are on different scales\n* there are some missing values","metadata":{}},{"cell_type":"code","source":"train_set.Transported.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:07.816854Z","iopub.execute_input":"2022-07-16T11:45:07.817646Z","iopub.status.idle":"2022-07-16T11:45:07.824695Z","shell.execute_reply.started":"2022-07-16T11:45:07.817618Z","shell.execute_reply":"2022-07-16T11:45:07.824010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we can see the distribution is almost equal","metadata":{}},{"cell_type":"markdown","source":"[List of content](#sections)\n<a id=\"3\"></a>\n## **<span style=\"color:#F26300;\">3. Data Visualization</span>**","metadata":{}},{"cell_type":"markdown","source":"We are going to look at two types of plots:\n*  Univariate plots to better understand each attribute.\n*  Multivariate plots to better understand the relationships between attributes.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"3.1\"></a>\n## **<span style=\"color:#F26300;\">3.1 Univariate Plots for Numerical Features</span>**","metadata":{}},{"cell_type":"code","source":"#this will create a box plot for numeric values\ntrain_set.plot(kind='box',subplots=True,layout=(2,3),sharex=False,sharey=False,figsize=(15,15))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:07.997927Z","iopub.execute_input":"2022-07-16T11:45:07.999109Z","iopub.status.idle":"2022-07-16T11:45:08.637044Z","shell.execute_reply.started":"2022-07-16T11:45:07.999057Z","shell.execute_reply":"2022-07-16T11:45:08.635853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#this will create a historgram for numeric values\ntrain_set.hist(figsize=(10,10))","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:08.638901Z","iopub.execute_input":"2022-07-16T11:45:08.639410Z","iopub.status.idle":"2022-07-16T11:45:09.568848Z","shell.execute_reply.started":"2022-07-16T11:45:08.639383Z","shell.execute_reply":"2022-07-16T11:45:09.567621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Only 'Age' has a Gaussian-like distribution, the rest are highly skewed!","metadata":{}},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"3.2\"></a>\n## **<span style=\"color:#F26300;\">3.2 Univariate Plots for Categorical Features</span>**","metadata":{}},{"cell_type":"code","source":"sns.set(rc={'figure.figsize':(10,10)})\nfig, axes = plt.subplots(2, 2)\nnames=['HomePlanet','CryoSleep','Destination','VIP']\n\nfor name, ax in zip(names, axes.flatten()):\n    sns.countplot(x=name,data=train_set,ax=ax)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:09.570077Z","iopub.execute_input":"2022-07-16T11:45:09.570372Z","iopub.status.idle":"2022-07-16T11:45:10.053762Z","shell.execute_reply.started":"2022-07-16T11:45:09.570347Z","shell.execute_reply":"2022-07-16T11:45:10.052898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"3.3\"></a>\n## **<span style=\"color:#F26300;\">3.3 Multivariate Plots for Numerical Features</span>**","metadata":{}},{"cell_type":"code","source":"#sns.heatmap(data=train_set.corr(),annot=True) \n#I won't create this heatmap because the default method for .corr() is pearson and it assums Gaussian Distribution and Outliers can heavily influence the outcomes\n#but the distribution for numerical features here are highly skewed and this heatmap can not give reliable answers!","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:10.055865Z","iopub.execute_input":"2022-07-16T11:45:10.056094Z","iopub.status.idle":"2022-07-16T11:45:10.059418Z","shell.execute_reply.started":"2022-07-16T11:45:10.056072Z","shell.execute_reply":"2022-07-16T11:45:10.058780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.pairplot(data=train_set.select_dtypes(['number','bool']),hue=\"Transported\",palette='CMRmap',kind='scatter')","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:10.060446Z","iopub.execute_input":"2022-07-16T11:45:10.060858Z","iopub.status.idle":"2022-07-16T11:45:30.421617Z","shell.execute_reply.started":"2022-07-16T11:45:10.060835Z","shell.execute_reply":"2022-07-16T11:45:30.420442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Apparantly there's not linear correlation between the numerical columns.\n\nLet's take a closer look at the diagonal plots...","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(2, 3,figsize=(15,15))\nnames=['Age','RoomService','FoodCourt','ShoppingMall','Spa','VRDeck']\n\nfor name, ax in zip(names, axes.flatten()):\n    sns.kdeplot(x=name,hue='Transported',data=train_set,ax=ax)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:30.423050Z","iopub.execute_input":"2022-07-16T11:45:30.423341Z","iopub.status.idle":"2022-07-16T11:45:32.749477Z","shell.execute_reply.started":"2022-07-16T11:45:30.423314Z","shell.execute_reply":"2022-07-16T11:45:32.748030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(2, 3,figsize=(18,15))\nnames=['Age','RoomService','FoodCourt','ShoppingMall','Spa','VRDeck']\n\nfor name, ax in zip(names, axes.flatten()):\n    sns.stripplot(y=name,x='Transported',data=train_set,ax=ax)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:32.751360Z","iopub.execute_input":"2022-07-16T11:45:32.751738Z","iopub.status.idle":"2022-07-16T11:45:33.671455Z","shell.execute_reply.started":"2022-07-16T11:45:32.751702Z","shell.execute_reply":"2022-07-16T11:45:33.670243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the last two plots, we can see:\n\n1-**Age:** \n* Children (up to 10-12) have a higher chance of being Transported\n* Adults (20-40) have a higher chance of **not** being Transported\n* Rest have almost an equal chance of being Transported\n\n2-**RoomService:**\n* People who spent no money on RoomService have a higher chance of being Transported; actually, as the amount of expenditure goes higher, fewer are among the Transported people and for the extreme expenditures, there are no Transported people.\n\n3-**FoodCourt:**\n* People who spent nothing on FoodCourt are less likely to be Transported, as the amount of expenditure goes higher, the chance of being Transported is almost equal, till we reach about 17000, from there on, we have a few people; but all of them are Transported.\n\n4-**ShoppingMall:**\n* There seems to be little to no distinction between Transported and not Transported people, based on expenditure in the ShoppingMall, this feature doesn't seem to be very informative\n\n5-**Spa:**\n* People who spent little to no money on Spa, have a higher chance of being Transported\n\n6-**VRDeck:**\n* Same as Spa","metadata":{}},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"3.4\"></a>\n## **<span style=\"color:#F26300;\">3.4 Multivariate Plots for Categorical Features</span>**","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(2, 2)\nnames=['HomePlanet','CryoSleep','Destination','VIP']\n\nfor name, ax in zip(names, axes.flatten()):\n    sns.barplot(x=name,y='Transported',data=train_set,ax=ax)\n    ax.set( ylabel=\"Transportation Probability\")","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:33.673140Z","iopub.execute_input":"2022-07-16T11:45:33.674240Z","iopub.status.idle":"2022-07-16T11:45:35.454066Z","shell.execute_reply.started":"2022-07-16T11:45:33.674207Z","shell.execute_reply":"2022-07-16T11:45:35.453171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'll explore the relevance of each categorical feature to the target, using chi2 test later, but what we can tell so far is:\n\n1- People in VIP section, have less chance of being Transported\n\n2- People in CryoSleep have much higher chance of survival\n\nI think this is because, VIP people were awake and probably scattered in different parts of the spaceship; shopping or eating or whatever, so they were less safe in case of collision, whereas people in cryosleep, they probably were in a strong container, which would keep them safe.\n\n3-For the destination and homeplanet, we can see how the Transportation Probability changes among them, in the plot.","metadata":{}},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"4\"></a>\n## **<span style=\"color:#F26300;\">4. Handling Missing Values and Outliers</span>**\n<a id=\"4.1\"></a>\n## **<span style=\"color:#F26300;\">4.1 Outliers</span>**","metadata":{}},{"cell_type":"markdown","source":"Using below function, I'll detect rows with n number of outliers, I'll drop rows with more than 5 outliers","metadata":{}},{"cell_type":"code","source":"from collections import Counter\ndef outlier_detect(df,n,cols):\n    rows,to_drop=[],[]\n    for col in cols:\n        Q1=np.nanpercentile(df[col],25)\n        Q3=np.nanpercentile(df[col],75)\n        IQR=Q3-Q1\n        outlier_point=1.5*IQR\n        rows.extend(df[(df[col]<Q1-outlier_point)|(df[col]>Q3+outlier_point)].index)\n    for r,c in Counter(rows).items():\n        if c>=n: to_drop.append(r)\n    return to_drop","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:35.455387Z","iopub.execute_input":"2022-07-16T11:45:35.456636Z","iopub.status.idle":"2022-07-16T11:45:35.464701Z","shell.execute_reply.started":"2022-07-16T11:45:35.456596Z","shell.execute_reply":"2022-07-16T11:45:35.462918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"to_drop=outlier_detect(train_set,5,train_set.select_dtypes('float').columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:35.469342Z","iopub.execute_input":"2022-07-16T11:45:35.469770Z","iopub.status.idle":"2022-07-16T11:45:35.489483Z","shell.execute_reply.started":"2022-07-16T11:45:35.469744Z","shell.execute_reply":"2022-07-16T11:45:35.488578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_set.drop(to_drop,inplace=True,axis=0)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:35.490979Z","iopub.execute_input":"2022-07-16T11:45:35.491718Z","iopub.status.idle":"2022-07-16T11:45:35.497941Z","shell.execute_reply.started":"2022-07-16T11:45:35.491692Z","shell.execute_reply":"2022-07-16T11:45:35.496831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_set.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:35.499555Z","iopub.execute_input":"2022-07-16T11:45:35.500136Z","iopub.status.idle":"2022-07-16T11:45:35.510239Z","shell.execute_reply.started":"2022-07-16T11:45:35.500111Z","shell.execute_reply":"2022-07-16T11:45:35.509300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"4.2\"></a>\n## **<span style=\"color:#F26300;\">4.2 Missing and Duplicates</span>**","metadata":{}},{"cell_type":"code","source":"#there are no rows with all null values\ntrain_set.isna().all(axis=1).unique()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:35.511633Z","iopub.execute_input":"2022-07-16T11:45:35.512878Z","iopub.status.idle":"2022-07-16T11:45:35.525444Z","shell.execute_reply.started":"2022-07-16T11:45:35.512809Z","shell.execute_reply":"2022-07-16T11:45:35.524767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#there are no duplicates\ntrain_set.duplicated().any()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:35.526354Z","iopub.execute_input":"2022-07-16T11:45:35.527120Z","iopub.status.idle":"2022-07-16T11:45:35.545375Z","shell.execute_reply.started":"2022-07-16T11:45:35.527095Z","shell.execute_reply":"2022-07-16T11:45:35.544315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_prcnt=[train_set[col].isna().sum()/train_set.shape[0] *100 for col in train_set.columns]","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:35.546299Z","iopub.execute_input":"2022-07-16T11:45:35.546543Z","iopub.status.idle":"2022-07-16T11:45:35.557547Z","shell.execute_reply.started":"2022-07-16T11:45:35.546521Z","shell.execute_reply":"2022-07-16T11:45:35.556625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"miss_tbl=pd.DataFrame(missing_prcnt,columns=['%missing'],index=train_set.columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:35.559005Z","iopub.execute_input":"2022-07-16T11:45:35.559552Z","iopub.status.idle":"2022-07-16T11:45:35.568053Z","shell.execute_reply.started":"2022-07-16T11:45:35.559518Z","shell.execute_reply":"2022-07-16T11:45:35.567114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"miss_tbl","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:35.569252Z","iopub.execute_input":"2022-07-16T11:45:35.569515Z","iopub.status.idle":"2022-07-16T11:45:35.589860Z","shell.execute_reply.started":"2022-07-16T11:45:35.569492Z","shell.execute_reply":"2022-07-16T11:45:35.588716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's see if there's a pattern to missing values","metadata":{}},{"cell_type":"code","source":"mno.matrix(train_set, figsize = (20, 6))","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:35.591094Z","iopub.execute_input":"2022-07-16T11:45:35.592418Z","iopub.status.idle":"2022-07-16T11:45:35.943300Z","shell.execute_reply.started":"2022-07-16T11:45:35.592371Z","shell.execute_reply":"2022-07-16T11:45:35.942550Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"They seem pretty random...","metadata":{}},{"cell_type":"markdown","source":"We'll get back to the missing values and impute them.","metadata":{}},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"5\"></a>\n## **<span style=\"color:#F26300;\">5. Exploring Features</span>**\n<a id=\"5.1\"></a>\n## **<span style=\"color:#F26300;\">5.1 Names</span>**","metadata":{}},{"cell_type":"markdown","source":"I noticed people with same last names, I'm going to extract the last names and see if I can extract a feature named: \"Family\" from them!","metadata":{}},{"cell_type":"code","source":"train_set[['Name','Last']] = train_set.Name.str.split(\" \", expand=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:35.944364Z","iopub.execute_input":"2022-07-16T11:45:35.945069Z","iopub.status.idle":"2022-07-16T11:45:35.965364Z","shell.execute_reply.started":"2022-07-16T11:45:35.945039Z","shell.execute_reply":"2022-07-16T11:45:35.963666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_set.drop('Name',inplace=True,axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:35.966846Z","iopub.execute_input":"2022-07-16T11:45:35.967523Z","iopub.status.idle":"2022-07-16T11:45:35.976866Z","shell.execute_reply.started":"2022-07-16T11:45:35.967494Z","shell.execute_reply":"2022-07-16T11:45:35.975612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_set['Last'].ffill(axis=0,inplace=True)\n#I still haven't found a way to make use of the name column to improve the score, it needs more work!","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:35.978447Z","iopub.execute_input":"2022-07-16T11:45:35.979398Z","iopub.status.idle":"2022-07-16T11:45:35.988763Z","shell.execute_reply.started":"2022-07-16T11:45:35.979354Z","shell.execute_reply":"2022-07-16T11:45:35.987776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"5.2\"></a>\n## **<span style=\"color:#F26300;\">5.2 Cabin</span>**\n\nBased on competition's data explanations:  The cabin number where the passenger is staying. Takes the form deck/num/side, where side can be either P for Port or S for Starboard.\n\nI'll break cabin into 3 features:\n\n**CabinCap**: cabin's capacity based on the number each cabin name is repeated\n\n**CabinDeck**\n\n**CabinSide**","metadata":{}},{"cell_type":"code","source":"#This shows how many people were in each cabin\ncabin_cap= train_set.Cabin.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:35.990410Z","iopub.execute_input":"2022-07-16T11:45:35.990747Z","iopub.status.idle":"2022-07-16T11:45:36.000026Z","shell.execute_reply.started":"2022-07-16T11:45:35.990722Z","shell.execute_reply":"2022-07-16T11:45:35.999062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'm going to use the number of people in each cabin and replace the cabin names with number of people in it.","metadata":{}},{"cell_type":"code","source":"train_set['CabinCap']=train_set['Cabin'].map(cabin_cap)\ntrain_set['CabinCap']=train_set['CabinCap'].astype('object')","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:36.001215Z","iopub.execute_input":"2022-07-16T11:45:36.001882Z","iopub.status.idle":"2022-07-16T11:45:36.011799Z","shell.execute_reply.started":"2022-07-16T11:45:36.001854Z","shell.execute_reply":"2022-07-16T11:45:36.010845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g=sns.barplot(x='CabinCap',y='Transported',data=train_set)\ng.set( ylabel=\"Transportation Probability\")\n#seems like, people in cabins with moderate capacity had more chance of survival","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:36.013473Z","iopub.execute_input":"2022-07-16T11:45:36.013905Z","iopub.status.idle":"2022-07-16T11:45:36.481458Z","shell.execute_reply.started":"2022-07-16T11:45:36.013778Z","shell.execute_reply":"2022-07-16T11:45:36.480850Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# now let's extract CabinDeck and CabinSide\n\ntrain_set[['CabinDeck','CabinSide']]=train_set.Cabin.str.split('/',expand=True)[[0,2]]","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:36.482323Z","iopub.execute_input":"2022-07-16T11:45:36.482722Z","iopub.status.idle":"2022-07-16T11:45:36.504097Z","shell.execute_reply.started":"2022-07-16T11:45:36.482698Z","shell.execute_reply":"2022-07-16T11:45:36.502845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g=sns.barplot(x='CabinDeck',y='Transported',data=train_set)\ng.set( ylabel=\"Transportation Probability\")","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:36.510221Z","iopub.execute_input":"2022-07-16T11:45:36.510553Z","iopub.status.idle":"2022-07-16T11:45:36.975014Z","shell.execute_reply.started":"2022-07-16T11:45:36.510528Z","shell.execute_reply":"2022-07-16T11:45:36.973881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.barplot(x='CabinSide',y='Transported',data=train_set)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:36.976105Z","iopub.execute_input":"2022-07-16T11:45:36.976372Z","iopub.status.idle":"2022-07-16T11:45:37.255309Z","shell.execute_reply.started":"2022-07-16T11:45:36.976348Z","shell.execute_reply":"2022-07-16T11:45:37.254424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_set.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:37.256400Z","iopub.execute_input":"2022-07-16T11:45:37.257407Z","iopub.status.idle":"2022-07-16T11:45:37.276089Z","shell.execute_reply.started":"2022-07-16T11:45:37.257375Z","shell.execute_reply":"2022-07-16T11:45:37.275436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"5.3\"></a>\n## **<span style=\"color:#F26300;\">5.3 PassengerId</span>**\n\nBased on competition's data description, PassengerId is a unique Id for each passenger. Each Id takes the form gggg_pp where gggg indicates a group the passenger is travelling with and pp is their number within the group. People in a group are often family members, but not always.","metadata":{}},{"cell_type":"code","source":"train_set['GpId']=train_set.PassengerId.str.split('_',expand=True)[0]","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:37.277388Z","iopub.execute_input":"2022-07-16T11:45:37.278104Z","iopub.status.idle":"2022-07-16T11:45:37.300843Z","shell.execute_reply.started":"2022-07-16T11:45:37.278072Z","shell.execute_reply":"2022-07-16T11:45:37.299994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_set['GpSize']=train_set['GpId'].map(train_set['GpId'].value_counts())","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:37.302291Z","iopub.execute_input":"2022-07-16T11:45:37.303052Z","iopub.status.idle":"2022-07-16T11:45:37.312182Z","shell.execute_reply.started":"2022-07-16T11:45:37.303025Z","shell.execute_reply":"2022-07-16T11:45:37.311344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_set['GpSize']=train_set['GpSize'].astype('object')","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:37.313210Z","iopub.execute_input":"2022-07-16T11:45:37.313863Z","iopub.status.idle":"2022-07-16T11:45:37.322220Z","shell.execute_reply.started":"2022-07-16T11:45:37.313832Z","shell.execute_reply":"2022-07-16T11:45:37.321347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.barplot(x='GpSize',y='Transported',data=train_set)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:37.323147Z","iopub.execute_input":"2022-07-16T11:45:37.323693Z","iopub.status.idle":"2022-07-16T11:45:37.806128Z","shell.execute_reply.started":"2022-07-16T11:45:37.323649Z","shell.execute_reply":"2022-07-16T11:45:37.804999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_set.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:37.807408Z","iopub.execute_input":"2022-07-16T11:45:37.807695Z","iopub.status.idle":"2022-07-16T11:45:37.832663Z","shell.execute_reply.started":"2022-07-16T11:45:37.807671Z","shell.execute_reply":"2022-07-16T11:45:37.831238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"5.4\"></a>\n## **<span style=\"color:#F26300;\">5.4 Expenses</span>**","metadata":{}},{"cell_type":"code","source":"train_set[['RoomService','FoodCourt','ShoppingMall','Spa','VRDeck']]=train_set[['RoomService','FoodCourt','ShoppingMall','Spa','VRDeck']].fillna(0,axis=0)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:37.833530Z","iopub.execute_input":"2022-07-16T11:45:37.834651Z","iopub.status.idle":"2022-07-16T11:45:37.847155Z","shell.execute_reply.started":"2022-07-16T11:45:37.834606Z","shell.execute_reply":"2022-07-16T11:45:37.846325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#summing the total amount of expenditure\n\ntrain_set[\"TotalExp\"] = train_set[['RoomService','FoodCourt','ShoppingMall','Spa','VRDeck']].sum(axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:37.850660Z","iopub.execute_input":"2022-07-16T11:45:37.850943Z","iopub.status.idle":"2022-07-16T11:45:37.859276Z","shell.execute_reply.started":"2022-07-16T11:45:37.850920Z","shell.execute_reply":"2022-07-16T11:45:37.858132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.stripplot(y=\"TotalExp\", data=train_set)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:37.860772Z","iopub.execute_input":"2022-07-16T11:45:37.862034Z","iopub.status.idle":"2022-07-16T11:45:38.055922Z","shell.execute_reply.started":"2022-07-16T11:45:37.861994Z","shell.execute_reply":"2022-07-16T11:45:38.054982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.pairplot(data=train_set.select_dtypes(['number','bool']),hue=\"Transported\",palette='CMRmap',kind='reg')","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:45:38.057154Z","iopub.execute_input":"2022-07-16T11:45:38.057415Z","iopub.status.idle":"2022-07-16T11:46:28.452146Z","shell.execute_reply.started":"2022-07-16T11:45:38.057392Z","shell.execute_reply":"2022-07-16T11:46:28.450709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on the reg plots, we can see that there's a distinction between the lines for Transported True/False when 2 features are plotted. I'm going to use this and create some features, maybe they increase the score. ","metadata":{}},{"cell_type":"code","source":"for col in ['FoodCourt','Spa','VRDeck','ShoppingMall']:\n    colname='%'+col\n    train_set[colname]=[0 if train_set['TotalExp'][i]==0 else\n                        (train_set[col][i]/train_set['TotalExp'][i])*100 for i in train_set.index]\n","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:28.453535Z","iopub.execute_input":"2022-07-16T11:46:28.453859Z","iopub.status.idle":"2022-07-16T11:46:28.943329Z","shell.execute_reply.started":"2022-07-16T11:46:28.453830Z","shell.execute_reply":"2022-07-16T11:46:28.942178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_set['SPFC']=[0 if train_set['FoodCourt'][i]==0 else\n                    (train_set['Spa'][i]/train_set['FoodCourt'][i])*100 for i in train_set.index]\ntrain_set['VRFC']=[0 if train_set['FoodCourt'][i]==0 else\n                    (train_set['VRDeck'][i]/train_set['FoodCourt'][i])*100 for i in train_set.index]","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:28.944455Z","iopub.execute_input":"2022-07-16T11:46:28.944705Z","iopub.status.idle":"2022-07-16T11:46:29.141768Z","shell.execute_reply.started":"2022-07-16T11:46:28.944682Z","shell.execute_reply":"2022-07-16T11:46:29.140799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_set.drop(train_set.loc[train_set.TotalExp>=20000].index,inplace=True,axis=0)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.142978Z","iopub.execute_input":"2022-07-16T11:46:29.143344Z","iopub.status.idle":"2022-07-16T11:46:29.153557Z","shell.execute_reply.started":"2022-07-16T11:46:29.143312Z","shell.execute_reply":"2022-07-16T11:46:29.152656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"5.5\"></a>\n## **<span style=\"color:#F26300;\">5.5 Age</span>**\n\nI'm going to create age groups to acknowledge the peaks of the age's kde plot.","metadata":{}},{"cell_type":"code","source":"bins= [0,18,40,80]\nlabels = [1,2,3]\ntrain_set['AgeGp']=pd.cut(train_set['Age'], bins=bins, labels=labels,include_lowest=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.155577Z","iopub.execute_input":"2022-07-16T11:46:29.157139Z","iopub.status.idle":"2022-07-16T11:46:29.164809Z","shell.execute_reply.started":"2022-07-16T11:46:29.157109Z","shell.execute_reply":"2022-07-16T11:46:29.164062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"5.6\"></a>\n## **<span style=\"color:#F26300;\">5.6 Homeplanet-Destination</span>**\n\nBelow table was inspired by this [Notebook](https://www.kaggle.com/code/arootda/pycaret-visualization-optimization-0-81)\n","metadata":{}},{"cell_type":"code","source":"plt.subplots(figsize=(10, 5))\ng = sns.heatmap(train.pivot_table(index='HomePlanet', columns='Destination', values='Transported'), annot=True, cmap=\"YlGnBu\")\ng.set_title('Transportation Possibility by HomePlanet and Destination', weight='bold', size=15)\ng.set_xlabel('Destination', weight='bold', size=13)\ng.set_ylabel('HomePlanet', weight='bold', size=13)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.166277Z","iopub.execute_input":"2022-07-16T11:46:29.166933Z","iopub.status.idle":"2022-07-16T11:46:29.419080Z","shell.execute_reply.started":"2022-07-16T11:46:29.166898Z","shell.execute_reply":"2022-07-16T11:46:29.418122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_set['H-D']=train_set['HomePlanet']+\"-\"+train_set[\"Destination\"]","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.420211Z","iopub.execute_input":"2022-07-16T11:46:29.420486Z","iopub.status.idle":"2022-07-16T11:46:29.432145Z","shell.execute_reply.started":"2022-07-16T11:46:29.420462Z","shell.execute_reply":"2022-07-16T11:46:29.430342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"5.7\"></a>\n## **<span style=\"color:#F26300;\">5.7 Chi2 test</span>**","metadata":{}},{"cell_type":"markdown","source":"**Let's do a chi2 test to see if there are any features with a pvalue higher than 0.05**","metadata":{}},{"cell_type":"code","source":"from scipy.stats import chi2_contingency\ndef chi2_calc(df,target):\n    scores=[]\n    for col in df.columns:\n        ct=pd.crosstab(df[col],target)\n        stat,p,dof,expected=chi2_contingency(ct)\n        scores.append(p)\n    return pd.DataFrame(scores, index=df.columns, columns=['P value']).sort_values(by='P value')","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.436255Z","iopub.execute_input":"2022-07-16T11:46:29.436599Z","iopub.status.idle":"2022-07-16T11:46:29.444772Z","shell.execute_reply.started":"2022-07-16T11:46:29.436569Z","shell.execute_reply":"2022-07-16T11:46:29.443457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"chi2_calc(train_set.select_dtypes(['object','bool']),train_set.Transported)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.446418Z","iopub.execute_input":"2022-07-16T11:46:29.446754Z","iopub.status.idle":"2022-07-16T11:46:29.700066Z","shell.execute_reply.started":"2022-07-16T11:46:29.446718Z","shell.execute_reply":"2022-07-16T11:46:29.699188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**So all the categorical features have significant relevance to the target (pvalue <0.05)**","metadata":{}},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"5.8\"></a>\n## **<span style=\"color:#F26300;\">5.8 Dropping Columns and Defining The Target Column</span>**","metadata":{}},{"cell_type":"code","source":"# since based on the analysis of the graph, there was no difference\n#between transported/not transported kde plot, I'm dropping foodcourt and shoppingmall\ntrain_set.drop(['ShoppingMall','FoodCourt'],inplace=True,axis=1)\n\ntrain_set.drop(['Age','Cabin','Last','GpId','PassengerId'],inplace=True,axis=1)\n\ntrain_set.drop(['Destination','HomePlanet'],axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.701283Z","iopub.execute_input":"2022-07-16T11:46:29.701532Z","iopub.status.idle":"2022-07-16T11:46:29.715662Z","shell.execute_reply.started":"2022-07-16T11:46:29.701508Z","shell.execute_reply":"2022-07-16T11:46:29.714453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y=train_set.pop('Transported')","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.716887Z","iopub.execute_input":"2022-07-16T11:46:29.717449Z","iopub.status.idle":"2022-07-16T11:46:29.724205Z","shell.execute_reply.started":"2022-07-16T11:46:29.717410Z","shell.execute_reply":"2022-07-16T11:46:29.723250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y=y.map({False:0,True:1})","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.725550Z","iopub.execute_input":"2022-07-16T11:46:29.725796Z","iopub.status.idle":"2022-07-16T11:46:29.741543Z","shell.execute_reply.started":"2022-07-16T11:46:29.725772Z","shell.execute_reply":"2022-07-16T11:46:29.740440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.744996Z","iopub.execute_input":"2022-07-16T11:46:29.745882Z","iopub.status.idle":"2022-07-16T11:46:29.756332Z","shell.execute_reply.started":"2022-07-16T11:46:29.745831Z","shell.execute_reply":"2022-07-16T11:46:29.755328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"6\"></a>\n## **<span style=\"color:#F26300;\">6. PipeLine Implementation</span>**\n\nHere are the steps I'm going to take:\n\n* 1- Use 2 pipelines, one for categorical data and one for numerical data. in these pipelines, 2 things are going to happen:\n\n1-1 For numerical pipeline: imputing and scaling\n\n1-2 For categorical pipeline:imputing and encoding\n\n* 2-Use a ColumnTransformer to implement the functions in the pipeline on their respective data types\n\n\n* 3-Try a number of models to see which ones work better","metadata":{}},{"cell_type":"code","source":"train_set.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.757756Z","iopub.execute_input":"2022-07-16T11:46:29.758964Z","iopub.status.idle":"2022-07-16T11:46:29.775467Z","shell.execute_reply.started":"2022-07-16T11:46:29.758928Z","shell.execute_reply":"2022-07-16T11:46:29.774583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num=list(train_set.select_dtypes('float').columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.776510Z","iopub.execute_input":"2022-07-16T11:46:29.777494Z","iopub.status.idle":"2022-07-16T11:46:29.785953Z","shell.execute_reply.started":"2022-07-16T11:46:29.777458Z","shell.execute_reply":"2022-07-16T11:46:29.784933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.787565Z","iopub.execute_input":"2022-07-16T11:46:29.788295Z","iopub.status.idle":"2022-07-16T11:46:29.795508Z","shell.execute_reply.started":"2022-07-16T11:46:29.788223Z","shell.execute_reply":"2022-07-16T11:46:29.794488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat=list(train_set.select_dtypes(['object','bool','category']).columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.796744Z","iopub.execute_input":"2022-07-16T11:46:29.797098Z","iopub.status.idle":"2022-07-16T11:46:29.809273Z","shell.execute_reply.started":"2022-07-16T11:46:29.797076Z","shell.execute_reply":"2022-07-16T11:46:29.807460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.810854Z","iopub.execute_input":"2022-07-16T11:46:29.811301Z","iopub.status.idle":"2022-07-16T11:46:29.819912Z","shell.execute_reply.started":"2022-07-16T11:46:29.811243Z","shell.execute_reply":"2022-07-16T11:46:29.819299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ordinal_cat=['CabinCap','GpSize','AgeGp']","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.820779Z","iopub.execute_input":"2022-07-16T11:46:29.821389Z","iopub.status.idle":"2022-07-16T11:46:29.831858Z","shell.execute_reply.started":"2022-07-16T11:46:29.821364Z","shell.execute_reply":"2022-07-16T11:46:29.830671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"nominal_cat=[i for i in cat if i not in ordinal_cat]","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.833031Z","iopub.execute_input":"2022-07-16T11:46:29.834293Z","iopub.status.idle":"2022-07-16T11:46:29.841863Z","shell.execute_reply.started":"2022-07-16T11:46:29.834213Z","shell.execute_reply":"2022-07-16T11:46:29.840771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"nominal_cat","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.843504Z","iopub.execute_input":"2022-07-16T11:46:29.843978Z","iopub.status.idle":"2022-07-16T11:46:29.854792Z","shell.execute_reply.started":"2022-07-16T11:46:29.843950Z","shell.execute_reply":"2022-07-16T11:46:29.853639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_set[nominal_cat].info()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.856014Z","iopub.execute_input":"2022-07-16T11:46:29.856311Z","iopub.status.idle":"2022-07-16T11:46:29.873439Z","shell.execute_reply.started":"2022-07-16T11:46:29.856277Z","shell.execute_reply":"2022-07-16T11:46:29.872239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_set","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.874776Z","iopub.execute_input":"2022-07-16T11:46:29.875585Z","iopub.status.idle":"2022-07-16T11:46:29.901939Z","shell.execute_reply.started":"2022-07-16T11:46:29.875557Z","shell.execute_reply":"2022-07-16T11:46:29.901325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Note: OrdinalEncoder (without defining an order) is like LabelEncoder, but we can apply it to multiple columns at once\nwhereas LabelEncoder, is for one column transformation. The order in which they encode, is ascending, meaning A will be 1 and B will be 2**","metadata":{}},{"cell_type":"code","source":"#Importing Models\nfrom sklearn.ensemble import RandomForestClassifier,ExtraTreesClassifier,GradientBoostingClassifier,AdaBoostClassifier,ExtraTreesRegressor\nfrom sklearn.neural_network import MLPClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom xgboost import XGBClassifier\nfrom lightgbm import LGBMClassifier\nfrom catboost import CatBoostClassifier\n\n########################################################################\nfrom sklearn.preprocessing import FunctionTransformer\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.preprocessing import RobustScaler,StandardScaler,MinMaxScaler\nfrom sklearn.preprocessing import OrdinalEncoder,OneHotEncoder\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.experimental import enable_iterative_imputer\nfrom sklearn.impute import IterativeImputer,KNNImputer,SimpleImputer\nfrom category_encoders import MEstimateEncoder,PolynomialEncoder,BackwardDifferenceEncoder,LeaveOneOutEncoder,QuantileEncoder,WOEEncoder\nfrom sklearn.discriminant_analysis import LinearDiscriminantAnalysis\nfrom sklearn.preprocessing import PowerTransformer,QuantileTransformer\nfrom sklearn.preprocessing import KBinsDiscretizer\nfrom sklearn.tree import DecisionTreeRegressor, DecisionTreeClassifier\nrandom_state=0\n\n#setting up the imputer for numerical features\nNImputer=IterativeImputer(estimator=ExtraTreesRegressor(max_depth=7,min_samples_leaf=200,random_state=random_state),random_state=random_state,tol=1e-5,max_iter=20)\n\n#setting up the imputer for categorical features (We can explore with this imputer too, I'll let it be here just in case)\nCImputer=IterativeImputer(estimator=ExtraTreesClassifier(max_depth=10,min_samples_leaf=100,random_state=random_state),random_state=random_state,tol=1e-5,max_iter=20)\n\n#target encoder: it's usually used for high cardinality features. \n# t_encoder = MEstimateEncoder(m=10, random_state=random_state,handle_missing='return_nan')\n\n#PolynomialEncoder\n# p_ecnoder=PolynomialEncoder()\n\n#BackwardDifferenceCoding\n# b_encoder=BackwardDifferenceEncoder()\n\n#QuantileEncoder\n# q_encoder=QuantileEncoder(m=12,handle_missing='return_nan')\n\n#Weight of Evidence Encoder\nwe=WOEEncoder(drop_invariant=True,handle_missing='return_nan')\n\n\n#creating log function that has fit and fit_transform methods because numerical columns are mostly skewed\n# def log_transform(x):\n#     return np.log(x + 1)\n# log_pip=FunctionTransformer(log_transform)\n\n#Using power transformer class for standardizing data distirbution in numerical columns\n# pt=PowerTransformer()\n\n#Using quantile transformer\n# qt=QuantileTransformer(n_quantiles=500, random_state=0,output_distribution='uniform')\n\n#Using kbinsdiscretizer for skewed numerical features\n# kbins = KBinsDiscretizer(n_bins=2,encode='ordinal', strategy='uniform')\n\n#creating two preprocessing pipelines for categorical and numerical data types\nnumeric_pip=Pipeline(steps=[('NImputer',SimpleImputer(strategy='median')),('Scale',RobustScaler(unit_variance=True))])\n\nnomcat_pip=Pipeline(steps=[('WeightOfEvidence',we),('CImputer',SimpleImputer(strategy='mean'))])\n\nordcat_pip=Pipeline(steps=[('CImputer',SimpleImputer(strategy='mean'))])\n\n#creating a column transformer to implement the transformation\nct=ColumnTransformer(transformers=[('num',numeric_pip,num),('nomcat',nomcat_pip,nominal_cat),('ordcat',ordcat_pip,ordinal_cat)])","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:29.903366Z","iopub.execute_input":"2022-07-16T11:46:29.903578Z","iopub.status.idle":"2022-07-16T11:46:29.915054Z","shell.execute_reply.started":"2022-07-16T11:46:29.903556Z","shell.execute_reply":"2022-07-16T11:46:29.913976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#This is for testing to see if the columntransformer works properly\ntst = Pipeline([\n('coltrns',ct), #COLUMN TRANSFORMER\n])\nxx=tst.fit_transform(train_set,y)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-16T11:46:29.916624Z","iopub.execute_input":"2022-07-16T11:46:29.916883Z","iopub.status.idle":"2022-07-16T11:46:30.055834Z","shell.execute_reply.started":"2022-07-16T11:46:29.916861Z","shell.execute_reply":"2022-07-16T11:46:30.054624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"xx=pd.DataFrame(xx)\nxx.head()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-16T11:46:30.057278Z","iopub.execute_input":"2022-07-16T11:46:30.057620Z","iopub.status.idle":"2022-07-16T11:46:30.079324Z","shell.execute_reply.started":"2022-07-16T11:46:30.057595Z","shell.execute_reply":"2022-07-16T11:46:30.078151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"6.1\"></a>\n## **<span style=\"color:#F26300;\">6.1 Models and Scoring</span>**","metadata":{}},{"cell_type":"code","source":"classifiers = []\nclassifiers.append(RandomForestClassifier(random_state=random_state))\nclassifiers.append(GradientBoostingClassifier(random_state=random_state)) \nclassifiers.append(XGBClassifier(random_state=random_state))\nclassifiers.append(LGBMClassifier(random_state=random_state))\nclassifiers.append(CatBoostClassifier(random_state=random_state,verbose=0))\n# CREATING A FOR LOOP FOR SCORING EACH MODEL\ncv_results = []\ncv = KFold(n_splits=5,shuffle=True, random_state=random_state)\nfor classifier in classifiers :\n    classif = Pipeline([\n('coltrns',ct), #COLUMN TRANSFORMER\n('classifier', classifier)])\n    cvs=cross_val_score(classif, train_set, y, scoring = \"accuracy\", cv = cv, n_jobs=-1)\n    cv_results.append(cvs)\n\ncv_means = []\ncv_std = []\nfor cv_result in cv_results:\n    cv_means.append(cv_result.mean())\n    cv_std.append(cv_result.std())\n\n#CREATING A DATAFRAME OF MODEL SCORES\ncv_res = pd.DataFrame({\"CrossValMeans\":cv_means,\"CrossValSDs\": cv_std,\"Algorithm\":[\"RandomForestClassifier\",\n                                                                                 \"GradientBoostingClassifier\",\n                                                                                  \"XGBClassifier\",\n                                                                                  \"LGBMClassifier\",\n                                                                                  \"CatBoostClassifier\",\n                                                                                  ]})\n#PLOTTING\ng = sns.barplot(x=\"CrossValMeans\",y=\"Algorithm\",data = cv_res.sort_values(by='CrossValMeans'), palette=\"twilight_shifted_r\",**{'xerr':cv_std})\ng.set_xlabel(\"Mean Accuracy\")\ng = g.set_title(\"Cross validation scores\")\ng=sns.set(rc={'figure.figsize':(5,5)})\n\ncv_res.sort_values(by='CrossValMeans')","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:46:30.081111Z","iopub.execute_input":"2022-07-16T11:46:30.081423Z","iopub.status.idle":"2022-07-16T11:47:03.946470Z","shell.execute_reply.started":"2022-07-16T11:46:30.081398Z","shell.execute_reply":"2022-07-16T11:47:03.945504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"7\"></a>\n## **<span style=\"color:#F26300;\">7. Learning Curve</span>**\n\nHere we're going to take a look at the learning curve, using below function, before tuning the models.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import learning_curve\ndef plot_learning_curve(my_model, title, X, y, ylim=None, cv=None,\n                        n_jobs=-1, train_sizes=[np.linspace(.1, 1.0, 5)]):\n    \"\"\"Generate a simple plot of the test and training learning curve\"\"\"\n    plt.figure()\n    plt.title(title)\n    if ylim is not None:\n        plt.ylim(*ylim)\n    plt.xlabel(\"Training examples\")\n    plt.ylabel(\"Score\")\n    train_sizes, train_scores, test_scores = learning_curve(\n        my_model, X, y, cv=cv, n_jobs=-1, train_sizes=train_sizes,scoring=\"accuracy\")\n    train_scores_mean = np.mean(train_scores, axis=1)\n    train_scores_std = np.std(train_scores, axis=1)\n    test_scores_mean = np.mean(test_scores, axis=1)\n    test_scores_std = np.std(test_scores, axis=1)\n\n\n    plt.fill_between(train_sizes, train_scores_mean - train_scores_std,\n                     train_scores_mean + train_scores_std, alpha=0.1,\n                     color=\"r\")\n    plt.fill_between(train_sizes, test_scores_mean - test_scores_std,\n                     test_scores_mean + test_scores_std, alpha=0.1, color=\"g\")\n    plt.plot(train_sizes, train_scores_mean, 'o-', color=\"r\",\n             label=\"Training score\")\n    plt.plot(train_sizes, test_scores_mean, 'o-', color=\"g\",\n             label=\"Cross-validation score\")\n\n    plt.legend(loc=\"best\")\n    plt.grid()\n    return plt","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:47:03.947880Z","iopub.execute_input":"2022-07-16T11:47:03.948097Z","iopub.status.idle":"2022-07-16T11:47:03.958646Z","shell.execute_reply.started":"2022-07-16T11:47:03.948075Z","shell.execute_reply":"2022-07-16T11:47:03.957472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in [1,2,3,4]: #Index of the last 4 highest scores\n    model = Pipeline([\n        ('coltrns',ct), #COLUMN TRANSFORMER\n        ('classifier', classifiers[i])])\n    plot_learning_curve(model,model[1],train_set,y)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:47:03.960189Z","iopub.execute_input":"2022-07-16T11:47:03.960855Z","iopub.status.idle":"2022-07-16T11:48:40.409516Z","shell.execute_reply.started":"2022-07-16T11:47:03.960814Z","shell.execute_reply":"2022-07-16T11:48:40.408425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The lesser the training curve changes, the more the model is overfit. because it's working well on the training data and learning it amazingly, but when it comes to the validation set, it can not generalize and performs poorly.","metadata":{}},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"8\"></a>\n## **<span style=\"color:#F26300;\">8. Model Tuning using GridSearchCV</span>**\n\n\nBecause of highly skewed data and the fact that if I were to delete all outliers, a big portion of the dataset would be deleted, I'm going to use tree-based models as final models for tuning. they have the highest scores in the chart and they are:\n\n* Random Forest Classifier\n* XGBClassifier\n* LGBMClassifier\n* Gradient Boosting Classifier\n","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:48:40.410880Z","iopub.execute_input":"2022-07-16T11:48:40.411143Z","iopub.status.idle":"2022-07-16T11:48:40.417253Z","shell.execute_reply.started":"2022-07-16T11:48:40.411116Z","shell.execute_reply":"2022-07-16T11:48:40.416210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"8.1\"></a>\n## **<span style=\"color:#F26300;\">8.1 Gradient Boost</span>**","metadata":{}},{"cell_type":"markdown","source":"[Reference](https://www.analyticsvidhya.com/blog/2016/02/complete-guide-parameter-tuning-gradient-boosting-gbm-python/) for tuning gradient boost","metadata":{}},{"cell_type":"code","source":"#GB\n# gb_model = Pipeline([\n# ('coltrns',ct), #COLUMN TRANSFORMER\n# ('classifier', GradientBoostingClassifier(random_state=random_state))])\n\n# #there are two types of parameter to be tuned here – tree based and boosting parameters,\n# #in general Lower the learning rate and increase the estimators proportionally to get more robust models.\n# gb_grid =  {\n    \n#     #Generally the default value of 0.1 works but somewhere between 0.05 to 0.2 should work for different problems\n#     \"classifier__learning_rate\":[0.1],\n    \n#     #This should range around 40-70. Remember to choose a value on which your system can work fairly fast.\n#     #This is because it will be used for testing various scenarios and determining the tree parameters.\n#     \"classifier__n_estimators\":[60,70], \n    \n#     #This should be ~0.5-1% of total values\n#     \"classifier__min_samples_split\":[40,50],\n    \n#     #Can be selected based on intuition. This is just used for preventing overfitting\n#     \"classifier__min_samples_leaf\" : [120,150],\n    \n#     #Should be chosen (5-8) based on the number of observations and predictors.\n#     \"classifier__max_depth\" :[8,10],\n    \n#     #.8 is a commonly used start value\n#     \"classifier__subsample\":[.8,.5],\n   \n#     \"classifier__max_features\":['log2']\n\n#     }\n\n# gsgb = GridSearchCV(gb_model,gb_grid , cv=cv, scoring=\"accuracy\", n_jobs= -1, verbose = 1)\n\n# gsgb.fit(train_set,y)\n# # gb_model.fit(train_set,y)\n\n# gb_best = gsgb.best_estimator_\n\n# # Best score\n# display(gsgb.best_score_)\n# display(gb_best)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:48:40.418407Z","iopub.execute_input":"2022-07-16T11:48:40.418664Z","iopub.status.idle":"2022-07-16T11:48:40.428325Z","shell.execute_reply.started":"2022-07-16T11:48:40.418637Z","shell.execute_reply":"2022-07-16T11:48:40.427479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"8.2\"></a>\n## **<span style=\"color:#F26300;\">8.2 LGBM</span>**","metadata":{}},{"cell_type":"code","source":"lgbm_model = Pipeline([\n('coltrns',ct), #COLUMN TRANSFORMER\n('classifier', LGBMClassifier(random_state=random_state))])\n\nlgbm_grid =  {\n    'classifier__num_leaves': [31, 127],\n    'classifier__reg_alpha': [0.1, 0.5],\n    'classifier__reg_lambda': [0,.5, 1],\n    }\n\ngslgbm = GridSearchCV(lgbm_model,lgbm_grid , cv=cv, scoring=\"accuracy\", n_jobs= -1, verbose = 1)\n\ngslgbm.fit(train_set,y)\n\nlgbm_best = gslgbm.best_estimator_\n\n# Best score\ndisplay(gslgbm.best_score_)\n\ndisplay(lgbm_best)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:48:40.429888Z","iopub.execute_input":"2022-07-16T11:48:40.430209Z","iopub.status.idle":"2022-07-16T11:48:53.301192Z","shell.execute_reply.started":"2022-07-16T11:48:40.430177Z","shell.execute_reply":"2022-07-16T11:48:53.300602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"8.3\"></a>\n## **<span style=\"color:#F26300;\">8.3 XGBClassifier</span>**","metadata":{}},{"cell_type":"code","source":"xgbc_best = Pipeline([\n('coltrns',ct), #COLUMN TRANSFORMER\n('classifier', XGBClassifier(random_state=random_state))])\nxgbc_best.fit(train_set,y)\n\n# xgbc_grid =  {\n              \n#               'classifier__learning_rate': [0.03,0.01], \n#               'classifier__max_depth': [8,10],\n#               'classifier__min_child_weight': [40,50],\n#               'classifier__subsample': [.5,.8],\n#               'classifier__colsample_bytree': [.5,.8],\n#               'classifier__n_estimators': [100] \n#               }\n# gsxgbc = GridSearchCV(xgbc_model,xgbc_grid , cv=cv, scoring=\"accuracy\", n_jobs= -1, verbose = 1)\n\n# gsxgbc.fit(train_set,y)\n\n# xgbc_best = gsxgbc.best_estimator_\n\n# # Best score\n# display(gsxgbc.best_score_)\n\n# display(xgbc_best)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:48:53.302022Z","iopub.execute_input":"2022-07-16T11:48:53.302378Z","iopub.status.idle":"2022-07-16T11:48:54.348460Z","shell.execute_reply.started":"2022-07-16T11:48:53.302355Z","shell.execute_reply":"2022-07-16T11:48:54.347067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"8.4\"></a>\n## **<span style=\"color:#F26300;\">8.4 CatBoost</span>**","metadata":{}},{"cell_type":"code","source":"cb_best = Pipeline([\n('coltrns',ct), #COLUMN TRANSFORMER\n('classifier', CatBoostClassifier(random_state=random_state,verbose=0))])\n\n# cb_grid =  {'classifier__depth': [7,10,15],\n#                  'classifier__learning_rate' : [0.01,0.03,0.1,0.3],\n#                   'classifier__iterations': [100,150,200]\n#                  }\n# gscb= GridSearchCV(cb_model,cb_grid , cv=cv, scoring=\"accuracy\", n_jobs= -1, verbose = 1)\n\ncb_best.fit(train_set,y)\n\n# cb_best = gscb.best_estimator_\n\n# Best score\n# display(gscb.best_score_)\n\n# display(cb_best)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:48:54.350289Z","iopub.execute_input":"2022-07-16T11:48:54.350665Z","iopub.status.idle":"2022-07-16T11:48:58.407327Z","shell.execute_reply.started":"2022-07-16T11:48:54.350623Z","shell.execute_reply":"2022-07-16T11:48:58.406318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"9\"></a>\n## **<span style=\"color:#F26300;\">9. Transforming and Predicting Test Set</span>**\n","metadata":{}},{"cell_type":"code","source":"#Creating Rel Feature For Test Set\ntest_set[['Name','Last']] = test_set.Name.str.split(\" \", expand=True)\ntest_set.drop('Name',inplace=True,axis=1)\ntest_set['Last'].ffill(axis=0,inplace=True) # I'll drop 'Last', becuase I still don't know how to use names to improve the score\n\n#Creating cabin capacity using the number of people in it\ntcabin_cap= test_set.Cabin.value_counts()\ntest_set['CabinCap']=test_set['Cabin'].map(tcabin_cap)\ntest_set['CabinCap']=test_set['CabinCap'].astype('object')\n\n#Creating cabin deck and cabin side features\ntest_set[['CabinDeck','CabinSide']]=test_set.Cabin.str.split('/',expand=True)[[0,2]]\ntest_set['GpId']=test_set.PassengerId.str.split('_',expand=True)[0]\ntest_set['GpSize']=test_set['GpId'].map(test_set['GpId'].value_counts())\ntest_set['GpSize']=test_set['GpSize'].astype('object')\n\n#Total expense\ntest_set[['RoomService','FoodCourt','ShoppingMall','Spa','VRDeck']]=test_set[['RoomService','FoodCourt','ShoppingMall','Spa','VRDeck']].fillna(0,axis=0)\ntest_set[\"TotalExp\"] = test_set[['RoomService','FoodCourt','ShoppingMall','Spa','VRDeck']].sum(axis=1)\n\nfor col in ['FoodCourt','Spa','VRDeck','ShoppingMall']:\n    colname='%'+col\n    test_set[colname]=[0 if test_set['TotalExp'][i]==0 else\n                        (test_set[col][i]/test_set['TotalExp'][i])*100 for i in test_set.index]\n#Expense ratio based on pairplot\ntest_set['SPFC']=[0 if test_set['FoodCourt'][i]==0 else\n                    (test_set['Spa'][i]/test_set['FoodCourt'][i])*100 for i in test_set.index]\ntest_set['VRFC']=[0 if test_set['FoodCourt'][i]==0 else\n                    (test_set['VRDeck'][i]/test_set['FoodCourt'][i])*100 for i in test_set.index]\n#Age\nbins= [0,18,40,80]\nlabels = [1,2,3]\ntest_set['AgeGp']=pd.cut(test_set['Age'], bins=bins, labels=labels,include_lowest=True)\n\n#H-D\ntest_set['H-D']=test_set['HomePlanet']+\"-\"+test_set[\"Destination\"]\n\n\n#dropping PassengerId\n\ntest_set.drop(['ShoppingMall','FoodCourt'],inplace=True,axis=1)\n\ntest_set.drop(['Age','Cabin','Last','GpId','PassengerId'],inplace=True,axis=1)\n\ntest_set.drop(['Destination','HomePlanet'],axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T12:01:14.035579Z","iopub.execute_input":"2022-07-16T12:01:14.035961Z","iopub.status.idle":"2022-07-16T12:01:14.399917Z","shell.execute_reply.started":"2022-07-16T12:01:14.035935Z","shell.execute_reply":"2022-07-16T12:01:14.398032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_cb = pd.Series(cb_best.predict(test_set), name=\"cb\")\ntest_lgbm = pd.Series(lgbm_best.predict(test_set), name=\"lgbm\")\n# test_gb = pd.Series(gb_best.predict(test_set), name=\"gb\")\ntest_xgbc = pd.Series(xgbc_best.predict(test_set), name=\"xgbc\")\n\n\n# Concatenate all classifier results\nensemble_results = pd.concat([test_cb,test_lgbm,test_xgbc],axis=1)\n\n\ng= sns.heatmap(ensemble_results.corr(),annot=True)\n\n#The results mainly agree with eachother, but in general it's better if we have strong models that are not highly correlated so that they can\n#cover eachother's flaws in prediction to a degree!","metadata":{"execution":{"iopub.status.busy":"2022-07-16T12:03:34.670069Z","iopub.execute_input":"2022-07-16T12:03:34.670444Z","iopub.status.idle":"2022-07-16T12:03:35.008554Z","shell.execute_reply.started":"2022-07-16T12:03:34.670418Z","shell.execute_reply":"2022-07-16T12:03:35.007312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.ensemble import VotingClassifier\nvotingC = VotingClassifier(estimators=[ ('CB', cb_best),('LGBM',lgbm_best),('XGBC',xgbc_best)], voting='soft', n_jobs=-1)\n\nvotingC.fit(train_set,y)\npredictions = pd.DataFrame(votingC.predict(test_set)).values","metadata":{"execution":{"iopub.status.busy":"2022-07-16T12:03:41.407083Z","iopub.execute_input":"2022-07-16T12:03:41.407431Z","iopub.status.idle":"2022-07-16T12:03:47.917149Z","shell.execute_reply.started":"2022-07-16T12:03:41.407405Z","shell.execute_reply":"2022-07-16T12:03:47.915828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output = pd.DataFrame({'PassengerId': test['PassengerId'], 'Transported': predictions.flatten()})","metadata":{"execution":{"iopub.status.busy":"2022-07-16T12:03:49.566014Z","iopub.execute_input":"2022-07-16T12:03:49.566591Z","iopub.status.idle":"2022-07-16T12:03:49.572448Z","shell.execute_reply.started":"2022-07-16T12:03:49.566562Z","shell.execute_reply":"2022-07-16T12:03:49.571060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output.Transported=output.Transported.map({1:True,0:False})","metadata":{"execution":{"iopub.status.busy":"2022-07-16T12:03:51.575730Z","iopub.execute_input":"2022-07-16T12:03:51.577176Z","iopub.status.idle":"2022-07-16T12:03:51.585586Z","shell.execute_reply.started":"2022-07-16T12:03:51.577111Z","shell.execute_reply":"2022-07-16T12:03:51.584470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T12:03:54.292362Z","iopub.execute_input":"2022-07-16T12:03:54.292704Z","iopub.status.idle":"2022-07-16T12:03:54.303172Z","shell.execute_reply.started":"2022-07-16T12:03:54.292678Z","shell.execute_reply":"2022-07-16T12:03:54.302079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"[List of Content](#sections)\n<a id=\"10\"></a>\n## **<span style=\"color:#F26300;\">10. Submission</span>**","metadata":{}},{"cell_type":"code","source":"output.to_csv('submission.csv', index=False)\nprint(\"Your submission was successfully saved!\")","metadata":{"execution":{"iopub.status.busy":"2022-07-16T12:03:57.930677Z","iopub.execute_input":"2022-07-16T12:03:57.931040Z","iopub.status.idle":"2022-07-16T12:03:57.949968Z","shell.execute_reply.started":"2022-07-16T12:03:57.931013Z","shell.execute_reply":"2022-07-16T12:03:57.949073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Thank you for reading!","metadata":{}},{"cell_type":"markdown","source":"We can explore more with feature, for example, maybe the homeplanet-destination pairs have a statistically significant relation with the target or we can experiment with tuning and model selection; there are a lot of things we can do and I hope this notebook gives someone some ideas for their project!\n\n**I'd be happy to recieve your feedbacks and suggestions on how to improve my work :)**","metadata":{}}]}