{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-17T09:39:34.210335Z","iopub.execute_input":"2022-07-17T09:39:34.210819Z","iopub.status.idle":"2022-07-17T09:39:34.219706Z","shell.execute_reply.started":"2022-07-17T09:39:34.210781Z","shell.execute_reply":"2022-07-17T09:39:34.218414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Table Of Content\n\n* [Introduction](#section-one)\n* [1. Data Understanding](#dataunderstanding)\n    -  [Nomenclature of Dataset](#nomenclature)\n    -  [Dataset Information](#info)\n    -  [Statistic Descriptive](#statdesc)\n* [2. EDA](#eda)\n    -  [Numerical Features](#numfeat)\n    -  [Categorical Features](#catfeat)\n    -  [Correlation matrix](#corr)\n    -  [Target Distribution](#target)\n    -  [Missing Values](#missing)\n* [3. Modelling](#modelling)\n    - [Data Partition](#datapar)\n    - [Preprocessing](#prep) \n    - [Spot Check Algorithm](#spot_check)\n    - [Model Tuning](#tune) ","metadata":{"execution":{"iopub.status.busy":"2022-07-09T07:56:28.184289Z","iopub.execute_input":"2022-07-09T07:56:28.184724Z","iopub.status.idle":"2022-07-09T07:56:28.190786Z","shell.execute_reply.started":"2022-07-09T07:56:28.18469Z","shell.execute_reply":"2022-07-09T07:56:28.18922Z"}}},{"cell_type":"code","source":"#import all package\nimport numpy as np \nimport pandas as pd \n\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\nfrom sklearn.model_selection import train_test_split\n\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.preprocessing import OrdinalEncoder,PowerTransformer,OneHotEncoder\nfrom sklearn.impute import SimpleImputer\n\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import SVC\nfrom sklearn.ensemble import RandomForestClassifier\nfrom xgboost import XGBClassifier\nfrom lightgbm import LGBMClassifier\nfrom catboost import CatBoostClassifier \n\nfrom sklearn.model_selection import StratifiedKFold\n\nfrom skopt.space.space import Real, Categorical, Integer \nfrom skopt import BayesSearchCV","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:34.229500Z","iopub.execute_input":"2022-07-17T09:39:34.230643Z","iopub.status.idle":"2022-07-17T09:39:34.240346Z","shell.execute_reply.started":"2022-07-17T09:39:34.230602Z","shell.execute_reply":"2022-07-17T09:39:34.239256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"section-one\"></a>\n# Introduction","metadata":{}},{"cell_type":"markdown","source":"In this notebook, we will make a prediction model about the Spaceship Titanic. This notebook is an introduction notebook regarding the machine learning competition at Kaggle. and this is my first notebook, so if there are any mistakes, feel free to comment\n\nThis dataset is made by Kaggle, inspired by the titanic dataset. Because at first glance, the two datasets have some similarities. If you often do projects regarding predictions, you must be familiar with these datasets.\n\nThe Spaceship Titanic dataset is a dataset that contains data on spaceship passengers. And this notebook is to find out with the available information we can predict whether the passenger is transported or not. In other words, this case is a prediction of classification.\n\nIn this notebook, I will do several steps before making a prediction, and this step is the step I usually do so that it may be different from the steps others took. These steps include:\n\n1. Data Understanding\n2. EDA & Cleaning\n4. Modelling","metadata":{}},{"cell_type":"markdown","source":"# 1. Data Understanding","metadata":{}},{"cell_type":"code","source":"# code for importing data\ndf = pd.read_csv(\"/kaggle/input/spaceship-titanic/train.csv\") #import train data\ndf_test = pd.read_csv(\"/kaggle/input/spaceship-titanic/test.csv\") #import test data ","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:34.250788Z","iopub.execute_input":"2022-07-17T09:39:34.251649Z","iopub.status.idle":"2022-07-17T09:39:34.316294Z","shell.execute_reply.started":"2022-07-17T09:39:34.251606Z","shell.execute_reply":"2022-07-17T09:39:34.314876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head() #preview 5 first row in df","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:34.318700Z","iopub.execute_input":"2022-07-17T09:39:34.319087Z","iopub.status.idle":"2022-07-17T09:39:34.345067Z","shell.execute_reply.started":"2022-07-17T09:39:34.319040Z","shell.execute_reply":"2022-07-17T09:39:34.343810Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"nomenclatur\"></a> \nNomenclature Of Dataset : \n<ul>\n    <li>PassengerId - A unique Id for each passenger. Each Id takes the form gggg_pp where gggg indicates a group the passenger is travelling with and pp is their number within the group. People in a group are often family members, but not always.</li>\n    <li>HomePlanet - The planet the passenger departed from, typically their planet of permanent residence.</li>\n    <li>CryoSleep - Indicates whether the passenger elected to be put into suspended animation for the duration of the voyage. Passengers in cryosleep are confined to their cabins.</li>\n    <li>Cabin - The cabin number where the passenger is staying. Takes the form deck/num/side, where side can be either P for Port or S for Starboard.</li>\n    <li>Destination - The planet the passenger will be debarking to.</li>\n    <li>Age - The age of the passenger.</li>\n    <li>VIP - Whether the passenger has paid for special VIP service during the voyage.</li>\n    <li>RoomService, FoodCourt, ShoppingMall, Spa, VRDeck - Amount the passenger has billed at each of the Spaceship Titanic's many luxury amenities.</li>\n    <li>Name - The first and last names of the passenger.</li>\n    <li>Transported - Whether the passenger was transported to another dimension. This is the target, the column we are trying to predict.</li>\n\n\n\n\n\n\n \n\n\n    \n    \n    \n    \n    \n    \n    ","metadata":{}},{"cell_type":"markdown","source":"<a id=\"info\"></a>\n## Dataset Information","metadata":{}},{"cell_type":"code","source":"df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:34.347179Z","iopub.execute_input":"2022-07-17T09:39:34.347695Z","iopub.status.idle":"2022-07-17T09:39:34.379804Z","shell.execute_reply.started":"2022-07-17T09:39:34.347644Z","shell.execute_reply":"2022-07-17T09:39:34.378451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the information above, we can find out:\n<ul>\n    <li> that this dataset consists of 14 columns.</li>\n    <li> We may use six columns of type float64 (numerical) to do scaling or normalization, depending on the model used </li>\n    <li>We will maybe do a label or one hot encoder of the seven columns of type Object (Categorical).</li>\n    <li>One column as a target of type boolean. </li>\n    <li> the dataset has a total of 8693 rows, and all columns/features have missing values ​​except PassengerId and Transported features; therefore, we need to impute missing values</li>\n   \n","metadata":{}},{"cell_type":"markdown","source":"<a id=\"statdesc\"></a>\n## Statistic Descriptive","metadata":{}},{"cell_type":"code","source":"df.describe().T","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:34.382279Z","iopub.execute_input":"2022-07-17T09:39:34.382656Z","iopub.status.idle":"2022-07-17T09:39:34.428808Z","shell.execute_reply.started":"2022-07-17T09:39:34.382621Z","shell.execute_reply":"2022-07-17T09:39:34.427727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the statistical description above, it can be seen that numeric features other than `Age` have a max value that is very far from the 75% value for each feature. This causes the data to tend to lean to the left, or it could be that this data has an outlier value. The data cleaning process will be carried out to overcome this in the next section.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"eda\"></a>\n# 2. EDA","metadata":{}},{"cell_type":"markdown","source":"<a id = \"numfeat\"></a>\n## Numerical Feature","metadata":{"execution":{"iopub.status.busy":"2022-07-05T12:50:31.586827Z","iopub.execute_input":"2022-07-05T12:50:31.587316Z","iopub.status.idle":"2022-07-05T12:50:31.593215Z","shell.execute_reply.started":"2022-07-05T12:50:31.587282Z","shell.execute_reply":"2022-07-05T12:50:31.591721Z"}}},{"cell_type":"code","source":"#fungsi untuk plotting\ndef num_eda(data,feature,target,bins ):\n    df_set_pos = data[data[target] == True].drop(target,axis = 1)\n    df_set_neg = data[data[target] == False].drop(target,axis = 1)\n    #plot numerical data (Classification Task)\n    fig, axes = plt.subplots(2,2,figsize = (18,4))\n\n    ax1 = sns.histplot(x = feature,data = data,ax = axes[0,0],bins = bins,kde = True,edgecolor = \"k\",color = \"orange\")\n    ax1.grid(linestyle='--', linewidth=0.5, color='gray')\n    ax1.set_title(f\"{feature} Distribution\")\n    \n    ax2 = sns.histplot(x = feature,data = df_set_pos,ax = axes[0,1],bins = bins,label = \"Transported\",kde = True,color = \"green\",linewidth = 0 )\n    ax2_1 = sns.histplot(x = feature,data = df_set_neg,ax = axes[0,1],label = \"Not - Transported\",bins = bins,kde = True,color = \"red\",linewidth = 0)\n    ax2.grid(linestyle='--', linewidth=0.1, color='gray')\n    ax2.set_title(f\"{feature} Distribution by Target Class\")\n    ax2.legend()\n    \n    ax3 = sns.boxplot(x = feature,data = data,ax = axes[1,0],color = \"orange\")\n    ax3.grid(linestyle='--', linewidth=0.5, color='gray')\n    \n    ax4 = sns.boxplot(x = feature,y = target,data = data,ax = axes[1,1],orient = \"h\",palette = [\"red\",\"green\"])\n    ax4.grid(linestyle='--', linewidth=0.1, color='gray')\n    ax4.legend()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:34.430391Z","iopub.execute_input":"2022-07-17T09:39:34.431517Z","iopub.status.idle":"2022-07-17T09:39:34.447683Z","shell.execute_reply.started":"2022-07-17T09:39:34.431477Z","shell.execute_reply":"2022-07-17T09:39:34.446442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"age\"></a>\n### Age","metadata":{}},{"cell_type":"code","source":"num_eda(df,\"Age\",\"Transported\",20)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:34.450206Z","iopub.execute_input":"2022-07-17T09:39:34.451119Z","iopub.status.idle":"2022-07-17T09:39:35.364921Z","shell.execute_reply.started":"2022-07-17T09:39:34.451077Z","shell.execute_reply":"2022-07-17T09:39:35.363541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the graph above is the distribution of the age of passengers, it can be seen that on the histogram graph on the right, more passengers aged 0-20 were transported, at the age of 20-40 more passengers were not transported, while those aged 40 and over tended to be the same.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"roomservice\"></a>\n### Room Service, SPA, VR Deck","metadata":{"execution":{"iopub.status.busy":"2022-07-05T09:27:27.626787Z","iopub.execute_input":"2022-07-05T09:27:27.627319Z","iopub.status.idle":"2022-07-05T09:27:27.633906Z","shell.execute_reply.started":"2022-07-05T09:27:27.627276Z","shell.execute_reply":"2022-07-05T09:27:27.631777Z"}}},{"cell_type":"code","source":"num_eda(df,\"RoomService\",\"Transported\",30)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:35.366405Z","iopub.execute_input":"2022-07-17T09:39:35.366781Z","iopub.status.idle":"2022-07-17T09:39:37.387762Z","shell.execute_reply.started":"2022-07-17T09:39:35.366748Z","shell.execute_reply":"2022-07-17T09:39:37.386596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_eda(df,\"Spa\",\"Transported\",30)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:37.389473Z","iopub.execute_input":"2022-07-17T09:39:37.389945Z","iopub.status.idle":"2022-07-17T09:39:38.222562Z","shell.execute_reply.started":"2022-07-17T09:39:37.389907Z","shell.execute_reply":"2022-07-17T09:39:38.221302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_eda(df,\"VRDeck\",\"Transported\",30)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:38.224200Z","iopub.execute_input":"2022-07-17T09:39:38.224578Z","iopub.status.idle":"2022-07-17T09:39:39.111755Z","shell.execute_reply.started":"2022-07-17T09:39:38.224544Z","shell.execute_reply":"2022-07-17T09:39:39.110477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the graph above, passengers who spend big on Room Service, Spa, and VR Deck, are passengers who are not transported. We can see from the boxplot that in these three features (RoomService, Spa, and VRDeck), passengers who are Transported have a max value that is smaller than the max value of passengers who are not Transported.","metadata":{}},{"cell_type":"markdown","source":"<a id =\"foodcourt\"></a>\n### Food Court, Shopping Mall","metadata":{"execution":{"iopub.status.busy":"2022-07-05T11:26:22.925763Z","iopub.execute_input":"2022-07-05T11:26:22.926293Z","iopub.status.idle":"2022-07-05T11:26:22.931936Z","shell.execute_reply.started":"2022-07-05T11:26:22.926251Z","shell.execute_reply":"2022-07-05T11:26:22.930916Z"}}},{"cell_type":"code","source":"num_eda(df,\"FoodCourt\",\"Transported\",30)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:39.113730Z","iopub.execute_input":"2022-07-17T09:39:39.114498Z","iopub.status.idle":"2022-07-17T09:39:39.990472Z","shell.execute_reply.started":"2022-07-17T09:39:39.114450Z","shell.execute_reply":"2022-07-17T09:39:39.989189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_eda(df,\"ShoppingMall\",\"Transported\",30)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:39.994103Z","iopub.execute_input":"2022-07-17T09:39:39.994519Z","iopub.status.idle":"2022-07-17T09:39:40.833837Z","shell.execute_reply.started":"2022-07-17T09:39:39.994480Z","shell.execute_reply":"2022-07-17T09:39:40.832604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the graph above, there is no significant difference between Transported and Not Transported passengers regarding their spending on Food Courts and Shopping Malls.","metadata":{}},{"cell_type":"markdown","source":"## Total Bills And Percentage","metadata":{}},{"cell_type":"code","source":"def BillFeature(df):\n    data = df.copy()\n    Bills = [\"RoomService\",\"Spa\",\"VRDeck\",\"ShoppingMall\",\"FoodCourt\"]\n    data[Bills] = data[Bills].fillna(0)\n    data[\"TotalBills\"] = data[Bills].sum(1)\n    for Bill in Bills:\n        data[Bill + \"Perc\"] = data[Bill].divide(data[\"TotalBills\"])\n        data[Bill + \"Perc\"] = data[Bill + \"Perc\"].fillna(0)\n        \n    return data","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:40.835349Z","iopub.execute_input":"2022-07-17T09:39:40.835726Z","iopub.status.idle":"2022-07-17T09:39:40.843196Z","shell.execute_reply.started":"2022-07-17T09:39:40.835691Z","shell.execute_reply":"2022-07-17T09:39:40.842166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = BillFeature(df)\ndf_test = BillFeature(df_test)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:40.844604Z","iopub.execute_input":"2022-07-17T09:39:40.845430Z","iopub.status.idle":"2022-07-17T09:39:40.876310Z","shell.execute_reply.started":"2022-07-17T09:39:40.845367Z","shell.execute_reply":"2022-07-17T09:39:40.875310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_eda(df,\"TotalBills\",\"Transported\",30)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:40.878021Z","iopub.execute_input":"2022-07-17T09:39:40.878636Z","iopub.status.idle":"2022-07-17T09:39:41.760871Z","shell.execute_reply.started":"2022-07-17T09:39:40.878600Z","shell.execute_reply":"2022-07-17T09:39:41.759247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After performing EDA on the numerical data, we find that five numerical features have asymmetrical distribution. Therefore, we will normalize the numerical data, which is expected to increase the score on the model. because there is a value of 0 in the data, we can't use the box-cox normalization so we will use the Yeo-Jonhson transformation method.","metadata":{}},{"cell_type":"markdown","source":"## Categorical Features","metadata":{"execution":{"iopub.status.busy":"2022-07-05T12:50:53.062806Z","iopub.execute_input":"2022-07-05T12:50:53.06326Z","iopub.status.idle":"2022-07-05T12:50:53.067662Z","shell.execute_reply.started":"2022-07-05T12:50:53.063216Z","shell.execute_reply":"2022-07-05T12:50:53.06661Z"}}},{"cell_type":"markdown","source":"<a id = \"unique\"></a>\n### Unique Values Count","metadata":{"execution":{"iopub.status.busy":"2022-07-05T12:52:31.583562Z","iopub.execute_input":"2022-07-05T12:52:31.584256Z","iopub.status.idle":"2022-07-05T12:52:31.588241Z","shell.execute_reply.started":"2022-07-05T12:52:31.584213Z","shell.execute_reply":"2022-07-05T12:52:31.587224Z"}}},{"cell_type":"code","source":"df.select_dtypes(include = [\"object\",\"bool\"]).nunique()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:41.762564Z","iopub.execute_input":"2022-07-17T09:39:41.763369Z","iopub.status.idle":"2022-07-17T09:39:41.791065Z","shell.execute_reply.started":"2022-07-17T09:39:41.763319Z","shell.execute_reply":"2022-07-17T09:39:41.789744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the information above, several categorical features have many unique values, namely `PassengerId`, `Cabin`, and `Name`. Because of that, we will do feature engineering on these features so that we can use these features with fewer unique numbers.\n\nWhile for other category features, because the number of unique values ​​is small, we can use a one-hot encoder or label encoder.","metadata":{}},{"cell_type":"code","source":"#plot categorical Data (classification Task)\ndef cat_eda(data,feature,target):\n    fig, axes = plt.subplots(1, 2, figsize=(15, 5))\n    \n    ax1 = sns.countplot(x=feature, data=data, ax=axes[0],color = \"orange\",edgecolor = \"k\")\n    ax1.bar_label(ax1.containers[0])\n    ax1.set_title(f\"{feature} Distribution\")\n    \n    ax2 = sns.countplot(x = feature,hue = target, data = data,ax = axes[1],palette = [\"r\",'lime'],edgecolor = \"k\")\n    ax2.set_title(f\"{feature} Distribution with target\")\n    ax2.bar_label(ax2.containers[0])\n    ax2.bar_label(ax2.containers[1])","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:41.793364Z","iopub.execute_input":"2022-07-17T09:39:41.793887Z","iopub.status.idle":"2022-07-17T09:39:41.803690Z","shell.execute_reply.started":"2022-07-17T09:39:41.793837Z","shell.execute_reply":"2022-07-17T09:39:41.802403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"homeplanet\"></a>\n### Home Planet","metadata":{"execution":{"iopub.status.busy":"2022-07-05T12:22:53.092282Z","iopub.execute_input":"2022-07-05T12:22:53.092725Z","iopub.status.idle":"2022-07-05T12:22:53.099573Z","shell.execute_reply.started":"2022-07-05T12:22:53.09269Z","shell.execute_reply":"2022-07-05T12:22:53.097941Z"}}},{"cell_type":"code","source":"cat_eda(df,\"HomePlanet\",\"Transported\")","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:41.805876Z","iopub.execute_input":"2022-07-17T09:39:41.806350Z","iopub.status.idle":"2022-07-17T09:39:42.184705Z","shell.execute_reply.started":"2022-07-17T09:39:41.806301Z","shell.execute_reply":"2022-07-17T09:39:42.183183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Passengers come from 3 planets, namely Europa, Earth, and Mars. Most passengers come from Earth. The passenger with the highest percentage transported is Europa, Earth has the lowest percentage of transported passengers, while Mars has the same rate of Not Transported and Transported.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"destination\"></a>\n## Destination","metadata":{"execution":{"iopub.status.busy":"2022-07-09T03:47:40.492774Z","iopub.execute_input":"2022-07-09T03:47:40.493915Z","iopub.status.idle":"2022-07-09T03:47:40.498601Z","shell.execute_reply.started":"2022-07-09T03:47:40.493869Z","shell.execute_reply":"2022-07-09T03:47:40.497736Z"}}},{"cell_type":"code","source":"cat_eda(df,\"Destination\",\"Transported\")","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:42.186880Z","iopub.execute_input":"2022-07-17T09:39:42.187654Z","iopub.status.idle":"2022-07-17T09:39:42.584135Z","shell.execute_reply.started":"2022-07-17T09:39:42.187602Z","shell.execute_reply":"2022-07-17T09:39:42.582866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are three destination planets, namely TRAPPIST-1e, PSO J318.5-22, and 55 Cancri e. Most passengers go to TRAPPIST-1e. Passengers with the highest percentage transported are Passengers who go to 55 Cancri e, while TRAPPIST-1e is the lowest percentage of passengers who are Transported. PSO J318.5-22 has the same Not Transported and Transported rates.\n","metadata":{}},{"cell_type":"markdown","source":"### Cryo Sleep","metadata":{"execution":{"iopub.status.busy":"2022-07-05T12:38:27.108444Z","iopub.execute_input":"2022-07-05T12:38:27.109827Z","iopub.status.idle":"2022-07-05T12:38:27.114507Z"}}},{"cell_type":"code","source":"cat_eda(df,\"CryoSleep\",\"Transported\")","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:42.585781Z","iopub.execute_input":"2022-07-17T09:39:42.586972Z","iopub.status.idle":"2022-07-17T09:39:42.982041Z","shell.execute_reply.started":"2022-07-17T09:39:42.586920Z","shell.execute_reply":"2022-07-17T09:39:42.981215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Many passengers choose not to do Cryo Sleep. However, passengers who have the highest percentage of Not Transported are passengers who do not choose to do Cryo Sleep.\n","metadata":{}},{"cell_type":"markdown","source":"### VIP","metadata":{}},{"cell_type":"code","source":"cat_eda(df,\"VIP\",\"Transported\")","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:42.983269Z","iopub.execute_input":"2022-07-17T09:39:42.983926Z","iopub.status.idle":"2022-07-17T09:39:43.387290Z","shell.execute_reply.started":"2022-07-17T09:39:42.983887Z","shell.execute_reply":"2022-07-17T09:39:43.386227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on this data, it turns out that not many passengers have VIP access. Only 199 people out of 8693 have VIP access. Regarding the percentage of Transported and Not Transported, there is not much difference between passengers who have VIP access and those who don't.\n","metadata":{}},{"cell_type":"markdown","source":"## Cabin","metadata":{}},{"cell_type":"markdown","source":"Previously we know that `Cabin` has many unique values. Therefore we need feature engineering in `Cabin`.\n\nThe nomenclature explains that `Cabin` consists of three parts of information: Deck, Number, and Side. A \"/\" sign separates all three parts of information, and We will create a new feature by splitting the Cabin into these three pieces of information.","metadata":{}},{"cell_type":"code","source":"df[[\"Deck\",\"Number\",\"Side\"]] = df.Cabin.str.split(\"/\",expand = True)\ndf_test[[\"Deck\",\"Number\",\"Side\"]] = df_test.Cabin.str.split('/',expand = True)\ndf[[\"Deck\",\"Number\",\"Side\"]].nunique()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:43.389014Z","iopub.execute_input":"2022-07-17T09:39:43.389699Z","iopub.status.idle":"2022-07-17T09:39:43.433542Z","shell.execute_reply.started":"2022-07-17T09:39:43.389661Z","shell.execute_reply":"2022-07-17T09:39:43.432315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After splitting into three parts, we now get a new feature called Deck, Number, and Side, which totalled eight unique values, Number summing 1817 unique values, and Side totalling two unique values. Because of that, we only use Deck and Side features.","metadata":{}},{"cell_type":"code","source":"cat_eda(df,\"Deck\",\"Transported\")","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:43.435245Z","iopub.execute_input":"2022-07-17T09:39:43.436117Z","iopub.status.idle":"2022-07-17T09:39:43.985124Z","shell.execute_reply.started":"2022-07-17T09:39:43.436073Z","shell.execute_reply":"2022-07-17T09:39:43.981048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Cabins filled with many passengers are Cabin F and G, which are filled with around 2500 passengers, while the other cabins are filled with less than 1000 passengers.\n\nCabins B and C have a higher number of transported passengers. Cabins F and E have a higher number of Not Transported passengers than the Transported ones. In contrast, the rest of the cabins (A, D, G, T) have almost the same proportion between Transported and Not Transported.","metadata":{}},{"cell_type":"code","source":"cat_eda(df,\"Side\",\"Transported\")","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:43.986827Z","iopub.execute_input":"2022-07-17T09:39:43.987184Z","iopub.status.idle":"2022-07-17T09:39:44.316688Z","shell.execute_reply.started":"2022-07-17T09:39:43.987152Z","shell.execute_reply":"2022-07-17T09:39:44.315415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Side P is for Port, and S is Starboard. Both have the same number of passengers. However, it has a different number of Not Transported and Transported. On Side P, the number of Not Transported is more than Transported, while on Side S, the number of Transported is more than Not Transported.","metadata":{}},{"cell_type":"markdown","source":"# Passenger Id","metadata":{}},{"cell_type":"markdown","source":"In `PassengerId`,  there are many unique values. Therefore we need feature engineering on this feature.\n\n`PassengerId` has the following value format, `gggg_pp` contains two pieces of information, namely 'gggg' for the passenger number and 'pp' as the number of groups for that passenger. We will take the value of `pp` as a new feature.\n\nAfter feature engineering, we find that `Group` or `pp` has several unique values of 8.","metadata":{}},{"cell_type":"code","source":"df[[\"Number\",\"Group\"]] = df.PassengerId.str.split(\"_\",expand = True)\ndf_test[[\"Number\",\"Group\"]] = df_test.PassengerId.str.split(\"_\",expand = True) \ndf[[\"Number\",\"Group\"]].nunique()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:44.318458Z","iopub.execute_input":"2022-07-17T09:39:44.318834Z","iopub.status.idle":"2022-07-17T09:39:44.361856Z","shell.execute_reply.started":"2022-07-17T09:39:44.318798Z","shell.execute_reply":"2022-07-17T09:39:44.360952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_eda(df,\"Group\",\"Transported\")","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:44.363170Z","iopub.execute_input":"2022-07-17T09:39:44.363694Z","iopub.status.idle":"2022-07-17T09:39:44.867788Z","shell.execute_reply.started":"2022-07-17T09:39:44.363662Z","shell.execute_reply":"2022-07-17T09:39:44.866808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Name","metadata":{}},{"cell_type":"code","source":"#get the name length \ndf[\"name_length\"] = df.Name.str.len()\ndf_test[\"name_length\"] = df_test.Name.str.len()\n#get the first name and last name length\ndf[[\"first_length\",\"last_length\"]] = df.Name.str.split(\" \",expand = True)\ndf[\"first_length\"] = df.first_length.str.len()\ndf[\"last_length\"] = df.last_length.str.len()\n\ndf_test[[\"first_length\",\"last_length\"]] = df_test.Name.str.split(\" \",expand = True)\ndf_test[\"first_length\"] = df_test.first_length.str.len()\ndf_test[\"last_length\"] = df_test.last_length.str.len()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:44.872976Z","iopub.execute_input":"2022-07-17T09:39:44.873627Z","iopub.status.idle":"2022-07-17T09:39:44.946098Z","shell.execute_reply.started":"2022-07-17T09:39:44.873586Z","shell.execute_reply":"2022-07-17T09:39:44.944766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_eda(df,\"last_length\",\"Transported\")","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:44.947553Z","iopub.execute_input":"2022-07-17T09:39:44.947929Z","iopub.status.idle":"2022-07-17T09:39:45.533280Z","shell.execute_reply.started":"2022-07-17T09:39:44.947892Z","shell.execute_reply":"2022-07-17T09:39:45.531896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_eda(df,\"first_length\",\"Transported\")","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:45.534892Z","iopub.execute_input":"2022-07-17T09:39:45.535927Z","iopub.status.idle":"2022-07-17T09:39:45.948889Z","shell.execute_reply.started":"2022-07-17T09:39:45.535885Z","shell.execute_reply":"2022-07-17T09:39:45.947984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_eda(df,\"last_length\",\"Transported\")","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:45.950172Z","iopub.execute_input":"2022-07-17T09:39:45.951041Z","iopub.status.idle":"2022-07-17T09:39:46.517716Z","shell.execute_reply.started":"2022-07-17T09:39:45.951003Z","shell.execute_reply":"2022-07-17T09:39:46.516470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"corr\"></a>\n## Correlation Matrix","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(7,7))\nsns.heatmap(df.corr(method=\"spearman\"),cmap = \"bwr\" ,vmin = -1,vmax=1,annot = True,cbar = False,fmt = \".2f\")\nplt.title(\"Correlation Matrix\");","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:46.519574Z","iopub.execute_input":"2022-07-17T09:39:46.520052Z","iopub.status.idle":"2022-07-17T09:39:47.851336Z","shell.execute_reply.started":"2022-07-17T09:39:46.520005Z","shell.execute_reply":"2022-07-17T09:39:47.850147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"target\"></a>\n## Target Distribution","metadata":{}},{"cell_type":"code","source":"df.Transported.value_counts().plot(kind = \"pie\",autopct='%1.0f%%',colors = [\"r\",\"lime\" ])","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:47.852892Z","iopub.execute_input":"2022-07-17T09:39:47.853262Z","iopub.status.idle":"2022-07-17T09:39:47.961008Z","shell.execute_reply.started":"2022-07-17T09:39:47.853227Z","shell.execute_reply":"2022-07-17T09:39:47.959476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The pie chart above shows that the proportion of Transported and Not Transported is the same. Because the ratio of positive and negative classes in this dataset is balanced, we can use metric accuracy","metadata":{}},{"cell_type":"markdown","source":"<a id = \"missing\"></a>\n## Missing Value Distribution","metadata":{"execution":{"iopub.status.busy":"2022-07-05T07:33:07.718398Z","iopub.execute_input":"2022-07-05T07:33:07.718876Z","iopub.status.idle":"2022-07-05T07:33:07.72413Z","shell.execute_reply.started":"2022-07-05T07:33:07.71884Z","shell.execute_reply":"2022-07-05T07:33:07.722847Z"}}},{"cell_type":"code","source":"plt.figure(figsize=(8,4))\nsns.displot(\n    data=df.isna().melt(value_name=\"missing\"),\n    y=\"variable\",\n    hue=\"missing\",\n    multiple=\"fill\",\n    aspect=2\n)\nplt.title(\"Missing Value Proportion Each Feature\");","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:47.962975Z","iopub.execute_input":"2022-07-17T09:39:47.964534Z","iopub.status.idle":"2022-07-17T09:39:49.175298Z","shell.execute_reply.started":"2022-07-17T09:39:47.964469Z","shell.execute_reply":"2022-07-17T09:39:49.174196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The graph above shows the proportion of missing values ​​in each feature. Because the ratio is not more than 0.1, each feature with a missing value will be imputed. However, because the data has an asymmetric distribution, imputation using the mean will not be used. Therefore imputation using the median will be more suitable.","metadata":{}},{"cell_type":"markdown","source":"At the End of EDA, we removed three built-in features from the dataset, namely `PassengerId`, `Cabin`, and `Name`. However, we added 12 new features from the `Cabin` and `PassengerId` features. So the final total of the features we will use is 22.","metadata":{}},{"cell_type":"markdown","source":"<a id = \"modelling\"></a>\n# 3. Modelling","metadata":{}},{"cell_type":"markdown","source":"<a id = \"datapar\"></a>\n## Create Train set and Test set\nThe data partition we will use here is the Train-Validation-Test split, where we will divide 80% of the data into Train Sets and 20% into Test Sets.\n\nWe will use The train set in this notebook for training and evaluation using the Stratified K-Fold Cross Validation method. Meanwhile, we use the test set for evaluation at the end of the modelling.","metadata":{}},{"cell_type":"code","source":"## stratified shuffle\nX = df.drop(columns=[\"Transported\",\"Number\",\"PassengerId\",\"Cabin\",\"Name\"])\ny = df.Transported\ny = y.apply(lambda x : 1 if x == True else 0)\n\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)\nX_train.shape, X_test.shape, y_train.shape, y_test.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:49.176542Z","iopub.execute_input":"2022-07-17T09:39:49.177472Z","iopub.status.idle":"2022-07-17T09:39:49.208274Z","shell.execute_reply.started":"2022-07-17T09:39:49.177434Z","shell.execute_reply":"2022-07-17T09:39:49.206845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"prep\"></a>\n## Preprocessing\nNext, we will create a Pipeline for the Preprocessing process. Here I made two pipelines, namely pipelines for numerical features and categorical features.\n\nIn the numerical pipeline, I do two steps: first, impute the missing values, and next, transform the features using the Yeo-Johnson method. I also did two phases in the categorical pipeline, imputation and the One-Hot-Encoder to change the categorical feature into a number.","metadata":{}},{"cell_type":"code","source":"def prep_pipeline(X_train,numerical_columns = \"default\",categorical_columns = \"default\"):\n    if numerical_columns == \"default\":\n        numerical_columns = X_train.select_dtypes(include = [\"float64\"]).columns\n\n    if categorical_columns == \"default\": \n        categorical_columns = X_train.select_dtypes(include = [\"object\"]).columns  \n        \n    numeric_transformer = Pipeline(steps = [\n        (\"impute\", SimpleImputer(strategy = \"median\")),\n        (\"Transform\",PowerTransformer())\n    ])\n\n    categorical_transformer = Pipeline(steps=[\n        (\"impute\",SimpleImputer(strategy=\"most_frequent\")),\n        (\"encoder\",OneHotEncoder())\n    ])\n\n    preprocessor = ColumnTransformer(transformers = [\n        (\"numerical\",numeric_transformer,numerical_columns),\n        (\"categorical\",categorical_transformer,categorical_columns)\n    ])\n    return preprocessor\npreprocessor = prep_pipeline(X_train)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:49.209773Z","iopub.execute_input":"2022-07-17T09:39:49.210141Z","iopub.status.idle":"2022-07-17T09:39:49.224598Z","shell.execute_reply.started":"2022-07-17T09:39:49.210106Z","shell.execute_reply":"2022-07-17T09:39:49.223603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"spot_check\"></a>\n## Spot Check Model\n\nFor the training model, I use several models, both simple models to complex models. The models that I will use in this notebook are Logistics Regression, KNN, SVM, Random Forest, LGM, XGBoost, and CatBoost. I will take three candidate models that have a high Cross Validation score, and then I will do Model Tuning.","metadata":{"execution":{"iopub.status.busy":"2022-07-08T11:06:40.168457Z","iopub.execute_input":"2022-07-08T11:06:40.1689Z","iopub.status.idle":"2022-07-08T11:06:40.189689Z","shell.execute_reply.started":"2022-07-08T11:06:40.168862Z","shell.execute_reply":"2022-07-08T11:06:40.188895Z"}}},{"cell_type":"code","source":"df_model = pd.DataFrame(columns = [\"model\",\"set_data\",\"score\"])\nset_data = [\"test\",\"train\"]\nmodels = {\n        \"KNN\" : KNeighborsClassifier(),\n        \"SVM\":SVC(),\n        \"Random Forest\":RandomForestClassifier(random_state = 42,n_jobs = -1),\n        \"Logistic Regression\" : LogisticRegression(random_state = 42), \n        \"LGBM\" : LGBMClassifier(random_state = 42),\n        \"XGB\" : XGBClassifier(random_state = 42),\n        \"CatB\" : CatBoostClassifier(random_state = 42,verbose = 0)\n        }\n\nscorer = \"accuracy\"\nnum_cv = 5\ncv = StratifiedKFold(n_splits = num_cv,shuffle = True,random_state = 42)\n\nfor m in models:\n    pipeline = Pipeline([  \n    ('prep', preprocessor), \n    ('algo', models[m])\n])\n    spot_check = cross_val_score(pipeline,X_train,y_train,cv = cv,scoring = scorer,n_jobs= -1 )\n    spot_check = spot_check.mean()\n    model = pipeline.fit(X_train,y_train)\n    score = pipeline.score(X_test,y_test)\n    model_list = [m] * 2\n    tes = pd.DataFrame(list(zip(model_list,set_data,[score,spot_check])),columns = [\"model\",\"set_data\",\"score\"])\n    df_model = pd.concat([df_model,tes],ignore_index = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:39:49.226182Z","iopub.execute_input":"2022-07-17T09:39:49.227283Z","iopub.status.idle":"2022-07-17T09:40:48.447571Z","shell.execute_reply.started":"2022-07-17T09:39:49.227240Z","shell.execute_reply":"2022-07-17T09:40:48.446248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#set figsize\nplt.figure(figsize=(10, 5))\nplots = sns.barplot(x=\"model\", y=\"score\", data=df_model, ci=None,hue = \"set_data\")\nplots.set_title(\"Mean Score of Model\")\nplots.bar_label(plots.containers[0],fmt = \"%.3f\")\nplots.bar_label(plots.containers[1],fmt = \"%.3f\")\nplt.yticks(np.arange(0,1.1,step = 0.1))\nplt.ylabel(\"accuracy\")\nplt.legend();","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:40:48.450028Z","iopub.execute_input":"2022-07-17T09:40:48.450826Z","iopub.status.idle":"2022-07-17T09:40:48.733248Z","shell.execute_reply.started":"2022-07-17T09:40:48.450764Z","shell.execute_reply":"2022-07-17T09:40:48.732222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The graph above shows that the average model has a score between 0.77 - 0.80. with the highest model equal to 0.808 for the Train score and 0.808 for the test score. We will take three models with a high score which we will then tune. The selected models are LGBM, XGBoost, and CatBoost.  ","metadata":{}},{"cell_type":"markdown","source":"<a id = \"tune\"></a>\n## Model Tuning Use Bayesian Search","metadata":{"execution":{"iopub.status.busy":"2022-07-08T12:18:07.292749Z","iopub.execute_input":"2022-07-08T12:18:07.29365Z","iopub.status.idle":"2022-07-08T12:18:07.298087Z","shell.execute_reply.started":"2022-07-08T12:18:07.293606Z","shell.execute_reply":"2022-07-08T12:18:07.297003Z"}}},{"cell_type":"markdown","source":"## LGBM and XGBost","metadata":{}},{"cell_type":"code","source":"models_tune = {\n    \"CatBoost\": CatBoostClassifier(random_state=42,verbose = 0),\n    \"LGBM\": LGBMClassifier(random_state=42),\n    \"XGB\": XGBClassifier(random_state=42)\n}\n\n#model parameters\nparams_catboost = {\n    \"algo__learning_rate\": Real(low=0.001, high=1, prior='log-uniform', transform='identity'),\n    \"algo__max_depth\": Integer(low=2, high=7, transform='identity'),\n    \"algo__l2_leaf_reg\": Real(low=0.001, high=100, prior='log-uniform', transform='identity'),\n}\n\nparams_lgbm = {\n    'algo__max_depth': Integer(low=3, high=12, transform='identity'),\n    'algo__learning_rate': Real(low=0.001, high=1, prior='log-uniform', transform='identity'),\n    'algo__colsample_bytree': Real(low=0.1, high=1, transform='identity'),\n    'algo__subsample': Real(low=0.2, high=0.8, transform='identity'),\n    'algo__num_leaves': Integer(low=20, high=3000, transform='identity'),\n    'algo__reg_alpha': Real(low=0.001, high=100, prior='log-uniform', transform='identity'),\n    'algo__reg_lambda': Real(low=0.001, high=100, prior='log-uniform', transform='identity')\n}\n\nparams_xgb = { \n 'algo__gamma': Integer(low=1, high=10, prior='uniform', transform='identity'),\n 'algo__learning_rate': Real(low=0.01, high=1, prior='log-uniform', transform='identity'),\n 'algo__max_depth': Integer(low=1, high=12, prior='uniform', transform='identity'),\n 'algo__reg_alpha': Real(low=0.001, high=100, prior='log-uniform', transform='identity'),\n 'algo__reg_lambda': Real(low=0.001, high=100, prior='log-uniform', transform='identity'),\n }\n\nmodels_params = dict(zip(models_tune, [params_catboost, params_lgbm, params_xgb]))","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:40:48.735289Z","iopub.execute_input":"2022-07-17T09:40:48.736743Z","iopub.status.idle":"2022-07-17T09:40:48.780077Z","shell.execute_reply.started":"2022-07-17T09:40:48.736663Z","shell.execute_reply":"2022-07-17T09:40:48.778834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nmodel_dict = {}\n\nfor model in models_tune:\n        pipeline = Pipeline([\n                (\"prep\" , preprocessor),\n                (\"algo\",models_tune[model])\n        ])\n\n        model_dict[model] = BayesSearchCV(pipeline,models_params[model], cv=5, scoring='accuracy', n_iter=20, n_jobs=-1, verbose=0, random_state=42)\n        model_dict[model].fit(X_train, y_train)\n\n        print(model)\n        print(\"Best parameters found on training set:\")\n        print(model_dict[model].best_params_)\n        print(\"Best score found on training set, validation set, and test set :\")\n        print(model_dict[model].score(X_train, y_train), model_dict[model].best_score_, model_dict[model].score(X_test, y_test))","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:40:48.781836Z","iopub.execute_input":"2022-07-17T09:40:48.782323Z","iopub.status.idle":"2022-07-17T09:51:47.286454Z","shell.execute_reply.started":"2022-07-17T09:40:48.782285Z","shell.execute_reply":"2022-07-17T09:51:47.285155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"prediction = model_dict[\"CatBoost\"].predict(df_test)\nprediction = (prediction == 1)\nmy_submission = pd.DataFrame({'PassengerId': df_test.PassengerId, 'Transported': prediction})\n# you could use any filename. We choose submission here\nmy_submission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:52:27.054614Z","iopub.execute_input":"2022-07-17T09:52:27.055637Z","iopub.status.idle":"2022-07-17T09:52:27.182906Z","shell.execute_reply.started":"2022-07-17T09:52:27.055582Z","shell.execute_reply":"2022-07-17T09:52:27.181950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"my_submission","metadata":{"execution":{"iopub.status.busy":"2022-07-17T09:52:29.606436Z","iopub.execute_input":"2022-07-17T09:52:29.607580Z","iopub.status.idle":"2022-07-17T09:52:29.622099Z","shell.execute_reply.started":"2022-07-17T09:52:29.607531Z","shell.execute_reply":"2022-07-17T09:52:29.620570Z"},"trusted":true},"execution_count":null,"outputs":[]}]}