{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Titanic - EDA and 8 Classification using Scikit-Learn","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"**CONTENTS:**\n* [Introduction](#1)\n* [Load and Check Data](#2)\n* [Variable Description](#3)\n* [Univariate Analysis](#4)\n* [Data Analysis](#5)\n* [Missing Value](#6)\n* [Visualization](#7)\n* [Feature Engineering](#8)\n* [Model Selection](#9)\n* [Generate Prediction Value File](#10)","metadata":{}},{"cell_type":"markdown","source":"<a id = \"1\"></a><br>\n# Introduction\n\nThis project is going to predict whether a person survived or did not survived from the Titanic shipwreck.\\\n**Exploratory Data Analysis (EDA)** method is used to describe the data, view the distribution of data, compare relationships between data, clean the data and so on.\\\nThen Machine Learning methods are used to create a model which can predict each passenger survived the Titanic shipwreck or not. In this notebook, 8 different machine learning Classifier are considered (**Decision Tree, Random Forest, KNN, Gaussian Naive Bayes, Support Vector Machines, Logistic Regression, Stochastic Gradient Descent, Gradient Boosting Classifier**), and finally higher accuracy score of whom can be selected to do the prediction.","metadata":{}},{"cell_type":"markdown","source":"<a id = \"2\"></a><br>\n# Load and Check Data","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\ntrain_data = pd.read_csv('/kaggle/input/titanic/train.csv')\n# print(train_data)\n\ntest_data = pd.read_csv('/kaggle/input/titanic/test.csv')\n# print(test_data)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:27:54.308696Z","iopub.execute_input":"2022-07-11T13:27:54.309177Z","iopub.status.idle":"2022-07-11T13:27:55.226463Z","shell.execute_reply.started":"2022-07-11T13:27:54.309077Z","shell.execute_reply":"2022-07-11T13:27:55.225457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train_data\ntrain_data.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:27:58.937837Z","iopub.execute_input":"2022-07-11T13:27:58.938237Z","iopub.status.idle":"2022-07-11T13:27:58.978761Z","shell.execute_reply.started":"2022-07-11T13:27:58.938203Z","shell.execute_reply":"2022-07-11T13:27:58.977649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# test_data\ntest_data.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:28:02.743287Z","iopub.execute_input":"2022-07-11T13:28:02.744085Z","iopub.status.idle":"2022-07-11T13:28:02.766164Z","shell.execute_reply.started":"2022-07-11T13:28:02.744028Z","shell.execute_reply":"2022-07-11T13:28:02.765197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data[\"SibSp\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:29:41.935515Z","iopub.execute_input":"2022-07-11T13:29:41.936141Z","iopub.status.idle":"2022-07-11T13:29:41.947872Z","shell.execute_reply.started":"2022-07-11T13:29:41.936104Z","shell.execute_reply":"2022-07-11T13:29:41.946545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.describe()\n# test_data.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:29:44.920980Z","iopub.execute_input":"2022-07-11T13:29:44.921600Z","iopub.status.idle":"2022-07-11T13:29:44.954868Z","shell.execute_reply.started":"2022-07-11T13:29:44.921549Z","shell.execute_reply":"2022-07-11T13:29:44.954111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"3\"></a><br>\n# Variable Description","metadata":{}},{"cell_type":"code","source":"train_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:29:48.869197Z","iopub.execute_input":"2022-07-11T13:29:48.869828Z","iopub.status.idle":"2022-07-11T13:29:48.891690Z","shell.execute_reply.started":"2022-07-11T13:29:48.869791Z","shell.execute_reply":"2022-07-11T13:29:48.890674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"4\"></a><br>\n# Univariate Analysis\n\n* Categorical Variable: Survived, Pclass, Sex, Cabin, Embarked, Name, Ticket, SibSp and Parch\n* Numerical Variable: Fare, age and passengerId","metadata":{}},{"cell_type":"markdown","source":"**Categorical Variable**\n\nFirstly, define **Survived, Pclass, Sex, Embarked, SibSp, Parch** as one kind of categorical variable, since these variables can be classified in limited groups (in this project less than 8 groups). The features of these variables can be described by visualized figures.","metadata":{}},{"cell_type":"code","source":"def bar_plot(variable):\n    \n    vari = train_data[variable]\n    variValue = vari.value_counts()\n    \n    plt.figure(figsize = (9,3))\n    plt.bar(variValue.index, variValue, width = 0.3, color=\"#87CEFA\")\n    plt.xticks(variValue.index, variValue.index.values)\n    plt.ylabel(\"Frequency\")\n    plt.title(variable)\n    \n    for a, b , label in zip(variValue.index, variValue, variValue):\n        plt.text(a, b, label, ha='center', va='bottom')\n    \n    plt.show()\n\n# A function to draw the bar plot for all categorical variables.","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:29:53.067554Z","iopub.execute_input":"2022-07-11T13:29:53.067936Z","iopub.status.idle":"2022-07-11T13:29:53.074556Z","shell.execute_reply.started":"2022-07-11T13:29:53.067906Z","shell.execute_reply":"2022-07-11T13:29:53.073588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical_vari1 = [\"Survived\", \"Pclass\", \"Sex\", \"Embarked\", \"SibSp\", \"Parch\"]\n\nfor x in categorical_vari1:\n    bar_plot(x)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:29:56.521282Z","iopub.execute_input":"2022-07-11T13:29:56.522740Z","iopub.status.idle":"2022-07-11T13:29:57.423446Z","shell.execute_reply.started":"2022-07-11T13:29:56.522676Z","shell.execute_reply":"2022-07-11T13:29:57.422608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Secondly, define **Cabin, Name, Ticket** as the other kind of categorical variable, since these variables cannot be classified in limited groups (in this project more than 100 groups).","metadata":{}},{"cell_type":"code","source":"categorical_vari2 = [\"Cabin\", \"Name\", \"Ticket\"]\n\nfor x in categorical_vari2:\n    print(\"{} \\n\".format(train_data[x].value_counts()))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:30:11.337311Z","iopub.execute_input":"2022-07-11T13:30:11.337731Z","iopub.status.idle":"2022-07-11T13:30:11.351800Z","shell.execute_reply.started":"2022-07-11T13:30:11.337698Z","shell.execute_reply":"2022-07-11T13:30:11.350849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Numerical Variable**\n\nUsing numerical variabe (Fare, age and passengerId) to draw hist plot.","metadata":{}},{"cell_type":"code","source":"def hist_plot(variable):\n    \n    plt.figure(figsize = (9,3))\n    plt.hist(train_data[variable], bins = 50, color=\"#87CEFA\")\n    plt.xlabel(variable)\n    plt.ylabel(\"Frequency\")\n    plt.title(\"{} distribution with hist\".format(variable))\n    \n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:30:17.081613Z","iopub.execute_input":"2022-07-11T13:30:17.082191Z","iopub.status.idle":"2022-07-11T13:30:17.090102Z","shell.execute_reply.started":"2022-07-11T13:30:17.082141Z","shell.execute_reply":"2022-07-11T13:30:17.088480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numerical_vari = [\"Fare\", \"Age\", \"PassengerId\"]\n\nfor y in numerical_vari:\n    hist_plot(y)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:30:21.247632Z","iopub.execute_input":"2022-07-11T13:30:21.248040Z","iopub.status.idle":"2022-07-11T13:30:21.878339Z","shell.execute_reply.started":"2022-07-11T13:30:21.247998Z","shell.execute_reply":"2022-07-11T13:30:21.877164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"5\"></a><br>\n# Data Analysis\n\nWe try to find the connection between the variable and the result of the survival.\n\n* Pclass - Survived\n* Sex - Survived\n* Embarked - Survived\n* SibSp - Survived\n* Parch - Survived","metadata":{}},{"cell_type":"code","source":"# Pclass vs Survived\nPcla_Sur = train_data[[\"Pclass\", \"Survived\"]].groupby(\"Pclass\", as_index = False).mean()\n# as_index = False       # is a SQL way to show the dataframe\n\nPcla_Sur.sort_values(by=\"Survived\",ascending = False)     # sort the value by \"Survived\"","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:30:26.181521Z","iopub.execute_input":"2022-07-11T13:30:26.181912Z","iopub.status.idle":"2022-07-11T13:30:26.203245Z","shell.execute_reply.started":"2022-07-11T13:30:26.181879Z","shell.execute_reply":"2022-07-11T13:30:26.202407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Sex vs Survived\nSex_Sur = train_data[[\"Sex\", \"Survived\"]].groupby(\"Sex\", as_index = False).mean()\nSex_Sur.sort_values(by=\"Survived\",ascending = False)     # sort the value by \"Survived\"","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:30:29.910525Z","iopub.execute_input":"2022-07-11T13:30:29.910943Z","iopub.status.idle":"2022-07-11T13:30:29.928007Z","shell.execute_reply.started":"2022-07-11T13:30:29.910909Z","shell.execute_reply":"2022-07-11T13:30:29.927013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Embarked vs Survived\nEmba_Sur = train_data[[\"Embarked\", \"Survived\"]].groupby(\"Embarked\", as_index = False).mean()\nEmba_Sur.sort_values(by=\"Survived\",ascending = False)     # sort the value by \"Survived\"","metadata":{"execution":{"iopub.status.busy":"2022-07-11T06:50:26.899135Z","iopub.execute_input":"2022-07-11T06:50:26.899511Z","iopub.status.idle":"2022-07-11T06:50:26.916312Z","shell.execute_reply.started":"2022-07-11T06:50:26.89948Z","shell.execute_reply":"2022-07-11T06:50:26.915387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# SibSp vs Survived\nSib_Sur = train_data[[\"SibSp\", \"Survived\"]].groupby(\"SibSp\", as_index = False).mean()\nSib_Sur.sort_values(by=\"Survived\",ascending = False)     # sort the value by \"Survived\"","metadata":{"execution":{"iopub.status.busy":"2022-07-11T06:50:30.019443Z","iopub.execute_input":"2022-07-11T06:50:30.019785Z","iopub.status.idle":"2022-07-11T06:50:30.035086Z","shell.execute_reply.started":"2022-07-11T06:50:30.019756Z","shell.execute_reply":"2022-07-11T06:50:30.034277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Parch vs Survived\nPar_Sur = train_data[[\"Parch\", \"Survived\"]].groupby(\"Parch\", as_index = False).mean()\nPar_Sur.sort_values(by=\"Survived\",ascending = False)     # sort the value by \"Survived\"","metadata":{"execution":{"iopub.status.busy":"2022-07-11T06:50:34.084797Z","iopub.execute_input":"2022-07-11T06:50:34.085179Z","iopub.status.idle":"2022-07-11T06:50:34.102319Z","shell.execute_reply.started":"2022-07-11T06:50:34.085147Z","shell.execute_reply":"2022-07-11T06:50:34.101283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"6\"></a><br>\n# Missing Value\n\n* Find Missing Value\n* Fill Missing Value","metadata":{}},{"cell_type":"code","source":"len(train_data)     # show the length of the train data\n#len(test_data)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:30:36.214288Z","iopub.execute_input":"2022-07-11T13:30:36.214732Z","iopub.status.idle":"2022-07-11T13:30:36.222437Z","shell.execute_reply.started":"2022-07-11T13:30:36.214695Z","shell.execute_reply":"2022-07-11T13:30:36.220988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Combine train data with test data, show in dataframe\ntrain_test_combi = pd.concat([train_data, test_data], ignore_index=True)\ntrain_test_combi","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:30:38.659723Z","iopub.execute_input":"2022-07-11T13:30:38.660097Z","iopub.status.idle":"2022-07-11T13:30:38.703449Z","shell.execute_reply.started":"2022-07-11T13:30:38.660066Z","shell.execute_reply":"2022-07-11T13:30:38.702257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Find Missing Value**","metadata":{}},{"cell_type":"code","source":"train_test_combi.columns[train_test_combi.isnull().any()]\n# train_test_combi.isnull().any()     # bool type, judge True/False","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:30:43.396573Z","iopub.execute_input":"2022-07-11T13:30:43.396973Z","iopub.status.idle":"2022-07-11T13:30:43.408280Z","shell.execute_reply.started":"2022-07-11T13:30:43.396942Z","shell.execute_reply":"2022-07-11T13:30:43.407125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_test_combi.isnull().sum()     # Show the number of \"null\" in each variable.","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:30:46.946469Z","iopub.execute_input":"2022-07-11T13:30:46.946874Z","iopub.status.idle":"2022-07-11T13:30:46.958452Z","shell.execute_reply.started":"2022-07-11T13:30:46.946843Z","shell.execute_reply":"2022-07-11T13:30:46.957280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Fill Missing Value**\n\n* Age has 263 missing value\n* Embarked has 2 missing value\n* Fare only has 1 missing value\n* Cabin has 1014 missing value (since this variable has less useful information, it will be handled later)","metadata":{}},{"cell_type":"markdown","source":"**1. Fill Embarked Value**","metadata":{}},{"cell_type":"code","source":"# Find the sample which has the missing value in Embarked\ntrain_test_combi[train_test_combi[\"Embarked\"].isnull()]","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:30:50.815879Z","iopub.execute_input":"2022-07-11T13:30:50.816339Z","iopub.status.idle":"2022-07-11T13:30:50.835031Z","shell.execute_reply.started":"2022-07-11T13:30:50.816296Z","shell.execute_reply":"2022-07-11T13:30:50.834234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The values of \"Pclass\", \"Ticket\", \"Fare\", \"Cabin\" of these two people are all the same, we suppose that they came from the same Port of Embarkation. Since \"Fare\" is a Numerical Variable, we use it to find the relationship between Embarked and Fare. From the plot below, the range of Fare from Cherbourg is large, it can contain the Fare 80.0, so we decided using **C** as the Missing Value of these two samples.","metadata":{}},{"cell_type":"code","source":"# use box plot to show the range of Fare from each Port of Embarkation\ntrain_test_combi.boxplot(column = \"Fare\", by = \"Embarked\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:30:55.200634Z","iopub.execute_input":"2022-07-11T13:30:55.201198Z","iopub.status.idle":"2022-07-11T13:30:55.376379Z","shell.execute_reply.started":"2022-07-11T13:30:55.201163Z","shell.execute_reply":"2022-07-11T13:30:55.375455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_test_combi[\"Embarked\"] = train_test_combi[\"Embarked\"].fillna(\"C\")     \n# fill the missing value\n\ntrain_test_combi[train_test_combi[\"Embarked\"].isnull()]     \n# no NaN value in Embarked","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:30:58.315629Z","iopub.execute_input":"2022-07-11T13:30:58.316447Z","iopub.status.idle":"2022-07-11T13:30:58.333737Z","shell.execute_reply.started":"2022-07-11T13:30:58.316368Z","shell.execute_reply":"2022-07-11T13:30:58.332557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**2. Fill Fare Value**","metadata":{}},{"cell_type":"code","source":"# Find the sample which has the missing value in Fare\ntrain_test_combi[train_test_combi[\"Fare\"].isnull()]","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:31:02.817235Z","iopub.execute_input":"2022-07-11T13:31:02.817695Z","iopub.status.idle":"2022-07-11T13:31:02.837052Z","shell.execute_reply.started":"2022-07-11T13:31:02.817657Z","shell.execute_reply":"2022-07-11T13:31:02.835950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The Fare of this sample is missing, we can replace with mean. To be precisely and rigorously, using the sample whose Pclass value is 3.","metadata":{}},{"cell_type":"code","source":"mean_Fare = np.mean(train_test_combi[train_test_combi[\"Pclass\"] == 3][\"Fare\"])     \n# choose Pclass value is 3\n\ntrain_test_combi[\"Fare\"] = train_test_combi[\"Fare\"].fillna(mean_Fare)    \n# fill the missing value\n\ntrain_test_combi[train_test_combi[\"Fare\"].isnull()]     \n# no NaN value in Embarked","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:31:05.896343Z","iopub.execute_input":"2022-07-11T13:31:05.896790Z","iopub.status.idle":"2022-07-11T13:31:05.932920Z","shell.execute_reply.started":"2022-07-11T13:31:05.896754Z","shell.execute_reply":"2022-07-11T13:31:05.931947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**3. Fill Age Value**\n\nSince there is more missing value for 'Age' variable, we will fill it with the help of further more analysis and data visualization.","metadata":{}},{"cell_type":"markdown","source":"<a id = \"7\"></a><br>\n# Visualization","metadata":{}},{"cell_type":"markdown","source":"**1. Correlation Between SibSp -- Parch -- Age -- Fare -- Survived**\n\nWe use heatmap to explore the correlation between numbers of variable to Survived.","metadata":{}},{"cell_type":"code","source":"list1 = [\"SibSp\", \"Parch\", \"Age\", \"Fare\", \"Survived\"]\nsns.heatmap(train_test_combi[list1].corr(), cmap = \"Oranges\", annot = True, fmt = \".2f\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:31:10.441305Z","iopub.execute_input":"2022-07-11T13:31:10.442184Z","iopub.status.idle":"2022-07-11T13:31:10.720595Z","shell.execute_reply.started":"2022-07-11T13:31:10.442137Z","shell.execute_reply":"2022-07-11T13:31:10.719712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the plot above, we can find that Fare feature seems to have correlation with Survived feature (0.26).","metadata":{}},{"cell_type":"markdown","source":"**2. Correlation Between SibSp -- Survived**","metadata":{}},{"cell_type":"code","source":"# use catplot() to describe the classification data.\nSibSp_plot = sns.catplot(x = \"SibSp\", y = \"Survived\", data = train_test_combi, \\\n                         kind = \"bar\", height = 3.5)\nSibSp_plot.set_ylabels(\"Survived Probability\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:31:14.507922Z","iopub.execute_input":"2022-07-11T13:31:14.508368Z","iopub.status.idle":"2022-07-11T13:31:14.866536Z","shell.execute_reply.started":"2022-07-11T13:31:14.508327Z","shell.execute_reply":"2022-07-11T13:31:14.865541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the bar plot above, we found that\n* Having a lot of SibSp have less chance to survive.\n* If sibsp == 0 or 1 or 2, passenger has more chance to survive.","metadata":{}},{"cell_type":"markdown","source":"**3. Correlation Between Parch -- Survived**","metadata":{}},{"cell_type":"code","source":"Parch_plot = sns.catplot(x = \"Parch\", y = \"Survived\", data = train_test_combi, \\\n                         kind = \"bar\", height = 3.5)\nParch_plot.set_ylabels(\"Survived Probability\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:31:18.632068Z","iopub.execute_input":"2022-07-11T13:31:18.632505Z","iopub.status.idle":"2022-07-11T13:31:19.118473Z","shell.execute_reply.started":"2022-07-11T13:31:18.632468Z","shell.execute_reply":"2022-07-11T13:31:19.117467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the bar plot above, we found that\n* Small familes have more chance to survive.\n* There is a std in survival of passenger with parch = 3.","metadata":{}},{"cell_type":"markdown","source":"**4. Correlation Between Pclass -- Survived**","metadata":{}},{"cell_type":"code","source":"Pclass_plot = sns.catplot(x = \"Pclass\", y = \"Survived\", data = train_test_combi, \\\n                          kind = \"bar\", height = 3.5)\nPclass_plot.set_ylabels(\"Survived Probability\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:31:22.917068Z","iopub.execute_input":"2022-07-11T13:31:22.917562Z","iopub.status.idle":"2022-07-11T13:31:23.209027Z","shell.execute_reply.started":"2022-07-11T13:31:22.917519Z","shell.execute_reply":"2022-07-11T13:31:23.207738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the bar plot above, we found that\n\n* Upper class have more chance to survive.","metadata":{}},{"cell_type":"markdown","source":"**5. Correlation Between Sex -- Survived**","metadata":{}},{"cell_type":"code","source":"Sex_plot = sns.catplot(x = \"Sex\", y = \"Survived\", data = train_test_combi, \\\n                       kind = \"bar\", height = 3.5)\nSex_plot.set_ylabels(\"Survived Probability\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:31:26.369139Z","iopub.execute_input":"2022-07-11T13:31:26.369552Z","iopub.status.idle":"2022-07-11T13:31:26.596187Z","shell.execute_reply.started":"2022-07-11T13:31:26.369518Z","shell.execute_reply":"2022-07-11T13:31:26.595267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the bar plot above, we found that\n\n* Female has a high survival rate.","metadata":{}},{"cell_type":"markdown","source":"**6. Correlation Between Age -- Survived**","metadata":{}},{"cell_type":"code","source":"Age_plot = sns.FacetGrid(train_test_combi, col = \"Survived\")     \n# Number of the grid, here is 2 \n\nAge_plot.map(sns.histplot, \"Age\", kde = True, bins = 30)     \n# map(function, iterable), use function on each item\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:31:29.805879Z","iopub.execute_input":"2022-07-11T13:31:29.806300Z","iopub.status.idle":"2022-07-11T13:31:30.249539Z","shell.execute_reply.started":"2022-07-11T13:31:29.806256Z","shell.execute_reply":"2022-07-11T13:31:30.248523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the hist plot above, we found that\n* age <= 10 has a high survival rate,\n* oldest passengers (80) survived,\n* large number of 20 years old did not survive,\n* most passengers are in 15-35 age range,\n* use age feature in training,\n* use age distribution for missing value of age.","metadata":{}},{"cell_type":"markdown","source":"**7. Correlation Amoung Pclass -- Survived -- Age**","metadata":{}},{"cell_type":"code","source":"Pclass_Age_plot = sns.FacetGrid(train_test_combi, col = \"Survived\", row = \"Pclass\",height = 3)\n# Number of the grid, here is 6.\n# row and column show the classification of \"Pclass\"(1,2,3) and \"Survived\"(0,1) respectively.\n\nPclass_Age_plot.map(sns.histplot, \"Age\", kde = True, bins = 25)     \n# map(function, iterable), use function on each item\n\nPclass_Age_plot.add_legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:31:33.550050Z","iopub.execute_input":"2022-07-11T13:31:33.550504Z","iopub.status.idle":"2022-07-11T13:31:35.159221Z","shell.execute_reply.started":"2022-07-11T13:31:33.550464Z","shell.execute_reply":"2022-07-11T13:31:35.158092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the hist plot above, we found that\n* pclass is important feature for model training.","metadata":{}},{"cell_type":"markdown","source":"**8. Correlation Amoung Embarked -- Sex -- Pclass -- Survived**","metadata":{}},{"cell_type":"code","source":"Embarked_Sex_Pclass_plot = sns.FacetGrid(train_test_combi, row = \"Embarked\", height = 2.5)\n# Number of the grid, here is 3.\n\nEmbarked_Sex_Pclass_plot.map(sns.pointplot, \"Pclass\", \"Survived\", \"Sex\")     \n# map(function, iterable), use function on each item.\n# sns.pointplot(x = \"Pclass\", y = \"Survived\", hue = \"Sex\"), \n# hue shows the variable in different colour.\n\nEmbarked_Sex_Pclass_plot.add_legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:31:39.053720Z","iopub.execute_input":"2022-07-11T13:31:39.054126Z","iopub.status.idle":"2022-07-11T13:31:40.184466Z","shell.execute_reply.started":"2022-07-11T13:31:39.054093Z","shell.execute_reply":"2022-07-11T13:31:40.183586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the hist plot above, we found that\n* Female passengers have a high survival rate from Southampton and Queenstown than male,\n* Males have better survival rate from Cherbourg,\n* Embarked and Sex is important feature for model training.","metadata":{}},{"cell_type":"markdown","source":"**9. Correlation Amoung Embarked -- Sex -- Fare -- Survived**","metadata":{}},{"cell_type":"code","source":"Embarked_Sex_Fare_plot = sns.FacetGrid(train_test_combi, col = \"Survived\", \\\n                                       row = \"Embarked\", height = 3)\n# Number of the grid, here is 6.\n# row and column show the classification of \"Embarked\"(S,C,Q) and \"Survived\"(0,1) respectively.\n\nEmbarked_Sex_Fare_plot.map(sns.barplot, \"Sex\", \"Fare\")     \n# map(function, iterable), use function on each item\n\nEmbarked_Sex_Fare_plot.add_legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:31:44.025938Z","iopub.execute_input":"2022-07-11T13:31:44.026353Z","iopub.status.idle":"2022-07-11T13:31:45.215088Z","shell.execute_reply.started":"2022-07-11T13:31:44.026316Z","shell.execute_reply":"2022-07-11T13:31:45.214151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the barplot above, we found that\n* Passsengers who pay higher fare have better survival.\n* Fare can be used as categorical for training.","metadata":{}},{"cell_type":"markdown","source":"**Fill Missing Value: Age Feature**\n\nAge has 263 missing value, we can use age distribution for missing value of age.","metadata":{}},{"cell_type":"code","source":"train_test_combi[train_test_combi[\"Age\"].isnull()]","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:31:49.209051Z","iopub.execute_input":"2022-07-11T13:31:49.209476Z","iopub.status.idle":"2022-07-11T13:31:49.243570Z","shell.execute_reply.started":"2022-07-11T13:31:49.209440Z","shell.execute_reply":"2022-07-11T13:31:49.242291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.catplot(x = \"Sex\", y = \"Age\", data = train_test_combi, kind = \"box\", height = 3)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:31:53.018158Z","iopub.execute_input":"2022-07-11T13:31:53.018761Z","iopub.status.idle":"2022-07-11T13:31:53.193476Z","shell.execute_reply.started":"2022-07-11T13:31:53.018723Z","shell.execute_reply":"2022-07-11T13:31:53.192514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Sex feature is not informative for age prediction, the age distribution seems to be same.","metadata":{}},{"cell_type":"code","source":"sns.catplot(x = \"Sex\", y = \"Age\", hue = \"Pclass\", data = train_test_combi, \\\n            kind = \"box\", height = 4)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:31:55.785224Z","iopub.execute_input":"2022-07-11T13:31:55.787529Z","iopub.status.idle":"2022-07-11T13:31:56.162875Z","shell.execute_reply.started":"2022-07-11T13:31:55.787481Z","shell.execute_reply":"2022-07-11T13:31:56.161689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems like different Pclass has different range of passengers in age. 1st class passengers are older than the 2nd, and the passengers in 2nd are older than 3rd class.","metadata":{}},{"cell_type":"code","source":"sns.catplot(x = \"SibSp\", y = \"Age\", data = train_test_combi, kind = \"box\", height = 4)\nplt.show()\n\nsns.catplot(x = \"Parch\", y = \"Age\", data = train_test_combi, kind = \"box\", height = 4)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:31:59.104907Z","iopub.execute_input":"2022-07-11T13:31:59.105330Z","iopub.status.idle":"2022-07-11T13:31:59.668569Z","shell.execute_reply.started":"2022-07-11T13:31:59.105294Z","shell.execute_reply":"2022-07-11T13:31:59.667385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list2 = [ \"Age\", \"Sex\", \"SibSp\", \"Parch\", \"Pclass\"]\nsns.heatmap(train_test_combi[list2].corr(), cmap = \"Greens\", annot = True, fmt = \".2f\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:32:02.990980Z","iopub.execute_input":"2022-07-11T13:32:02.991689Z","iopub.status.idle":"2022-07-11T13:32:03.230303Z","shell.execute_reply.started":"2022-07-11T13:32:02.991644Z","shell.execute_reply":"2022-07-11T13:32:03.229215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the plot above, we can find that Sex feature is not correlated with Age feature, while Pclass, SibSp and Parch seem to have correlation with Age feature.","metadata":{}},{"cell_type":"code","source":"train_test_combi[train_test_combi[\"Age\"].isnull()].index","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:32:06.926694Z","iopub.execute_input":"2022-07-11T13:32:06.927141Z","iopub.status.idle":"2022-07-11T13:32:06.936458Z","shell.execute_reply.started":"2022-07-11T13:32:06.927098Z","shell.execute_reply":"2022-07-11T13:32:06.935544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# list(train_test_combi[train_test_combi[\"Age\"].isnull()].index)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T14:08:08.015363Z","iopub.execute_input":"2022-07-11T14:08:08.016005Z","iopub.status.idle":"2022-07-11T14:08:08.019920Z","shell.execute_reply.started":"2022-07-11T14:08:08.015968Z","shell.execute_reply":"2022-07-11T14:08:08.018944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_test_combi['Age'][(train_test_combi['SibSp'] == train_test_combi.iloc[5]['SibSp']) \\\n& (train_test_combi['Parch'] == train_test_combi.iloc[5]['Parch']) \\\n& (train_test_combi['Pclass'] == train_test_combi.iloc[5]['Pclass'])]","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:32:14.942534Z","iopub.execute_input":"2022-07-11T13:32:14.943305Z","iopub.status.idle":"2022-07-11T13:32:14.956698Z","shell.execute_reply.started":"2022-07-11T13:32:14.943262Z","shell.execute_reply":"2022-07-11T13:32:14.955494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"index_nan_age = list(train_test_combi[train_test_combi['Age'].isnull()].index)\n# show the index of the sample which has NaN Age value, transform datatype to list.\n\nfor i in index_nan_age:\n    age_pred = train_test_combi['Age'] \\\n               [(train_test_combi['SibSp'] == train_test_combi.iloc[i]['SibSp']) & \\\n                (train_test_combi['Parch'] == train_test_combi.iloc[i]['Parch']) & \\\n                (train_test_combi['Pclass'] ==train_test_combi.iloc[i]['Pclass'])].median()\n    age_med = train_test_combi['Age'].median()\n    if not np.isnan(age_pred):\n        train_test_combi['Age'].iloc[i] = age_pred\n    else:\n        train_test_combi['Age'].iloc[i] = age_med","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:32:18.067507Z","iopub.execute_input":"2022-07-11T13:32:18.067932Z","iopub.status.idle":"2022-07-11T13:32:18.766593Z","shell.execute_reply.started":"2022-07-11T13:32:18.067896Z","shell.execute_reply":"2022-07-11T13:32:18.765553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_test_combi[train_test_combi['Age'].isnull()]","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:32:33.860590Z","iopub.execute_input":"2022-07-11T13:32:33.861208Z","iopub.status.idle":"2022-07-11T13:32:33.875180Z","shell.execute_reply.started":"2022-07-11T13:32:33.861169Z","shell.execute_reply":"2022-07-11T13:32:33.873956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_test_combi.isnull().sum() ","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:32:40.703618Z","iopub.execute_input":"2022-07-11T13:32:40.704359Z","iopub.status.idle":"2022-07-11T13:32:40.714884Z","shell.execute_reply.started":"2022-07-11T13:32:40.704298Z","shell.execute_reply":"2022-07-11T13:32:40.713968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"8\"></a><br>\n# Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"The dataset still contains categorical data, but modeling needs something numerical, so we need to do some conversion to get numerical value of these data.","metadata":{}},{"cell_type":"markdown","source":"**Cabin Feature**\n\nAt first, we'll drop the 'Cabin' feature, since not a lot more useful information can be extracted from it, and it has lots of missing data.","metadata":{}},{"cell_type":"code","source":"# Drop the 'Cabin' feature, not a lot more useful information can be extracted from it, \n# and it has lots of missing data.\n\ntrain_test_combi = train_test_combi.drop(['Cabin'], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:32:47.382400Z","iopub.execute_input":"2022-07-11T13:32:47.383036Z","iopub.status.idle":"2022-07-11T13:32:47.389441Z","shell.execute_reply.started":"2022-07-11T13:32:47.382998Z","shell.execute_reply":"2022-07-11T13:32:47.388491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Ticket Feature**\n\nWe'll also drop the 'Ticket' feature, since not a lot more useful information can be extracted from it.","metadata":{}},{"cell_type":"code","source":"# Drop the 'Ticket' feature, not a lot more useful information can be extracted from it.\n\ntrain_test_combi = train_test_combi.drop(['Ticket'], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:32:50.970708Z","iopub.execute_input":"2022-07-11T13:32:50.971544Z","iopub.status.idle":"2022-07-11T13:32:50.979355Z","shell.execute_reply.started":"2022-07-11T13:32:50.971479Z","shell.execute_reply":"2022-07-11T13:32:50.978084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_test_combi","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:32:54.416564Z","iopub.execute_input":"2022-07-11T13:32:54.417353Z","iopub.status.idle":"2022-07-11T13:32:54.442141Z","shell.execute_reply.started":"2022-07-11T13:32:54.417309Z","shell.execute_reply":"2022-07-11T13:32:54.440972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Name Feature**\n\nIn this section, let's handle the information about 'Name' Feature. 'Name' Feature contains categorical data, we need at first extract a title for each name in the dataset (such as 'Mr', 'Miss', 'Mrs'...), then in order to doing classification reasonable and limited, we need to replace various titles with more common names. Finally, all the categorical titles should be converted into numerical value to support further analysis.","metadata":{}},{"cell_type":"code","source":"# Extract a title for each Name in the train and test datasets\n\ntrain_test_combi['Title'] = train_test_combi.Name.str.extract(' ([A-Za-z]+)\\.', expand=False)\n\npd.crosstab(train_test_combi['Title'], train_test_combi['Sex'])","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:32:59.837658Z","iopub.execute_input":"2022-07-11T13:32:59.838493Z","iopub.status.idle":"2022-07-11T13:32:59.878819Z","shell.execute_reply.started":"2022-07-11T13:32:59.838430Z","shell.execute_reply":"2022-07-11T13:32:59.877665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Replace various titles with more common names\n\ntrain_test_combi['Title'] = train_test_combi['Title'].replace(['Capt', 'Col', 'Don', \\\n                            'Dona', 'Dr', 'Jonkheer', 'Major', 'Rev'], 'Rare')\n    \ntrain_test_combi['Title'] = train_test_combi['Title'].replace(['Countess', 'Lady', 'Sir'],\\\n                                                              'Royal')\n\ntrain_test_combi['Title'] = train_test_combi['Title'].replace(['Mlle', 'Ms'], 'Miss')\n\ntrain_test_combi['Title'] = train_test_combi['Title'].replace('Mme', 'Mrs')\n\ntrain_test_combi[['Title', 'Survived']].groupby(['Title'], as_index = False).mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:33:04.217894Z","iopub.execute_input":"2022-07-11T13:33:04.218638Z","iopub.status.idle":"2022-07-11T13:33:04.244797Z","shell.execute_reply.started":"2022-07-11T13:33:04.218591Z","shell.execute_reply":"2022-07-11T13:33:04.243470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Map each of the title groups to a numerical value\n\ntitle_mapping = {'Mr': 1, 'Miss': 2, 'Mrs': 3, 'Master': 4, 'Royal': 5, 'Rare': 6}\n\ntrain_test_combi['Title'] = train_test_combi['Title'].map(title_mapping)\n#train_test_combi['Title'] = train_test_combi['Title'].fillna(0)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:33:08.396779Z","iopub.execute_input":"2022-07-11T13:33:08.397565Z","iopub.status.idle":"2022-07-11T13:33:08.405730Z","shell.execute_reply.started":"2022-07-11T13:33:08.397514Z","shell.execute_reply":"2022-07-11T13:33:08.404848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_test_combi.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:33:11.713390Z","iopub.execute_input":"2022-07-11T13:33:11.714149Z","iopub.status.idle":"2022-07-11T13:33:11.733945Z","shell.execute_reply.started":"2022-07-11T13:33:11.714105Z","shell.execute_reply":"2022-07-11T13:33:11.732474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_test_combi[train_test_combi[\"Title\"].isnull()]","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:33:16.898174Z","iopub.execute_input":"2022-07-11T13:33:16.898580Z","iopub.status.idle":"2022-07-11T13:33:16.912166Z","shell.execute_reply.started":"2022-07-11T13:33:16.898546Z","shell.execute_reply":"2022-07-11T13:33:16.911350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We can drop the name feature now that we've extracted the titles\n\ntrain_test_combi = train_test_combi.drop(['Name'], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:33:19.605327Z","iopub.execute_input":"2022-07-11T13:33:19.605912Z","iopub.status.idle":"2022-07-11T13:33:19.611389Z","shell.execute_reply.started":"2022-07-11T13:33:19.605876Z","shell.execute_reply":"2022-07-11T13:33:19.610558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Sex Feature**\n\nWe need to map each of the Sex groups to a numerical value.","metadata":{}},{"cell_type":"code","source":"# Map each of the Sex groups to a numerical value, Male is 0, Female is 1\n\nsex_mapping = {'male': 0, 'female': 1}\n\ntrain_test_combi['Sex'] = train_test_combi['Sex'].map(sex_mapping)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:33:23.202545Z","iopub.execute_input":"2022-07-11T13:33:23.203168Z","iopub.status.idle":"2022-07-11T13:33:23.210701Z","shell.execute_reply.started":"2022-07-11T13:33:23.203116Z","shell.execute_reply":"2022-07-11T13:33:23.209643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Embarked Feature**\n\nWe need to map each of the Embarked groups to a numerical value.","metadata":{}},{"cell_type":"code","source":"# Map each of the Embarked groups to a numerical value, 'S' is 1, 'C' is 2, 'Q' is 3\n\nembarked_mapping = {'S': 1, 'C': 2, 'Q': 3}\n\ntrain_test_combi['Embarked'] = train_test_combi['Embarked'].map(embarked_mapping)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:33:27.624658Z","iopub.execute_input":"2022-07-11T13:33:27.625228Z","iopub.status.idle":"2022-07-11T13:33:27.631333Z","shell.execute_reply.started":"2022-07-11T13:33:27.625190Z","shell.execute_reply":"2022-07-11T13:33:27.630528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_test_combi.head(15)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:33:30.458099Z","iopub.execute_input":"2022-07-11T13:33:30.458559Z","iopub.status.idle":"2022-07-11T13:33:30.478226Z","shell.execute_reply.started":"2022-07-11T13:33:30.458520Z","shell.execute_reply":"2022-07-11T13:33:30.476942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Divided the processed dataset into two datasets, train data and test data.\npro_train_data = train_test_combi[0:891]\npro_test_data = train_test_combi[891:1309]\npro_train_data.info()\npro_test_data","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:33:34.723979Z","iopub.execute_input":"2022-07-11T13:33:34.724364Z","iopub.status.idle":"2022-07-11T13:33:34.757583Z","shell.execute_reply.started":"2022-07-11T13:33:34.724328Z","shell.execute_reply":"2022-07-11T13:33:34.756538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"9\"></a><br>\n# Model Selection","metadata":{}},{"cell_type":"markdown","source":"**Testing Different Models**\n\nThe accuracy of these following models will be tested using training data:\n\n* Decision Tree Classifier\n* Random Forest Classifier\n* KNN or k-Nearest Neighbors\n* Gaussian Naive Bayes\n* Support Vector Machines\n* Logistic Regression\n* Stochastic Gradient Descent\n* Gradient Boosting Classifier","metadata":{}},{"cell_type":"markdown","source":"**Splitting the Training Data**\n\nWe will use part of the training data (at about 30%) to test the accuracy of the models.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nX_Variable = pro_train_data.drop(['Survived', 'PassengerId'], axis = 1)\ny_Target = pro_train_data['Survived']\n\nX_train, X_valid, y_train, y_valid = train_test_split(X_Variable, y_Target, \\\n                                                      test_size = 0.3, random_state = 0)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:33:41.431997Z","iopub.execute_input":"2022-07-11T13:33:41.432442Z","iopub.status.idle":"2022-07-11T13:33:41.669582Z","shell.execute_reply.started":"2022-07-11T13:33:41.432392Z","shell.execute_reply":"2022-07-11T13:33:41.668665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**The accuracy of each model**\n\nFor each model, we set the model classifier, and then fit it with 70% of our training data, predict for 30% of the validation data, test and check the accuracy.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import accuracy_score","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:33:44.952847Z","iopub.execute_input":"2022-07-11T13:33:44.953280Z","iopub.status.idle":"2022-07-11T13:33:44.958626Z","shell.execute_reply.started":"2022-07-11T13:33:44.953243Z","shell.execute_reply":"2022-07-11T13:33:44.957623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Decision Tree Classifier\nfrom sklearn.tree import DecisionTreeClassifier\n\nmodel_DT = DecisionTreeClassifier()\nmodel_DT.fit(X_train, y_train)\ny_pred_DT = model_DT.predict(X_valid)\naccuracy_DT = round(accuracy_score(y_pred_DT, y_valid) * 100, 2)\nprint('The accuracy score of Decision Tree model is:', accuracy_DT)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:33:47.609353Z","iopub.execute_input":"2022-07-11T13:33:47.610167Z","iopub.status.idle":"2022-07-11T13:33:47.824072Z","shell.execute_reply.started":"2022-07-11T13:33:47.610125Z","shell.execute_reply":"2022-07-11T13:33:47.821871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Random Forest Classifier\nfrom sklearn.ensemble import RandomForestClassifier\n\nmodel_RF = RandomForestClassifier()\nmodel_RF.fit(X_train, y_train)\ny_pred_RF = model_RF.predict(X_valid)\naccuracy_RF = round(accuracy_score(y_pred_RF, y_valid) * 100, 2)\nprint('The accuracy score of Random Forest model is:', accuracy_RF)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:33:50.848909Z","iopub.execute_input":"2022-07-11T13:33:50.849716Z","iopub.status.idle":"2022-07-11T13:33:51.265247Z","shell.execute_reply.started":"2022-07-11T13:33:50.849655Z","shell.execute_reply":"2022-07-11T13:33:51.263999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# KNN or k-Nearest Neighbors\nfrom sklearn.neighbors import KNeighborsClassifier\n\nmodel_KNN = KNeighborsClassifier()\nmodel_KNN.fit(X_train, y_train)\ny_pred_KNN = model_KNN.predict(X_valid)\naccuracy_KNN = round(accuracy_score(y_pred_KNN, y_valid) * 100, 2)\nprint('The accuracy score of KNN model is:', accuracy_KNN)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:33:53.826764Z","iopub.execute_input":"2022-07-11T13:33:53.828484Z","iopub.status.idle":"2022-07-11T13:33:53.854020Z","shell.execute_reply.started":"2022-07-11T13:33:53.828396Z","shell.execute_reply":"2022-07-11T13:33:53.852861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Gaussian Naive Bayes\nfrom sklearn.naive_bayes import GaussianNB\n\nmodel_NB = GaussianNB()\nmodel_NB.fit(X_train, y_train)\ny_pred_NB = model_NB.predict(X_valid)\naccuracy_NB = round(accuracy_score(y_pred_NB, y_valid) * 100, 2)\nprint('The accuracy score of Gaussian Naive Bayes model is:', accuracy_NB)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:33:57.213637Z","iopub.execute_input":"2022-07-11T13:33:57.214375Z","iopub.status.idle":"2022-07-11T13:33:57.230901Z","shell.execute_reply.started":"2022-07-11T13:33:57.214329Z","shell.execute_reply":"2022-07-11T13:33:57.229800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Support Vector Machines\nfrom sklearn.svm import SVC\n\nmodel_svc = SVC(kernel = 'linear')\nmodel_svc.fit(X_train, y_train)\ny_pred_svc = model_svc.predict(X_valid)\naccuracy_svc = round(accuracy_score(y_pred_svc, y_valid) * 100, 2)\nprint('The accuracy score of Support Vector Machines model is:', accuracy_svc)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:34:00.155440Z","iopub.execute_input":"2022-07-11T13:34:00.155850Z","iopub.status.idle":"2022-07-11T13:34:01.017939Z","shell.execute_reply.started":"2022-07-11T13:34:00.155817Z","shell.execute_reply":"2022-07-11T13:34:01.016649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Logistic Regression\nfrom sklearn.linear_model import LogisticRegression\n\nmodel_LR = LogisticRegression()\nmodel_LR.fit(X_train, y_train)\ny_pred_LR = model_LR.predict(X_valid)\naccuracy_LR = round(accuracy_score(y_pred_LR, y_valid) * 100, 2)\nprint('The accuracy score of Logistic Regression model is:', accuracy_LR)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:34:03.852503Z","iopub.execute_input":"2022-07-11T13:34:03.852910Z","iopub.status.idle":"2022-07-11T13:34:03.899258Z","shell.execute_reply.started":"2022-07-11T13:34:03.852876Z","shell.execute_reply":"2022-07-11T13:34:03.898348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Stochastic Gradient Descent\nfrom sklearn.linear_model import SGDClassifier\n\nmodel_SGD = SGDClassifier()\nmodel_SGD.fit(X_train, y_train)\ny_pred_SGD = model_SGD.predict(X_valid)\naccuracy_SGD = round(accuracy_score(y_pred_SGD, y_valid) * 100, 2)\nprint('The accuracy score of Stochastic Gradient Descent model is:', accuracy_SGD)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:34:07.317017Z","iopub.execute_input":"2022-07-11T13:34:07.317436Z","iopub.status.idle":"2022-07-11T13:34:07.334767Z","shell.execute_reply.started":"2022-07-11T13:34:07.317386Z","shell.execute_reply":"2022-07-11T13:34:07.333620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Gradient Boosting Classifier\nfrom sklearn.ensemble import GradientBoostingClassifier\n\nmodel_GB = GradientBoostingClassifier()\nmodel_GB.fit(X_train, y_train)\ny_pred_GB = model_GB.predict(X_valid)\naccuracy_GB = round(accuracy_score(y_pred_GB, y_valid) * 100, 2)\nprint('The accuracy score of Gradient Boosting Classifier model is:', accuracy_GB)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:34:12.947265Z","iopub.execute_input":"2022-07-11T13:34:12.948080Z","iopub.status.idle":"2022-07-11T13:34:13.067249Z","shell.execute_reply.started":"2022-07-11T13:34:12.948035Z","shell.execute_reply":"2022-07-11T13:34:13.065991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Above are the accuracy scores of all selected models, let's compare the accuracy of each model.","metadata":{}},{"cell_type":"code","source":"models = pd.DataFrame({\n    'Model': ['Decision Tree', 'Random Forest', 'KNN', 'Gaussian Naive Bayes', \\\n              'Support Vector Machines', 'Logistic Regression', \\\n              'Stochastic Gradient Descent', 'Gradient Boosting Classifier'],\n    'Score': [accuracy_DT, accuracy_RF, accuracy_KNN, accuracy_NB, accuracy_svc, \\\n              accuracy_LR, accuracy_SGD, accuracy_GB]})\nmodels.sort_values(by = 'Score', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:34:19.207441Z","iopub.execute_input":"2022-07-11T13:34:19.207881Z","iopub.status.idle":"2022-07-11T13:34:19.224945Z","shell.execute_reply.started":"2022-07-11T13:34:19.207833Z","shell.execute_reply":"2022-07-11T13:34:19.223654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the table above, it can be find that the accuracy score of **Gradient Boosting Classifier model** is the highest (sometimes is **Random Forest model**), so we'll try to use these two models to predict the survival value for testing data, to see which one can get more accuracy result.","metadata":{}},{"cell_type":"markdown","source":"<a id = \"10\"></a><br>\n# Generate Prediction Value File\n\nIn this section, we can generate a prediction value file about the survival information in the testing dataset.","metadata":{}},{"cell_type":"code","source":"# Gradient Boosting Classifier model\nmodel_pred_GB = GradientBoostingClassifier()\nmodel_pred_GB.fit(X_Variable, y_Target)\ny_predictions_GB = model_pred_GB.predict(pro_test_data.drop(['Survived', 'PassengerId'], \\\n                                                      axis = 1)).astype(int)\nprint(y_predictions_GB)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:52:57.243550Z","iopub.execute_input":"2022-07-11T13:52:57.243888Z","iopub.status.idle":"2022-07-11T13:52:57.375887Z","shell.execute_reply.started":"2022-07-11T13:52:57.243861Z","shell.execute_reply":"2022-07-11T13:52:57.374881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Random Forest model\nmodel_pred_RF = RandomForestClassifier()\nmodel_pred_RF.fit(X_Variable, y_Target)\ny_predictions_RF = model_pred_RF.predict(pro_test_data.drop(['Survived', 'PassengerId'], \\\n                                                      axis = 1)).astype(int)\n# print(y_predictions_RF)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#set the output as a dataframe and convert to csv file named submission.csv\n\noutput1 = pd.DataFrame({'PassengerId': test_data['PassengerId'], \\\n                        'Survived': y_predictions_GB})\n#output2 = pd.DataFrame({'PassengerId': test_data['PassengerId'], \\\n#                        'Survived': y_predictions_RF})\n\noutput1.to_csv('submission.csv', index=False)\noutput1","metadata":{"execution":{"iopub.status.busy":"2022-07-11T13:53:01.977916Z","iopub.execute_input":"2022-07-11T13:53:01.978268Z","iopub.status.idle":"2022-07-11T13:53:01.991749Z","shell.execute_reply.started":"2022-07-11T13:53:01.978237Z","shell.execute_reply":"2022-07-11T13:53:01.990711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Well, we predicted one version of the survival value for people listed in the testing dataset. A part of the result is shown above.","metadata":{}},{"cell_type":"markdown","source":"<font size=\"+2\" color=purple ><b><u> Upvote If you like it !!! 😄🔝</u></b></font>","metadata":{}}]}