{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"I am new to the world of machine learning. I'm trying to improve myself by starting with the Titanic survival prediction. I would appreciate it if you could post any shortcomings in the notebook or any issues where I can improve myself as a comment.\n\nNotebook Contents:\n<a id=\"toc\"></a>\n\n1. [Import Necessary Libraries](#1)\n\n2. [Loading Data](#2)\n\n3. [Exploratory Data Analysis](#3)\n\n    3.1 `Pclass`\n    \n    3.2 `Age`\n    \n    3.3 `Sex`\n    \n    3.4 `Embarked`\n    \n    3.5 `Fare`\n    \n    3.6 `ibsp` and `parch`\n    \n    3.7 `Name `\n    \n4. [Cleaning Data](#4)\n\n5. [Models](#5)\n\n    5.1 Adaboost\n    \n    5.2 Random Forest Model\n    \n    5.3 Gradient Tree Boosting\n    \n    5.4 Ensemble Learning\n    \n    5.5 Voting Classifier\n    \n    5.6 Bagging Classifier\n\n6. [Submission](#6)","metadata":{}},{"cell_type":"markdown","source":"# **Part 1** <a id=\"1\"></a>\n# Importing necessary libraries","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport warnings\nwarnings.simplefilter(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:14.424423Z","iopub.execute_input":"2022-08-12T10:11:14.424831Z","iopub.status.idle":"2022-08-12T10:11:15.581127Z","shell.execute_reply.started":"2022-08-12T10:11:14.424735Z","shell.execute_reply":"2022-08-12T10:11:15.580056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Part 2** <a id=\"2\"></a>\n\n# Loading Data","metadata":{}},{"cell_type":"markdown","source":"Analyze both training and testing data to get better understanding of the data but first let's read the train data as `df_train` and test data as `df_test` and concat these dataframes to `df` to get the full data.","metadata":{}},{"cell_type":"code","source":"df_train = pd.read_csv(\"../input/titanic/train.csv\")\ndf_test = pd.read_csv(\"../input/titanic/test.csv\")\n\ndf_test_copy = df_test.copy()\n\ndf = pd.concat([df_train, df_test])","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:15.583245Z","iopub.execute_input":"2022-08-12T10:11:15.583665Z","iopub.status.idle":"2022-08-12T10:11:15.632370Z","shell.execute_reply.started":"2022-08-12T10:11:15.583605Z","shell.execute_reply":"2022-08-12T10:11:15.628093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:15.633723Z","iopub.execute_input":"2022-08-12T10:11:15.634672Z","iopub.status.idle":"2022-08-12T10:11:15.662616Z","shell.execute_reply.started":"2022-08-12T10:11:15.634590Z","shell.execute_reply":"2022-08-12T10:11:15.661351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:15.664888Z","iopub.execute_input":"2022-08-12T10:11:15.665457Z","iopub.status.idle":"2022-08-12T10:11:15.694920Z","shell.execute_reply.started":"2022-08-12T10:11:15.665406Z","shell.execute_reply":"2022-08-12T10:11:15.693985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have 1309 rows and 12 data columns in of `df` dataframe.","metadata":{}},{"cell_type":"markdown","source":"# **Part 3**  <a id=\"3\"></a>\n# Exploratory Data Analysis","metadata":{}},{"cell_type":"markdown","source":"`Pclass` is Ticket class which has three values 1, 2 or 3.\n`Name`, `Sex` and `Age` features explain themselves.\n`SibSp` is the number of the siblings / spouses aboard the Titanic.\n`Parch` is the number of the parents / children aboard the Titanic.\n`Ticket` is the ticket number of the passenger.\n`Fare` is the passenger fare which is a numerical feature.\n`Cabin` is the cabin number of the passenger.\n`Embarked` is port of embarkation and it is a categorical feature which has 3 values C, Q or S.","metadata":{}},{"cell_type":"code","source":"df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:15.696393Z","iopub.execute_input":"2022-08-12T10:11:15.697382Z","iopub.status.idle":"2022-08-12T10:11:15.709374Z","shell.execute_reply.started":"2022-08-12T10:11:15.697335Z","shell.execute_reply":"2022-08-12T10:11:15.708301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First, let's take a look at the summary of train and test data separately. Immediately, we see that `Age`, `Cabin` and `Embarked` have missing values in the train data that we'll have to deal with.","metadata":{}},{"cell_type":"code","source":"# Visualizing the Number of missing values in train data\nplt.figure(figsize = (15,5))\nplt.bar(df_train.columns, df_train.isna().sum(),color='red' )\nplt.xlabel(\"Columns name\")\nplt.ylabel(\"Number of missing values in data\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:15.711053Z","iopub.execute_input":"2022-08-12T10:11:15.711292Z","iopub.status.idle":"2022-08-12T10:11:16.009574Z","shell.execute_reply.started":"2022-08-12T10:11:15.711265Z","shell.execute_reply":"2022-08-12T10:11:16.008678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Visualizing the Number of missing values in test data\nplt.figure(figsize = (15,5))\nplt.bar(df_test.columns, df_test.isna().sum())\nplt.xlabel(\"Columns name\")\nplt.ylabel(\"Number of missing values in data\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:16.010887Z","iopub.execute_input":"2022-08-12T10:11:16.011753Z","iopub.status.idle":"2022-08-12T10:11:16.247540Z","shell.execute_reply.started":"2022-08-12T10:11:16.011711Z","shell.execute_reply":"2022-08-12T10:11:16.246846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have also see `Age` , `Fare` and `Cabin` columns have null values in the test data.","metadata":{}},{"cell_type":"markdown","source":"## 3.1 Pclass","metadata":{}},{"cell_type":"markdown","source":"Now, let's investigate the the importance of the `Pclass` feature.","metadata":{}},{"cell_type":"code","source":"ax = sns.countplot('Pclass',hue='Survived',data=df)\nfor p in ax.patches:\n   ax.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0.25, p.get_height()+0.01))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:16.248790Z","iopub.execute_input":"2022-08-12T10:11:16.249573Z","iopub.status.idle":"2022-08-12T10:11:16.499702Z","shell.execute_reply.started":"2022-08-12T10:11:16.249514Z","shell.execute_reply":"2022-08-12T10:11:16.498853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Although the rescue number seems to be high for Pclass 3, when we look at the rate, the rescue rate for passengers with Pclass 1 is higher. 136/216 for Pclass 1 which is about 63%, for Pclass 2 87/184 which is about 47% and lastly Pclass 3 passengers have a survival rate of about 25%.","metadata":{}},{"cell_type":"markdown","source":"## 3.2 Age","metadata":{}},{"cell_type":"markdown","source":"We will fill the missing `Age` values with its mean and by doing `Sex` separation.","metadata":{}},{"cell_type":"code","source":"df_train.groupby('Sex')['Age'].mean() ","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:16.501035Z","iopub.execute_input":"2022-08-12T10:11:16.501920Z","iopub.status.idle":"2022-08-12T10:11:16.513085Z","shell.execute_reply.started":"2022-08-12T10:11:16.501879Z","shell.execute_reply":"2022-08-12T10:11:16.512218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.loc[(df_train.Age.isnull())&(df_train.Sex=='female'),'Age']= 27.915709\ndf_train.loc[(df_train.Age.isnull())&(df_train.Sex=='male'),'Age']= 30.726645","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:16.516096Z","iopub.execute_input":"2022-08-12T10:11:16.516569Z","iopub.status.idle":"2022-08-12T10:11:16.526446Z","shell.execute_reply.started":"2022-08-12T10:11:16.516480Z","shell.execute_reply":"2022-08-12T10:11:16.525843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.Age.isnull().any() #So there are null values left in Age feature of the train data","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:16.527664Z","iopub.execute_input":"2022-08-12T10:11:16.528482Z","iopub.status.idle":"2022-08-12T10:11:16.538522Z","shell.execute_reply.started":"2022-08-12T10:11:16.528441Z","shell.execute_reply":"2022-08-12T10:11:16.537845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test.groupby('Sex')['Age'].mean() ","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:16.540022Z","iopub.execute_input":"2022-08-12T10:11:16.540462Z","iopub.status.idle":"2022-08-12T10:11:16.552281Z","shell.execute_reply.started":"2022-08-12T10:11:16.540367Z","shell.execute_reply":"2022-08-12T10:11:16.551268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test.loc[(df_test.Age.isnull())&(df_test.Sex=='female'),'Age']= 30.272362\ndf_test.loc[(df_test.Age.isnull())&(df_test.Sex=='male'),'Age']= 30.272732","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:16.553976Z","iopub.execute_input":"2022-08-12T10:11:16.554520Z","iopub.status.idle":"2022-08-12T10:11:16.565693Z","shell.execute_reply.started":"2022-08-12T10:11:16.554478Z","shell.execute_reply":"2022-08-12T10:11:16.565029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test.Age.isnull().any() #So there are null values left in Age feature of the test data","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:16.567221Z","iopub.execute_input":"2022-08-12T10:11:16.567505Z","iopub.status.idle":"2022-08-12T10:11:16.578070Z","shell.execute_reply.started":"2022-08-12T10:11:16.567462Z","shell.execute_reply":"2022-08-12T10:11:16.577270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3.3 Sex","metadata":{}},{"cell_type":"markdown","source":"Visualizing the number of survived passenger according to `Sex`, as we can see from the below figure the most of the survivors is from females so the `Sex` of the passengers is an important feature. We will map the females as `0` and males as `1`.","metadata":{}},{"cell_type":"code","source":"ax = sns.countplot('Sex',hue='Survived',data=df_train)\nfor p in ax.patches:\n   ax.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0.25, p.get_height()+0.01))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:16.579184Z","iopub.execute_input":"2022-08-12T10:11:16.579419Z","iopub.status.idle":"2022-08-12T10:11:16.796471Z","shell.execute_reply.started":"2022-08-12T10:11:16.579389Z","shell.execute_reply":"2022-08-12T10:11:16.795576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['Sex'] = df_train['Sex'].map( {'female': 0, 'male': 1} ).astype(int)\ndf_test['Sex'] = df_test['Sex'].map( {'female': 0, 'male': 1} ).astype(int)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:16.798006Z","iopub.execute_input":"2022-08-12T10:11:16.798321Z","iopub.status.idle":"2022-08-12T10:11:16.808316Z","shell.execute_reply.started":"2022-08-12T10:11:16.798278Z","shell.execute_reply":"2022-08-12T10:11:16.807429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:16.809344Z","iopub.execute_input":"2022-08-12T10:11:16.809930Z","iopub.status.idle":"2022-08-12T10:11:16.830841Z","shell.execute_reply.started":"2022-08-12T10:11:16.809897Z","shell.execute_reply":"2022-08-12T10:11:16.829932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:16.831985Z","iopub.execute_input":"2022-08-12T10:11:16.832294Z","iopub.status.idle":"2022-08-12T10:11:16.847217Z","shell.execute_reply.started":"2022-08-12T10:11:16.832264Z","shell.execute_reply":"2022-08-12T10:11:16.846580Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3.4 Embarked","metadata":{}},{"cell_type":"markdown","source":"`Embarked` feature has missing values only on training data. We will replace the missing values with common used ones. The most common Port of Embarkation is 'S'. ","metadata":{}},{"cell_type":"code","source":"df_train['Embarked'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:16.848138Z","iopub.execute_input":"2022-08-12T10:11:16.848684Z","iopub.status.idle":"2022-08-12T10:11:16.856223Z","shell.execute_reply.started":"2022-08-12T10:11:16.848639Z","shell.execute_reply":"2022-08-12T10:11:16.855329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax= sns.countplot('Embarked',hue='Survived',data=df)\nfor p in ax.patches:\n   ax.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0.25, p.get_height()+0.01))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:16.857489Z","iopub.execute_input":"2022-08-12T10:11:16.858255Z","iopub.status.idle":"2022-08-12T10:11:17.111463Z","shell.execute_reply.started":"2022-08-12T10:11:16.858219Z","shell.execute_reply":"2022-08-12T10:11:17.110669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.loc[(df_train.Embarked.isnull()),'Embarked']= 'S'","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.112932Z","iopub.execute_input":"2022-08-12T10:11:17.113153Z","iopub.status.idle":"2022-08-12T10:11:17.119559Z","shell.execute_reply.started":"2022-08-12T10:11:17.113126Z","shell.execute_reply":"2022-08-12T10:11:17.118364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.Embarked.isnull().any() #So no null values left finally in the Embarked feature","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.120910Z","iopub.execute_input":"2022-08-12T10:11:17.121243Z","iopub.status.idle":"2022-08-12T10:11:17.133864Z","shell.execute_reply.started":"2022-08-12T10:11:17.121213Z","shell.execute_reply":"2022-08-12T10:11:17.132925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['Embarked'] = df_train['Embarked'].map( {'S': 0, 'C': 1, 'Q': 2} ).astype(int)\ndf_test['Embarked'] = df_test['Embarked'].map( {'S': 0, 'C': 1, 'Q': 2} ).astype(int)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.135122Z","iopub.execute_input":"2022-08-12T10:11:17.135352Z","iopub.status.idle":"2022-08-12T10:11:17.147955Z","shell.execute_reply.started":"2022-08-12T10:11:17.135326Z","shell.execute_reply":"2022-08-12T10:11:17.147240Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3.5 Fare","metadata":{}},{"cell_type":"markdown","source":"Filling missing Fare values with median in test data. ","metadata":{}},{"cell_type":"code","source":"df_test[\"Fare\"].fillna(df_test[\"Fare\"].median(), inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.149030Z","iopub.execute_input":"2022-08-12T10:11:17.149375Z","iopub.status.idle":"2022-08-12T10:11:17.160059Z","shell.execute_reply.started":"2022-08-12T10:11:17.149346Z","shell.execute_reply":"2022-08-12T10:11:17.159384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['FareBand'] = pd.qcut(df_train['Fare'], 4, labels = [1, 2, 3, 4])\ndf_test['FareBand'] = pd.qcut(df_test['Fare'], 4, labels = [1, 2, 3, 4])","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.161259Z","iopub.execute_input":"2022-08-12T10:11:17.162008Z","iopub.status.idle":"2022-08-12T10:11:17.177135Z","shell.execute_reply.started":"2022-08-12T10:11:17.161975Z","shell.execute_reply":"2022-08-12T10:11:17.176187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3.6 ibsp and parch","metadata":{}},{"cell_type":"markdown","source":"Let's examine ibsp and parch features. ibsp is # of siblings / spouses aboard the Titanic and parch feature is # of parents / children aboard the Titanic. ","metadata":{}},{"cell_type":"code","source":"df_train['FamilySize'] = df_train['SibSp'] + df_train['Parch'] + 1\ndf_test['FamilySize'] = df_test['SibSp'] + df_test['Parch'] + 1","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.178678Z","iopub.execute_input":"2022-08-12T10:11:17.179081Z","iopub.status.idle":"2022-08-12T10:11:17.188378Z","shell.execute_reply.started":"2022-08-12T10:11:17.179014Z","shell.execute_reply":"2022-08-12T10:11:17.187453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Adding the column named 'IsAlone' if the 'FamilySize' is equal to 1 the 'IsAlone' is 1.","metadata":{}},{"cell_type":"code","source":"df_train['IsAlone'] = 0\ndf_train.loc[df_train['FamilySize'] == 1, 'IsAlone'] = 1\n\ndf_test['IsAlone'] = 0\ndf_test.loc[df_test['FamilySize'] == 1, 'IsAlone'] = 1","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.189619Z","iopub.execute_input":"2022-08-12T10:11:17.190282Z","iopub.status.idle":"2022-08-12T10:11:17.203977Z","shell.execute_reply.started":"2022-08-12T10:11:17.190241Z","shell.execute_reply":"2022-08-12T10:11:17.203268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax= sns.countplot('IsAlone',hue='Survived',data=df_train)\nfor p in ax.patches:\n   ax.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0.25, p.get_height()+0.01))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.209998Z","iopub.execute_input":"2022-08-12T10:11:17.210721Z","iopub.status.idle":"2022-08-12T10:11:17.429915Z","shell.execute_reply.started":"2022-08-12T10:11:17.210685Z","shell.execute_reply":"2022-08-12T10:11:17.428937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3.7 Name","metadata":{}},{"cell_type":"code","source":"#create a combined group of both datasets\ndf = [df_train, df_test]\n\nfor data in df:\n    data['Title'] = data.Name.str.extract(' ([A-Za-z]+)\\.', expand=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.431233Z","iopub.execute_input":"2022-08-12T10:11:17.432115Z","iopub.status.idle":"2022-08-12T10:11:17.440868Z","shell.execute_reply.started":"2022-08-12T10:11:17.432077Z","shell.execute_reply":"2022-08-12T10:11:17.439743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for data in df:\n    data['Title'] = data['Title'].replace(['Lady', 'Capt', 'Col',\n    'Don', 'Dr', 'Major', 'Rev', 'Jonkheer', 'Dona'], 'Rare')\n    data['Title'] = data['Title'].replace(['Countess', 'Lady', 'Sir'], 'Royal')\n    data['Title'] = data['Title'].replace('Mlle', 'Miss')\n    data['Title'] = data['Title'].replace('Ms', 'Miss')\n    data['Title'] = data['Title'].replace('Mme', 'Mrs')","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.442290Z","iopub.execute_input":"2022-08-12T10:11:17.442598Z","iopub.status.idle":"2022-08-12T10:11:17.459922Z","shell.execute_reply.started":"2022-08-12T10:11:17.442568Z","shell.execute_reply":"2022-08-12T10:11:17.458896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"titles= {\"Mr\": 1, \"Miss\": 2, \"Mrs\": 3, \"Master\": 4, \"Royal\": 5, \"Rare\": 6}\nfor data in df:\n    data['Title'] = data['Title'].map(titles)\n    data['Title'] = data['Title'].fillna(0)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.461064Z","iopub.execute_input":"2022-08-12T10:11:17.461914Z","iopub.status.idle":"2022-08-12T10:11:17.469408Z","shell.execute_reply.started":"2022-08-12T10:11:17.461876Z","shell.execute_reply":"2022-08-12T10:11:17.468822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.470709Z","iopub.execute_input":"2022-08-12T10:11:17.471068Z","iopub.status.idle":"2022-08-12T10:11:17.498906Z","shell.execute_reply.started":"2022-08-12T10:11:17.471033Z","shell.execute_reply":"2022-08-12T10:11:17.498004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Part 4**  <a id=\"4\"></a>\n# Cleaning Data","metadata":{}},{"cell_type":"markdown","source":"Dropping some non-useful columns.","metadata":{}},{"cell_type":"code","source":"df_train.drop(columns=['PassengerId','Cabin', 'Ticket','Name','SibSp','Parch','Cabin','FamilySize','Age'], axis=1, inplace = True)\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.500708Z","iopub.execute_input":"2022-08-12T10:11:17.501259Z","iopub.status.idle":"2022-08-12T10:11:17.517344Z","shell.execute_reply.started":"2022-08-12T10:11:17.501210Z","shell.execute_reply":"2022-08-12T10:11:17.516425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test.drop(columns=['PassengerId','Cabin', 'Ticket','Name','SibSp','Parch','Cabin','FamilySize','Age'], axis=1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.518979Z","iopub.execute_input":"2022-08-12T10:11:17.519495Z","iopub.status.idle":"2022-08-12T10:11:17.535353Z","shell.execute_reply.started":"2022-08-12T10:11:17.519451Z","shell.execute_reply":"2022-08-12T10:11:17.534553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.536730Z","iopub.execute_input":"2022-08-12T10:11:17.537401Z","iopub.status.idle":"2022-08-12T10:11:17.556933Z","shell.execute_reply.started":"2022-08-12T10:11:17.537360Z","shell.execute_reply":"2022-08-12T10:11:17.556141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Part 5**  <a id=\"5\"></a>\n# Models","metadata":{}},{"cell_type":"code","source":"df_train_X = df_train.drop(['Survived'],axis=1)\ndf_train_y = df_train['Survived']","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.558300Z","iopub.execute_input":"2022-08-12T10:11:17.558784Z","iopub.status.idle":"2022-08-12T10:11:17.565843Z","shell.execute_reply.started":"2022-08-12T10:11:17.558743Z","shell.execute_reply":"2022-08-12T10:11:17.564976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Train test split\n","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nX_train, X_test, y_train, y_test = train_test_split(df_train_X, \n                                                    df_train_y, test_size=0.20, \n                                                    random_state=101, stratify = df_train_y)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.567203Z","iopub.execute_input":"2022-08-12T10:11:17.568032Z","iopub.status.idle":"2022-08-12T10:11:17.787465Z","shell.execute_reply.started":"2022-08-12T10:11:17.567982Z","shell.execute_reply":"2022-08-12T10:11:17.786516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.1 Adaboost","metadata":{}},{"cell_type":"markdown","source":"Boosting: Ensemble method combining several weak learners to form a strong learner.","metadata":{}},{"cell_type":"markdown","source":"Most popular boosting methods:\n1. AdaBoost\n1. Gradient Boosting","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import AdaBoostClassifier\nada = AdaBoostClassifier(n_estimators=110, random_state=0)\nada.fit(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:17.788760Z","iopub.execute_input":"2022-08-12T10:11:17.789025Z","iopub.status.idle":"2022-08-12T10:11:18.318910Z","shell.execute_reply.started":"2022-08-12T10:11:17.788993Z","shell.execute_reply":"2022-08-12T10:11:18.317886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = ada.predict(X_test)\npredictions_ada = ada.predict(df_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.320136Z","iopub.execute_input":"2022-08-12T10:11:18.320382Z","iopub.status.idle":"2022-08-12T10:11:18.372174Z","shell.execute_reply.started":"2022-08-12T10:11:18.320352Z","shell.execute_reply":"2022-08-12T10:11:18.371150Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Adaboost_Test_score = ada.score(X_test, y_test)\n\nprint('Test set accuracy of ada: {:.3f}'.format(Adaboost_Test_score))","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.373462Z","iopub.execute_input":"2022-08-12T10:11:18.373737Z","iopub.status.idle":"2022-08-12T10:11:18.404314Z","shell.execute_reply.started":"2022-08-12T10:11:18.373702Z","shell.execute_reply":"2022-08-12T10:11:18.403344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import classification_report\nprint(classification_report(y_test,predictions))","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.405869Z","iopub.execute_input":"2022-08-12T10:11:18.406721Z","iopub.status.idle":"2022-08-12T10:11:18.419456Z","shell.execute_reply.started":"2022-08-12T10:11:18.406674Z","shell.execute_reply":"2022-08-12T10:11:18.418010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.2 Random Forest Model","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\n\nrf = RandomForestClassifier(n_estimators=100)\n\nrf.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.422942Z","iopub.execute_input":"2022-08-12T10:11:18.423196Z","iopub.status.idle":"2022-08-12T10:11:18.615475Z","shell.execute_reply.started":"2022-08-12T10:11:18.423163Z","shell.execute_reply":"2022-08-12T10:11:18.614596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random_forest_Test_score = rf.score(X_test, y_test)\n\nprint('Test set accuracy of rf: {:.3f}'.format(random_forest_Test_score))","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.616856Z","iopub.execute_input":"2022-08-12T10:11:18.617168Z","iopub.status.idle":"2022-08-12T10:11:18.643678Z","shell.execute_reply.started":"2022-08-12T10:11:18.617126Z","shell.execute_reply":"2022-08-12T10:11:18.642774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = rf.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.644934Z","iopub.execute_input":"2022-08-12T10:11:18.645314Z","iopub.status.idle":"2022-08-12T10:11:18.665668Z","shell.execute_reply.started":"2022-08-12T10:11:18.645281Z","shell.execute_reply":"2022-08-12T10:11:18.664968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(classification_report(y_test,predictions))","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.666859Z","iopub.execute_input":"2022-08-12T10:11:18.667247Z","iopub.status.idle":"2022-08-12T10:11:18.676423Z","shell.execute_reply.started":"2022-08-12T10:11:18.667210Z","shell.execute_reply":"2022-08-12T10:11:18.675456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.3 Gradient Tree Boosting","metadata":{}},{"cell_type":"markdown","source":"`GradientBoostingClassifier` supports both binary and multi-class classification. The following example shows how to fit a gradient boosting classifier with 100 decision stumps as weak learners.","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import GradientBoostingClassifier\nclf = GradientBoostingClassifier(n_estimators=90, learning_rate=1,max_depth=2, random_state=0)\nclf.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.677463Z","iopub.execute_input":"2022-08-12T10:11:18.677776Z","iopub.status.idle":"2022-08-12T10:11:18.756425Z","shell.execute_reply.started":"2022-08-12T10:11:18.677746Z","shell.execute_reply":"2022-08-12T10:11:18.755585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"GradientBoosting_Test_score = clf.score(X_test, y_test)\n\nprint('Test set accuracy of clf: {:.3f}'.format(GradientBoosting_Test_score))","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.757471Z","iopub.execute_input":"2022-08-12T10:11:18.757712Z","iopub.status.idle":"2022-08-12T10:11:18.766317Z","shell.execute_reply.started":"2022-08-12T10:11:18.757684Z","shell.execute_reply":"2022-08-12T10:11:18.765368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = clf.predict(X_test)\npredictions_clf = clf.predict(df_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.767514Z","iopub.execute_input":"2022-08-12T10:11:18.768145Z","iopub.status.idle":"2022-08-12T10:11:18.781988Z","shell.execute_reply.started":"2022-08-12T10:11:18.768098Z","shell.execute_reply":"2022-08-12T10:11:18.780929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(classification_report(y_test,predictions))","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.783091Z","iopub.execute_input":"2022-08-12T10:11:18.783940Z","iopub.status.idle":"2022-08-12T10:11:18.798826Z","shell.execute_reply.started":"2022-08-12T10:11:18.783901Z","shell.execute_reply":"2022-08-12T10:11:18.797844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.4 Ensemble Learning","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.neighbors import KNeighborsClassifier as KNN\n\n# Set seed for reproducibility\nSEED = 1\n\n# Instantiate lr\nlr = LogisticRegression(random_state=SEED)\n\n# Instantiate knn\nknn = KNN(n_neighbors=3)\n\n# Instantiate dt\ndt = DecisionTreeClassifier(min_samples_leaf=0.26, random_state=SEED)\n\n# Define the list classifiers\nclassifiers = [\n    ('Logistic Regression', lr),\n    ('K Nearest Neighbors', knn),\n    ('Classification Tree', dt)\n]","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.801835Z","iopub.execute_input":"2022-08-12T10:11:18.802110Z","iopub.status.idle":"2022-08-12T10:11:18.808412Z","shell.execute_reply.started":"2022-08-12T10:11:18.802080Z","shell.execute_reply":"2022-08-12T10:11:18.807488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import accuracy_score\n\n# Iterate over the pre-defined list of classifiers\nfor clf_name, clf in classifiers:\n    # Fit clf to the training set\n    clf.fit(X_train, y_train)\n    \n    # Predict y_pred\n    y_pred = clf.predict(X_test)\n    \n    # Calculate accuracy\n    accuracy = accuracy_score(y_test, y_pred)\n    \n    # Evaluate clf's accuracy on the test set\n    print('{:s} : {:.3f}'.format(clf_name, accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.809770Z","iopub.execute_input":"2022-08-12T10:11:18.810011Z","iopub.status.idle":"2022-08-12T10:11:18.884017Z","shell.execute_reply.started":"2022-08-12T10:11:18.809981Z","shell.execute_reply":"2022-08-12T10:11:18.883114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.5 Voting Classifier","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import VotingClassifier\n\n# VotingClassifier vc\nvc = VotingClassifier(estimators=classifiers)\n\n# Fit\nvc.fit(X_train, y_train)\n\n# Evaluate the test set predictions\ny_pred = vc.predict(X_test)\n\n# Calculate accuracy score\naccuracy = accuracy_score(y_test, y_pred)\nprint('Voting Classifier: {:.3f}'.format(accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.885234Z","iopub.execute_input":"2022-08-12T10:11:18.885465Z","iopub.status.idle":"2022-08-12T10:11:18.943479Z","shell.execute_reply.started":"2022-08-12T10:11:18.885437Z","shell.execute_reply":"2022-08-12T10:11:18.942421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5.6 Bagging Classifier","metadata":{}},{"cell_type":"code","source":"from sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import BaggingClassifier\n\n# Instantiate dt\ndt = DecisionTreeClassifier(random_state=1)\n\n# Instantiate bc\nbc = BaggingClassifier(base_estimator=dt, n_estimators=70, random_state=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.944807Z","iopub.execute_input":"2022-08-12T10:11:18.945125Z","iopub.status.idle":"2022-08-12T10:11:18.950952Z","shell.execute_reply.started":"2022-08-12T10:11:18.945082Z","shell.execute_reply":"2022-08-12T10:11:18.950096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import accuracy_score\n\n# Fit bc to the training set\nbc.fit(X_train, y_train)\n\n# Predict test set labels\ny_pred = bc.predict(X_test)\npredictions_bc = bc.predict(df_test)\n\n\n# Evaluate acc_test\nacc_test = accuracy_score(y_test, y_pred)\nprint('Test set accuracy of bc: {:.3f}'.format(acc_test))","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:18.952063Z","iopub.execute_input":"2022-08-12T10:11:18.952352Z","iopub.status.idle":"2022-08-12T10:11:19.150303Z","shell.execute_reply.started":"2022-08-12T10:11:18.952312Z","shell.execute_reply":"2022-08-12T10:11:19.149390Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Part 6**  <a id=\"6\"></a>\n# Submission","metadata":{}},{"cell_type":"code","source":"submission = pd.concat([df_test_copy['PassengerId'],pd.Series(predictions_ada,name=\"Survived\")],axis=1)\nsubmission.to_csv(\"./submission.csv\",index=False,header=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:19.151701Z","iopub.execute_input":"2022-08-12T10:11:19.151957Z","iopub.status.idle":"2022-08-12T10:11:19.161452Z","shell.execute_reply.started":"2022-08-12T10:11:19.151926Z","shell.execute_reply":"2022-08-12T10:11:19.160369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T10:11:19.162682Z","iopub.execute_input":"2022-08-12T10:11:19.162897Z","iopub.status.idle":"2022-08-12T10:11:19.178898Z","shell.execute_reply.started":"2022-08-12T10:11:19.162871Z","shell.execute_reply":"2022-08-12T10:11:19.177944Z"},"trusted":true},"execution_count":null,"outputs":[]}]}